czbenchmarks.datasets.single_cell ================================= .. py:module:: czbenchmarks.datasets.single_cell Attributes ---------- .. autoapisummary:: czbenchmarks.datasets.single_cell.logger Classes ------- .. autoapisummary:: czbenchmarks.datasets.single_cell.SingleCellDataset Module Contents --------------- .. py:data:: logger .. py:class:: SingleCellDataset(dataset_type_name: str, path: pathlib.Path, organism: czbenchmarks.datasets.types.Organism, task_inputs_dir: Optional[pathlib.Path] = None) Bases: :py:obj:`czbenchmarks.datasets.dataset.Dataset` Abstract base class for single cell datasets containing gene expression data. Handles loading and validation of AnnData objects with the following requirements: - Must have gene names in `adata.var['ensembl_id']` or `adata.var_names`. - Gene names must start with the organism prefix (e.g., "ENSG" for human). - Must contain raw counts in `adata.X` (non-negative integers). - Should be stored in H5AD format. .. attribute:: adata Loaded AnnData object containing gene expression data. :type: ad.AnnData Initialize a SingleCellDataset instance. :param dataset_type_name: Name of the dataset type (used for directory naming). :type dataset_type_name: str :param path: Path to the dataset file. :type path: Path :param organism: Enum value indicating the organism. :type organism: Organism :param task_inputs_dir: Directory for storing task-specific inputs. :type task_inputs_dir: Optional[Path] .. py:attribute:: adata :type: anndata.AnnData .. py:method:: load_data(backed: Literal['r', 'r+'] | bool | None = None) -> None Load the dataset from the path. This method reads the dataset file in H5AD format and loads it into the `adata` attribute as an AnnData object. :param backed: Whether to load the dataset into memory or use backed mode. Memory: False or None. Default is None. Backed: True, 'r' for read-only, 'r+' for read-write :type backed: Literal['r', 'r+'] | bool | None Populates: adata (ad.AnnData): Loaded AnnData object containing gene expression data.