Distributed self-adaptive multi-target genetic algorithm for performing single cell clustering and marker prediction by using high-dimensional data

The distributed adaptive multi-objective genetic algorithm directly processed single-cell oscillary high-dimensional data, which solved the problems of information loss and model overfitting caused by dimensionality reduction technology, and achieved effective detection and biological interpretation of complex and rare cell populations and gene markers.

CN119968641APending Publication Date: 2025-05-09MOHAMMED BIN RASHID UNIV OF MEDICINE & HEALTH SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380067246.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-26
Filing Date
2023-09-18
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art often uses dimensionality reduction techniques in single-cell omics data analysis, resulting in information loss and model overfitting, making it difficult to effectively detect complex and rare cell populations and gene markers.

Method used

A distributed adaptive multi-objective genetic algorithm is used to directly process high-dimensional data without data reduction, and single-cell clustering and feature extraction are optimized through fitness functions and genetic operators (cross, mutations, selection).

Benefits of technology

It effectively overcomes the problems of information loss and model overfitting, and can detect biologically interpretable cell types and related markers from high-dimensional data, improving the reproducibility and biological interpretability of the results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119968641A_ABST
    Figure CN119968641A_ABST
Patent Text Reader

Abstract

A method and system for detecting cell types and related markers, the method comprising: a) running a genetic algorithm in a distributed system, where each node is configured to: i) receive a matrix of feature expressions of cells; ii) initializing a set of solutions with a plurality of solutions derived from the matrix, each solution having a set of randomly allocated clusters of single cells; iii) calculating the fitness of each solution in the set relative to a plurality of targets; iv) applying a genetic operator to the set to produce a set of filial generations; v) calculating the fitness of each solution in the filial generation set; vi) comparing the fitness of the parent set and the fitness of the child set, and reserving the solution with the highest fitness for the next generation; vii) repeating the steps ii-vi by a predetermined algebra; b) selecting a solution set (180) with the highest fitness; c) extracting single cell clusters and features from the selected solution set (190).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to methods and systems for detecting cell types and markers associated therewith. Background Art

[0002] In recent years, single-cell sequencing has greatly advanced our understanding of biological systems. The revolution that has taken place in high-throughput technologies has provided us with deeper insights into the transcriptome at single-cell resolution, enabling a better understanding of cellular heterogeneity. For example, single-cell RNA sequencing (scRNA-seq) can identify complex and rare cell populations, visualize the trajectories of developing different cell lineages, and even reveal gene regulatory networks. However, known methods are applied to low-dimensional spaces, which will result in a certain amount of data loss when compressed to a specific number of dimensions. Single-cell omics (OMICs) have identified a large number of disease-related cell types in the past few years. With the advent of high-throughput omics, researchers are faced with the challenge of processing and analyzing multidimensional data sets obtained from DNA, RNA or peptide sequencing and their analysis techniques. It includes information about nucleotide / protein sequences and their expression, interactions and other related phenotypes.

[0003] Recent state-of-the-art approaches involve dimensionality reduction techniques, which result in loss of information about the variance present in single-cell omics data. Typically, low-dimensional data are employed when single-cell omics (scOMICs) data are used to detect complex and rare cell populations, cell lineages, and gene markers.

[0004] In contrast, preferred embodiments of the present invention disclosed herein use high-dimensional data without applying any data reduction methods to overcome problems associated with information loss.

[0005] In the current era of big data, there are many challenges in analyzing the high-dimensional data obtained. In fact, when the number of variables (i.e., number of transcripts, expression levels) exceeds the sample size (i.e., cells, tissues), the problem amplifies exponentially. This phenomenon is often referred to as the "curse of dimensionality."

[0006] In high-dimensional data, the observed "zero values" are increasingly sparse, which is difficult to overcome and deal with. For example, scRNA-seq measurements may contain cells that lack a certain molecular marker or a specific mapped read, making it more difficult to obtain representative data about the population. In addition, high correlations can be observed between two or several predictor variables, resulting in perfect multicollinearity. Other disadvantages of dimensionality are model overfitting, estimation instability, local convergence, and huge estimation errors, which adversely impair the reproducibility of the results. These challenges, as well as some others, require us to adopt advanced computational methods to deal with the scale and complexity of high-dimensional datasets.

[0007] The application of computational methods on massive and heterogeneous data requires compression and normalization before processing. It is known to use statistical methods to accurately and effectively reduce data through feature selection or dimensionality reduction. Feature selection provides the selection of the most relevant variables from the original data, while dimensionality reduction is done by projecting the low-dimensional components (k) into a new subspace as representatives of the original high-dimensional data (p). Dimensionality reduction methods such as principal component analysis (PCA), linear discriminant analysis (LDA), generalized discriminant analysis (GDA), etc. have been introduced to solve difficult problems involving high-dimensional data, each of which is applicable to the data mining process and various types of high-dimensional data.

[0008] A widely used linear transformation method for reducing the dimension in multi-feature data is PCA, which aims to present data in a new coordinate system where the limited variables of the data set can be expressed without significant errors. It transforms the data into principal components by establishing linear correlations between the original features of the data set. However, in attempts to compress and standardize data, the PCA method results in information loss, low readability and interpretability of the original features. In addition, one of the shortcomings of PCA is the use of normalized data instead of raw data or scaled data, otherwise the algorithm will not be able to find the best principal component and will be biased towards components with high variance. In addition, outliers in the data must be removed before processing. The assumption of orthogonality is another limitation of PCA application. Due to the orthogonal design between the principal components, the technique cannot decompose data with non-orthogonal axes. In practice, most studies consider high-variance PCA and ignore most low-variance PCA that leads to biased high-value signals (i.e., clustering of highly expressed features), and may not be able to capture low signals from the data (i.e., clustering of small rare low-expressing cells in scRNA-seq or other omics data).

[0009] Another problem with using PCA is that the resulting cell clusters are not biologically interpretable, i.e., they are not interpretable in terms of molecular biology. PCA-derived methods only cluster cells because they have similar variance but do not explain why this variance exists in the data using knowledge of molecular biology. Scientists apply many downstream analyses to identify the causes of cell clustering and try to establish correlations between cell clustering and biology.

[0010] Since applying PCA for dimensionality reduction before clustering reduces the clusters, dimensionality reduction methods are usually avoided in prior art methods to prevent the consequence and complexity of losing information from the obtained data and to avoid significant bias in the interpretation of the results.

[0011] It would be desirable to alleviate at least some of the problems discussed above. Summary of the invention

[0012] The present invention provides a method for detecting cell types and related markers thereof, the method comprising:

[0013] a. Run multiple instances of the algorithm in a distributed system, where each instance is configured according to the following steps:

[0014] i. receiving a two-dimensional input matrix M, wherein each vertical vector of the input matrix represents a single cell, and each value in the vector represents a feature expression corresponding to the single cell;

[0015] ii. Initialize the parent solution set S1, S2, ..., S from the matrix n , where n is a predefined finite number, wherein the solution set includes a plurality of solutions S1, S2, ..., S derived from the matrix n , where each solution S i A set of randomly assigned clusters C1, C2, ..., C with single cells k , where k is the total number of clusters in a given solution;

[0016] iii. Use the fitness function to calculate each solution S in the solution set i A fitness function derived from a plurality of objectives;

[0017] iv. applying a genetic operator to the solution set to produce a set of offspring solutions;

[0018] v. using the fitness function to calculate the fitness of each solution in the offspring set;

[0019] vi. Compare the fitness of each solution in the solution set and the offspring set, where each solution with the highest fitness is retained in the new solution set for the next generation;

[0020] vii. Repeating steps iii, iv, v and vi for a predetermined number of generations;

[0021] b. Selecting the solution set with the highest fitness among the multiple instances;

[0022] c. Extract single-cell clusters and features from the selected solution set.

[0023] The present invention also provides a system for detecting cell types and related markers thereof, the system comprising a device for performing the following steps:

[0024] a. Run multiple instances of the algorithm in a distributed system, where each instance is configured according to the following steps:

[0025] i. receiving a two-dimensional input matrix, wherein each vertical vector of the input matrix represents a single cell, and each value in the vector represents a feature expression corresponding to the single cell;

[0026] ii. Initialize the parent solution set S1, S2, ..., S from the matrix n , where n is a predefined finite number, wherein the solution set includes a plurality of solutions S1, S2, ..., S derived from the matrix n , where each solution S i A set of randomly assigned clusters C1, C2, ..., C with single cells k , where n is the total number of clusters in a given solution;

[0027] iii. Use the fitness function to calculate each solution S in the solution set i A fitness function derived from a plurality of objectives;

[0028] iv. applying a genetic operator to the solution set to produce a set of offspring solutions;

[0029] v. Use the fitness function to calculate the fitness of each solution in the offspring set;

[0030] vi. Compare the fitness of each solution in the solution set and the offspring set, where each solution with the highest fitness is retained in the new solution set for the next generation;

[0031] vii. Repeating steps iii, iv, v and vi for a predetermined number of generations;

[0032] b. Selecting the solution set with the highest fitness among the multiple instances;

[0033] c. Extract single-cell clusters and features from the selected solution set.

[0034] In a preferred embodiment, the fitness function is derived from two objectives.

[0035] In preferred embodiments, the matrix comprises raw count or log2 transcriptomic data or proteomic data.

[0036] In a preferred embodiment, the relative trajectory of a cell type is calculated as the product of marker gene regulation and its correlation, where the trajectory of a cluster preferably defines whether the cluster's regulation is up or down compared to other cell types.

[0037] In another preferred embodiment, the solution i Each cluster C in i The trajectory is

[0038] In another preferred embodiment, the solution set selected in step (b) is selected for trajectory analysis or trajectory prediction.

[0039] In a preferred embodiment, the solution i Each cluster C i With variable number of gene features G i .

[0040] In a preferred embodiment, each solution S i Is multidimensional and of variable length.

[0041] In a preferred embodiment, the genetic operator includes a crossover operator, wherein the crossover operator is used in cluster C. i…k Shuffle single cells.

[0042] In a preferred embodiment, the genetic operator includes a mutation operator, wherein the mutation operator is in cluster C i Replace or add a single cell or feature G i .

[0043] In preferred embodiments, the signature expression represents mRNA or protein.

[0044] In a preferred embodiment, the genetic operator performs an adaptive operation to avoid local optima.

[0045] In another preferred embodiment, in step vi, each solution with the highest fitness is retained in the set of new solutions for the next generation using a selection operation that calls an adaptive operation to avoid local optimality by removing other solutions from the solution group S and replacing them with randomly generated solutions.

[0046] In a preferred embodiment, all adaptive operations are performed by a distributed system by monitoring all instances of the algorithm.

[0047] In a preferred embodiment, the marker is a regulatory marker.

[0048] In preferred embodiments, the systems set forth above may be adapted to perform the method steps of the embodiments set forth above. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Embodiments of the present invention will now be described by way of example and with reference to the accompanying drawings, in which like reference numerals are used to refer to like parts, and in which:

[0050] Figure 1A and Figure 1B is a schematic diagram showing different components of a distributed multi-objective genetic algorithm on high-dimensional data (scaled raw counts) for clustering, landmark detection and trajectory analysis;

[0051] Figure 2A , Figure 2C and Figure 2C Depicts the sequential steps from sequencing to quality control (QC) ( Figure 2A ), the generation of matrix M ( Figure 2B ) and the randomly generated initial solution set S n ( Figure 2C );

[0052] Figure 3 The adaptation of a running instance of an algorithm in a distributed computing setting is shown. The adaptation node monitors the evolvability of the other x (usually x=100) running instances of the algorithm. The adaptation node can write an instruction file for each running instance to command the adaptation step to maintain the evolvability of the solution;

[0053] Figure 4 shows adaptation of a running instance of the algorithm in a distributed computing setting; and

[0054] FIG. 5A to FIG. 5F Depicted are fitness plots for single-cell RNAseq and protein-seq data that were processed for quality control (QC) and log-transformed as described in the algorithm steps;

[0055] Best fit runs using PBMC, brain and protein single cell data show that clusters generated by methods according to preferred embodiments Fig. 6A (PBMC), Figure 6B (brain) and Figure 6C (Protein).

[0056] Figure 7 Graphs showing PBMC data.

[0057] Figure 8 Graphs showing PBMC data.

[0058] Fig. 9A and Fig. 9B Shown are the frequencies where the mean standard deviation in the Cell Tesseract cluster is lower (CT < random / leiden) or higher (CT > random / leiden) than the median of the random and Leiden clusters, respectively. DETAILED DESCRIPTION

[0059] The human body is composed of many types of cells. It is still impossible to identify all major cell types and their subtypes. One of the main problems in disease diagnosis and treatment in precision medicine is the identification of disease-associated cell types. For example, tumors have their own cellular heterogeneity and ultimately affect the entire body. In order to identify which cell in the tumor contributes to the pathogenesis of cancer, clustering methods are applied. In order to identify disease-associated cell types, it is important to use biologically meaningful criteria to identify cell clusters. Traditional methods lack the use of the entire data variability and reduce its dimensionality to identify cell clusters. A preferred embodiment of the present invention proposes a method that does not reduce data variability or dimensionality in order to detect cell clusters in a more meaningful way by introducing meaningful biological targets. For tumors, blood, or dysplastic brain tissues of epilepsy patients, clustering techniques can be applied to decipher the cell biological insights / targets of the disease, which can be further analyzed for future drug design.

[0060] In order to solve problems related to single-cell omics clustering, variable number feature selection, and trajectory analysis, this paper describes methods and devices using a distributed adaptive multi-objective genetic algorithm.

[0061] According to a preferred embodiment of the present invention, high-dimensional data is used without applying any data reduction methods to overcome problems associated with information loss. High-dimensional data retains all observable variance within the data, keeping both high-value signals and low-value signals, which enables the preferred embodiment to detect common or rare clusters from scOMICs data. Preferred embodiments of the present invention apply multiple targets to define cell clusters and their potential gene features. These targets are derived from molecular biological knowledge, and cell clusters are interpretable based on these defined targets. Therefore, the method and system according to the preferred embodiment are driven by multiple targets and detect biologically interpretable cell types.

[0062] Inspired by the above-mentioned dimensionality reduction problem, the method and system disclosed herein solve these problems by introducing a simple metric that accelerates the clustering of original high-dimensional data and has lossless data reduction. The challenges of data clustering include feature loss during dimensionality reduction and clustering based on unknown parameters, such as in the case of available data clustering tools. This article discloses an adaptive genetic algorithm that can be better explained, which has multiple specific goal-driven (i.e., low variance of gene regulation, high expression correlation, relative trajectory of cell clustering) data clustering, with the flexibility of adding additional goals of research interest. The preferred embodiment of the present invention will identify naturally occurring cell clusters, the variable number of marker genes (i.e., features) associated with each cluster, and the relative trajectory of the cell type that will be calculated as the product of marker gene regulation and its correlation. The trajectory of the cluster defines whether the cluster is upregulated in a given data set, and how the cluster is correlated in terms of different gene regulation, i.e., whether the regulation is up or down compared with other cell types. However, the single cell trajectory cannot be effectively derived by reducing the dimension. The trajectory is mathematically defined in equation 9.

[0063] It should be noted that marker genes are referred to as features in the model. Each cell type typically has a set of marker genes (also referred to as gene markers or features) that define the unique molecular activity of the cell type. The terms gene marker, marker gene, and feature are used synonymously herein. In a preferred embodiment, the marker is a regulatory marker.

[0064] Genetic algorithms (GAs) are agnostic to the dimensionality of the data and can traverse complex exponentially scaled search spaces to produce meaningful solutions for one or more objectives. Evolutionary computation is a broad subfield in artificial intelligence that includes genetic programming, genetic algorithms (GAs), evolutionary strategies, and other sub-branches of optimization algorithms inspired by nature. Over the past few decades, a lot of research has been done on proposing new and more efficient GAs (including adaptive genetic algorithms, multi-objective GAs, and neuroevolutionary algorithms) to solve optimization-related problems. There have also been many advances in proposing complex genetic operator methods (i.e., variable length chromosomes, adaptive mutations). These advances are key improvements in traversing the complex search space of the problem. However, it is impractical for traditional algorithms to implement and find solutions in a meaningful time. One of the most important factors for genetic algorithms is how to encode precision medicine problems into their chromosomes, and how these encoded chromosomes evolve to produce meaningful solutions based on the fitness function of the quantified objectives. The preferred embodiment solves the problem of detecting cell types and their associated markers (or features) from scOMICs data sets. For genetic diseases, cell type detection is an important advancement in the implementation of precision medicine. These cell types and their associated markers may be disease relevant, which can be targeted for future therapeutic options.

[0065] There are many computational problems in precision medicine, which are NP-difficult and require approximate or heuristic-based methods to produce optimal or suboptimal solutions. Traditional diagnostic and therapeutic options lack molecular resolution, and traditional methods do not accurately identify only pathogenic molecules, but target a wide range of molecules and biological pathways. For example, chemotherapy lacks the ability to target the pathogenic cells of cancer, instead it targets all cells indiscriminately. One of the main goals of precision medicine is to detect and target specific cell molecules (i.e., cell types, features) to accelerate disease diagnosis and treatment compared with traditional methods. For example, single-cell transcriptomics is at the forefront of discovering new cell types and markers. Clustering single cells using RNA transcriptome data and identifying cell type-specific markers is a major challenge. Each cell type generally has a set of marker genes (or features) that define the unique molecular activity of the cell type. According to a preferred method, the number of marker genes of the cell type is not fixed, and each cell type can have a variable number of clusters, i.e., the marker genes of each cluster do not have a fixed number, to simplify the problem. Due to the exponential search space, it is highly complex to identify these variable marker sets and their cell types. It is known that a similar clustering problem, k-nearest neighbor, is NP-hard. The number of possible clusters and different features associated with each cluster is exponential, and traditional algorithms will not be able to identify a solution in a feasible time. Therefore, an optimization algorithm will be a better way to find the best solution or a solution close to the best solution. Genetic algorithms are a viable alternative for optimization-related problems in high-dimensional data and have an impact in adapting to the multi-dimensionality of omics data (such as clustering of transcriptome data). In recent years, NP-hard problem examples have shown that genetic algorithms perform the same or even better than reinforcement learning. This result enforces the idea of ​​applying genetic algorithms as a meaningful alternative to traditional algorithms and is often used to reinforce machine learning algorithms.

[0066] Some people criticize GA for producing suboptimal solutions due to premature convergence. However, the preferred embodiment of the method described herein uses an updated version of genetic operators (i.e., crossover, mutation and selection) to configure so that they can effectively traverse the search space and produce a near-optimal solution or an optimal solution. This requires careful design of GA, including fitness function and its relationship with the problem target. Adaptive GA is a good alternative to some examples of traditional GA and reinforcement learning algorithms. In particular, in order to reduce the risk of stagnation in local optimum, adaptive genetic algorithms can introduce required diversity into the solution pool through the manipulation of genetic operators. In human evolution, mutation rates are different in genomes, and there are hot spots that mutate more frequently than other regions in the genome. According to this phenomenon, in GA, mutation operators are realized to produce improved results by guiding mutations to specific chromosome regions or targeted targets. Genetic algorithms have previously been used for feature selection on low-dimensional data or high-dimensional data.

[0067] In this paper, a distributed adaptive multi-objective genetic algorithm is proposed to solve problems related to single-cell omics clustering, feature selection, and trajectory analysis.

[0068] The distributed adaptive multi-objective genetic algorithm according to the preferred embodiment is:

[0069] i) Unsupervised clustering of single cells using high-dimensional transcriptome or proteome data;

[0070] ii) identifying clusters and features (i.e., genes) based on multiple targets in a biologically interpretable high-dimensional space;

[0071] iii) the number of clusters in the solution is not user defined, but rather the algorithm of the preferred method evolves to the best solution with the number of clusters that represents the target;

[0072] iv) Each cluster has its own variable length features (or genes) that define the cluster based on a set objective;

[0073] v) Find the best solution through genetic algorithm operators (crossover, mutation, selection) and adaptation of solution sets to ensure evolvability;

[0074] vi) implementing adaptivity within a cluster computing setting, where each instance of the algorithm is checked for evolvability;

[0075] vii) Apply distributed computing to find the best possible solutions and effective hyperparameters for genetic algorithm operators (crossover, mutation, and selection);

[0076] viii) Adaptive parameters are implemented in a distributed setting to keep the algorithm evolvable across running instances (in real time);

[0077] ix) detecting a single cell trajectory of each cell cluster, wherein the single cell trajectory is defined as the product of the marker gene regulation level of each cluster and the correlation between marker genes in the high-dimensional space.

[0078] Single-cell technology is only a decade old, and methods are still developing. Variants of genetic algorithms have previously been proposed mainly on low-dimensional single-cell data for clustering or feature (gene expression) selection. There is a huge gap for algorithms that can handle high-dimensional data from single-cell omics in order to cluster and detect interpretable related features. Single-cell trajectories cannot be effectively derived by reducing the dimension. It is necessary to use high-dimensional values ​​for each gene expression because it is calculated from RNA-seq. In high-dimensional space, each value represents the actions of many biological machines, and reducing the dimension will result in the loss of valuable biological insights from the data. For example, dimensionality reduction can remove subtle variability that exists in splicing patterns but not in high-dimensional space. PCA cannot handle zero or extreme downregulation of genes, while high-dimensional data will retain zero or low mRNA seq or proteomic expression data. Due to the importance of detecting subtle but important rare transcriptional dynamics (i.e., rare use of certain transcripts or low-level expression of certain gene markers), it is also necessary to implement unsupervised clustering algorithms, which may not be detected on low-dimensional data using supervised machine learning methods. The preferred distributed multi-objective algorithm can detect clusters and their associated features (gene expression) from single-cell transcriptomic (RNA sequencing (RNA-seq)) or proteomic datasets.

[0079] In a preferred embodiment, the method is a computer-implemented method. In addition, in a preferred method, multiple instances are run in multiple computing clusters in a distributed system. A typical distributed system configured with a computing cluster can be used for implementation. The method and system can be implemented using any appropriately configured hardware device and / or software device.

[0080] Figure 1A and Figure 1BIt is a schematic diagram showing the different components of the exemplary method according to the preferred embodiment of the present invention, and the exemplary method implements a distributed multi-objective genetic algorithm to high-dimensional data (raw counts of scaling), for clustering, marker detection and trajectory analysis. Initial data are usually extracted from single-cell RNA or proteomics experiments, wherein the total RNA or protein of each cell from the tissue is extracted according to the laboratory scheme, and RNA is subsequently sequenced or the protein abundance of thousands or millions of cells is measured. Matrix M is composed of raw counts or log2 transcriptome data or proteome data. The algorithm initializes the parent solution set S, and applies genetic operators (that is, mutations, crossovers) to produce daughter solution sets (O). Fitness is calculated for both S and O sets, and a new solution set is selected for the next generation (gen). The process continues until termination conditions are met.

[0081] In particular, Figure 1A and Figure 1B The steps of an exemplary method 100 for detecting cell types and their associated markers are depicted. At step 110, a two-dimensional input matrix M is received, wherein each vertical vector of the input matrix represents a single cell, and each value in the vector represents a feature expression corresponding to the single cell. Figure 1A and Figure 1B In the exemplary embodiments shown for illustrative purposes, the single-cell data is a transcriptome or a proteome, a high-dimensional raw count or a log2. Other types of single-cell data (e.g., single-cell non-coding RNA or metabolites) can be processed in the same manner. Therefore, the present invention is not limited to a transcriptome or a proteome.

[0082] At step 120, the parent solution set S1, S2, ..., S from the matrix is ​​initialized. n , where n is a predefined finite number, wherein the solution set includes a plurality of solutions S1, S2, ..., S derived from the matrix n , where each solution Si has a set of randomly assigned clusters C1, C2, ..., C k , where k is the total number of clusters in a given solution.

[0083] At step 130, each solution S in the solution set is calculated using a fitness function derived from multiple objectives. i The fitness F(S i ). In this exemplary embodiment, two targets are used. More targets may be used depending on the situation.

[0084] At step 140, genetic operators (crossover, mutation, selection) are applied to the solution set to generate a set of offspring solutions for the next generation. In this exemplary method, crossover and mutation operators are used to generate a new set of offspring (O), and adaptive parameters A and B are also updated.

[0085] At step 150, the fitness of each solution in the offspring set (O) is calculated using a fitness function.

[0086] At step 160, the fitness of each solution in the solution set (S) is compared to those solutions in the offspring set (O), and each solution with the highest fitness is retained in the new solution set for the next generation.

[0087] This process continues until a termination condition is met 170 and the method terminates 199. Figure 1A 140, wherein the generation counter gen is incremented.

[0088] like Figure 1B As depicted in Figure 1A The processes depicted in FIG. 1 run in a distributed environment.

[0089] From Figure 1A to Figure 1B The arrow reflects Figure 1A The process of is implemented on multiple distributed computing instances. In an exemplary embodiment, Figure 1A The process shown is run on multiple computing instances. In the exemplary embodiment 100, such computing instances are used. More or fewer instances may be used as appropriate. Once the process as depicted terminates a predetermined number of generations on each of the multiple computing instances, the solution set (S ) with the highest fitness among the 180 or more instances is selected. i ). From the selected solution set (S i ) to extract 190 single-cell clusters and features.

[0090] Methods embodying the present invention (including the exemplary methods disclosed herein) are typically computer-implemented. In the exemplary method, the following distributed computing architecture is adopted: wherein x (e.g., x=100) implemented algorithm instances are run in many computing clusters. The adaptive node will check the evolvability of each algorithm instance by reading the fitness function value from each generation to measure the progress of the algorithm instance. If any instance of the algorithm has not evolved compared to its previous generation fitness value, it is because the algorithm has converged to a suboptimal solution too early. In order to remove it from the suboptimal solution, the adaptive node signals the instance of the algorithm to change its existing solution pool by adding a new solution, thereby improving the algorithm's evolvability.

[0091] It should be understood that other appropriately configured distributed computing architectures or other hardware devices and / or software devices may also be used to implement the embodiments of the present invention.

[0092] Quality Control for Single-cell Omics

[0093] The mean counts of the raw sequencing matrix data were calculated and features were retained if at least 10% of the cells had the mean transcript count, and these features were then normalized using a log2 transformation. No single cells were excluded from the analysis.

[0094] Initializing solution pool S

[0095] The solution space of the genetic algorithm according to a preferred embodiment consists of a two-dimensional matrix M, in which each vertical vector (e.g. Figure 2B ) represents a single cell, and each value in the vector represents the feature (gene) expression (read counts obtained from mRNA sequencing or protein expression) corresponding to the single cell. In each generation, the solution set S will have n solutions S1, S2, ..., S obtained from the matrix M. n , where n is a finite number defined by the user and is typically < 1000. In GA terminology, these solutions are called "chromosomes", and in this GA, each chromosome in the solution set S has single-cell clusters and features of variable length.

[0096] Since this is an unsupervised implementation, each solution S derived from the matrix M i A set of randomly assigned clusters C1, C2, ..., C with single cells k , where n is the solution S i The total number of clusters in each C i There are two sets of variables, G - a set of L features (or genes) and R - clusters C i The feature set in G has a maximum bound of 25,000 (the number of genes in the human genome), where each feature is a dimension in a single cell. i , the number of features L is randomly drawn from the set G. The minimum number of feature sets is L = 25, and the maximum number of features is L = 300. In large-scale single-cell transcriptome studies of low-dimensional RNA-seq data, the maximum number of features for clustering is 300. i The number of single cells R is randomly assigned from the minimum number of single cells in the cluster min(R) = 25 and the maximum number, which is max(R) = the total number of single cells - ((n-1)XR min ) definition.

[0097] Figure 2 depicts the sequential steps from sequencing to quality control (QC) and the generation of the matrix M. The algorithm takes an input mapping file and generates a random solution set S n QC is implemented before. Each initial solution S of the algorithm i is derived from M and has a random number of clusters assigned to it. Each cluster (C i ) has a variable number of traits (or genes) (G i ), that is, the variable “number of clusters” and the variable “feature set (G)”.

[0098] In Figure 2a, step 210 depicts mRNA sequencing or protein expression from a laboratory experiment to generate single-cell omics. The single-cell omics data is mapped and processed in step 220. The matrix M is generated in step 230, and an initial solution set S is randomly generated in step 240. n .

[0099] One of the particularly advantageous features according to the preferred embodiment is the design of each solution Si. A solution Si has a plurality of clusters C assigned to it. k , where k is chosen randomly. Solution S i Each cluster C i Characterized by a variable number of genes (G i )(See Figure 2C ). Therefore, each cluster within the solution set can have a different number of independent features. In order to have biologically meaningful features for clustering, the length of the features |G| should have a user-defined minimum value (e.g., 20 or 25). Cell types with fixed gene feature methods may not allow some of the important markers to be selected due to limitations. This variable number of gene features will allow detection of cell types with the correct number of marker genes and will improve the accuracy of cell clustering. Therefore, for cell cluster C1, the length of |G| may be 20, while for C2, 25 gene features may be included. This ensures that the single cell clusters and features extracted from the selected solution set are biologically interpretable. In particular, by having the correct number of marker genes rather than applying dimensionality reduction, the extracted single cell clusters and features are biologically interpretable. Therefore, Figure 2A , Figure 2B and Figure 2C This reflects how, for example, mRNA sequencing or protein expression occurs in laboratory experiments, and how the resulting input data is processed by an automated method / system according to a preferred embodiment of the present invention to generate single-cell omics. Since the resulting extracted single-cell clusters and features are biologically interpretable, medical diagnosis and treatment are facilitated.

[0100] In addition, data variability is not reduced by the method described in the preferred embodiment, but remains as it is. Single-cell omics variability is intrinsic and multidimensional. The variability of the data exists due to the active participation of numerous biological processes involving cellular molecules, and traditional clustering algorithms only generate clusters based on the potential variance from the reduction technology without interpolating any biological knowledge into it. The method described in the preferred embodiment uses high-dimensional data from RNA sequencing, and the algorithm achieves a specific goal to cluster scOMICs data. For example, if we set the goal of applying biological knowledge and want to observe the clustering of single-cell omics with respect to the similarity of cellular gene expression, the method according to the preferred embodiment clusters the scOMICs data based on this specific goal.

[0101] Calculating fitness

[0102] For each solution S i , the fitness is calculated in two layers. First, for each cluster C i Assign a fitness function derived from two main objectives. Second, assign a i,..k The fitness value of each solution S is calculated i The fitness function is essentially a minimization function where both goals will have low variance (distance) for clustering tightly linked single cells. In a preferred embodiment, both goals are very important for real world applications, particularly identifying cell types that behave in a similar way molecularly. For example, a goal can be derived where cell types are clustered based on cell type. However, in an embodiment, the goal is more about molecular similarity.

[0103] Goal 1: The first objective in the fitness function defines the average correlation of each single cell calculated using equations 1, 2, and 3, where cluster C i The feature vectors of the single cells in the cluster are used to calculate the correlation coefficient r. i,rmean The mean r of .

[0104] Collection G i A cluster C with R cells and L features in i The average correlation r of each single cell in is calculated as:

[0105]

[0106] Mean correlation of clusters

[0107] Goal 2: The second goal is to calculate the mean expression "exp" of each single cell and its relationship with the cluster. This goal calculates (Equations 4 and 5) the mean expression expmean(Ci) of the cluster.

[0108] Expression differences

[0109] Mean expression of clusters

[0110] Clusters F(C) are assigned by dividing by the maximum possible distance calculated for the values ​​(0,0) and (1, maximum feature expression (raw read count or log2)) (Equations 6 and 7) i )’s fitness. Next, solve for the fitness F(S i ) is the average fitness of all cluster fitnesses (Equation 8). The greater the fitness, the greater the chance that it will be selected for reproduction. The maximum fitness also has an inverse relationship with the distance function derived from goals 1 and 2. In other words, the higher the fitness of the solution means the smaller the distance between the single cell and its cluster mean for goals 1 and 2.

[0111] Distance between two points

[0112] Cluster fitness

[0113] Solution fitness

[0114] The trajectory T of a cell cluster is calculated based on two determined objectives (Eq. 9), and it tracks the state of gene expression and the correlation of single cells in a cluster.

[0115] Solution S i Each cluster C in i The trajectory is

[0116]

[0117] Preferred embodiments accommodate two goals. The exemplary embodiments described herein are described with two purposes; however, in other embodiments, additional goals are also provided. According to such embodiments, multiple types of goals can be added and achieved. For example: 1) clustering cell types based on dysregulation of a set of disease-related marker genes (or features); or 2) clustering cell types based only on upregulated marker genes or downregulated marker genes.

[0118] Select a parent for the new child

[0119] An elitist approach has been implemented to retain the most suitable solutions for the next generation. Both the parent (S) set and the offspring (O) set, which have an equal number of solutions, compete to successfully enter the new solution set. For each S from the parent set and the offspring set, i and O i The fitness of is ranked, and the top n = |S| sets with the highest fitness are retained for the next generation of reproduction. In order to select the next generation, the offspring (from crossover and mutation) compete with the parent generation, and the top n solutions successfully become the next generation of S.

[0120] In a preferred embodiment, the adaptive parameter A will check whether the fitness function has converged after a certain (user-defined) number of generations gen (gen=100 in this exemplary embodiment). If the fluctuation of the average fitness is 0 in these previous generations, the algorithm will update the adaptive parameter A=1 (default value is 0). If A=1, the algorithm will take measures to remove existing solutions from S to move the search away from premature convergence positions in the search space (see the adaptive section for details). The algorithm will only keep the best solution best (S) with the highest fitness in the parent generation, i.e. the set of parents (S). i ), and all other S i will be replaced by a random set of n-1 solutions and will subsequently be used to generate another n children for the set child (O). If adaptive A=0, the algorithm will move to the next step without replacing any solution in the set parent (S). The application of adaptive parameters speeds up the evolution of the algorithm and keeps the algorithm away from generating suboptimal solutions.

[0121] Crossover operator and mutation operator

[0122] Two selection procedures have been implemented for the crossover operation. i is multidimensional and of variable length. Due to this complexity, the crossover operator is not very useful in clustering C. i Two selection criteria are applied to cluster C i ......shuffle the single cells between k. The first choice will pick a random solution S i , and randomly select S i In the next step, single cells between clusters will be exchanged without exchanging their associated features G.

[0123] The second selection criterion selects the top five solutions (with the highest fitness) S iA solution in the set is chosen, and the cluster with the weakest fitness will be exchanged for the cluster with the fitness closest to the average. Although this is a probabilistic approach, it is more direct and more influential. The probability of using one of the two choices in each generation is equal (p = 0.50). If the algorithm detects premature convergence and the adaptive parameter A = 1, the probability of using the second selection criterion for crossover will increase by 80% or (probability will be assigned, p = 0.80).

[0124] Figure 3 The exemplary crossover operation depicted in operates as follows: Select a random solution S from the solution pool. Compare the distance between each Ci,j and the cluster mean Corr and the mean Exp. Ci,j is placed to the closest C, and it adopts G that belongs to the new C. Crossover also has an adaptive switch. If A is on, instead of a random solution S, it will increase the probability of picking the most suitable solution in the pool. Recall that Figure 1A Schematic diagram of using crossover and mutation after calculating fitness. These operations are used to produce the next generation of solutions.

[0125] The mutation operator is designed to (randomly) i Replace or add a single cell or feature G i First, for single-cell mutation operations, the method is configured with a directed approach, where selection will be made from the solution S i Randomly select cluster C from i , and only the weakest (lowest fitness) single cell in the cluster is deterministically selected for replacement. i The weakest single cell will be clustered by another cluster C p The weakest single-cell replacement, where i≠p and C p The distance from this specific single cell is small. Due to the importance of each single cell in the data, no single cell is deleted.

[0126] Secondly, the mutation operation starts from cluster C i The feature G i Randomly remove or add features in G. This operation involves removing existing features and adding new features to G i In , 25,000 features are randomly sampled from the genome. If the algorithm detects premature convergence and becomes A = 1, the mutation rate is increased by 50% to introduce more diversity into the solution set S.

[0127] Distributed computing and adaptation to ensure evolvability of solutions:

[0128] The algorithm consists of several parameters (mutation rate, number of generations (gen), selection probability), and it is difficult to run a test set to find the best values ​​for these parameters. The exemplary method and system are preferably implemented using a distributed computing method. According to an exemplary embodiment, 100 instances of the algorithm are run by assigning different genetic algorithm parameter values ​​in each instance. According to a preferred embodiment, all instances of the algorithm are monitored. The algorithm updates the adaptive variable A or B for the set of parameter values ​​in a single running instance I. The "best solution" will be determined by comparing the fitness values ​​of each best solution produced by each instance of the genetic algorithm. The "best solution" is used to obtain results related to clustering, feature detection and trajectory analysis ( Figure 4 ). According to an exemplary embodiment, the method is configured to perform adaptation within a distributed operating environment using the following specific steps:

[0129] Distributed computing architecture ( Figure 4 ) within the adaptive node starts x running instances of the algorithm I x (x=100 in the exemplary embodiment, but x can be adjusted according to preference), and each running instance is monitored by an adaptive node to ensure that the fitness is evolving according to the specified goals. The node can send real-time instructions to the running instance of the algorithm to change its internal parameters or stop the run and replace it with a new instance.

[0130] To enhance evolvability, the adaptive node parameter A (set to 0 by default) checks convergence of the average fitness every x generations. If the adaptive node finds a premature fitness stagnation point for the program instance, it sets A=1 (instruction file). If A=1, then this particular instance of the program removes all solutions from the set S except for the best fitting solution. The remaining solutions are randomly generated for the next generation to compete. Therefore, the next generation of solutions will have the exact same number of solutions as before, and will only include the best parents and children.

[0131] Due to poor evolvability (no fitness progress), if the termination parameter of the program instance is B=1 (default setting is 0), the adaptive node stops the program and restarts a new instance of the program with a new population by increasing the mutation rate. In the end, the adaptive node will have x=100 runs.

[0132] Figure 4 The adaptation of a running instance of an algorithm in a distributed computing setting is shown. The adaptation node monitors the evolvability of the other x (x=100 in the exemplary embodiment, but x can be adjusted according to preference) running instances of the algorithm. The adaptation node can write an instruction file for each running instance to command the adaptation step to maintain the evolvability of the solution.

[0133] Single cell datasets:

[0134] Dataset 1: 2,700 single cells from PBMC (peripheral blood mononuclear cells). (https: / / satijalab.org / seurat / archive / v3.1 / pbmc3k_tutorial.html) Obtain raw RNAseq data.

[0135] Dataset 2: 7,283 single cells from the ACC (anterior cingulate cortex) region of the brain. From the Allen Brain Atlas (https: / / portal.brain-map.org / atlases-and-data / rnaseq) Obtain raw RNAseq data.

[0136] Dataset 3: Single cells from CD8+ naive T cells. From GSE100501 (https: / / www.ncbi.nlm.nih.gov / geo / query / acc.cgi?acc=GSM2685258) Obtain raw protein sequencing analysis data.

[0137] These datasets have been used in prior art PCA-based methods. Using these datasets enables comparison with these prior art methods.

[0138] For comparison purposes, several clusters were randomly created and single cells were randomly assigned to those clusters with a random number of features in each cluster.

[0139] Comparison Test:

[0140] Cell Tesseract (CT) is an exemplary implementation of a distributed adaptive multi-objective genetic algorithm used in a preferred embodiment of the present invention. The analysis has been performed using the most frequently used single-cell transcriptome software described in Seurat, which implements the unsupervised clustering algorithm "Leiden", which implements a network-based algorithm that partitions data based on its network membership. In two comparisons, the dispersion (standard deviation (stdev)) of the feature (gene) expression of the clusters identified by Cell Tesseract is compared with the stdev distribution of features in the clusters randomly assigned or detected by the Leiden algorithm. A proportional Z test can be used to estimate significance.

[0141] result

[0142] Single cell data clustering using Cell Tesseract

[0143] Single-cell RNAseq and protein-seq data were processed for quality control (QC) and log-transformed as described in the algorithm steps. Each dataset was run with a different set of parameters to identify the best solution. i The fitness of each run is monitored between generations to confirm that the algorithm is evolving within the search high-dimensional space. The results are FIG. 5A to FIG. 5F Shown in.

[0144] Figure 5A Fitness plots of PBMCs are shown (1000 generations, 60 runs).

[0145] Figure 5B Fitness plots of PBMCs are shown (5000 generations, 40 runs).

[0146] Figure 5C Fitness plots of the brain are shown (1000 generations, 93 runs).

[0147] FIG5D shows the fitness plot of the brain (5000 generations, 7 runs).

[0148] Figure 5E The fitness graph of protein single-cell data (1000 generations, 1 run) is shown.

[0149] Fig. 5F A fitness plot of protein single-cell data is shown (10,000 generations, 1 run). The fitness of each run is represented by a different color, which can be identified as different shades in the black and white rendering.

[0150] The evolvability of the algorithm runs shows a multiple increase in fitness from the starting point. In an exemplary embodiment, results from 100 runs of PBMC single cell data were used for comparison purposes. From the distributed runs, the best solution with the highest fitness was selected for data clustering, feature selection, and trajectory prediction. The clusters generated by the method according to the preferred embodiment using the best fit runs of PBMC, brain, and protein single cell data are shown in Fig. 6A (PBMC), Figure 6B (brain) and Figure 6C (Protein).

[0151] The clusters were visualized using plotly of the Rgl package in R, with each cluster being numbered and represented by a different color, which can be identified as different shades in a black and white rendering. XC, YE, and Zmean E represent correlation, average mean expression, and track T, respectively. Track T on the Z axis is the product of the correlation of the mRNA expression level of the cluster and the marker gene. Due to the variable number of marker genes, each cell type defines its unique track. Track T is mathematically defined in equation (9).

[0152] To test the performance of the Cell Tesseract algorithm employed in the preferred embodiment, the mean standard deviation (stdev) of the features (gene expression) in each cluster generated by CellTesseract was compared to the corresponding distribution of standard deviations in all clusters generated by random cluster assignment. This allowed us to compare the standard deviation of each cluster of CT to the distribution of standard deviations of randomly generated clusters and their associated features. The plots of the PBMC data are shown in Figure 7 middle.

[0153] Using PBMC data, the mean standard deviation (y-axis) calculated for features in Cell Tesseract (CT) (red dots) for each cluster (CT1 to CT18) is compared to all clusters using random cluster assignment (distribution shown as box plot). The results show that the mean and deviation of Cell Tesseract are significantly different from those generated randomly.

[0154] Furthermore, the mean standard deviation of the features in the individual clusters generated by Cell Tesseract is compared to the corresponding distribution of standard deviations across all clusters generated (in low-dimensional space) by one of the gold standard clustering algorithms, Leiden ( Figure 8 ). The results show that the standard deviation of the Cell Tesseract cluster features is not significantly different compared to Leiden. Next, the results of the two comparisons (randomly assigned clusters and Leiden) were quantified. The results show that the overall Cell Tesseract clusters show a significantly smaller standard deviation of cluster expression compared to the randomly assigned clusters (P < 7.3X10 -5 ), while there was no significant difference between Cell Tesseract and Leiden (P<0.16)( Figure 8 ).

[0155] Figure 8 Shown are the mean standard deviations (y-axis) of features in Cell Tesseract (red dots) for each cluster (CT1 to CT18) using PBMC data compared to all clusters using the Leiden algorithm (distributions shown as box plots).

[0156] Fig. 9A and Fig. 9BThe frequency of cases where the mean standard deviation in the Cell Tesseract clustering is lower (CT < random / leiden) or higher (CT > random / leiden) than the median of the random and Leiden clusterings, respectively, is shown. The mean standard deviation of CT is significantly different from the randomly generated cell clusters and features; however; the mean standard deviation is similar to the Leiden clustering algorithm. Therefore, the method according to the preferred embodiment performs significantly better than randomly generated results and is comparable to the commonly used Leiden clustering algorithm.

[0157] discuss

[0158] The method according to the preferred embodiment is configured using a genetic algorithm, which can process both single-cell transcriptome data and proteome data, for clustering of biologically driven targets. The implemented target enables detection of cell types in which the level of marker genes is similarly regulated (i.e., similar mRNA expression levels) and the correlation of mRNA expression between cells is higher. The target also enables identification of the track of each cell type, which means that compared with other cell types, regulation is upward or downward. The method according to this preferred embodiment searches for the best solution in high-dimensional single-cell omics data. In an exemplary specific implementation, it has been applied to both transcriptome data and proteome data to obtain target-driven clustering that can be explained according to the target. Therefore, this preferred method can be used to identify pathogenic cell types and their related marker genes. The first target in the exemplary method studies the correlation between single cells in clustering. The co-expression dynamics level of single cells in the second target quantitative clustering. These two targets are later used to identify the track of clustering, which defines whether clustering is upregulated in a given data set and how the consistency of clustering is related in terms of different gene regulation. The number of clusters and the number of features in the clustering are unknown. This unsupervised nature of the search for the best solution makes it flexible to identify clusters with different quantitative measures.

[0159] Since no dimensionality reduction techniques are applied, the exemplary method is configured to process thousands of dimensions to find target-driven interpretable clusters from single-cell omics data. Although the high dimensionality of the data introduces huge complexity, CellTesseract (an exemplary embodiment of the preferred method) evolves to find biologically meaningful solutions. Each solution represents a group of cell clusters, in which their marker genes are similarly regulated in terms of mRNA expression levels and regulatory correlations. Unlike traditional dimensionality reduction methods, the preferred method provides information that helps to explain clusters with meaningful molecular insights. Compared with randomly assigned clusters and the most frequently used single-cell clustering algorithm in Seurat, clustering has been verified. The exemplary method according to the preferred embodiment of the present invention is significantly (p<7.3X10 -5) outperformed random cluster assignment, which showed clusters composed of highly correlated single cells, and detected similarly co-expressed feature sets. The resulting data also produced comparable (p<0.16) results relative to Seurat in terms of mean standard deviation. Therefore, this paper describes a preferred method that introduces variable solution or chromosome length, adaptive parameters, and directed genetic operators to find clusters, features, and trajectories from single-cell omics datasets.

[0160] The method and system of the exemplary genetic algorithm configuration described herein provide solutions for single cell clustering algorithms, which can effectively detect the clustering of omics (i.e., transcriptome or proteomics) data sets and the features of their related variable quantities. This GA acts on high-dimensional data that maintain observable changes in omics molecules. The exemplary GA applies target-driven clustering, which is interpretable according to the set target and can be applied to many single cell omics data, including conditional (processed or unprocessed) data. For example, the preferred method can be applied to case-control scOMICs data, with the goal of finding pathogenic cell types in cases. From patient tissue and control tissue, the algorithm can identify cell types and their marker genes that are differently regulated in patient tissue rather than in control (i.e., triple negative breast cancer). The preferred method can also be applied to single cell proteomics data obtained from patient tissue to identify protein markers of diseases (e.g., epilepsy). Therefore, the GA described in the context of the preferred embodiment improves the two-dimensional clustering based on PCA, as shown in the results.

[0161] Thus, the preferred embodiments avoid creating undesired linear correlations between variables, which can lead to failures when the data set cannot be defined by means and covariances. In addition, low-dimensional methods must ignore certain variables during feature selection. The high-dimensional methods described in the preferred embodiments overcome the problem of selecting features that depend on univariate statistics and can lead to suboptimal subsets.

[0162] In addition, the high-dimensional methods described in the preferred embodiments overcome the problem of bias towards high-value signals (e.g., clusters of highly expressed features). In contrast, the preferred embodiments capture low signals from the data (i.e., small rare low-expressing cell clusters in scRNA-seq or other omics data).

[0163] One of the limitations found in Scanpy's output is that cells of the same type are not always grouped together, and cells of different types are not well separated. One of the general concerns with supervised clustering methods such as k-means is that the user must provide the number of clusters or the resolution at which the number of clusters is determined. These complexities highlight the need for better unsupervised solutions to detect clusters and their associated features (i.e., genes).

[0164] Specific applications of the preferred embodiments

[0165] As mentioned above, single cell is a very new technology that is only ten years old. The methods and systems according to the preferred embodiments can affect the diagnosis and treatment of many genetic diseases that are currently resistant to traditional drug treatment. The following is a list of high-risk applications of the preferred embodiments:

[0166] 1. The methods and systems according to the preferred embodiments use high-dimensional single-cell omics (scOMICs) data (e.g., transcriptome or proteome) to identify cell-specific markers. These markers belonging to specific cell types are the main regulators of a group of cells. The markers identified from our algorithm can be tested as diagnostic predictors of diseases. Such markers can be applied to clinical trials to quantify diagnostic sensitivity and specificity. Such markers can be widely used, and the method can be applied to a variety of diseases. For example, there are currently no markers for most rare genetic diseases and cancers. The methods and systems according to the preferred embodiments can be applied to such unmedicated diseases to identify scOMICs-based disease markers and their associated cell types. Diseases to which the preferred embodiments can be easily applied are epilepsy and breast cancer. In both cases, the preferred methods and systems can be applied to find predictive markers for diagnosis and treatment.

[0167] 2. The methods and systems according to the preferred embodiments can be applied to gene therapy. For example:

[0168] a. CRISPR / Cas9 can be designed in a cell type-specific manner, and methods and systems according to preferred embodiments can identify such cell types and target molecules for CRISPR / Cas9 design.

[0169] b. mRNA drug technology based on antisense oligonucleotides (ASO): The methods and systems according to the preferred embodiments can target cell type specific markers that are upregulated in genetic diseases. Marker detection will be the key to targeting cells with ASO mRNA drugs. For example, if we use breast cancer tissue to apply scOMICs, the methods and systems according to the preferred embodiments will be able to detect cell types and their associated markers that are the main regulators of the phenotype. ASO mRNA technology can then be applied to target these markers to block or reverse regulation. This specific implementation can be widely applied to most genetic diseases.

[0170] 3. Discovery of cell type specific markers can bring great opportunities for collaboration with drugs. Recently, the drug design industry has adopted single cell technology for single cell level molecular quantification. The method and system according to the preferred embodiment can be applied to pharmacogenetic testing experiments of treated and untreated types in cell models for various drug compounds to identify marker sensitivity and specificity for chemical compounds or drugs.

[0171] 4. As can be seen from the above, the methods and systems according to the preferred embodiments can be widely applied. Therefore, research collaborations with many genetic and cell biology research groups are obvious opportunities. Additional prediction functions can be added for different biological problems involving single-cell omics.

[0172] The invention is not limited to the embodiments described herein but may be modified or altered without departing from the scope of the invention as defined by the appended claims.

Claims

1. A method for detecting cell types and related markers thereof, the method comprising: a. Run multiple instances of the algorithm in a distributed system, where each instance is configured according to the following steps: i. receiving a two-dimensional input matrix M, wherein each vertical vector of the input matrix represents a single cell, and each value in the vector represents a feature expression corresponding to the single cell; ii. Initialize the parent solution set S1, S2, ..., S from the matrix n , wherein n is a predefined finite number, wherein the solution set includes a plurality of solutions S1, S2, ..., S derived from the matrix n , where each solution S i A set of randomly assigned clusters C1, C2, ..., C with single cells k , where k is the total number of clusters in a given solution; iii. Use the fitness function to calculate each solution S in the solution set i A fitness function, wherein the fitness function is derived from a plurality of objectives; iv. applying a genetic operator to the solution set to generate a set of offspring solutions; v. using the fitness function to calculate the fitness of each solution in the offspring set; vi. comparing the fitness of each solution in the solution set and the offspring set, wherein each solution with the highest fitness is retained in the new solution set for the next generation; vii. Repeating steps iii, iv, v and vi for a predetermined number of generations; b. selecting the solution set having the highest fitness among the multiple instances; c. Extract single-cell clusters and features from the selected solution set.

2. The method of claim 1, wherein the fitness function is derived from two objectives.

3. The method of claim 1 or 2, wherein the matrix comprises raw counts or log2 transcriptomic data or proteomic data.

4. A method as claimed in any preceding claim, wherein the relative trajectory of cell types is calculated as the product of marker gene regulation and its correlation.

5. The method of claim 4, wherein the trajectory of a cluster defines whether the regulation of the cluster is up or down compared to other cell types.

6. The method according to any one of claims 4 or 5, wherein the solution S i Each cluster C in i The trajectory is 7. The method of claim 4, wherein the solution set selected in step (b) is selected for trajectory analysis or trajectory prediction.

8. A method as claimed in any preceding claim, wherein the solution S i Each cluster C i With variable number of gene features G i .

9. A method as claimed in any preceding claim, wherein each solution S i Is multidimensional and of variable length.

10. A method as claimed in any preceding claim, wherein the genetic operator comprises a crossover operator, wherein the crossover operator is used in cluster C i…k The single cells are shuffled in between.

11. A method as claimed in any preceding claim, wherein the genetic operator comprises a mutation operator, wherein the mutation operator is applied to cluster C i Replace or add a single cell or feature G i .

12. A method as claimed in any preceding claim, wherein the characteristic expression represents mRNA or protein.

13. A method as claimed in any preceding claim, wherein the genetic operator performs an adaptive operation to avoid local optima.

14. The method of claim 13, wherein in step vi, each solution with the highest fitness is retained in a set of new solutions for the next generation using a selection operation that calls an adaptive operation to avoid local optimality by removing other solutions from the solution group S and replacing them with randomly generated solutions.

15. A method as claimed in any one of claims 1 to 5, wherein all adaptive operations are performed by a distributed system by monitoring all instances of the algorithm.

16. The method of any preceding claim, wherein the marker is a regulatory marker.

17. A system for detecting cell types and markers associated therewith, the system comprising means for performing the following steps: a. Run multiple instances of the algorithm in a distributed system, where each instance is configured according to the following steps: i. receiving a two-dimensional input matrix, wherein each vertical vector of the input matrix represents a single cell, and each value in the vector represents a feature expression corresponding to the single cell; ii. Initialize the parent solution set S1, S2, ..., S from the matrix n , wherein n is a predefined finite number, wherein the solution set includes a plurality of solutions S1, S2, ..., S derived from the matrix n , where each solution S i A set of randomly assigned clusters C1, C2, ..., C with single cells k , where n is the total number of clusters in a given solution; iii. Use the fitness function to calculate each solution S in the solution set i wherein the fitness function is derived from a plurality of objectives; iv. applying a genetic operator to the solution set to generate a set of offspring solutions; v. using the fitness function to calculate the fitness of each solution in the offspring set; vi. comparing the fitness of each solution in the solution set and the offspring set, wherein each solution with the highest fitness is retained in the new solution set for the next generation; vii. Repeating steps iii, iv, v and vi for a predetermined number of generations; b. selecting the solution set having the highest fitness among the multiple instances; c. Extract single-cell clusters and features from the selected solution set.

Citation Information

Patent Citations

  • SE100501C1