Cell characterization model pre-training method, cell downstream task processing method, device, storage medium and program product

By identifying characteristic genes and gene pathways in the cell characterization model and constructing a pre-trained model, the problem of high noise in cell characterization information is solved, and the accuracy and discriminative power of cell characterization are improved.

CN119832987BActive Publication Date: 2025-11-18PEKING UNIVERSITY CHENGDU ACADEMY FOR ADVANCED INTERDISCIPLINARY BIOTECHNOLOGIES +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411742754.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-11-18
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

In existing technologies, the cell characterization information output by pre-trained models has a high noise content, and there is still room for improvement in the accuracy of cell characterization.

Method used

By identifying multiple characteristic genes in a cell sample set, the first characteristic gene sequence of the sample cells is obtained, and multiple gene pathways are matched based on the multiple characteristic genes. The cell characterization model is then adjusted until the training termination condition is met, thus constructing a pre-trained cell characterization model.

Benefits of technology

It reduces the noise content in cell-encoded information, improves the signal-to-noise ratio, increases the distinguishability of cell-encoded information for cells, and optimizes the accuracy of cell characterization results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832987B_ABST
    Figure CN119832987B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of biology, and provides a cell characterization model pre-training method, a cell downstream task processing method, equipment, a readable storage medium and a program product, which can improve cell characterization accuracy. The cell characterization model pre-training method comprises the following steps: determining a plurality of characteristic genes of a cell sample set, and obtaining a first characteristic gene sequence corresponding to each sample cell in the cell sample set; a plurality of gene pathways for realizing different biological functions are matched according to the plurality of characteristic genes; cell coding information corresponding to each sample cell is determined by a cell characterization model to be trained according to each first characteristic gene sequence and each gene pathway, a second characteristic gene sequence corresponding to each sample cell is predicted according to the cell coding information; the cell characterization model is adjusted according to the difference between the second characteristic gene sequence and the first characteristic gene sequence of the sample cell until a training end condition is met, and a pre-trained cell characterization model is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biotechnology, and in particular to a method for pre-training a cell characterization model, a method for processing downstream cell tasks, an apparatus, a computer device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] In fields such as biology and medicine, the accuracy of cell characterization information is crucial for improving performance in various cell-related downstream tasks. In the practice of cell characterization, ribonucleic acid (RNA) can be extracted from cell or tissue samples, and then combined with omics sequencing technology to determine the expression levels of specific genes, thereby characterizing cells using gene expression levels. However, the expression data obtained in this way are highly dependent on experimental methods, have high data noise, and pose many challenges to subsequent analysis.

[0003] With the rise of artificial intelligence models, model-based, AI-driven scientific research is becoming a hot research topic. Feature vectors generated by pre-trained models, with their excellent discriminative and representational capabilities, provide more refined distinctions and descriptions of cells. Providing high-quality cell feature vectors through pre-trained models effectively improves the performance of various downstream tasks related to cells.

[0004] However, the inventors found in practice that the cell characterization information output by the pre-trained model in the relevant technology has a high noise content, and there is still room for improvement in the accuracy of cell characterization. Summary of the Invention

[0005] Based on this, it is necessary to provide a cell characterization model pre-training method, a cell downstream task processing method, an apparatus, a computer device, a computer-readable storage medium, and a computer program product to address the above-mentioned technical problems.

[0006] In a first aspect, this application provides a method for pre-training a cell characterization model, including:

[0007] Identify multiple characteristic genes in a cell sample set, and obtain the first characteristic gene sequence corresponding to each of the multiple sample cells in the cell sample set;

[0008] Multiple gene pathways for achieving different biological functions are obtained by matching the multiple characteristic genes; each gene pathway includes various characteristic genes involved in achieving the corresponding biological function of the gene pathway;

[0009] The cell characterization model to be trained determines the cell coding information corresponding to each of the sample cells based on each of the first feature gene sequences and each of the gene pathways, and predicts the second feature gene sequence corresponding to each of the sample cells based on the cell coding information;

[0010] Based on the difference between the second characteristic gene sequence and the first characteristic gene sequence corresponding to the sample cells, the cell representation model is adjusted until the training termination condition is met, thus obtaining a pre-trained cell representation model.

[0011] In one embodiment, the gene pathway is obtained through the following steps:

[0012] The multiple feature genes are input into multiple pathway construction networks; each of the multiple pathway construction networks corresponds to a different pathway construction rule, and the different pathway construction rules characterize the realization process of biological functions based on different dimensions.

[0013] Each pathway is used to construct a network, which determines multiple characteristic genes that perform related biological functions among the multiple characteristic genes according to the corresponding pathway construction rules, and outputs a corresponding undirected graph based on the multiple characteristic genes that perform the same biological function.

[0014] Based on the output of each undirected graph, the gene pathways of the multiple characteristic genes are obtained.

[0015] In one embodiment, the process of each pathway construction network determining multiple characteristic genes that perform associated biological functions from among the multiple characteristic genes according to corresponding pathway construction rules includes:

[0016] If the plurality of pathway construction networks includes a first pathway construction network, then the first pathway construction network determines the plurality of characteristic genes that realize the associated biological functions among the plurality of characteristic genes according to the pathway construction rules related to the gene ontology;

[0017] If the plurality of pathway construction networks includes a second pathway construction network, then the second pathway construction network determines the plurality of characteristic genes that realize the associated biological functions among the plurality of characteristic genes according to the pathway construction rules related to cell biological processes;

[0018] If the multiple pathway construction network includes a third pathway construction network, then the third pathway construction network determines multiple characteristic genes that realize related biological functions among the multiple characteristic genes according to the pathway construction rules related to immunity.

[0019] If the plurality of pathway construction networks include a fourth pathway construction network, the fourth pathway construction network determines the plurality of characteristic genes that realize the associated biological functions among the plurality of characteristic genes according to the pathway construction rules related to cell transcription.

[0020] In one embodiment, the cell characterization model to be trained determines the cell coding information corresponding to each of the sample cells based on each of the first feature gene sequences and each of the gene pathways, including:

[0021] A first matrix is ​​constructed based on the first characteristic gene sequence corresponding to each of the multiple sample cells and the gene expression level of each characteristic gene in each of the pre-acquired first characteristic gene sequences; each row of the first matrix corresponds to one sample cell, each column corresponds to one characteristic gene, and each matrix element represents the gene expression level of the characteristic gene of the sample cell.

[0022] A second matrix is ​​constructed based on each of the gene pathways; each row of the second matrix corresponds to a characteristic gene, each column corresponds to a gene pathway, and each matrix element represents whether the characteristic gene is included in the gene pathway;

[0023] The cell characterization model to be trained extracts features from the product of the first matrix and the second matrix, and obtains the cell coding information corresponding to each sample cell based on the feature extraction results.

[0024] In one embodiment, the second matrix includes multiple matrices, and the gene pathways in the same second matrix are constructed based on the same pathway construction rules, while the gene pathways in different second matrices correspond to different pathway construction rules.

[0025] The feature extraction from the product of the first matrix and the second matrix by the cell representation model to be trained includes:

[0026] The product of the first matrix and each of the second matrices is obtained from the cell representation model to be trained, resulting in multiple third matrices;

[0027] The cell characterization model uses multiple feature extraction modules to extract features from each of the third matrices based on a self-attention mechanism, resulting in sub-feature extraction results for each of the multiple feature extraction modules. The multiple feature extraction modules are used to extract features from the third matrices constructed based on different pathway construction rules.

[0028] The feature extraction result is obtained by fusing the results of each sub-feature extraction.

[0029] In one embodiment, the cell characterization model to be trained determines the cell coding information corresponding to each of the sample cells based on each of the first feature gene sequences and each of the gene pathways, including:

[0030] A partial feature gene in each of the first feature gene sequences is masked to obtain the corresponding feature gene mask sequence. The cell characterization model to be trained determines the cell coding information corresponding to each sample cell based on each feature gene mask sequence and each gene pathway.

[0031] The step of predicting the second characteristic gene sequence corresponding to each sample cell based on the cell coding information includes:

[0032] Based on the cell coding information, feature gene prediction is performed on the masked portion of each feature gene mask sequence, and the second feature gene sequence corresponding to each sample cell is obtained based on the prediction results.

[0033] Secondly, this application also provides a method for downstream cell task processing, the method comprising:

[0034] Determine the downstream tasks for the target cell, obtain the characteristic genes corresponding to the target cell, and determine the characteristic gene sequence and multiple gene pathways corresponding to the target cell based on the characteristic genes corresponding to the target cell;

[0035] The target cell encoding information of the target cell is obtained by a pre-trained cell characterization model based on the characteristic gene sequence and multiple gene pathways corresponding to the target cell; the cell characterization model is trained based on any of the cell characterization model pre-training methods described above.

[0036] The downstream tasks targeting the target cell are processed based on the target cell's encoded information.

[0037] Thirdly, this application also provides a pre-training device for a cell characterization model, comprising:

[0038] The first sequence acquisition module is used to determine multiple characteristic genes of the cell sample set and acquire the first characteristic gene sequence corresponding to each of the multiple sample cells in the cell sample set.

[0039] A gene pathway matching module is used to match multiple gene pathways for achieving different biological functions based on the multiple characteristic genes; each gene pathway includes various characteristic genes involved in achieving the biological function corresponding to the gene pathway;

[0040] The characterization module is used to determine the cell coding information corresponding to each of the sample cells by the cell characterization model to be trained based on each of the first feature gene sequences and each of the gene pathways, and to predict the second feature gene sequence corresponding to each of the sample cells based on the cell coding information.

[0041] The model adjustment module is used to adjust the cell representation model according to the difference between the second feature gene sequence and the first feature gene sequence corresponding to the sample cells until the training termination condition is met, thereby obtaining a pre-trained cell representation model.

[0042] Fourthly, this application also provides a cell downstream task processing device, the device comprising:

[0043] The task determination module is used to determine the downstream tasks for the target cell, obtain the characteristic genes corresponding to the target cell, and determine the characteristic gene sequence and multiple gene pathways corresponding to the target cell based on the characteristic genes corresponding to the target cell.

[0044] The target encoding information acquisition module is used to acquire the target cell encoding information of the target cell by a pre-trained cell characterization model based on the characteristic gene sequence and multiple gene pathways corresponding to the target cell; the cell characterization model is trained based on the cell characterization model pre-training method described above.

[0045] The task processing module is used to process the downstream tasks targeting the target cell based on the target cell encoding information.

[0046] Fifthly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the cell characterization model pre-training method or the steps of the cell downstream task processing method described above.

[0047] Sixthly, this application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the cell characterization model pre-training method or the steps of the cell downstream task processing method described above.

[0048] In a seventh aspect, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the cell characterization model pre-training method or the steps of the cell downstream task processing method described in any of the preceding claims.

[0049] The aforementioned cell characterization model pre-training method, cell downstream task processing method, device, computer equipment, computer-readable storage medium, and computer program product can determine multiple characteristic genes of a cell sample set and obtain the first characteristic gene sequences corresponding to each of the multiple sample cells in the cell sample set. Then, based on the multiple characteristic genes, multiple gene pathways for achieving different biological functions are obtained, wherein each gene pathway includes various characteristic genes involved in achieving the corresponding biological function. Subsequently, the cell characterization model to be trained can determine the cell coding information corresponding to each sample cell based on each first characteristic gene sequence and each gene pathway, and predict the second characteristic gene sequence corresponding to each sample cell based on the cell coding information. Based on the difference between the second characteristic gene sequence and the first characteristic gene sequence corresponding to the sample cell, the cell characterization model is adjusted until the training termination condition is met, and a pre-trained cell characterization model is obtained. In this embodiment, multiple gene pathways for achieving different biological functions are constructed based on multiple characteristic genes. The sequences of each first characteristic gene and each gene pathway are input into the cell representation model to be trained. The cell representation model determines the cell coding information corresponding to each sample cell based on the input information. This effectively combines the cell's own characteristic genes and gene pathways related to biological functions, which can reduce the noise content in the generated cell coding information and improve the signal-to-noise ratio, increase the distinguishability of cell coding information for cells, effectively optimize the representation effect of cell coding information, and improve the accuracy of cell representation results. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 This is a flowchart illustrating a pre-training method for a cell characterization model in one embodiment;

[0052] Figure 2 This is a flowchart illustrating another cell characterization model pre-training method in one embodiment;

[0053] Figure 3 This is an architectural diagram of a cell characterization model in one embodiment;

[0054] Figure 4 This is a schematic diagram of a gene pathway in one embodiment;

[0055] Figure 5 This is a flowchart illustrating a downstream task processing method for cells in one embodiment;

[0056] Figure 6 This is a flowchart illustrating another downstream task processing method for cells in one embodiment;

[0057] Figure 7 This is a structural block diagram of a cell characterization model pre-training device in one embodiment;

[0058] Figure 8 This is a structural block diagram of a cell downstream task processing device in one embodiment;

[0059] Figure 9 This is an internal structural diagram of a computer device in one embodiment;

[0060] Figure 10 This is an internal structural diagram of another computer device in one embodiment. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0062] To enable those skilled in the art to better understand this application, the relevant technologies are described below.

[0063] In fields such as biology and medicine, performance improvements in various cell-related downstream tasks often depend on high-quality cell characterization information.

[0064] Traditionally, cell characterization involves extracting ribonucleic acid (RNA) from cell or tissue samples using biochemical methods. Then, techniques such as real-time quantitative polymerase chain reaction (PCR), including reverse transcription quantitative polymerase chain reaction (RT-qPCR), digital PCR, or conventional PCR, such as RNA sequencing (RNA-Seq), are used to measure the expression levels of specific genes, thereby obtaining cell characterization information. Taking RT-qPCR as an example, this technology measures gene expression levels in real time by monitoring the fluorescence signal during deoxyribonucleic acid (DNA) amplification. This method can provide absolute quantification of gene expression. Furthermore, by comparing the cycle threshold (Ct value) of the target gene and a reference gene (such as a housekeeping gene), the relative expression level of the target gene can be calculated. Combined with relevant analysis methods (such as the standard curve method and the 2^-ΔΔCt method), researchers can analyze changes in gene expression and thus infer the gene's function and biological significance. However, these traditional methods for obtaining gene expression levels are often limited by experimental conditions and detection sensitivity, highly dependent on the experimental methods used, and generate significant data noise, making them inconvenient to use.

[0065] With the development of artificial intelligence (AI) technology, model-based, AI-driven scientific research has become a hot topic, finding widespread application in various tasks such as cell type classification, gene perturbation prediction, drug sensitivity prediction, drug response prediction, and cell target discovery. For example, graph-based representation methods construct cell networks to model cell or gene expression data as graph structures, where nodes represent individual cells or genes, and edges represent interactions between cells or genes. Algorithms such as Graph Neural Networks (GNNs) can learn low-dimensional representations of nodes in the graph, capturing complex relationships and network structures between cells, helping to understand the role of cells in biological processes. Among related technologies, Graph Convolutional Networks (GCNs) and Graph Attention Networks (GATs) have been used to analyze single-cell RNA sequencing data to identify cell subpopulations and infer cell state transitions. Furthermore, with the rise of AI models, large models that have achieved significant results in natural language processing are also being applied to cell data analysis. By pre-training these large models, cell feature vectors that allow for more refined cell differentiation and characterization can be obtained.

[0066] However, the inventors found in practice that the cell characterization information output by the model in the related technology has a high noise content, and the quality of cell characterization still needs to be improved.

[0067] Based on this, it is necessary to provide a cell characterization model pre-training method, a cell downstream task processing method, an apparatus, a computer device, a computer-readable storage medium, and a computer program product to address the above-mentioned technical problems.

[0068] In one embodiment, such as Figure 1 As shown, a pre-training method for a cell characterization model is provided. This embodiment illustrates the method using a terminal as an example. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0069] S101, determine multiple characteristic genes of the cell sample set, and obtain the first characteristic gene sequence corresponding to each of the multiple sample cells in the cell sample set.

[0070] In practical applications, a cell sample set consisting of cells of the target species can be obtained. The cell sample set includes multiple sample cells, each of which can be a single cell. The target species can be any species to be analyzed.

[0071] In one embodiment, the cell sample set may contain a large amount of single-cell gene sequencing data of the target species. The cell sample set includes multiple sample cells, each with corresponding sequencing data, which can characterize multiple genes possessed by the sample cell corresponding to the sequencing data. The representational characteristics of the dataset are related to the cell line itself. Some related technologies use a single dataset to train cell representation models. However, the data in a single dataset often only shows the gene expression levels of some genes, and non-expressed genes will not appear. Therefore, the transferability of models trained on a single dataset is poor. In this regard, in some examples of this application, the number of sample cells in the cell sample set can be at a preset level or higher, such as reaching tens of millions. By using massive cell data for subsequent pre-training of the cell representation model, the generalization ability of the model can be increased, which helps to more effectively construct single-cell representation methods applicable to various situations. This makes the single-cell representation results more stable, reduces the noise contained in the cell representation results, improves the discrimination of different cell types, and provides a reliable foundation for subsequent downstream cell tasks based on cell representation. It also helps to more efficiently handle downstream tasks that require high generalization ability, such as perturbation tasks like cell perturbation prediction.

[0072] In some exemplary embodiments, the cell sample set can be constructed by selecting non-repeating cell data from a publicly available dataset. For each sample cell provided in the publicly available dataset, a sample cell with corresponding original data can be selected. This original data includes the gene expression levels of the sample cell and related metadata. By selecting sample cells with original data, the overall integrity of the cell sample set can be improved. Then, the original data corresponding to the sample cells is populated to unify the data dimensions of each sample cell in the cell sample set.

[0073] After obtaining the cell sample set, characteristic genes within the set can be identified. Specifically, some genes exhibit relatively stable expression levels across different cells; that is, the expression levels of the same gene are the same or similar in all or most cells. These genes are expressed in all cells to similar degrees, and are also known as housekeeping genes. Housekeeping genes lack uniqueness and cannot be used to characterize cell samples. Other genes, however, show significant differences in expression levels across different cells. In this step, characteristic genes can be identified. Characteristic genes can be understood as specific genes whose expression levels differ across different cell types. Characterizing cells by identifying characteristic genes in the cell sample set helps improve the effectiveness of cell characterization and enhances the distinguishability between different cells.

[0074] In one embodiment, multiple characteristic genes of a cell sample set can be determined by the following steps: determining the cellular genes possessed by each sample cell in the cell sample set; determining hypervariable genes among multiple sample cell genes based on the gene expression level of each cellular gene in different sample cells; and obtaining multiple characteristic genes of the cell sample set based on the multiple hypervariable genes.

[0075] Among them, the gene expression level of hypervariable genes differed more than the threshold in different sample cells.

[0076] In practice, the original data of the cell sample set may contain a massive number of genes. For example, in one embodiment, the cell sample set contains 60,664 genes. During model training, the genes are converted into word embeddings, but the number of word embeddings provided in the model is often much smaller than the number of genes. In addition, if a massive number of genes are processed at the same time, more noise will be introduced into the cell characterization results, reducing the distinguishability of the cell characterization results for different cells.

[0077] In some embodiments, such as Figure 2 As shown, highly variable genes can be screened from a cell sample set. The cell specificity of highly variable genes often stems from functional differences in genes, and these genes are frequently associated with specific cell functions, making them ideal gene targets. In some examples, the Highly Variable Genes (HVG) algorithm can be used to select highly variable genes from the cell sample set. Then, multiple highly variable genes can be retained as feature genes; for example, a preset number of highly variable genes can be selected. Through highly variable gene screening, genes with significant expression variations across different cells will be retained, while housekeeping genes will be removed. Selecting these highly variable genes as characterization helps improve the effectiveness of cell characterization and enhances the distinguishability between different cells.

[0078] After identifying multiple characteristic genes corresponding to the cell sample set, the raw data of the sample cells can be filtered. For each sample cell, characteristic genes can be selected from the multiple genes corresponding to that sample cell. Then, based on the selected multiple characteristic genes, a characteristic gene sequence of that sample cell can be constructed. For ease of differentiation, this characteristic gene sequence is also called the first characteristic gene sequence. The first characteristic gene sequence can include the gene identifier (gene ID) and gene expression level of each of the multiple characteristic genes.

[0079] S102, based on matching multiple characteristic genes, multiple gene pathways are obtained to achieve different biological functions; each gene pathway includes various characteristic genes involved in achieving the corresponding biological function of the gene pathway.

[0080] On the other hand, after obtaining multiple characteristic genes from a cell sample set, multiple gene pathways can be matched based on these characteristic genes. For example, multiple characteristic genes can be matched with multiple pre-obtained gene pathways. For each pre-obtained gene pathway, if the multiple characteristic genes contain all the genes present in that gene pathway, then the gene pathway can be considered successfully matched. Each gene pathway contains multiple characteristic genes, and different gene pathways have at least some different characteristic genes. The multiple characteristic genes contained in the same gene pathway can be characteristic genes used to achieve related or similar biological functions. In other words, multiple characteristic genes that can achieve the same or similar biological functions can constitute a gene pathway.

[0081] In this step, by matching the combination of feature genes with multiple pre-acquired gene pathways, multiple feature gene sets can be obtained. Each feature gene set corresponds to a matched gene pathway, and multiple feature genes in the same feature gene set work together to achieve the same or related biological functions. In some exemplary embodiments, multiple pre-acquired gene pathways are embedded in the cell characterization model. By inputting multiple feature genes into the cell characterization model to be trained, the cell characterization model can match and obtain the relevant multiple gene pathways.

[0082] S103, the cell characterization model to be trained determines the cell coding information corresponding to each sample cell based on each first characteristic gene sequence and each gene pathway, and predicts the second characteristic gene sequence corresponding to each sample cell based on the cell coding information.

[0083] In practical applications, after obtaining the first characteristic gene sequence and multiple gene pathways corresponding to each of the multiple sample cells, the cell characterization model can characterize a single cell by combining the first characteristic gene sequence and gene pathways, thereby obtaining the cell encoding information (cell embedding) corresponding to the sample cell.

[0084] Since gene pathways describe the realization of biological functions and involve the coordination of multiple genes, they can characterize the collaborative effects between different characteristic genes in sample cells. Furthermore, cellular heterogeneity is often reflected through differences in gene pathway expression. Compared to related technologies that characterize cells solely based on gene sequences, this embodiment introduces gene pathways for cell characterization on top of the first characteristic gene sequence. This allows for the combination of different pathway information, reducing noise content in traditional input data, improving the signal-to-noise ratio, and characterizing sample cells in a differentiated manner based on their biological functions. This effectively enhances the discriminative power of the obtained cell coding information for sample cells, thus improving the training effect of model pre-training.

[0085] After obtaining cell coding information, the cell characterization model can re-predict the characteristic gene sequence of the sample cells based on the cell coding information to obtain the corresponding characteristic gene sequence. For easy differentiation, the characteristic gene sequence output by the cell characterization model is also called the second characteristic gene sequence.

[0086] Specifically, a cell representation model can include an encoder and a decoder. The input information of the cell representation model includes a first feature gene sequence, and multiple pre-acquired gene pathways can be embedded into the cell representation model. When the input information is input into the cell representation model, the encoder can obtain multiple gene pathways related to the sample based on the input first feature gene sequence and the matching results between the feature genes in the first feature gene sequence and the pre-embedded multiple gene pathways. Then, it can encode the first feature gene sequence and the matched multiple gene pathways to obtain the cell coding information corresponding to each sample cell. The decoder then decodes the cell coding information and transforms it into a second feature gene sequence via a fully connected network. In some exemplary embodiments, the cell representation model can be built based on the BERT (Bidirectional Encoder Representation from Transformers) model architecture, such as building a cell representation model based on the open-source model Geneformer. Figure 3 An architecture for a cell representation model is shown, which may include a self-attention layer, layer normalization, and a feedforward network.

[0087] S104. Based on the difference between the second characteristic gene sequence and the first characteristic gene sequence corresponding to the sample cells, adjust the cell representation model until the training termination condition is met, and obtain the pre-trained cell representation model.

[0088] After obtaining the second feature gene sequence output by the model, the second feature gene sequence can be compared with the first feature gene sequence to obtain the difference between the two sequences. Based on this difference, the model loss value is determined. For example, the cross-entropy can be calculated based on the difference between the second and first feature gene sequences corresponding to the sample cells, and this cross-entropy can be used as the model loss value. Then, the model parameters of the cell representation model are adjusted based on the model loss value to optimize the parameters. Repeating the above steps until the training termination condition is met yields a pre-trained cell representation model. After training is complete, a model file corresponding to the cell representation model can be generated and saved, containing the necessary information.

[0089] In the aforementioned pre-training method for cell characterization models, multiple characteristic genes of a cell sample set can be identified, and the first characteristic gene sequences corresponding to each of the multiple sample cells in the cell sample set can be obtained. Then, multiple gene pathways for achieving different biological functions are obtained based on the matching of multiple characteristic genes, wherein each gene pathway includes various characteristic genes involved in achieving the biological function corresponding to the gene pathway. Afterwards, the cell characterization model to be trained can determine the cell coding information corresponding to each sample cell based on each first characteristic gene sequence and each gene pathway, and predict the second characteristic gene sequence corresponding to each sample cell based on the cell coding information. Based on the difference between the second characteristic gene sequence and the first characteristic gene sequence corresponding to the sample cell, the cell characterization model is adjusted until the training termination condition is met, and the pre-trained cell characterization model is obtained. In this embodiment, multiple gene pathways for achieving different biological functions are constructed based on multiple characteristic genes. The sequences of each first characteristic gene and each gene pathway are input into the cell representation model to be trained. The cell representation model determines the cell coding information corresponding to each sample cell based on the input information. This effectively combines the cell's own characteristic genes and gene pathways related to biological functions, which can reduce the noise content in the generated cell coding information and improve the signal-to-noise ratio, increase the distinguishability of cell coding information for cells, effectively optimize the representation effect of cell coding information, and improve the accuracy of cell representation results.

[0090] In one exemplary embodiment, multiple gene pathways can be obtained in advance through the following steps:

[0091] S201 involves inputting multiple characteristic genes into multiple pathways to construct networks; each pathway network corresponds to a different pathway construction rule, and the different pathway construction rules characterize the realization process of biological functions based on different dimensions.

[0092] In practice, gene pathways can be constructed based on a single dimension or multiple different gene pathways based on multiple dimensions. The construction of gene pathways can be achieved through pathway construction networks. In some exemplary embodiments, the pathway construction network can be a graph attention network (GAN).

[0093] In this embodiment, multiple pathway construction networks can be pre-trained or specified. Each pathway construction network corresponds to a different pathway construction rule. These multiple pathway construction rules can characterize the realization process of biological functions from different dimensions. For example, pathway construction rules can be set from the dimensions of cellular components, molecular functions, and biological processes, or from the dimension of cell signal transduction to metabolic pathways, or from the dimension of cellular immunity or cellular transcription. In specific implementation, pathway construction rules can be set according to the actual situation and the corresponding pathway construction network can be obtained.

[0094] S201, each pathway constructs a network to determine multiple characteristic genes that realize related biological functions among multiple characteristic genes according to the corresponding pathway construction rules, and outputs the corresponding undirected graph based on multiple characteristic genes that realize the same biological function.

[0095] After obtaining multiple pathway construction networks, for each pathway construction network, multiple feature genes can be input into the pathway construction network. The pathway construction network then determines multiple feature genes that are used to achieve the same or related biological functions from the multiple feature genes according to the corresponding pathway construction rules, and then constructs and outputs the corresponding undirected graph.

[0096] In this undirected graph, feature genes are the nodes, and the edges depict the connections between these nodes. In some examples, the weights of the edges in the undirected graph can be trainable variables for constructing the network. For example... Figure 4 As shown, four undirected graphs were constructed using networks built through four pathways for five characteristic genes (g1, g2, g3, g4, and g5). These four undirected graphs represent the relationships between the five characteristic genes from different dimensions. It is understandable that... Figure 4 The undirected graph shown is only an example. An undirected graph does not necessarily contain all the characteristic genes. For example, if a characteristic gene is not used to achieve a certain biological function, then the characteristic gene may not appear on the undirected graph.

[0097] S203, based on the output undirected graphs, obtain the gene pathways of multiple characteristic genes.

[0098] After obtaining an undirected graph output by a network constructed from multiple pathways, multiple gene pathways can be derived from these undirected graphs. Specifically, for an undirected graph, gene pathways can be determined based on the characteristic genes corresponding to the nodes in the graph and the relationships between the nodes.

[0099] In this embodiment, on the one hand, the gene pathways between feature genes are generated by the trained pathway construction network, which helps to integrate gene pathway knowledge into the cell characterization process. Gene pathways are quickly generated based on the pathway construction rules learned by the pathway construction network. On the other hand, by using multiple pathway construction rules to generate different gene pathways, the realization process of biological functions can be described from multiple dimensions, the diversity of gene pathway information can be improved, and it helps to better utilize multiple gene pathways to characterize cells.

[0100] In one embodiment, step S201, where each pathway construction network determines multiple characteristic genes that perform associated biological functions from among multiple characteristic genes according to corresponding pathway construction rules, may include:

[0101] If the multiple pathway construction network includes a first pathway construction network, then the first pathway construction network determines the multiple characteristic genes that realize the associated biological function from among the multiple characteristic genes according to the pathway construction rules related to gene ontology; if the multiple pathway construction network includes a second pathway construction network, then the second pathway construction network determines the multiple characteristic genes that realize the associated biological function from among the multiple characteristic genes according to the pathway construction rules related to cell biological processes; if the multiple pathway construction network includes a third pathway construction network, then the third pathway construction network determines the multiple characteristic genes that realize the associated biological function from among the multiple characteristic genes according to the pathway construction rules related to immunity; if the multiple pathway construction network includes a fourth pathway construction network, then the fourth pathway construction network determines the multiple characteristic genes that realize the associated biological function from among the multiple characteristic genes according to the pathway construction rules related to cell transcription.

[0102] In this embodiment, four pathway construction networks can be pre-acquired, allowing the integration of four types of gene pathways as knowledge into the cell representation model. These four pathways are: Gene Ontology, Reactome, Immune, and Transcription Factors. For ease of distinction, the pathway construction network for the Gene Ontology pathway is referred to as the first pathway construction network, the network for the Reactome pathway as the second pathway construction network, the network for the Immune pathway as the third pathway construction network, and the network for the Transcription Factors pathway as the fourth pathway construction network.

[0103] Furthermore, when multiple characteristic genes are input into the first pathway construction network, the first pathway construction network can construct pathways from three directions—biological processes, molecular functions, and cellular components—according to pathway construction rules related to the gene ontology, thereby identifying multiple characteristic genes that realize associated biological functions among the multiple characteristic genes.

[0104] When multiple characteristic genes are input into the second pathway construction network, the second pathway construction network can identify characteristic genes related to various biological processes such as cell signal transduction and metabolism from multiple characteristic genes according to the pathway construction rules related to cell biological processes. It can construct pathways that can characterize cell biological processes and demonstrate cell heterogeneity by characterizing processes such as metabolism, cell cycle and apoptosis.

[0105] When multiple characteristic genes are input into the third pathway construction network, the third pathway construction network can identify multiple characteristic genes that realize related biological functions among the multiple characteristic genes according to the immune-related pathway construction rules. This network mainly involves constructing pathways related to the activation, proliferation, differentiation or effector functions of immune cells. The pathway information output by the third pathway construction network can effectively distinguish immune cells from other cells.

[0106] When multiple characteristic genes are input into the fourth pathway construction network, the fourth pathway construction network can identify multiple characteristic genes that realize related biological functions based on the pathway construction rules related to cell transcription, thereby constructing gene pathways related to cell proliferation, differentiation and signal response functions. These gene pathways are also related to the specificity exhibited by cells.

[0107] In this embodiment, by constructing networks through the first pathway, the second pathway, the third pathway, and the fourth pathway, we can obtain various pathway knowledge related to gene pathways and quickly construct the corresponding gene pathways.

[0108] In one embodiment, in step S103, the cell characterization model to be trained determines the cell coding information corresponding to each sample cell based on each first feature gene sequence and each gene pathway, which may include the following steps:

[0109] S301, construct a first matrix based on the first characteristic gene sequence corresponding to each of the multiple sample cells and the gene expression level of each characteristic gene in each pre-acquired first characteristic gene sequence; each row of the first matrix corresponds to a sample cell, each column corresponds to a characteristic gene, and each matrix element represents the gene expression level of the characteristic gene of the sample cell.

[0110] In practical applications, the gene expression level corresponding to the gene in each sample cell can be obtained. Then, after obtaining the first characteristic gene sequences corresponding to multiple sample cells, a first matrix can be constructed based on each first characteristic gene sequence and the gene expression level of each characteristic gene within that sequence. The first matrix consists of M rows and N columns, where M corresponds to the total number of sample cells (each row corresponds to one sample cell), N corresponds to the total number of characteristic genes, and each column corresponds to one characteristic gene. Each matrix element represents the gene expression level of a specific sample cell on its corresponding characteristic gene. For example, the matrix element (M1, N2) represents the specific gene expression level of sample cell M1 on characteristic gene N2.

[0111] S302, construct a second matrix based on each gene pathway; each row in the second matrix corresponds to a feature gene, each column corresponds to a gene pathway, and each matrix element represents whether the feature gene is included in the gene pathway.

[0112] On the other hand, a second matrix can be constructed based on each gene pathway. This second matrix consists of P rows and Q columns, where P corresponds to the total number of characteristic genes, with each row corresponding to one characteristic gene, and Q corresponds to the total number of gene pathways, with each column corresponding to one gene pathway. Each matrix element indicates whether a characteristic gene is included in a certain gene pathway. For example, the matrix element (P1, Q2) indicates whether characteristic gene P1 is in gene pathway Q2. In some examples, the values ​​of matrix elements can be 0 or 1, where 0 indicates that the gene does not appear in the gene pathway, and 1 indicates that the gene pathway does not appear. Accordingly, the second matrix can be a matrix consisting entirely of 0s and 1s.

[0113] S303: The cell representation model to be trained extracts features from the product of the first matrix and the second matrix, and obtains the cell coding information corresponding to each sample cell based on the feature extraction results.

[0114] After obtaining the first and second matrices, they can be input into the cell representation model to be trained. The cell representation model obtains the product of the first and second matrices. It can be understood that by multiplying the first and second matrices, a matrix composed of sample cells and gene pathways can be obtained. That is, sample cells can be represented through gene pathways. Then, the cell representation model can extract features from the matrix product and obtain the cell coding information corresponding to each sample cell based on the feature extraction results.

[0115] In this embodiment, sample cells combining gene pathway characterization can be obtained based on the product of the first matrix and the second matrix, effectively introducing multiple gene pathways for cell characterization and effectively increasing the distinguishability of cell coding information for sample cells.

[0116] In one embodiment, there are multiple second matrices. Gene pathways in the same second matrix are constructed based on the same pathway construction rules, while gene pathways in different second matrices correspond to different pathway construction rules. For example, for gene pathways output by multiple pathway construction networks, a second matrix can be constructed based on multiple gene pathways output by the same pathway construction network, thereby obtaining multiple second matrices corresponding to different pathway construction networks.

[0117] Accordingly, in S303, feature extraction of the product of the first and second matrices by the cell representation model to be trained may include the following steps:

[0118] The cell characterization model obtains the product of the first matrix and each second matrix to obtain multiple third matrices; multiple feature extraction modules in the cell characterization model extract features from each third matrix according to the self-attention mechanism to obtain the sub-feature extraction results of each feature extraction module; multiple feature extraction modules are used to extract features from the third matrices constructed based on different pathway construction rules; the feature extraction result is obtained based on the fusion result of each sub-feature extraction result.

[0119] In this embodiment, a multi-head attention mechanism can be used to extract features from the product of the first and second matrices. Specifically, after obtaining multiple second matrices, the first matrix can be multiplied by each of the multiple second matrices to obtain multiple third matrices. For example, for the first matrix A, the second matrix B1, and the second matrix B2, multiplying the first matrix A and the second matrix B1 yields the third matrix C1, and multiplying the first matrix A and the second matrix B2 yields the third matrix C2. It can be understood that the third matrices C1 and C2 are constructed based on different path construction rules.

[0120] The cell conditioning model can include multiple feature extraction modules, which are used to extract features from the third matrix constructed based on different pathway construction rules. Furthermore, each feature extraction module can separately extract features from its corresponding third matrix, resulting in sub-feature extraction results. For example, four feature extraction modules can be set up to extract features from the third matrix corresponding to the Gene Ontology pathway, the Reactome pathway, the Immune pathway, and the Transcription Factors pathway, respectively, yielding four sub-feature extraction results.

[0121] Furthermore, the results of multiple sub-feature extractions can be fused to obtain feature extraction results that integrate information from multiple pathways.

[0122] In this embodiment, multiple feature extraction modules can be used to extract features from the third matrix constructed based on different pathway construction rules in parallel. This enables the cell representation model to learn different types of pathway knowledge, capture richer contextual information, and improve the global understanding of sample cells and gene pathways. By fusing the results of multiple sub-feature extractions to obtain the final feature extraction result, the limitations of local view in single feature extraction can be avoided, allowing the model to handle gene pathways of different cells more flexibly and enhancing the model's adaptability and generalization performance.

[0123] In one embodiment, in step S103, the cell characterization model to be trained determines the cell coding information corresponding to each sample cell based on each first feature gene sequence and each gene pathway, which may include the following steps:

[0124] A portion of the characteristic genes in each first characteristic gene sequence are masked to obtain the corresponding characteristic gene mask sequence. The cell characterization model to be trained then determines the cell coding information corresponding to each sample cell based on the characteristic gene mask sequence and each gene pathway.

[0125] In specific implementations, partial feature genes in each first feature gene sequence can be masked to obtain a feature gene mask sequence corresponding to each first feature gene sequence. In some embodiments, partial feature genes in each first feature gene sequence can be randomly masked, thereby improving the generalization ability of the cell characterization model and enhancing its robustness. The order of each feature gene in the feature gene mask sequence is the same as the order of each feature gene in the first feature gene sequence, but some feature genes are masked, such as... Figure 4 As shown, during random masking, both the gene identifier and gene expression level of the characteristic gene are masked.

[0126] Once the feature gene mask sequence is obtained, the cell characterization model to be trained can determine the cell coding information corresponding to each sample cell based on each feature gene mask sequence and each gene pathway. Specifically, the feature gene mask sequence can be regarded as the first feature gene sequence mentioned above, and then the cell coding information can be determined based on each first feature gene sequence and gene pathway in one or more of the previous embodiments, which will not be elaborated here.

[0127] Accordingly, in step S103, the second characteristic gene sequence corresponding to each sample cell is predicted based on the cell coding information, including the following steps:

[0128] Based on the cell coding information, the masked portion of each characteristic gene mask sequence is used to predict the characteristic genes, and the second characteristic gene sequence corresponding to each sample cell is obtained based on the prediction results.

[0129] The cell characterization model can determine the cell coding information corresponding to each sample cell based on the feature gene mask sequence and each gene pathway. The specific determination method can be referred to the aforementioned embodiments, and will not be repeated here. Then, the cell characterization model predicts the masked feature gene based on the cell coding information and the context information of the masked part in the feature gene mask sequence, thereby predicting the masked feature gene based on the prediction result and obtaining the second feature gene sequence.

[0130] In this embodiment, by masking the first feature gene sequence and predicting the masked portion based on cell coding information to obtain the second feature gene sequence, the cell representation model can make fuller use of the contextual information in the feature gene mask sequence to predict the masked feature gene, thereby learning a more accurate cell representation.

[0131] In one embodiment, such as Figure 5 As shown, this application also provides a method for downstream task processing in cells. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0132] S501, determine the downstream tasks for the target cell, obtain the characteristic genes corresponding to the target cell, and determine the characteristic gene sequence and multiple gene pathways corresponding to the target cell based on the characteristic genes corresponding to the target cell.

[0133] In this step, after receiving the downstream task for the target cell, the initial characteristic gene sequence corresponding to the target cell can be obtained. In practical applications, for the downstream task, after obtaining the cell data of the target cell, the characteristic genes of the target cell can be extracted first, and data alignment can be performed. Specifically, the data of the target cell is first aligned with the data in the cell sample set. The alignment process includes determining the characteristic genes of the target cell. These characteristic genes are the characteristic genes determined when processing the cell sample set. In other words, for each gene of the target cell, if the gene matches any characteristic gene in the cell sample set, it will be retained; if the gene does not match any characteristic gene in the cell sample set, it will be removed. Then, the initial characteristic gene sequence corresponding to the target cell is obtained based on the retained genes.

[0134] After identifying the characteristic genes of the target cell, such as Figure 6 As shown, the characteristic gene sequence of the target cell can be obtained based on the characteristic gene of the target cell, and matched with multiple pre-obtained gene pathways to determine multiple gene pathways composed of the characteristic genes of the target cell. The method for determining the characteristic gene sequence of the target cell and the multiple gene pathways can be the same as the acquisition method in the previous embodiment. The specific implementation method can be referred to the previous text, and will not be repeated here.

[0135] S502, a trained cell characterization model, obtains the target cell coding information based on the characteristic gene sequence and multiple gene pathways corresponding to the target cell.

[0136] The cell characterization model is trained based on the cell characterization model pre-training method of any of the above embodiments.

[0137] In this step, a pre-trained cell representation model can generate target cell coding information based on the characteristic gene sequence and multiple gene pathways corresponding to the target cell. In some embodiments, since the original output of the pre-trained cell representation model is the second characteristic gene sequence, but the downstream task processing result may not be limited to the characteristic gene sequence, the last linear layer used for linear mapping in the cell representation model can be removed, and the target cell coding information determined by the cell representation model can be directly output as the processing result. Thus, different cells can be effectively represented by the target cell coding information output by the cell representation model. Of course, in other embodiments, considering the adjustment of model parameters, the pre-trained cell representation model can be redefined and combined with new neural network layers after loading to directly generate downstream output.

[0138] S503 processes downstream tasks targeting the target cell based on the target cell's encoded information.

[0139] After obtaining the target cell encoding information, the noise data contained in the target cell encoding information is significantly reduced, and more effective information is preserved or supplemented, so that the target cell encoding information can effectively characterize a single cell, and then the target cell encoding information can be used to process downstream tasks targeting the target cell.

[0140] In some exemplary embodiments, the cell representation model can be structurally adjusted. For example, if the cell representation model includes transformer layers, a fully connected layer or embedding layer can be added after the transformer layers. The fully connected layer or embedding layer can then transform the target cell encoding information to meet the specific output requirements of downstream tasks, such as obtaining feature gene sequences or cell classification results. For drug response and cell perturbation prediction tasks, an initial state representing single-cell changes can be constructed by combining inputs, and then predictions can be made based on the definition of the downstream model.

[0141] In the aforementioned downstream task processing method for cells, downstream tasks targeting the target cell can be determined, the characteristic genes corresponding to the target cell can be obtained, and the characteristic gene sequences and multiple gene pathways corresponding to the target cell can be determined based on the characteristic genes. Then, a pre-trained cell representation model obtains the target cell coding information based on the characteristic gene sequences and multiple gene pathways corresponding to the target cell. This cell representation model is trained using the aforementioned cell representation model pre-training method, and then the downstream tasks targeting the target cell are processed based on the target cell coding information. In this embodiment, the pre-trained cell representation model can obtain target cell coding information that can accurately represent and distinguish the target cell based on the characteristic gene sequences and multiple gene pathways corresponding to the target cell. By applying this target cell coding information to the downstream task, the task processing results of the downstream task can be effectively optimized.

[0142] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0143] Based on the same inventive concept, this application also provides a cell characterization model pre-training device for implementing the cell characterization model pre-training method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more embodiments of the cell characterization model pre-training device provided below can be found in the limitations of the cell characterization model pre-training method described above, and will not be repeated here.

[0144] In one exemplary embodiment, such as Figure 7 As shown, a cell characterization model pre-training device is provided, comprising:

[0145] The first sequence acquisition module 701 is used to determine multiple characteristic genes of the cell sample set and acquire the first characteristic gene sequence corresponding to each of the multiple sample cells in the cell sample set.

[0146] The gene pathway matching module 702 is used to obtain multiple gene pathways for achieving different biological functions based on the multiple feature genes; each gene pathway includes various feature genes involved in achieving the biological function corresponding to the gene pathway;

[0147] The characterization module 703 is used to determine the cell coding information corresponding to each of the sample cells by the cell characterization model to be trained based on each of the first feature gene sequences and each of the gene pathways, and to predict the second feature gene sequence corresponding to each of the sample cells based on the cell coding information.

[0148] The model adjustment module 704 is used to adjust the cell representation model according to the difference between the second feature gene sequence and the first feature gene sequence corresponding to the sample cells until the training termination condition is met, so as to obtain a pre-trained cell representation model.

[0149] In one embodiment, the apparatus further includes a gene pathway acquisition module, the gene pathway acquisition module being used for:

[0150] The multiple feature genes are input into multiple pathway construction networks; each of the multiple pathway construction networks corresponds to a different pathway construction rule, and the different pathway construction rules characterize the realization process of biological functions based on different dimensions.

[0151] Each pathway is used to construct a network, which determines multiple characteristic genes that perform related biological functions among the multiple characteristic genes according to the corresponding pathway construction rules, and outputs a corresponding undirected graph based on the multiple characteristic genes that perform the same biological function.

[0152] Based on the output of each undirected graph, the gene pathways of the multiple characteristic genes are obtained.

[0153] In one embodiment, the gene pathway acquisition module is used for:

[0154] If the plurality of pathway construction networks includes a first pathway construction network, then the first pathway construction network determines the plurality of characteristic genes that realize the associated biological functions among the plurality of characteristic genes according to the pathway construction rules related to the gene ontology;

[0155] If the plurality of pathway construction networks includes a second pathway construction network, then the second pathway construction network determines the plurality of characteristic genes that realize the associated biological functions among the plurality of characteristic genes according to the pathway construction rules related to cell biological processes;

[0156] If the multiple pathway construction network includes a third pathway construction network, then the third pathway construction network determines multiple characteristic genes that realize related biological functions among the multiple characteristic genes according to the pathway construction rules related to immunity.

[0157] If the plurality of pathway construction networks include a fourth pathway construction network, the fourth pathway construction network determines the plurality of characteristic genes that realize the associated biological functions among the plurality of characteristic genes according to the pathway construction rules related to cell transcription.

[0158] In one embodiment, the characterization module 703 is used for:

[0159] A first matrix is ​​constructed based on the first characteristic gene sequence corresponding to each of the multiple sample cells and the gene expression level of each characteristic gene in each of the pre-acquired first characteristic gene sequences; each row of the first matrix corresponds to one sample cell, each column corresponds to one characteristic gene, and each matrix element represents the gene expression level of the characteristic gene of the sample cell.

[0160] A second matrix is ​​constructed based on each of the gene pathways; each row of the second matrix corresponds to a characteristic gene, each column corresponds to a gene pathway, and each matrix element represents whether the characteristic gene is included in the gene pathway;

[0161] The cell characterization model to be trained extracts features from the product of the first matrix and the second matrix, and obtains the cell coding information corresponding to each sample cell based on the feature extraction results.

[0162] In one embodiment, the second matrix includes multiple matrices, and the gene pathways in the same second matrix are constructed based on the same pathway construction rules. The gene pathways in different second matrices correspond to different pathway construction rules.

[0163] The characterization module 703 is used for:

[0164] The product of the first matrix and each of the second matrices is obtained from the cell characterization model to obtain multiple third matrices;

[0165] The cell characterization model uses multiple feature extraction modules to extract features from each of the third matrices based on a self-attention mechanism, resulting in sub-feature extraction results for each of the multiple feature extraction modules. The multiple feature extraction modules are used to extract features from the third matrices constructed based on different pathway construction rules.

[0166] The feature extraction result is obtained by fusing the results of each sub-feature extraction.

[0167] In one embodiment, the represented module 703 is used for:

[0168] A partial feature gene in each of the first feature gene sequences is masked to obtain the corresponding feature gene mask sequence. The cell characterization model to be trained determines the cell coding information corresponding to each sample cell based on each feature gene mask sequence and each gene pathway.

[0169] Based on the cell coding information, feature gene prediction is performed on the masked portion of each feature gene mask sequence, and the second feature gene sequence corresponding to each sample cell is obtained based on the prediction results.

[0170] Based on the same inventive concept, this application also provides a downstream cell task processing apparatus for implementing the aforementioned downstream cell task processing method. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the downstream cell task processing apparatus provided below can be found in the limitations of the downstream cell task processing method described above, and will not be repeated here.

[0171] In one exemplary embodiment, such as Figure 8 As shown, a cell-based downstream task processing device is provided, comprising:

[0172] The task determination module 801 is used to determine the downstream task for the target cell, obtain the characteristic gene corresponding to the target cell, and determine the characteristic gene sequence and multiple gene pathways corresponding to the target cell based on the characteristic gene corresponding to the target cell.

[0173] The target encoding information acquisition module 802 is used to acquire the target cell encoding information of the target cell by a pre-trained cell characterization model based on the characteristic gene sequence and multiple gene pathways corresponding to the target cell; the cell characterization model is trained based on the cell characterization model pre-training method described above.

[0174] The task processing module 803 is used to process the downstream tasks targeting the target cell based on the target cell encoding information.

[0175] Each module in the aforementioned cell characterization model pre-training device and cell downstream task processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0176] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores cell data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network. When executed by the processor, the computer program implements a cell characterization model pre-training method or a cell downstream task processing method.

[0177] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 10 As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a cell characterization model pre-training method or a cell downstream task processing method. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0178] Those skilled in the art will understand that Figure 9 and Figure 10The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0179] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0180] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0181] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0182] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0183] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0184] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0185] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for pre-training a cell characterization model, characterized in that, The method includes: Identify multiple characteristic genes in a cell sample set, and obtain the first characteristic gene sequence corresponding to each of the multiple sample cells in the cell sample set; Multiple gene pathways for achieving different biological functions are obtained by matching the multiple feature genes. These gene pathways are acquired through the following steps: inputting the multiple feature genes into multiple pathway construction networks; each pathway construction network corresponds to a different pathway construction rule, and these different rules characterize the implementation process of biological functions based on different dimensions; each pathway construction network determines multiple feature genes that achieve associated biological functions from among the multiple feature genes according to the corresponding pathway construction rule, and outputs a corresponding undirected graph based on these multiple feature genes; based on the output undirected graphs, each gene pathway of the multiple feature genes is obtained; each gene pathway includes the various feature genes involved in achieving the corresponding biological function of the gene pathway. The cell characterization model to be trained determines the cell coding information corresponding to each of the sample cells based on each of the first feature gene sequences and each of the gene pathways, and predicts the second feature gene sequence corresponding to each of the sample cells based on the cell coding information; Based on the difference between the second feature gene sequence and the first feature gene sequence corresponding to the sample cells, the cell representation model is adjusted until the training termination condition is met, thus obtaining a pre-trained cell representation model. Wherein, the cell characterization model to be trained determines the cell coding information corresponding to each of the sample cells based on each of the first feature gene sequences and each of the gene pathways, including: A first matrix is ​​constructed based on the first characteristic gene sequences corresponding to multiple sample cells and the gene expression levels of each characteristic gene in each of the pre-acquired first characteristic gene sequences. Each row of the first matrix corresponds to one sample cell, each column corresponds to one characteristic gene, and each matrix element represents the gene expression level of the characteristic gene in the sample cell. A second matrix is ​​constructed based on each gene pathway. Each row of the second matrix corresponds to one characteristic gene, each column corresponds to one gene pathway, and each matrix element represents whether the characteristic gene is included in the gene pathway. The cell representation model to be trained extracts features from the product of the first matrix and the second matrix, and obtains the cell coding information corresponding to each sample cell based on the feature extraction results. The process of determining cell coding information corresponding to each sample cell based on each of the first feature gene sequences and each of the gene pathways by the cell characterization model to be trained, and predicting the second feature gene sequence corresponding to each sample cell based on the cell coding information, includes: A portion of the characteristic genes in each of the first characteristic gene sequences are masked to obtain the corresponding characteristic gene mask sequence. The cell characterization model to be trained determines the cell coding information corresponding to each of the sample cells based on the characteristic gene mask sequence and the gene pathway. Based on the cell coding information, the masked portion of each characteristic gene mask sequence is predicted, and the second characteristic gene sequence corresponding to each of the sample cells is obtained based on the prediction result.

2. The method according to claim 1, characterized in that, The process of constructing a network from each pathway, according to corresponding pathway construction rules, determines multiple characteristic genes among the multiple characteristic genes that realize associated biological functions, including: If the plurality of pathway construction networks includes a first pathway construction network, then the first pathway construction network determines the plurality of characteristic genes that realize the associated biological functions among the plurality of characteristic genes according to the pathway construction rules related to the gene ontology; If the plurality of pathway construction networks includes a second pathway construction network, then the second pathway construction network determines the plurality of characteristic genes that realize the associated biological functions among the plurality of characteristic genes according to the pathway construction rules related to cell biological processes; If the multiple pathway construction network includes a third pathway construction network, then the third pathway construction network determines multiple characteristic genes that realize related biological functions among the multiple characteristic genes according to the pathway construction rules related to immunity. If the plurality of pathway construction networks include a fourth pathway construction network, the fourth pathway construction network determines the plurality of characteristic genes that realize the associated biological functions among the plurality of characteristic genes according to the pathway construction rules related to cell transcription.

3. The method according to claim 1, characterized in that, The second matrix includes multiple matrices. Gene pathways in the same second matrix are constructed based on the same pathway construction rules, while gene pathways in different second matrices correspond to different pathway construction rules. The feature extraction from the product of the first matrix and the second matrix by the cell representation model to be trained includes: The product of the first matrix and each of the second matrices is obtained from the cell representation model to be trained, resulting in multiple third matrices; The cell characterization model uses multiple feature extraction modules to extract features from each of the third matrices based on a self-attention mechanism, resulting in sub-feature extraction results for each of the multiple feature extraction modules. The multiple feature extraction modules are used to extract features from the third matrices constructed based on different pathway construction rules. The feature extraction result is obtained by fusing the results of each sub-feature extraction.

4. A method for downstream cellular task processing, characterized in that, The method includes: Determine the downstream tasks for the target cell, obtain the characteristic genes corresponding to the target cell, and determine the characteristic gene sequence and multiple gene pathways corresponding to the target cell based on the characteristic genes corresponding to the target cell; The target cell encoding information of the target cell is obtained by a pre-trained cell characterization model based on the characteristic gene sequence and multiple gene pathways corresponding to the target cell; the cell characterization model is trained based on the cell characterization model pre-training method according to any one of claims 1 to 3. The downstream tasks targeting the target cell are processed based on the target cell's encoded information.

5. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the cell characterization model pre-training method according to any one of claims 1 to 3 or the steps of the cell downstream task processing method according to claim 4.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the cell characterization model pre-training method according to any one of claims 1 to 3 or the steps of the cell downstream task processing method according to claim 4.

7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the cell characterization model pre-training method according to any one of claims 1 to 3 or the steps of the cell downstream task processing method according to claim 4.