A transcriptome data analysis method, device, equipment and readable storage medium

CN122531497APending Publication Date: 2026-08-07JINAN UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JINAN UNIVERSITY
Filing Date
2026-05-12
Publication Date
2026-08-07

AI Technical Summary

Benefits of technology

[0014]As can be seen from the above technical solutions, the transcriptome data analysis method, apparatus, device, and readable storage medium provided in this application include: providing raw transcriptome data and sample batch source information of biological samples; processing the raw transcriptome data to obtain a gene expression matrix; obtaining standardized model input data based on the gene expression matrix and sample batch source information; inputting the standardized model input data into a transcriptome data analysis model to obtain a universal transcriptome hidden representation; the transcriptome data analysis model is configured to process the standardized model input data to obtain the universal transcriptome hidden representation, and to obtain the predicted results of masked gene expression values ​​based on the universal transcriptome hidden representation; and to obtain the prediction results corresponding to the pre-set analysis task using the universal transcriptome hidden representation based on a pre-set analysis task. This application achieves transcriptome data analysis by processing raw transcriptome data to obtain standardized model input data, inputting the standardized model input data into a transcriptome data analysis model to obtain a universal transcriptome hidden representation within the model, and finally obtaining the prediction results corresponding to the pre-set analysis task using the universal transcriptome hidden representation according to the pre-set analysis task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531497A_ABST
    Figure CN122531497A_ABST
Patent Text Reader

Abstract

The application discloses a transcriptome data analysis method, device, equipment and readable storage medium, comprising: providing transcriptome original data and sample batch source information, processing the transcriptome original data to obtain a gene expression matrix, obtaining standardized model input data based on the gene expression matrix and the sample batch source information, inputting the standardized model input data into a transcriptome data analysis model to obtain a general transcriptome hidden representation, and obtaining a prediction result corresponding to a pre-set analysis task based on the general transcriptome hidden representation. The application processes transcriptome data to obtain standardized model input data, inputs the model input data into a transcriptome data analysis model to obtain a general transcriptome hidden representation inside the model, and finally obtains a prediction result corresponding to a pre-set analysis task based on the general transcriptome hidden representation, thereby realizing analysis on transcriptome data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data analysis technology, and in particular to a transcriptome data analysis method, apparatus, device, and readable storage medium. Background Technology

[0002] Transcriptome data, as the core carrier of gene expression information in organisms, can accurately reflect gene expression levels and regulatory states. It has irreplaceable and important applications in many bioinformatics fields, such as disease subtyping, drug sensitivity prediction, cell composition analysis, and gene regulatory mechanism research. It serves as a crucial bridge connecting gene sequences and biological phenotypes, and is of great significance for advancing life science research, clinical diagnosis, and drug development. In recent years, neural network models have rapidly developed in data mining and analysis across various fields due to their powerful feature extraction, nonlinear fitting, and global dependency capture capabilities. Compared to traditional statistical analysis methods, neural network models can effectively handle high-dimensional, sparse, and heterogeneous complex data, uncovering deep correlations hidden within the data and significantly improving analytical accuracy and efficiency. Therefore, how to utilize neural network models to analyze transcriptome data has been a topic of ongoing interest. Summary of the Invention

[0003] In view of this, this application provides a method, apparatus, device, and readable storage medium for transcriptome data analysis to facilitate the analysis of transcriptome data.

[0004] To achieve the above objectives, the following solution is proposed: A transcriptome data analysis method, comprising: Provide raw transcriptome data and batch origin information for biological samples; The raw transcriptome data were processed to obtain the gene expression matrix; Based on the gene expression matrix and sample batch source information, standardized model input data is obtained. The standardized model input data includes: the masked standardized gene expression matrix, mask position index information, gene-level protein sequence embedding features, and batch identification information. The standardized model input data is input into the transcriptome data analysis model to obtain a universal transcriptome hidden representation. The transcriptome data analysis model is configured to process the standardized model input data to obtain a universal transcriptome hidden representation and, based on the universal transcriptome hidden representation, obtain the predicted expression values ​​of the masked genes. Based on a pre-defined analysis task, the prediction results corresponding to the pre-defined analysis task are obtained using a universal transcriptome hidden representation.

[0005] Optional transcriptome data analysis models include: a cascaded input layer, an embedding layer, a feature modeling layer, and an output layer; By inputting the standardized model data into the transcriptome data analysis model, a universal transcriptome hidden representation is obtained, including: The standardized model input data is obtained through the input layer; By using an embedding layer, the masked standardized gene expression matrix, gene-level protein sequence embedding features, and batch identification information are processed to obtain fused features; By modeling the fused features through the feature modeling layer, a general transcriptome hidden representation is obtained; The output layer, based on the universal transcriptome hidden representation and mask position index information, yields the predicted expression values ​​of the masked genes.

[0006] Optionally, the feature modeling layer includes: a static graph structure modeling sublayer and a multi-layer global latent attention sublayer. Through the feature modeling layer, the fused features are modeled to obtain a general transcriptome hidden representation, including: By modeling sub-layers using a static graph structure, a normalized adjacency matrix is ​​obtained based on a pre-constructed gene relationship graph. The normalized adjacency matrix is ​​then used to propagate the fused features after linear mapping. Combined with residual connections and normalization, an initial representation with enhanced structure is obtained. By using multiple global latent attention sublayers, interactive modeling is performed on the initial representation after structural enhancement to obtain a universal transcriptome hidden representation.

[0007] Optional, the process of pre-constructing a gene relationship map includes: Determine a unified set of gene nodes based on a pre-defined gene sequence list; Iterate through any two genes, calculate the correlation between any gene pair, and obtain the original set of gene pair correlations; Based on the original set of gene pair correlations, a weighted edge set is constructed. A weighted adjacency matrix is ​​constructed based on a unified set of gene nodes and a set of weighted edges. A gene relationship graph is obtained by using a unified set of gene nodes, a set of weighted edges, and a weighted adjacency matrix.

[0008] Optionally, interactive modeling is performed on the structurally enhanced initial representation to obtain a universal transcriptome hidden representation, including: The initial representation after structural enhancement is processed to obtain the attention output; By nonlinearly enhancing the attention output, a universal transcriptome hidden representation is obtained.

[0009] Optionally, when the number of genes exceeds a preset value, the initial representation after structural enhancement is processed to obtain attention output, including: Divide the gene dimensions according to the preset block size; For each block, the initial representation after structural enhancement is processed to obtain the corresponding attention output; At the gene level, the attention outputs corresponding to each segment are spliced ​​together to obtain the final attention output.

[0010] Optionally, the pre-defined analysis task is a classification task. Based on the pre-defined analysis task, using a universal transcriptome hidden representation, the prediction results corresponding to the pre-defined analysis task are obtained, including: At the gene level, global average pooling is applied to the general transcriptome hidden representation to obtain the sample-level hidden representation; By utilizing sample-level hidden representations and mapping them through a feedforward network, classification prediction results are obtained.

[0011] A transcriptome data analysis device, comprising: The data provision module is used to provide raw transcriptomic data and sample batch origin information for biological samples; The transcriptome data processing module is used to process raw transcriptome data to obtain gene expression matrices; The model input data acquisition module is used to obtain standardized model input data based on the gene expression matrix and sample batch source information. The standardized model input data includes: the masked standardized gene expression matrix, mask position index information, gene-level protein sequence embedding features, and batch identification information. The general representation acquisition module takes the standardized model input data and inputs it into the transcriptome data analysis model to obtain the general transcriptome hidden representation. The transcriptome data analysis model is configured to process the standardized model input data to obtain the general transcriptome hidden representation and, based on the general transcriptome hidden representation, obtain the predicted results of the masked gene expression values. The results prediction module is used to obtain the prediction results corresponding to the pre-defined analysis tasks by utilizing the general transcriptome hidden representation.

[0012] A transcriptome data analysis device includes: a memory and a processor; Memory, used to store programs; A processor is used to execute programs that implement the steps of transcriptome data analysis methods as described above.

[0013] A readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the transcriptome data analysis method as described in any of the preceding claims.

[0014] As can be seen from the above technical solutions, the transcriptome data analysis method, apparatus, device, and readable storage medium provided in this application include: providing raw transcriptome data and sample batch source information of biological samples; processing the raw transcriptome data to obtain a gene expression matrix; obtaining standardized model input data based on the gene expression matrix and sample batch source information; inputting the standardized model input data into a transcriptome data analysis model to obtain a universal transcriptome hidden representation; the transcriptome data analysis model is configured to process the standardized model input data to obtain the universal transcriptome hidden representation, and to obtain the predicted results of masked gene expression values ​​based on the universal transcriptome hidden representation; and to obtain the prediction results corresponding to the pre-set analysis task using the universal transcriptome hidden representation based on a pre-set analysis task. This application achieves transcriptome data analysis by processing raw transcriptome data to obtain standardized model input data, inputting the standardized model input data into a transcriptome data analysis model to obtain a universal transcriptome hidden representation within the model, and finally obtaining the prediction results corresponding to the pre-set analysis task using the universal transcriptome hidden representation according to the pre-set analysis task.

[0015] Furthermore, transcriptome data analysis models can enhance the ability to model and impute missing gene expression values ​​by learning the correlation features between gene expression in transcriptome data, and generate universal transcriptome hidden representations with universal expression capabilities, thereby providing unified and reusable feature input support for various downstream bioinformatics analysis tasks. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0017] Figure 1 A flowchart of a transcriptome data analysis method provided in this application embodiment; Figure 2 An internal architecture diagram of an optional transcriptome data analysis model provided in this application embodiment; Figure 3 An internal architecture diagram of another optional transcriptome data analysis model provided in this application embodiment; Figure 4 A flowchart illustrating a standardized model input data acquisition process provided in this application embodiment; Figure 5 This is a schematic diagram of a transcriptome data analysis device provided in an embodiment of this application; Figure 6 This is a hardware structure block diagram of a transcriptome data analysis device provided in an embodiment of this application. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] Figure 1 A flowchart of a transcriptome data analysis method provided in this application embodiment is included, which may include the following steps: Step S100: Provide raw transcriptome data and sample batch source information for the biological sample.

[0020] Specifically, batch origin information is used to characterize the experimental batch, sequencing batch, or data source batch of the sample, etc.

[0021] Step S101: Process the raw transcriptome data to obtain the gene expression matrix.

[0022] Specifically, to ensure data quality and consistency, the raw transcriptome data can first be deduplicated. For samples with duplicate sequencing or submissions, only the original gene count version is retained, and the standardized or secondary processed expression matrices are removed to obtain the deduplicated raw transcriptome data, ensuring the originality and consistency of the data source. Then, based on the deduplicated raw transcriptome data, a pre-defined gene set is selected, and a gene expression matrix is ​​constructed. The gene expression matrix can be arranged with samples as rows and genes as columns, and the matrix elements can represent the expression level of the corresponding gene in the corresponding sample.

[0023] Step S102: Based on the gene expression matrix and sample batch source information, obtain standardized model input data.

[0024] Specifically, the standardized model input data may include: a masked standardized gene expression matrix, mask position index information, gene-level protein sequence embedding features, and batch identification information. The data format of the standardized model input data is adapted to the input requirements of the transcriptome data analysis model.

[0025] Step S103: Input the standardized model data into the transcriptome data analysis model to obtain the universal transcriptome hidden representation.

[0026] Specifically, the universal transcriptome hidden representation integrates gene expression numerical features, gene sequence inherent properties, global expression context of samples, batch information, gene structure relationship propagation information, and global dependency modeling information. It can be used as a high-quality gene representation for subsequent downstream pre-defined analysis tasks, such as mask expression value reconstruction tasks, sample classification tasks, disease classification or disease risk prediction tasks, gene necessity score prediction tasks, gene expression change prediction tasks under compound perturbation conditions, drug sensitivity or drug response prediction tasks, biological age assessment tasks, cell composition deconvolution analysis tasks, and biological phenotypic association analysis tasks.

[0027] The transcriptome data analysis model is configured to process standardized model input data to obtain a universal transcriptome hidden representation, and based on the universal transcriptome hidden representation, to obtain the predicted expression values ​​of masked genes.

[0028] During the training of transcriptome data analysis models, a masked regression loss can be constructed based on the predicted expression values ​​and actual expression values ​​of the masked genes, using a loss function. This masked regression loss can then be used to update the model parameters. The loss function can include mean squared error loss, mean absolute error loss, cross-entropy loss, or a combination thereof.

[0029] Step S104: Based on the pre-defined analysis task, the prediction results corresponding to the pre-defined analysis task are obtained using the universal transcriptome hidden representation.

[0030] Specifically, through the aforementioned structural design, the transcriptome data analysis model can comprehensively model gene expression information, gene intrinsic attribute information, overall sample context information, and batch difference information within a unified embedding representation space. This enhances the model's representation and generalization capabilities for complex transcriptome data. Simultaneously, it supports self-supervised expression reconstruction and supervised classification tasks, and possesses flexible task extension capabilities, thereby improving the model's versatility, transferability, and practical application value. Furthermore, it provides a stable and universal transcriptome hidden representation for various downstream analysis tasks. It should be noted that this step can be performed independently of the transcriptome data analysis model, or it can be embedded in the output layer of the transcriptome data analysis model as an output head for another downstream supervised task.

[0031] When the model is applied to downstream supervised tasks, other downstream supervised task output heads can be selectively enabled according to pre-defined analysis task requirements. For example, when a classification task is enabled, the sample-level embedding represents the classification prediction result obtained through a classification mapping network. Furthermore, classification loss can be calculated based on the true class labels to supervise and optimize the model parameters, enabling the model to learn sample-level discriminative features.

[0032] This application provides a transcriptome data analysis method, which may include: providing raw transcriptome data and sample batch source information of a biological sample; processing the raw transcriptome data to obtain a gene expression matrix; obtaining standardized model input data based on the gene expression matrix and sample batch source information; inputting the standardized model input data into a transcriptome data analysis model to obtain a universal transcriptome hidden representation; the transcriptome data analysis model is configured to process the standardized model input data to obtain the universal transcriptome hidden representation, and to obtain the predicted results of masked gene expression values ​​based on the universal transcriptome hidden representation; and, based on a pre-set analysis task, using the universal transcriptome hidden representation to obtain the prediction results corresponding to the pre-set analysis task. This application achieves transcriptome data analysis by processing raw transcriptome data to obtain standardized model input data, inputting the standardized model input data into a transcriptome data analysis model to obtain a universal transcriptome hidden representation within the model, and finally obtaining the prediction results corresponding to the pre-set analysis task using the universal transcriptome hidden representation.

[0033] Furthermore, transcriptome data analysis models can enhance the ability to model and impute missing gene expression values ​​by learning the correlation features between gene expression in transcriptome data, and generate universal transcriptome hidden representations with universal expression capabilities, thereby providing unified and reusable feature input support for various downstream bioinformatics analysis tasks.

[0034] In some embodiments of this application, reference is made to Figure 2 As shown, Figure 2 This application provides an optional transcriptome data analysis model internal architecture diagram. The transcriptome data analysis model may include: a cascaded input layer, an embedding layer, a feature modeling layer, and an output layer. Based on this, step S103, inputting standardized model input data into the transcriptome data analysis model to obtain a general transcriptome hidden representation, may include: The input layer obtains standardized model input data.

[0035] Specifically, the input data for the standardized model may include: a standardized gene expression matrix after masking, mask position index information, gene-level protein sequence embedding features, and batch identification information.

[0036] By using an embedding layer, the standardized gene expression matrix after masking, gene-level protein sequence embedding features, and batch identification information are processed to obtain fused features.

[0037] Specifically, the masked standardized gene expression matrix, gene-level protein sequence embedding features, and batch identification information are fused to obtain fused features. By combining these multi-dimensional fused features within the model, the model can comprehensively consider inherent gene properties, the overall expression state of the sample, and systematic differences introduced by different experimental batches or data sources while learning gene expression patterns. This improves the model's robustness and generalization ability under complex transcriptome data conditions.

[0038] By using a feature modeling layer, the fused features are modeled to obtain a general transcriptome hidden representation.

[0039] Specifically, by integrating the input features into the feature modeling layer to perform structural dependency modeling and global interaction modeling, a general transcriptome hidden representation is obtained. This representation can not only be used to predict the expression values ​​of masked genes, but also for prediction in other analysis tasks. This avoids the need for repeated feature engineering design or model training for different tasks, effectively improving the efficiency, consistency, and scalability of transcriptome data analysis.

[0040] The output layer, based on the universal transcriptome hidden representation and mask position index information, yields the predicted expression values ​​of the masked genes.

[0041] Specifically, the general transcriptome hidden representation output by the feature modeling layer is mapped at the task level to generate the predicted expression values ​​of the masked genes.

[0042] Through the above structural design, the transcriptome data analysis model can comprehensively model gene expression information, gene intrinsic attribute information, overall sample context information, and batch difference information under a unified fusion embedding representation space. This enhances the model's ability to represent and generalize complex transcriptome data. It also supports self-supervised expression reconstruction and supervised classification tasks, enabling high-quality, generalized feature representation modeling of transcriptome data. Furthermore, it has flexible task extension capabilities, thereby improving the model's versatility, transferability, and practical application value. Finally, it provides a stable and reusable general transcriptome hidden representation for various downstream analysis tasks.

[0043] In some embodiments of this application, reference is made to Figure 3 As shown, Figure 3The following is an internal architecture diagram of an alternative transcriptome data analysis model provided in this application embodiment. The feature modeling layer may include a static graph structure modeling sublayer and a multi-layer global latent attention sublayer. This sublayer can perform structural dependency modeling and global interaction modeling on the fusion features output by the embedding layer, thereby enhancing the model's ability to characterize inter-gene regulatory relationships and global expression patterns. Ultimately, it outputs a structurally enhanced gene-level representation, i.e., a universal transcriptome hidden representation. The feature modeling module includes a static graph structure modeling sublayer and a multi-layer global latent attention sublayer. These two types of structures work together to achieve joint modeling of structural relationships and global dependencies.

[0044] Based on this, a feature modeling layer is used to model the fused features to obtain a general transcriptome hidden representation, which may include: By modeling sub-layers using a static graph structure, a normalized adjacency matrix is ​​obtained based on a pre-constructed gene relationship graph. The normalized adjacency matrix is ​​then used to propagate the fused features after linear mapping through graph structure propagation. Combined with residual connections and normalization processing, an initial representation with enhanced structure is obtained.

[0045] Specifically, based on a pre-constructed gene relationship graph, graph structure propagation modeling is performed on the regulation or correlation between genes to obtain a normalized adjacency matrix. The gene relationship graph can be: , in, A set of gene nodes, Let be the set of edges connecting genes. The weights are the edge weights corresponding to the weighted adjacency matrix W.

[0046] To ensure numerical stability, the edge weights can be processed using absolute values, and the node degree and lower bound can be calculated and truncated. , , , in, For genes With genes The strength of the correlation between them This is a preset non-zero positive number used to perform minimum value constraint processing on node degree.

[0047] By following the steps above, a symmetric normalized adjacency matrix can be constructed: .

[0048] The normalized adjacency matrix is ​​finally obtained: .

[0049] Fusion features By performing a linear mapping, we can obtain the fused features after linear mapping processing: , in, This is the weight matrix. This is a bias term.

[0050] Using normalized adjacency matrix Fusion features after linear mapping Propagating the graph structure: .

[0051] After performing residual connections and layer normalization, we can obtain the initial representation after structural enhancement: .

[0052] By using multiple global latent attention sublayers, interactive modeling is performed on the initial representation after structural enhancement to obtain a universal transcriptome hidden representation.

[0053] Specifically, the initial representation after structural enhancement is processed to obtain the attention output, assuming the first... The initial representation of the layered structure after enhancement is as follows: .

[0054] For the first The initial representation after the layer's structural enhancement is projected onto the query to obtain query features: , in, This is the weight matrix.

[0055] Rearranging the query into a multi-head form yields the rearranged query characteristics: , in, M represents the batch size, and M represents the number of attention heads. For the number of genes, The feature dimension corresponding to each attention head, d represents the total dimension of the model's hidden features.

[0056] After completing query projection and multi-head rearrangement, to improve training stability and avoid feature scales that are too large or too small, the query features can be further normalized using Root Mean Square Normalization (RMSNorm). RMSNorm normalization does not rely on mean centering; it scales the features based solely on their root mean square magnitude. This allows for scale standardization while preserving feature orientation information, thereby enhancing numerical stability and saving computational resources.

[0057] Query characteristics after rearrangement For each feature vector Calculate its root mean square value: , in, For feature vectors The component values ​​in the k-th dimension To prevent division by zero constant.

[0058] Perform normalization and learnable scaling: , in, The k-th dimension feature after normalization These are learnable parameters.

[0059] Thus, the normalized Query feature representation is obtained: .

[0060] Subsequently, a low-rank compression strategy is applied to the key features and value features to reduce the feature dimension while preserving the inter-gene association information, thereby improving the model's computational efficiency. This yields: , in, This is a low-rank latent representation matrix. This is the weight matrix.

[0061] By performing an upprojection reconstruction using the shared low-rank space, we can obtain: , in, The key feature matrix, The eigenvalue matrix, This is the weight matrix.

[0062] The above steps enable key features and value features to share the same low-rank latent representation, thus reducing computational complexity while maintaining expressive power. Furthermore, scaling the dot product attention calculation: , Then, by connecting the linear mapping with the residual, the attention output is obtained: , in, This is the weight matrix.

[0063] Furthermore, the attention output can be nonlinearly enhanced to obtain a universal transcriptome hidden representation.

[0064] Specifically, normalize the attention output: , The normalized attention output is nonlinearly transformed by a feedforward neural network module to enhance the model's expressive power.

[0065] right Performing a linear dimension upscaling transformation, mapping the feature dimension d to 4d through a linear layer, yields an intermediate representation: , in, , .

[0066] Dividing the intermediate feature into two sub-feature parts along the last dimension yields: , in, , .

[0067] The Swish activation function is applied to a portion of the features, and this signal is used as a gate signal to modulate another portion of the features, thus forming the Swiglu gated output: , , in, For the activation function, in this embodiment it can be the Sigmoid activation function. This is element-wise multiplication.

[0068] The gated features are dimensionality-reduced using a linear layer to restore their original dimensions, and then a general transcriptome hidden representation is obtained through residual connections. : , , Wherein, parameter matrix , .

[0069] Through the above gating mechanism, the model can adaptively control the information flow according to the input features, thereby enhancing important features and suppressing redundant information, improving feature expression ability, improving the model's ability to model complex transcriptome expression patterns, and maintaining high computational efficiency and training stability.

[0070] Multi-head attention interaction modeling is performed on the initial representation after structural enhancement. By query mapping, generating keys and values ​​from low-rank latent representations, scaling dot product attention computation, residual connections and feedforward nonlinear enhancement operations, global dependencies across genes are captured layer by layer.

[0071] This application, through a feature modeling layer, combines a static graph structure modeling sublayer and a multi-layer global potential attention sublayer to reveal information about structural relationships between genes, capture global dependency features across genes, reduce computational complexity while maintaining high expression capacity, and improve the stability and generalization performance of the transcriptome data analysis model on large-scale transcriptome data.

[0072] Furthermore, the final universal transcriptome hidden representation can be obtained by stacking the results of static graph structure modeling and multi-layer global latent attention modeling.

[0073] In some embodiments of this application, the process of pre-constructing a gene relationship map may include: S11. Determine a unified set of gene nodes based on a pre-defined gene sequence list.

[0074] Specifically, the process involves acquiring a gene set and its corresponding prior biological relationships. These prior biological relationships include at least one or more of the following: gene regulatory relationships, protein-protein interaction relationships, pathway co-existence relationships, or co-expression correlations. Using each gene in the gene set as a graph node, a node set is constructed and rearranged according to a predefined gene sequence list. The resulting unified gene node set can be: , Where N is the size of the uniform gene set, and its order is strictly consistent with the arrangement order of the gene expression matrix and gene embedding features.

[0075] S12. Traverse any two genes, calculate the correlation between any gene pair, and obtain the original correlation set of gene pairs.

[0076] Specifically, based on gene co-expression networks, the Pearson correlation coefficient between any pair of genes can be calculated to obtain the original set of correlations: , in, For genes With genes The strength of the correlation between them.

[0077] S13. Construct a weighted edge set based on the original set of gene pair correlations.

[0078] Specifically, an edge set is constructed based on the biological relationships between genes. When a preset relationship exists between two genes, a connection edge is established between the corresponding nodes, and edge weights are assigned to the connection edges. The edge weights are determined based on the correlation strength, interaction confidence, regulatory strength, or statistical correlation coefficient between genes.

[0079] To improve the stability and sparsity of the graph structure, correlations are filtered and processed to remove self-loops: For the retained gene pairs, construct a weighted edge set: , in, .

[0080] If the gene pair does not meet the preset threshold condition, such as the threshold being τ, then the corresponding edge is not constructed, thus obtaining a sparse graph structure.

[0081] S14. Construct a weighted adjacency matrix based on a unified set of gene nodes and a set of weighted edges.

[0082] Specifically, the edge weights are standardized by taking the absolute value of the edge weights, calculating the degree value of each node and performing symmetric normalization to obtain a normalized adjacency structure representation. The normalized adjacency structure representation is stored as static graph structure parameters for modeling gene-level features for structure propagation and relationship enhancement during model training and inference.

[0083] Based on the unified set of gene nodes and the set of weighted edges mentioned above, construct a weighted adjacency matrix: .

[0084] The elements of a weighted adjacency matrix are defined as follows: .

[0085] To ensure the preservation of node information, a unit self-loop term can be added to the adjacency matrix, resulting in: , in, It is an identity matrix.

[0086] S15. Using a unified set of gene nodes, a set of weighted edges, and a weighted adjacency matrix, obtain the gene relationship graph.

[0087] Specifically, using a unified set of gene nodes Weighted edge set and weighted adjacency matrix To obtain a gene relationship map .

[0088] The gene relationship graph is constructed under a unified gene indexing system and can be aligned node-by-node with gene expression matrices and gene-protein value sequence embedding features, serving as a fixed topological foundation for static graph structure modeling. The pre-constructed gene relationship graph built through the above steps transforms external biological prior knowledge into a computable graph structure, enabling the model to utilize known relationships between genes when modeling transcriptome data, thereby enhancing structural expression capabilities and biological consistency.

[0089] In some embodiments of this application, when the number of genes is greater than a preset value, the complexity is... To reduce computational complexity, a block-based latent attention mechanism can be introduced. Based on this, the initial representation after structural enhancement is processed to obtain the attention output, which may include: S31. Divide the gene dimensions according to the preset block size.

[0090] Specifically, let the block size be... Then the gene dimension can be divided into: , in, For the number of genes, The number of blocks after partitioning, if Not divisible If so, then fill in.

[0091] S32. For each block, process the initial representation after structural enhancement to obtain the corresponding attention output.

[0092] Specifically, the j-th block can be represented as: , Global latent attention interaction modeling is performed independently within each block to obtain the attention output corresponding to each block: , S33. At the gene dimension, the attention outputs corresponding to each segment are spliced ​​together to obtain the final attention output.

[0093] Specifically, by splicing the data at the gene level, the final attention output can be obtained: , Through this embodiment, the complexity can be reduced from Reduced to ,in, This structure significantly reduces computational costs while maintaining the ability to model local dependencies, enabling the model to run stably at a scale of approximately 20,000 genes. When N≤C, no block processing is required.

[0094] In some embodiments of this application, the pre-defined analysis task can be a classification task. Based on this, step S104, obtaining the prediction result corresponding to the pre-defined analysis task using a universal transcriptome hidden representation, may include: S41. At the gene level, the universal transcriptome hidden representation is subjected to global average pooling to obtain the sample-level hidden representation.

[0095] Specifically, the normalized universal transcriptome hidden representation can be globally averaged and pooled at the gene level: , We obtain the sample-level hidden representation: , in, This is the sample-level hidden representation corresponding to the b-th sample. Let be the normalized hidden representation vector corresponding to the i-th gene in the b-th sample.

[0096] S42. Using sample-level hidden representations, the classification prediction results are obtained through a feedforward network mapping.

[0097] Specifically, the sample-level hidden representation is used to map the input sample classification header to categories: , , Obtain the classification prediction output: , in, For activation function, , This is the weight matrix. , For bias terms, This represents the number of categories.

[0098] The sample classification head can include a first fully connected layer, a non-linear activation layer, and a second fully connected layer connected in sequence, which are used to output the predicted probability distribution of each category and obtain the classification prediction result.

[0099] Furthermore, when the pre-defined analysis task can be a classification task, in order to optimize the model in a targeted manner, a classification loss can be constructed based on the cross-entropy between the classification prediction result and the true classification result label, and the model parameters can be optimized and updated using the classification loss; alternatively, the mask regression loss and the classification loss can be weighted or directly summed to obtain the total model loss, which can be used to jointly optimize the model parameters.

[0100] For classification tasks, the classification loss can be constructed using the following formula based on the cross-entropy between the classification prediction result and the true classification result label: , in, To categorize the true results, This represents the classification prediction results.

[0101] In multi-task training scenarios, the following formula can be used to calculate the mask regression loss. With classification loss By performing weighted summation or direct summation, the total model loss can be obtained. : .

[0102] In the model pre-training stage, this embodiment primarily uses the masked expression value prediction task as the self-supervised training objective. The model parameters are optimized using the masked regression loss function, enabling the model to learn the intrinsic relationships between gene expressions, thereby obtaining a universal transcriptome hidden representation with general characterization capabilities. During this stage, other downstream supervision task output heads, such as classification output heads, can be constructed and integrated into the model structure, but they do not participate in loss calculation; their corresponding classification loss weights are set to zero or disabled. When other downstream supervision task output heads, such as classification output heads, are pre-integrated into the model architecture as predefined structural components of the basic large model, the model possesses the ability to directly support downstream classification tasks. In the basic large model pre-training stage, the classification module does not participate in loss calculation; it is only enabled and participates in training in specific downstream classification application tasks, thereby achieving seamless transfer from self-supervised pre-training to downstream supervised tasks.

[0103] The model in this embodiment supports the following three training modes in practical applications: First, only the mask expression value prediction task is enabled. In this mode, the model only calculates the mask regression loss for self-supervised pre-training to learn a general transcriptome representation. Second, only the sample-level classification task is enabled. In this mode, the model only calculates the classification loss for supervised learning scenarios, such as disease classification or sample type identification tasks. Third, both the mask expression value prediction task and the sample-level classification task are enabled. In this mode, the model adopts a multi-task joint training approach, optimizing both the mask regression loss and the classification loss simultaneously, so that the model has both expression reconstruction ability and classification discrimination ability, thereby further improving the generalization and discrimination ability of the embedded representation.

[0104] In the above embodiments, the training process of the transcriptome data analysis model used in the analysis of the pre-set analysis task can be carried out in a variety of ways. The training process of one of the transcriptome data analysis models is described below.

[0105] A method for training a multidimensional embedding fusion transcriptome data analysis model may include: S51. Provide raw transcriptome data and batch source information for biological samples.

[0106] Specifically, to construct high-quality transcriptome datasets for building foundational large-scale models for transcriptome data characterization, publicly available large-scale transcriptome databases, such as the GEO and ARCHS4 databases, can be selected as sources of raw transcriptome sequencing data, and raw gene expression count data of the corresponding samples can be obtained. These databases are maintained by internationally authoritative bioinformatics institutions and contain a large amount of standardized and organized transcriptome sequencing data and corresponding sample annotation information, providing high-throughput expression data support from multiple sources, tissues, and disease states for transcriptome data analysis model training.

[0107] Batch origin information is used to characterize the experimental batch, sequencing batch, or data source batch of the sample, etc.

[0108] S52. Process the raw transcriptome data to obtain the gene expression matrix.

[0109] S53. Based on the gene expression matrix and sample batch source information, obtain the standardized model input data and the true expression value labels of the masked genes.

[0110] Specifically, standardized model input data can include: a masked standardized gene expression matrix, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information. The standardized model input data format accurately adapts to the input requirements of transcriptome data analysis models, ensuring a high degree of match between the data structure, parameter dimensions, and information integrity and the model's acceptance standards. This provides a data foundation for the smooth progress of subsequent transcriptome data characterization processes and the accuracy of results.

[0111] S54. Input the standardized model input data into the transcriptome data analysis model to be trained to obtain the predicted expression values ​​of the masked genes.

[0112] Specifically, the transcriptome data analysis model to be trained is used to characterize transcriptome-related information based on standardized model input data. Multi-dimensional features are integrated within the model, enabling it to learn gene expression patterns while comprehensively considering inherent gene attributes, overall sample expression status, and systematic differences introduced by different experimental batches or data sources. This improves the model's robustness and generalization ability under complex transcriptome data conditions, provides a unified feature representation foundation, and allows the transcriptome data analysis model not only to predict masked gene expression values ​​but also to be directly used for prediction in other analysis tasks. This avoids repeated feature engineering design or model training for different tasks, effectively improving the efficiency, consistency, and scalability of transcriptome data analysis.

[0113] During model training, a masking prediction strategy is adopted to guide the model to reconstruct and learn the expression values ​​of masked genes using the contextual information of unmasked gene expression. This enables the model to automatically learn the potential association patterns and global expression structure between gene expressions without relying on specific downstream task labels.

[0114] S55. Update the parameters of the transcriptome data analysis model to be trained until convergence.

[0115] Specifically, the training objective is to make the predicted expression values ​​of masked genes as close as possible to the actual expression values ​​of the masked genes. The parameters of the transcriptome data analysis model to be trained are updated until convergence. Some gene expression values ​​in the standardized model input data are masked using a pre-defined masking strategy. During model training, the expression context information of unmasked genes can be used to predict the expression values ​​of masked genes. Based on the error between the prediction results and the corresponding actual expression value labels, a loss function is used to train the transcriptome data analysis model to obtain the target transcriptome data analysis model. The loss function can include mean squared error loss, mean absolute error loss, cross-entropy loss, or a combination thereof.

[0116] Transcriptome data analysis models can learn the correlation features between gene expression in transcriptome data through multidimensional embedding mechanisms, enhance the ability to model and impute missing gene expression values, and generate transcriptome feature embeddings with universal expression capabilities, thereby providing unified and reusable feature input support for various downstream bioinformatics analysis tasks.

[0117] Furthermore, when training the transcriptome data analysis model, the standardized model input data can be divided into training, validation, and test datasets. The training set can be used for model parameter updates, the validation set for model performance monitoring and optimal model selection, and the test set for final evaluation of the model's generalization ability. The training, validation, and test sets can be divided in a 9:0.5:0.5 ratio. Data can be stored in HDF5 format and read in batches using a custom data loading module, thus supporting efficient reading of large-scale samples.

[0118] The standardized model input data includes at least gene expression features, gene-level protein sequence embedding feature matrix, mask position index information, the true expression value of the masked gene, and batch identification information corresponding to the sample, as well as other sample-level or gene-level auxiliary features introduced according to the specific application scenario.

[0119] During model training, the training status of the model is monitored using a validation dataset. When preset conditions are met, the model parameters are saved or updated to obtain the trained target transcriptome data analysis model. For example, the preset condition could be a preset loss value. When the model loss value reaches the preset loss value, the current model parameters are saved, resulting in the trained target transcriptome data analysis model.

[0120] In some embodiments of this application, the transcriptome data analysis model to be trained may include: an input layer, an embedding layer, a feature modeling layer, and an output layer cascaded in sequence; Based on this, the training process for the transcriptome data analysis model to be trained may include: The input layer obtains standardized model input data.

[0121] Specifically, the input data for the standardized model may include: a standardized gene expression matrix after masking, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information.

[0122] Through the embedding layer, fused features are obtained by using the masked standardized gene expression matrix, gene-level protein sequence embedding feature matrix, and batch identification information from the input data of the standardized model.

[0123] Specifically, the masked standardized gene expression matrix, gene-level protein sequence embedding feature matrix, and batch identification information in the input data of the standardized model are mapped into a vector representation of a unified dimension and then fused.

[0124] By using a feature modeling layer, the fused features are modeled to obtain a general transcriptome hidden representation.

[0125] Specifically, the fused features are input into the feature modeling layer to obtain a general transcriptome hidden representation. This universal transcriptome hidden representation integrates gene expression numerical features, gene sequence intrinsic properties, global expression context of samples, batch information, gene structure relationship propagation information, and global dependency modeling information. Therefore, it can be used as a high-quality gene characterization for subsequent downstream application analyses, such as masked expression value reconstruction tasks, sample classification tasks, disease classification or disease risk prediction tasks, gene necessity score prediction tasks, gene expression change prediction tasks under compound perturbation conditions, drug sensitivity or drug response prediction tasks, biological age assessment tasks, cell composition deconvolution analysis tasks, and biological phenotypic association analysis tasks.

[0126] Depending on the different downstream application analysis scenarios, corresponding input formats are supported. By constructing a universal transcriptome hidden representation, a large-scale transcriptome basic model framework is achieved, featuring both gene-level and sample-level dual-layer representation, self-supervised pre-training and unified modeling for downstream tasks, and transferability and scalability, supporting various bioinformatics application scenarios. Compared with traditional models trained based on a single task, the universal transcriptome hidden representation output by this application has stronger generalization and expression capabilities, and can serve as the core representation form for transcriptome data analysis models.

[0127] The predicted expression values ​​of the masked genes are obtained through the output layer, based on the universal transcriptome hidden representation and mask position index information.

[0128] Specifically, the general transcriptome hidden representation output by the feature modeling layer is task-level mapped to generate predicted expression values ​​of masked genes. This general transcriptome hidden representation characterizes the overall feature distribution of a sample at the transcriptome level and can serve as a universal transcriptome hidden representation for various downstream analysis tasks. These downstream tasks may include: disease classification, disease risk prediction, gene necessity score prediction, compound-perturbed gene expression level prediction, drug sensitivity prediction, biological age assessment, cell composition deconvolution, and biological phenotypic association analysis.

[0129] The training process of the aforementioned transcriptome data analysis model aims to make the predicted expression values ​​of the masked genes approach the labels of their actual expression values, thereby updating the model's parameters. A masked regression loss can be constructed based on the mean squared error between the predicted and actual expression values ​​of the masked genes, serving as a self-supervised training objective. This allows the model to learn the intrinsic dependencies between gene expressions and optimize the general hidden representation of the transcriptome.

[0130] Furthermore, during model training, parallel training can be employed to synchronously update model parameters across multiple computational units, thereby improving training efficiency. The model parameter update process can accumulate gradient information from multiple training batches and then uniformly update the model parameters after reaching a preset number of accumulation steps, achieving equivalent large-batch training.

[0131] Simultaneously, model parameters can be updated based on adaptive optimization strategies, and a learning rate scheduling mechanism can be used to perform learning rate warm-up in the early stages of training and learning rate decay in subsequent training stages to improve the stability and convergence of model training. A mixed-precision computation method is used for forward and backward propagation calculations, and a gradient scaling mechanism is introduced during low-precision gradient calculations to reduce computational resource consumption and ensure numerical stability. When the model's performance on the validation dataset surpasses historical best results, the corresponding model parameters are saved as the target transcriptome data analysis model.

[0132] In some embodiments of this application, the AdamW optimizer can be used to update model parameters. This optimizer, based on traditional adaptive optimization methods, introduces a weight decay mechanism, which helps suppress overfitting and improve generalization ability. Simultaneously, a learning rate scheduling strategy of warm-up followed by linear decay is employed during training. The learning rate is gradually increased in the early stages of training to allow the model to smoothly enter the convergence phase, and then gradually decreased in subsequent stages to achieve stable convergence. This scheduling method effectively avoids unstable oscillations in the early stages of training and enables refined optimization in later stages.

[0133] In some embodiments of this application, there are multiple options for the internal network architecture of the embedding layer. The following describes the training process of the model using one of the internal architectures of the embedding layer.

[0134] The embedding layer may include: an expression embedding sublayer, a gene embedding sublayer, a sample context embedding sublayer, a batch embedding sublayer, and a feature fusion layer. Based on this, through the embedding layer, using the masked normalized gene expression matrix, the gene-level protein sequence embedding feature matrix, and batch identification information from the normalized model input data, fused features are obtained, which may include: By using the expression embedding sublayer, the expression embedding matrix is ​​obtained based on the masked, standardized gene expression matrix.

[0135] Specifically, the masked normalized gene expression matrix undergoes multi-scale periodic encoding mapping to preserve gene expression intensity information and quantitative variation characteristics. Assume the masked normalized gene expression matrix is ​​as follows: , in, For batch size, This represents the number of genes.

[0136] A set of preset frequency reference vectors is constructed to implement multi-scale periodic coding mapping. Assuming the embedding dimension is d, we can set d=640. Then the frequency reference vector can be defined as: , in, For frequency index, .

[0137] At this point, the frequency set can be: , Among them, the frequency parameter can be a preset constant parameter, which is used to construct the expression change response at different frequency scales, thereby enhancing the model's ability to perceive expression changes of different orders of magnitude.

[0138] By performing frequency expansion mapping on each expression value in the masked, normalized gene expression matrix, we can obtain: , in, For the first In the nth sample The expression value of each gene, For frequency index, .

[0139] Further performing sine and cosine periodic transformations on the angle value yields: ; .

[0140] By concatenating the above data, a preliminary gene expression embedding representation is obtained: .

[0141] When the gene expression value is set to a preset mask identifier value m, and m = -10, a learnable mask embedding vector can be introduced. The gene expression embedding representation at the corresponding position is replaced, and the mask indicator function can be defined as follows: .

[0142] Through the above steps, the final gene expression embedding representation can be obtained: .

[0143] Thus, the embedding matrix can be obtained: .

[0144] By performing dimensionality transformation on the gene-level protein sequence embedding feature matrix through the gene embedding sublayer, the gene-level embedding feature matrix is ​​obtained.

[0145] Specifically, the gene-level protein sequence embedding feature matrix can be: , in, The total number of feature vectors embedded in gene-level protein sequences. For the embedded dimension.

[0146] To match the gene-level protein sequence embedding features with the unified feature space dimension within the model, a projection network consisting of two layers of linear mapping and nonlinear transformation is used to transform the dimension of the gene-level protein sequence embedding feature matrix, namely: , , in, For intermediate mapping features, It is a gene-level embedding feature. For non-linear activation functions, in this embodiment, the ReLU activation function can be used. For the first The gene-level protein sequence embedding feature vector of each gene. , This is the weight matrix. , This is a bias term.

[0147] Based on the above steps, the gene-level embedding feature matrix can be obtained: .

[0148] Furthermore, to adapt to batch sample input, the gene-level embedding feature matrix can be expanded in the sample dimension. Assuming the batch sample size is B, the expanded gene-level embedding feature matrix can be: .

[0149] Through the above design, the gene embedding sublayer can introduce the inherent sequence-level properties of genes into the unified feature space of the model and participate in the subsequent modeling process as static prior information, thereby enhancing the model's ability to express gene functional characteristics and potential regulatory relationships.

[0150] By using the sample context embedding sublayer, sample context embedding features are generated based on the masked, standardized gene expression matrix.

[0151] Specifically, the normalized gene expression matrix after masking can be: , in, For batch size, This represents the number of genes.

[0152] The process involves a first linear mapping, a nonlinear activation transformation, and a second linear mapping, namely: , , in, For intermediate mapping features, For the first The sample context embedding vector of each sample is used to characterize the potential state of the overall transcriptome expression of that sample. For non-linear activation functions, in this embodiment, the GELU activation function can be used. , This is the weight matrix. , This is a bias term.

[0153] To enable sample context to participate in gene-level modeling, the sample context embedding vector can be extended to gene-dimensional alignment. The sample context embedding features can be: .

[0154] Extending the sample context embedding features by expanding them along the gene dimension yields the expanded features: .

[0155] Thus, all genes within the same sample share the same context embedding vector, with replication and expansion occurring only at the gene dimension.

[0156] Based on the standardized gene expression matrix after masking, a nonlinear mapping is performed on the overall gene expression distribution of a single sample to generate a sample-level context representation, which is then extended in the gene dimension so that all genes within the same sample share the context embedding vector.

[0157] The sample context embedding sublayer is used to characterize the overall transcriptional state at the sample level, providing a global representation of the sample from the perspective of whole-gene expression distribution. This enables the model to perceive the overall expression pattern within a single sample. The sample context embedding sublayer uses a multilayer perceptron structure to perform nonlinear mapping on the sample-level expression vector, generating sample context embedding features. Through this sample context embedding feature construction method, the overall expression distribution of the sample is compressed into a low-dimensional latent representation, enabling the model to perceive the overall state of the sample during gene-level modeling. This improves the model's ability to express global changes at the sample level and provides prior global contextual knowledge for subsequent processing.

[0158] The batch embedding sublayer generates a batch embedding tensor based on batch identification information.

[0159] Specifically, assuming the number of batches is M, and each sample corresponds to a discrete batch identifier, that is: , in, For the first Batch number of each sample For batch size, .

[0160] Furthermore, define the batch embedding matrix: , Among them, the m-th row For the learnable batch embedding tensor corresponding to the m-th batch, Parameters can be updated during model training.

[0161] Through the above steps, discrete batch numbers can be mapped to continuous vector representations: .

[0162] The batch embedding tensor can be obtained: .

[0163] Since the model is modeled in the form of gene-level sequences, the sample-level batch embeddings are extended to the gene dimension. All genes within the same sample share the same batch embedding vector. The extended batch embedding tensor can be: .

[0164] The batch identification information corresponding to the samples is embedded to construct a representation to explicitly model the systematic biases introduced by different experimental batches, sequencing platforms or data sources. By introducing a learnable batch embedding tensor, the model can automatically learn and correct the impact of batch effects on gene expression patterns.

[0165] Through the feature fusion layer, the representation embedding matrix is ​​processed. Gene-level embedding feature matrix Sample context embedding features and batch embedding tensors The fusion process is performed to obtain a unified fusion feature representation.

[0166] Specifically, the fusion process can be performed using an additive approach, and the fusion features can be: .

[0167] Furthermore, before modeling the fusion features to obtain a general transcriptome hidden representation, the fusion features can be subjected to layer normalization and layer discarding to obtain the final fusion features.

[0168] Specifically, performing layer normalization and layer discarding on the fused features can improve numerical stability and convergence speed. The processing can be achieved using the following formula: , , in, For layer normalization, This is a discard layer.

[0169] Furthermore, regarding Perform two-layer feedforward mapping: , in, , This is the weight matrix. The activation function is non-linear; in this embodiment, it can be the ReLU activation function.

[0170] The final fusion features can be obtained: .

[0171] Based on this, modeling the fusion features to obtain a universal transcriptome hidden representation can include: modeling the final fusion features to obtain a universal transcriptome hidden representation.

[0172] In some embodiments of this application, obtaining the predicted expression value of the masked gene based on universal transcriptome hiding representation and mask location index information may include: S61. Perform layer normalization on the general transcriptome hidden representation to obtain the gene-level embedding matrix.

[0173] Specifically, the general transcriptome hidden representation output by the feature modeling layer. Perform layer normalization to obtain .

[0174] For the The nth sample, the nth The embedding representation of a gene can be defined as: .

[0175] This allows us to obtain the gene-level embedding matrix: .

[0176] To obtain an overall representation of the sample, global pooling is performed on the gene dimension, for example, using global max pooling: .

[0177] The sample-level embedding matrix can be obtained: .

[0178] in, This sample-level embedding matrix can characterize the potential representation of the overall transcriptional state of a sample.

[0179] S62. Based on the gene-level embedding matrix, extract the corresponding hidden representation according to the mask position index information.

[0180] Specifically, based on the mask position index information, which can be a mask position index matrix, the hidden representation of the corresponding masked gene position is extracted from the gene-level embedding matrix.

[0181] Let the mask position index matrix be: , in, The number of masked genes for each sample.

[0182] Using the mask position index matrix, the hidden representation corresponding to the masked gene position is extracted from the gene-level embedding matrix: , in, For the first The first sample The location index of the masked gene.

[0183] The final result is: .

[0184] S63. Based on the hidden representation, a fully connected network is used for regression prediction to obtain the predicted expression value of the masked gene.

[0185] Specifically, a subset of masked features can be constructed using the hidden representation of the masked gene location, and then input into a mask prediction head for regression mapping. The mask prediction head comprises a first fully connected layer, a non-linear activation layer, and a second fully connected layer connected in sequence, which output the predicted expression value of the corresponding masked gene.

[0186] Regression prediction is performed using a two-layer fully connected network: , , in, , This is the weight matrix. , This is a bias term.

[0187] Finally, the predicted expression values ​​of the masked genes are obtained: .

[0188] To enable the model to learn the intrinsic dependencies between expression values ​​and thus obtain high-quality gene embedding representations, the mask regression loss can be defined as the mean squared error. Based on this, the mask regression loss can be: , in, This represents the true expression value of the masked gene. The predicted expression value of the masked gene.

[0189] In some embodiments of this application, the standardized model input data may include: a masked standardized gene expression matrix, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information. The data format of the standardized model input data is adapted to the input requirements of the transcriptome model to be trained. (Reference) Figure 4 As shown, Figure 4 This application provides a flowchart for obtaining standardized model input data. Based on this, step S53, based on the gene expression matrix and sample batch source information, obtains standardized model input data and the true expression value labels of the masked genes, which may include: S71. Based on the preset gene annotation information, the gene expression matrix is ​​standardized to obtain standardized gene expression data.

[0190] Specifically, standardized expression data serves as the basis for subsequent gene sequence alignment and feature construction processing.

[0191] S72. Based on a preset gene sequence list, the standardized gene expression data are reordered to obtain a standardized gene expression matrix.

[0192] Specifically, based on a pre-defined gene sequence list, standardized gene expression data are reordered, and zeros are padded at the positions corresponding to missing genes to obtain a standardized gene expression matrix with consistent gene order. The pre-defined gene sequence list is constructed based on an authoritative gene annotation database and serves as the sole standard for aligning gene expression data during the training of the multidimensional embedding fusion transcriptome data analysis model. By constructing a unified gene sequence representation framework, transcriptome data from different sources and under different experimental conditions can be aligned into structurally consistent gene expression representations.

[0193] S73. Based on the preset gene sequence list, construct the corresponding gene-level protein sequence embedding feature matrix.

[0194] S74. Based on a preset masking strategy, mask the expression values ​​of some genes in the standardized gene expression matrix to obtain the masked standardized gene expression matrix, masking position index information, and the true expression value label of the masked genes.

[0195] Specifically, the preset masking strategy can be as follows: From the standardized gene expression matrix, a portion of gene expression values ​​are randomly selected as masking targets according to a preset ratio. These selected gene expression values ​​are then masked, making them invisible in the model input or replaced with preset placeholder values. The location index of the masked genes and their actual expression values ​​are recorded. The masked standardized gene expression matrix is ​​used as input features, the actual expression values ​​of the masked genes are used as prediction targets, and the masked location index information is used as auxiliary input information. This allows the model to utilize the remaining unmasked gene expression features as contextual information to reconstruct and predict the masked gene expression values.

[0196] The preset ratio is any ratio within a preset range of the total number of genes. Masking can be implemented in various ways, such as: replacing the selected gene expression value with zero; replacing the selected gene expression value with a preset random noise value; replacing the selected gene expression value with a mask identifier value used to identify missing states; keeping the gene expression value unchanged but marking it as a state to be predicted in the model, etc.

[0197] S75. Encode the source information of the sample batch to obtain batch identification information.

[0198] Specifically, batch origin information is used to characterize the experimental batch, sequencing batch, or data source batch of the sample. The sample batch origin information is encoded, mapping different batches to different discrete batch identifier values. These batch identifier values ​​are then further converted into continuous vector representations to obtain batch identification information. This information, along with the standardized gene expression matrix, mask position index information, and gene-level protein sequence embedding feature matrix, constitutes standardized model input data for subsequent model input. For samples that cannot be matched to a known batch origin, a pre-defined unknown batch identifier value is uniformly assigned as batch identification information.

[0199] The processing flow of this embodiment can, to a certain extent, ensure that transcriptome data from different sources and batches have a unified data structure and comparability at the input level, providing a data foundation for the model to learn a stable and generalizable transcriptome representation.

[0200] In some embodiments of this application, the preset gene annotation information may include: a unique gene identifier and corresponding gene length information. Based on this, S71, the gene expression matrix is ​​standardized based on the preset gene annotation information to obtain standardized gene expression data, which may include: S81. Based on the unique gene identifier, determine the gene length information corresponding to each gene in the gene expression matrix.

[0201] S82. For each gene, based on the corresponding gene length information, convert the original gene expression value into the number of reads per million transcripts.

[0202] Specifically, based on gene length information, the original gene expression values ​​are normalized. This can include converting the original gene expression values ​​into Transcripts Per Million (TPM) expression values ​​based on gene length information to eliminate the impact of differences in sequencing depth and gene length.

[0203] S83. Transform the TPM expression values ​​to obtain standardized gene expression data.

[0204] Specifically, the TPM expression value is transformed by log2(TPM+1) to compress the dynamic range of the expression value and enhance the numerical stability, thereby obtaining standardized transcriptome expression data at the gene level.

[0205] In some embodiments of this application, S73, constructing a corresponding gene-level protein sequence embedding feature matrix based on a preset gene sequence list, may include: S91. Based on a preset gene sequence list, determine the protein sequence information corresponding to each gene.

[0206] Specifically, the protein sequence information corresponding to each gene can be determined based on a preset gene sequence list.

[0207] S92. Input the protein sequence information into the pre-trained biological sequence representation model to obtain the corresponding gene-level protein sequence embedding feature vector.

[0208] Specifically, the biological sequence representation model is trained using protein sequence information as training samples and the gene-level protein sequence embedding feature vectors corresponding to the protein sequence information as training labels.

[0209] S93. Based on the gene-level protein sequence embedding feature vector, obtain the gene-level protein sequence embedding feature matrix, and construct the gene-level protein sequence embedding feature matrix that corresponds one-to-one with the standardized gene expression matrix.

[0210] Specifically, the gene-level protein sequence embedding feature matrix is ​​associated with the standardized gene expression matrix to construct a gene-level protein sequence embedding feature matrix that corresponds one-to-one with the standardized gene expression matrix. For genes lacking protein sequence information, zero vectors or preset default vectors are used to fill in the missing information to ensure that the feature dimensions are fixed.

[0211] In some embodiments of this application, a multi-GPU distributed data parallel training method can be adopted to support high-dimensional feature training on large-scale datasets.

[0212] Specifically, before training begins, process groups are initialized via a distributed communication module. Each GPU corresponds to an independent training process, and each process processes only a subset of the training data. To ensure data consistency across different GPUs, the training set is partitioned using a distributed sampler. At the start of each training cycle, the sampling random seed is reset, ensuring a different data order in each cycle and improving model generalization ability. After backpropagation, each process automatically synchronizes parameters through an efficient communication mechanism, ensuring consistent model parameters across all GPUs.

[0213] In some embodiments of this application, since the model dimension is large and the number of samples in a single batch is limited by the GPU memory, a gradient accumulation mechanism can be introduced.

[0214] Specifically, during actual training, after each mini-batch's forward computation and backpropagation, only the gradients are accumulated without immediately updating the model parameters. Once a preset number of steps are reached, a unified parameter update operation is performed. This approach achieves equivalent large-batch training results without increasing GPU memory usage, improves model convergence stability, and reduces training instability caused by gradient fluctuations.

[0215] In some embodiments of this application, an automatic mixed precision training mechanism can be used to improve training efficiency and reduce memory usage.

[0216] Specifically, during the forward propagation phase, the model is executed in low-precision floating-point format, thereby reducing memory usage, improving matrix multiplication speed, and increasing overall training throughput. When using half-precision floating-point format, a gradient scaling mechanism is introduced to avoid numerical underflow. When using bfloat16 format, good numerical stability can be maintained without gradient scaling. This strategy significantly improves training efficiency while ensuring that training accuracy is largely unaffected.

[0217] After each training cycle, the model's performance can be evaluated on the validation set, and the validation loss can be calculated. If the validation loss of the current training cycle is better than the historical best result, the current model parameters are saved as the best model; at the same time, periodic checkpoint files are saved at preset intervals so that training can be resumed in case of abnormal interruption.

[0218] Through the above training method, the model of this invention can support high-dimensional feature training at the scale of tens of thousands of genes, efficiently scale in a multi-GPU environment, improve memory utilization and computational efficiency, ensure the stability of the training process, and obtain embedding representations with good generalization ability. The training method described in this embodiment is applicable to large-scale biological transcriptome data and can also be extended to other high-dimensional structured biological data scenarios.

[0219] The present application provides a transcriptome data analysis device according to an embodiment. The transcriptome data analysis device described below corresponds to the transcriptome data analysis method described above.

[0220] Figure 5 This is a schematic diagram of a transcriptome data analysis device provided in an embodiment of this application, with reference to... Figure 5 As shown, the transcriptome data analysis device may include: Data providing module 10 is used to provide raw transcriptome data and sample batch origin information for biological samples; Transcriptome data processing module 20 is used to process raw transcriptome data to obtain gene expression matrix; The model input data acquisition module 30 is used to obtain standardized model input data based on the gene expression matrix and sample batch source information. The standardized model input data includes: the masked standardized gene expression matrix, mask position index information, gene-level protein sequence embedding features, and batch identification information. The general representation acquisition module 40 takes the standardized model input data and inputs it into the transcriptome data analysis model to obtain the general transcriptome hidden representation. The transcriptome data analysis model is configured to process the standardized model input data to obtain the general transcriptome hidden representation and obtain the predicted results of the masked gene expression values ​​based on the general transcriptome hidden representation. The result prediction module 50 is used to obtain the prediction results corresponding to the pre-defined analysis task by using a general transcriptome hidden representation based on the pre-defined analysis task.

[0221] This application provides a transcriptome data analysis device, which may include: a data providing module 10 for providing raw transcriptome data and sample batch source information of biological samples; a transcriptome data processing module 20 for processing the raw transcriptome data to obtain a gene expression matrix; a model input data acquisition module 30 for obtaining standardized model input data based on the gene expression matrix and sample batch source information; a universal representation acquisition module 40 for inputting the standardized model input data into a transcriptome data analysis model to obtain a universal transcriptome hidden representation, wherein the transcriptome data analysis model is configured to have the ability to process the standardized model input data to obtain a universal transcriptome hidden representation and obtain the predicted result of the masked gene expression value based on the universal transcriptome hidden representation; and a result prediction module 50 for obtaining the prediction result corresponding to the pre-set analysis task based on the universal transcriptome hidden representation. This application processes raw transcriptome data to obtain standardized model input data, and then inputs the standardized model input data into a transcriptome data analysis model to obtain a general transcriptome hidden representation within the model. Based on a pre-defined analysis task, the general transcriptome hidden representation is used to obtain the prediction results corresponding to the pre-defined analysis task, thereby realizing the analysis of transcriptome data.

[0222] Furthermore, transcriptome data analysis models can enhance the ability to model and impute missing gene expression values ​​by learning the correlation features between gene expression in transcriptome data, and generate universal transcriptome hidden representations with universal expression capabilities, thereby providing unified and reusable feature input support for various downstream bioinformatics analysis tasks.

[0223] Optional transcriptome data analysis models include: a cascaded input layer, an embedding layer, a feature modeling layer, and an output layer; The universal representation acquisition module 40 performs the process of inputting standardized model data into the transcriptome data analysis model to obtain universal transcriptome hidden representations, which may include: The standardized model input data is obtained through the input layer; By using an embedding layer, the masked standardized gene expression matrix, gene-level protein sequence embedding features, and batch identification information are processed to obtain fused features; By modeling the fused features through the feature modeling layer, a general transcriptome hidden representation is obtained; The output layer, based on the universal transcriptome hidden representation and mask position index information, yields the predicted expression values ​​of the masked genes.

[0224] Optionally, the feature modeling layer includes a static graph structure modeling sublayer and a multi-layer global latent attention sublayer. The general representation acquisition module 40 performs a process of modeling the fused features through the feature modeling layer to obtain a general transcriptome hidden representation, which may include: By modeling sub-layers using a static graph structure, a normalized adjacency matrix is ​​obtained based on a pre-constructed gene relationship graph. The normalized adjacency matrix is ​​then used to propagate the fused features after linear mapping. Combined with residual connections and normalization, an initial representation with enhanced structure is obtained. By using multiple global latent attention sublayers, interactive modeling is performed on the initial representation after structural enhancement to obtain a universal transcriptome hidden representation.

[0225] Optionally, the transcriptome data analysis device also includes a gene mapping module, the process of which pre-constructs a gene mapping module may include: Determine a unified set of gene nodes based on a pre-defined gene sequence list; Iterate through any two genes, calculate the correlation between any gene pair, and obtain the original set of gene pair correlations; Based on the original set of gene pair correlations, a weighted edge set is constructed. A weighted adjacency matrix is ​​constructed based on a unified set of gene nodes and a set of weighted edges. A gene relationship graph is obtained by using a unified set of gene nodes, a set of weighted edges, and a weighted adjacency matrix.

[0226] Optionally, the universal representation acquisition module 40 performs interactive modeling on the structure-enhanced initial representation to obtain a universal transcriptome hidden representation, which may include: The initial representation after structural enhancement is processed to obtain the attention output; By nonlinearly enhancing the attention output, a universal transcriptome hidden representation is obtained.

[0227] Optionally, when the number of genes exceeds a preset value, the general representation acquisition module 40 performs a process of processing the structure-enhanced initial representation to obtain the attention output, which may include: Divide the gene dimensions according to the preset block size; For each block, the initial representation after structural enhancement is processed to obtain the corresponding attention output; At the gene level, the attention outputs corresponding to each segment are spliced ​​together to obtain the final attention output.

[0228] Optionally, the pre-defined analysis task is a classification task. The result prediction module 50 performs a process based on the pre-defined analysis task, using a universal transcriptome hidden representation, to obtain the prediction result corresponding to the pre-defined analysis task. This process may include: At the gene level, global average pooling is applied to the general transcriptome hidden representation to obtain the sample-level hidden representation; By utilizing sample-level hidden representations and mapping them through a feedforward network, classification prediction results are obtained.

[0229] This application also provides a transcriptome data analysis device. Figure 6 A hardware block diagram of a transcriptome data analysis device is shown, with reference to... Figure 6 The hardware structure of a transcriptome data analysis device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4; In this embodiment, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4. Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. Memory 3 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device; The memory stores a program, which the processor can call. The program is used to implement the various processing steps in the aforementioned transcriptome data analysis method.

[0230] This application embodiment also provides a storage medium that can store a program suitable for execution by a processor, the program being used to implement various processing flows in the aforementioned transcriptome data analysis method.

[0231] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0232] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined with each other, and the same or similar parts can be referred to each other.

[0233] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for transcriptome data analysis, characterized in that, include: Provide raw transcriptome data and batch origin information for biological samples; The raw transcriptome data is processed to obtain a gene expression matrix; Based on the gene expression matrix and the sample batch source information, standardized model input data is obtained, which includes: a masked standardized gene expression matrix, mask position index information, gene-level protein sequence embedding features, and batch identification information. The standardized model input data is input into the transcriptome data analysis model to obtain a general transcriptome hidden representation. The transcriptome data analysis model is configured to process the standardized model input data to obtain a general transcriptome hidden representation and, based on the general transcriptome hidden representation, obtain the predicted expression value of the masked gene. Based on a pre-defined analysis task, the prediction results corresponding to the pre-defined analysis task are obtained using the general transcriptome hidden representation.

2. The method according to claim 1, characterized in that, The transcriptome data analysis model includes: an input layer, an embedding layer, a feature modeling layer, and an output layer, which are cascaded in sequence. The step of inputting the standardized model into the transcriptome data analysis model to obtain a universal transcriptome hidden representation includes: Standardized model input data is obtained through the input layer; Through the embedding layer, the masked standardized gene expression matrix, gene-level protein sequence embedding features, and batch identification information are processed to obtain fused features; The fused features are modeled using the feature modeling layer to obtain a general transcriptome hidden representation; Through the output layer, based on the universal transcriptome hidden representation and mask position index information, the predicted expression values ​​of the masked genes are obtained.

3. The method according to claim 2, characterized in that, The feature modeling layer includes a static graph structure modeling sublayer and a multi-layer global latent attention sublayer. The fused features are modeled through the feature modeling layer to obtain a general transcriptome hidden representation, including: Sub-layers are modeled using a static graph structure. Based on a pre-constructed gene relationship graph, a normalized adjacency matrix is ​​obtained. The normalized adjacency matrix is ​​then used to propagate the fusion features after linear mapping. Combined with residual connections and normalization, an initial representation with enhanced structure is obtained. By using multiple global latent attention sublayers, interactive modeling is performed on the initial representation after structural enhancement to obtain a universal transcriptome hidden representation.

4. The method according to claim 3, characterized in that, The process of preconstructing the gene relationship map includes: Determine a unified set of gene nodes based on a pre-defined gene sequence list; Iterate through any two genes, calculate the correlation between any gene pair, and obtain the original set of gene pair correlations; Based on the original set of correlations between the gene pairs, a weighted edge set is constructed. Based on the unified set of gene nodes and the set of weighted edges, a weighted adjacency matrix is ​​constructed; A gene relationship graph is obtained using the unified set of gene nodes, the set of weighted edges, and the weighted adjacency matrix.

5. The method according to claim 3, characterized in that, The interaction modeling of the initial representation after structural enhancement to obtain a universal transcriptome hidden representation includes: The initial representation after structural enhancement is processed to obtain the attention output; The attention output is nonlinearly enhanced to obtain a general transcriptome hidden representation.

6. The method according to claim 5, characterized in that, When the number of genes exceeds a preset value, the process of processing the initial representation after structural enhancement to obtain attention output includes: Divide the gene dimensions according to the preset block size; For each block, the initial representation after structural enhancement is processed to obtain the corresponding attention output; At the gene level, the attention outputs corresponding to each segment are spliced ​​together to obtain the final attention output.

7. The method according to claim 2, characterized in that, The pre-defined analysis task is a classification task. The prediction results corresponding to the pre-defined analysis task, obtained using the universal transcriptome hidden representation, based on the pre-defined analysis task, include: At the gene level, the general transcriptome hidden representation is subjected to global average pooling to obtain the sample-level hidden representation; By utilizing sample-level hidden representations and mapping them through a feedforward network, classification prediction results are obtained.

8. A transcriptome data analysis device, characterized in that, include: The data provision module is used to provide raw transcriptomic data and sample batch origin information for biological samples; The transcriptome data processing module is used to process the raw transcriptome data to obtain a gene expression matrix; The model input data acquisition module is used to obtain standardized model input data based on the gene expression matrix and the sample batch source information. The standardized model input data includes: a masked standardized gene expression matrix, mask position index information, gene-level protein sequence embedding features, and batch identification information. The general representation acquisition module inputs the standardized model input data into the transcriptome data analysis model to obtain a general transcriptome hidden representation. The transcriptome data analysis model is configured to process the standardized model input data to obtain the general transcriptome hidden representation and, based on the general transcriptome hidden representation, obtain the predicted expression value of the masked gene. The result prediction module is used to obtain the prediction results corresponding to the pre-defined analysis task by utilizing the general transcriptome hidden representation.

9. A transcriptome data analysis device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement the steps of the transcriptome data analysis method as described in any one of claims 1-7.

10. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the transcriptome data analysis method as described in any one of claims 1-7.