Multi-dimensional embedding fusion transcriptome-based model training method and device, and related equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-12
- Publication Date
- 2026-08-11
AI Technical Summary
现有转录组模型存在多源异构特征难以实现统一融合,基因表达值、蛋白质序列先验、样本全局分布、批次来源等信息维度割裂,缺乏专门针对转录组数据特性设计的嵌入框架,大多采用简单的特征拼接或单一维度嵌入方式,无法实现多源信息的深度融合与协同作用,导致模型无法充分利用多维度生物信息,提取的特征质量低下,难以精准刻画基因表达规律与生物特性
[0015]As can be seen from the above technical solutions, the multidimensional embedding fusion transcriptome basic model training method, apparatus, and related equipment provided in this application embodiment include: providing raw transcriptome data and sample batch source information of biological samples; processing the raw transcriptome data to obtain a gene expression matrix; obtaining structured training data and masked gene true expression value labels based on the gene expression matrix and the sample batch source information; the structured training data includes: a masked standardized gene expression matrix, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information; inputting the structured training data into the transcriptome basic model to be trained to obtain the predicted expression value of the masked gene; the transcriptome basic model to be trained is configured to have the ability to process the structured training data to obtain fusion features and obtain the predicted expression value of the masked gene based on the fusion features; and updating the parameters of the transcriptome basic model to be trained until convergence, with the predicted expression value of the masked gene approaching the true expression value label of the masked gene as the training objective. This application processes raw transcriptome data to obtain structured training data including a standardized gene expression matrix after masking, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information. During the training of the transcriptome basic model, the structured training data is processed to obtain fusion features, and based on the fusion features, the predicted expression values of the masked genes are obtained. This achieves unified fusion of multi-source features of transcriptome data, thereby constructing a robust transcriptome basic model with excellent generalization ability.
Smart Images

Figure CN122551902A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data analysis technology, and in particular to a method, apparatus and related equipment for training a multidimensional embedded fusion transcriptome basic model. Background Technology
[0002] Transcriptome data, as the core carrier of gene expression information in organisms, contains multi-source heterogeneous information such as gene expression values, corresponding protein sequences, and sample batches. It is the core data foundation for analyzing gene regulatory networks, mining disease-related biomarkers, and exploring the molecular mechanisms of life activities, and has irreplaceable and important value in many fields such as life science research, clinical diagnosis, and drug development.
[0003] With the development of sequencing technology, the scale of transcriptome data has increased dramatically, placing higher demands on the training of general representation models based on deep learning. Existing transcriptome models suffer from difficulties in achieving unified fusion of multi-source heterogeneous features, fragmented information dimensions such as gene expression values, protein sequence priors, global sample distribution, and batch origin, and a lack of embedding frameworks specifically designed for the characteristics of transcriptome data. Most models employ simple feature splicing or single-dimensional embedding methods, failing to achieve deep fusion and synergistic effects of multi-source information. This results in models that cannot fully utilize multi-dimensional biological information, extracting low-quality features that are difficult to accurately characterize gene expression patterns and biological characteristics. Therefore, there is an urgent need for a model training method that can achieve unified fusion of multi-source features from transcriptome data, construct robust and highly generalizable basic transcriptome models, and promote the in-depth mining and application of transcriptome data. Summary of the Invention
[0004] In view of this, this application provides a method, apparatus and related equipment for training a multidimensional embedded fusion transcriptome basic model, so as to facilitate the training of a transcriptome basic model that achieves unified fusion of multi-source features of transcriptome data.
[0005] To achieve the above objectives, the following solution is proposed: A method for training a multidimensional embedding fusion transcriptome basic model includes: Provide raw transcriptome data and batch origin information for biological samples; The raw transcriptome data is processed to obtain a gene expression matrix; Based on the gene expression matrix and the sample batch source information, structured training data and the true expression value labels of the masked genes are obtained. The structured training data includes: a standardized gene expression matrix after masking, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information. Structured training data is input into the transcriptome base model to be trained to obtain the predicted expression values of the masked genes. The transcriptome base model to be trained is configured to process the structured training data to obtain fusion features and, based on the fusion features, obtain the predicted expression values of the masked genes. The training objective is to make the predicted expression value of the masked gene approach the actual expression value of the masked gene. The parameters of the transcriptome base model to be trained are updated until convergence.
[0006] Optionally, the basic transcriptome model to be trained includes: a cascaded input layer, a multi-dimensional fusion embedding layer, a gene structure and global interaction modeling layer, and an output layer; The training process of the transcriptome base model to be trained includes: Structured training data is obtained through the input layer; Through the multi-dimensional fusion embedding layer, fusion features are obtained by utilizing the masked standardized gene expression matrix, gene-level protein sequence embedding feature matrix, and batch identification information in the structured training data; The fusion features are modeled using the gene structure and global interaction modeling layer to obtain a general feature representation; The predicted expression value of the masked gene is obtained through the output layer based on the general feature representation and the mask position index information. The parameters of the transcriptome base model are updated until convergence, with the training objective being that the predicted expression value of the masked gene approaches the actual expression value of the masked gene.
[0007] Optionally, the structured training data further includes: a masked standardized gene expression matrix, a gene-level protein sequence embedding feature matrix, and batch identification information. The multi-dimensional fusion embedding layer includes: an expression embedding sublayer, a gene embedding sublayer, a sample context embedding sublayer, a batch embedding sublayer, and a feature fusion layer. The fused features are obtained through the multi-dimensional fusion embedding layer using the masked standardized gene expression matrix, gene-level protein sequence embedding feature matrix, and batch identification information in the structured training data, including: By using the expression embedding sublayer, the expression embedding matrix is obtained based on the masked normalized gene expression matrix; By transforming the dimension of the gene-level protein sequence embedding feature matrix through the gene embedding sublayer, the gene-level embedding feature matrix is obtained. By using the sample context embedding sublayer, sample context embedding features are generated based on the masked, standardized gene expression matrix. A batch embedding tensor is generated based on batch identification information through a batch embedding sublayer. The expression embedding matrix, the gene-level embedding feature matrix, the sample context embedding feature, and the batch embedding tensor are fused through the feature fusion layer to obtain fused features.
[0008] Optionally, obtaining the predicted expression value of the masked gene based on the general feature representation and the mask position index information includes: The general feature representation is subjected to layer normalization to obtain a gene-level embedding matrix; Based on the gene-level embedding matrix, the corresponding hidden representation is extracted according to the mask position index information; Based on the hidden representation, regression prediction is performed using a fully connected network to obtain the predicted expression values of the masked genes.
[0009] Optionally, the structured training data includes: a standardized gene expression matrix after masking, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information. The step of obtaining structured training data and the true expression value labels of the masked genes based on the gene expression matrix and the sample batch source information includes: Based on the preset gene annotation information, the gene expression matrix is standardized to obtain standardized gene expression data. Based on a preset gene sequence list, the standardized gene expression data is reordered to obtain a standardized gene expression matrix; Based on a pre-defined list of gene sequences, a corresponding gene-level protein sequence embedding feature matrix is constructed. Based on a preset masking strategy, the expression values of some genes in the standardized gene expression matrix are masked to obtain the masked standardized gene expression matrix, masking position index information, and labels of the true expression values of the masked genes. The source information of the sample batch is encoded to obtain batch identification information.
[0010] Optionally, the preset gene annotation information includes: a unique gene identifier and corresponding gene length information. The standardization process of the gene expression matrix based on the preset gene annotation information to obtain standardized gene expression data includes: Based on the unique gene identifier, the gene length information corresponding to each gene in the gene expression matrix is determined; For each gene, based on the corresponding gene length information, the original gene expression value is converted into the number of reads per million transcripts. The expression value of reads per million transcripts was transformed to obtain standardized gene expression data.
[0011] Optionally, the step of constructing a corresponding gene-level protein sequence embedding feature matrix based on a preset gene sequence list includes: Based on a pre-defined gene sequence list, determine the protein sequence information corresponding to each gene; Protein sequence information is input into a pre-trained biological sequence representation model to obtain the corresponding gene-level protein sequence embedding feature vector. The biological sequence representation model is trained using protein sequence information as training samples and the gene-level protein sequence embedding feature vector corresponding to the protein sequence information as training labels. Based on the gene-level protein sequence embedding feature vector, the gene-level protein sequence embedding feature matrix is obtained, and a gene-level protein sequence embedding feature matrix corresponding one-to-one with the standardized gene expression matrix is constructed.
[0012] A multidimensional embedding fusion transcriptome basic model training device, comprising: The training data provision module provides raw transcriptome data and batch source information for biological samples. The raw data processing module is used to process the raw transcriptome data to obtain a gene expression matrix; The structured training data acquisition module is used to obtain structured training data and the true expression value labels of the masked genes based on the gene expression matrix and the sample batch source information. The model training module is used to input structured training data into the transcriptome basic model to obtain the predicted expression values of the masked genes. The transcriptome basic model to be trained is configured to process the structured training data to obtain fusion features and, based on the fusion features, obtain the predicted expression values of the masked genes. The model update module is used to update the parameters of the transcriptome base model until convergence, with the training objective being that the predicted expression value of the masked gene approaches the actual expression value label of the masked gene.
[0013] A training device for a multidimensional embedded fusion transcriptome basic model includes: a memory and a processor; The memory is used to store programs; The processor is configured to execute the program to implement the various steps of the multidimensional embedding fusion transcriptome basic model training method as described in any of the preceding claims.
[0014] A readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multidimensional embedding fusion transcriptome basic model training method as described in any of the preceding claims.
[0015] As can be seen from the above technical solutions, the multidimensional embedding fusion transcriptome basic model training method, apparatus, and related equipment provided in this application embodiment include: providing raw transcriptome data and sample batch source information of biological samples; processing the raw transcriptome data to obtain a gene expression matrix; obtaining structured training data and masked gene true expression value labels based on the gene expression matrix and the sample batch source information; the structured training data includes: a masked standardized gene expression matrix, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information; inputting the structured training data into the transcriptome basic model to be trained to obtain the predicted expression value of the masked gene; the transcriptome basic model to be trained is configured to have the ability to process the structured training data to obtain fusion features and obtain the predicted expression value of the masked gene based on the fusion features; and updating the parameters of the transcriptome basic model to be trained until convergence, with the predicted expression value of the masked gene approaching the true expression value label of the masked gene as the training objective. This application processes raw transcriptome data to obtain structured training data including a standardized gene expression matrix after masking, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information. During the training of the transcriptome basic model, the structured training data is processed to obtain fusion features, and based on the fusion features, the predicted expression values of the masked genes are obtained. This achieves unified fusion of multi-source features of transcriptome data, thereby constructing a robust transcriptome basic model with excellent generalization ability.
[0016] Furthermore, the transcriptome basic model can learn the correlation features between gene expression in transcriptome data through a multidimensional embedding mechanism, enhance the ability to model and fill in missing gene expression values, and generate transcriptome feature embeddings with universal expression capabilities, thereby providing unified and reusable feature input support for various downstream bioinformatics analysis tasks. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0018] Figure 1 A flowchart illustrating a training method for a multidimensional embedding fusion transcriptome basic model provided in this application embodiment; Figure 2 An internal architecture diagram of an optional transcriptome basic model provided in this application embodiment; Figure 3An internal architecture diagram of another optional transcriptome basic model provided in this application embodiment; Figure 4 A flowchart for acquiring structured training data is provided as an embodiment of this application; Figure 5 A schematic diagram of a multidimensional embedding fusion transcriptome basic model training device provided in this application embodiment; Figure 6 This is a hardware structure block diagram of a multidimensional embedded fusion transcriptome basic model training device provided in an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] Figure 1 A flowchart of a training method for a multidimensional embedding fusion transcriptome basic model provided in this application embodiment is included. The method may include the following steps: Step S100: Provide raw transcriptome data and sample batch source information for the biological sample.
[0021] Specifically, to construct high-quality transcriptome datasets for building foundational large-scale models for transcriptome data characterization, publicly available large-scale transcriptome databases, such as the GEO and ARCHS4 databases, can be selected as sources of raw transcriptome sequencing data, and raw gene expression count data of the corresponding samples can be obtained. These databases are maintained by internationally authoritative bioinformatics institutions and contain a large amount of standardized and organized transcriptome sequencing data and corresponding sample annotation information, providing high-throughput expression data support from multiple sources, tissues, and disease states for training basic transcriptome models.
[0022] Batch origin information is used to characterize the experimental batch, sequencing batch, or data source batch of the sample, etc.
[0023] Step S101: Process the raw transcriptome data to obtain the gene expression matrix.
[0024] Specifically, by identifying and processing erroneous, inconsistent, redundant, or incomplete information, and in conjunction with transcriptome data processing standards, raw transcriptome data can be quality controlled, quantitatively analyzed, and structured to ultimately generate high-quality, usable transcriptome data.
[0025] To ensure data quality and consistency, the raw transcriptome data can be deduplicated first. For samples with duplicate sequencing or submissions, only the original gene count version is retained, and the standardized or secondary processed expression matrices are removed to obtain the deduplicated raw transcriptome data, ensuring the originality and consistency of the data source. Based on the deduplicated raw transcriptome data, a gene set of a predetermined type is selected to construct a gene expression matrix. The gene expression matrix can be formatted with samples as rows and genes as columns, with matrix elements representing the expression level of the corresponding gene in the corresponding sample.
[0026] Step S102: Based on the gene expression matrix and sample batch source information, obtain structured training data and the true expression value labels of the masked genes.
[0027] Specifically, structured training data can include: a masked, standardized gene expression matrix, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information. The data format of structured training data can precisely adapt to the input requirements of basic transcriptome models, ensuring that the data structure, parameter dimensions, and information integrity are highly matched with the model's acceptance standards, providing a data foundation for the smooth progress of subsequent transcriptome data characterization processes and the accuracy of results.
[0028] Step S103: Input the structured training data into the transcriptome base model to be trained to obtain the predicted expression values of the masked genes.
[0029] Specifically, the training transcriptome baseline model is used to characterize transcriptome-related information based on structured training data. This model is configured to process the structured training data, obtain fused features, and, based on these fused features, predict the expression values of masked genes. By fusing multi-dimensional features within the model, it can comprehensively consider inherent gene attributes, the overall expression state of the sample, and systematic differences introduced by different experimental batches or data sources while learning gene expression patterns. This improves the model's robustness and generalization ability under complex transcriptome data conditions, providing a unified feature representation foundation. This allows the transcriptome baseline model not only to predict the expression values of masked genes but also to be directly used for predictions in other analytical tasks, avoiding repeated feature engineering or model training for different tasks and effectively improving the efficiency, consistency, and scalability of transcriptome data analysis.
[0030] During model training, a masking prediction strategy is adopted to guide the model to reconstruct and learn the expression values of masked genes using the contextual information of unmasked gene expression. This enables the model to automatically learn the potential association patterns and global expression structure between gene expressions without relying on specific downstream task labels.
[0031] Step S104: Update the parameters of the transcriptome base model to be trained until convergence.
[0032] Specifically, the training objective is to make the predicted expression values of the masked genes as close as possible to the actual expression values of the masked genes. The parameters of the transcriptome base model to be trained are updated until convergence. In the structured training data, some gene expression values are masked using a pre-defined masking strategy. During model training, the expression context information of the unmasked genes can be used to predict the expression values of the masked genes. Based on the error between the prediction results and the corresponding actual expression value labels, a loss function is used to train the transcriptome base model to obtain the target transcriptome base model. The loss function can include mean squared error loss, mean absolute error loss, cross-entropy loss, or a combination thereof.
[0033] Transcriptome basic models can learn the correlation features between gene expression in transcriptome data through multidimensional embedding mechanisms, enhance the ability to model and impute missing gene expression values, and generate transcriptome feature embeddings with universal expression capabilities, thereby providing unified and reusable feature input support for various downstream bioinformatics analysis tasks.
[0034] Furthermore, when training the transcriptome-based model, the structured training data can be divided into training, validation, and test datasets. The training set can be used for model parameter updates, the validation set for model performance monitoring and optimal model selection, and the test set for final generalization ability evaluation. The training, validation, and test sets can be divided in a 9:0.5:0.5 ratio. The data can be stored in HDF5 format and read in batches using a custom data loading module, thus supporting efficient reading of large-scale samples.
[0035] The structured training data includes at least gene expression features, gene-level protein sequence embedding feature matrix, mask position index information, the true expression value of the masked gene, and batch identification information corresponding to the sample, as well as other sample-level or gene-level auxiliary features introduced according to the specific application scenario.
[0036] During model training, the training status of the model is monitored using a validation dataset. When preset conditions are met, the model parameters are saved or updated to obtain the trained target transcriptome basic model. For example, the preset condition could be a preset loss value. When the model loss value reaches the preset loss value, the current model parameters are saved, resulting in the trained target transcriptome basic model.
[0037] As can be seen from the above embodiments, the multidimensional embedding fusion transcriptome basic model training method provided in this application includes: providing raw transcriptome data and sample batch source information of biological samples; processing the raw transcriptome data to obtain a gene expression matrix; obtaining structured training data and masked gene true expression value labels based on the gene expression matrix and sample batch source information; the structured training data includes: a masked standardized gene expression matrix, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information; inputting the structured training data into the transcriptome basic model to be trained to obtain the predicted expression value of the masked gene; the transcriptome basic model to be trained is configured to have the ability to process the structured training data to obtain fusion features and obtain the predicted expression value of the masked gene based on the fusion features; and updating the parameters of the transcriptome basic model to be trained until convergence, with the predicted expression value of the masked gene approaching the true expression value label of the masked gene as the training objective. This application processes raw transcriptome data to obtain structured training data, including a standardized gene expression matrix after masking, mask position index information, gene-level protein sequence embedding feature matrix, and batch identification information. During the training of the transcriptome basic model, the structured training data is processed to obtain fusion features, and based on the fusion features, the predicted expression values of the masked genes are obtained. This achieves unified fusion of multi-source features of transcriptome data, thereby constructing a robust transcriptome basic model with excellent generalization ability.
[0038] Furthermore, the transcriptome basic model can learn the correlation features between gene expression in transcriptome data through a multidimensional embedding mechanism, enhance the ability to model and fill in missing gene expression values, and generate transcriptome feature embeddings with universal expression capabilities, thereby providing unified and reusable feature input support for various downstream bioinformatics analysis tasks.
[0039] In some embodiments of this application, reference is made to Figure 2 As shown, Figure 2 The following is an internal architecture diagram of an optional transcriptome basic model provided in this application embodiment. The transcriptome basic model to be trained may include: an input layer, a multi-dimensional fusion embedding layer, a gene structure and global interaction modeling layer, and an output layer, which are cascaded in sequence. Based on this, the training process of the basic transcriptome model to be trained can include: Structured training data is obtained through the input layer.
[0040] Specifically, structured training data may include: a standardized gene expression matrix after masking, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information.
[0041] By using a multi-dimensional fusion embedding layer, fused features are obtained by utilizing the masked standardized gene expression matrix, gene-level protein sequence embedding feature matrix, and batch identification information in the structured training data.
[0042] Specifically, the masked standardized gene expression matrix, gene-level protein sequence embedding feature matrix, and batch identification information in the structured training data are mapped into a vector representation of a unified dimension and then fused.
[0043] By modeling the fusion features through gene structure and global interaction modeling layers, a general feature representation is obtained.
[0044] Specifically, the fusion feature input gene structure and global interaction modeling layer are used to model and obtain a general feature representation. This general feature representation integrates gene expression numerical characteristics, inherent gene sequence properties, global expression context of the sample, batch information, gene structure relationship propagation information, and global dependency modeling information. Therefore, it can be used as a high-quality gene characterization for subsequent downstream application analyses, such as mask expression value reconstruction tasks, sample classification tasks, disease classification or disease risk prediction tasks, gene necessity score prediction tasks, gene expression change prediction tasks under compound perturbation conditions, drug sensitivity or drug response prediction tasks, biological age assessment tasks, cell composition deconvolution analysis tasks, and biological phenotypic association analysis tasks.
[0045] Depending on the different downstream application analysis scenarios, it supports corresponding input formats. By constructing a general feature representation, it realizes a large-scale transcriptome basic model framework with dual-layer representation at the gene and sample levels, self-supervised pre-training and unified modeling for downstream tasks, and transferability and scalability, supporting various bioinformatics application scenarios. Compared with traditional models trained based on a single task, the general feature representation output by this application has stronger generalization and expressive power, and can serve as the core representation form of the basic transcriptome model.
[0046] The predicted expression values of the masked genes are obtained through the output layer, based on the general feature representation and mask position index information.
[0047] Specifically, task-level mapping is performed on the general feature representations output from the gene structure and global interaction modeling layer to generate predicted expression values for masked genes. These general feature representations characterize the overall feature distribution of samples at the transcriptome level and can serve as a universal feature representation for various downstream analysis tasks. These downstream tasks may include: disease classification, disease risk prediction, gene necessity score prediction, compound-perturbed gene expression level prediction, drug sensitivity prediction, biological age assessment, cell composition deconvolution, and biological phenotypic association analysis.
[0048] The training process of the aforementioned transcriptome baseline model aims to make the predicted expression values of the masked genes approach the labels of their actual expression values, thereby updating the parameters of the baseline transcriptome model. A masked regression loss can be constructed based on the mean squared error between the predicted and actual expression values of the masked genes, serving as a self-supervised training objective. This allows the model to learn the intrinsic dependencies between gene expressions and optimize general feature representations.
[0049] Furthermore, during model training, parallel training can be employed to synchronously update model parameters across multiple computational units, thereby improving training efficiency. The model parameter update process can accumulate gradient information from multiple training batches and then uniformly update the model parameters after reaching a preset number of accumulation steps, achieving equivalent large-batch training.
[0050] Simultaneously, model parameters can be updated based on adaptive optimization strategies, and combined with a learning rate scheduling mechanism, a learning rate warm-up process is performed in the early stages of training, and a learning rate decay process is performed in subsequent training stages to improve the stability and convergence of model training. A mixed-precision computation method is used for forward and backward propagation calculations, and a gradient scaling mechanism is introduced during low-precision gradient calculations to reduce computational resource consumption and ensure numerical stability. When the model's performance metrics on the validation dataset are better than the historical best results, the corresponding model parameters are saved as the base model for the target transcriptome.
[0051] In some embodiments of this application, the AdamW optimizer can be used to update model parameters. This optimizer, based on traditional adaptive optimization methods, introduces a weight decay mechanism, which helps suppress overfitting and improve generalization ability. Simultaneously, a learning rate scheduling strategy of warm-up followed by linear decay is employed during training. The learning rate is gradually increased in the early stages of training to allow the model to smoothly enter the convergence phase, and then gradually decreased in subsequent stages to achieve stable convergence. This scheduling method effectively avoids unstable oscillations in the early stages of training and enables refined optimization in later stages.
[0052] In some embodiments of this application, reference is made to Figure 3 As shown, Figure 3 The following is an alternative internal architecture diagram of the transcriptome-based model provided in this application embodiment. The multi-dimensional fusion embedding layer may include: an expression embedding sublayer, a gene embedding sublayer, a sample context embedding sublayer, a batch embedding sublayer, and a feature fusion layer. Based on this, through the multi-dimensional fusion embedding layer, using the masked standardized gene expression matrix, gene-level protein sequence embedding feature matrix, and batch identification information from the structured training data, fused features are obtained, which may include: By using the expression embedding sublayer, the expression embedding matrix is obtained based on the masked, standardized gene expression matrix.
[0053] Specifically, the masked normalized gene expression matrix undergoes multi-scale periodic encoding mapping to preserve gene expression intensity information and quantitative variation characteristics. Assume the masked normalized gene expression matrix is as follows: , in, For batch size, This represents the number of genes.
[0054] A set of preset frequency reference vectors is constructed to implement multi-scale periodic coding mapping. Assuming the embedding dimension is d, we can set d=640. Then the frequency reference vector can be defined as: , in, For frequency index, .
[0055] At this point, the frequency set can be: , Among them, the frequency parameter can be a preset constant parameter, which is used to construct the expression change response at different frequency scales, thereby enhancing the model's ability to perceive expression changes of different orders of magnitude.
[0056] By performing frequency expansion mapping on each expression value in the masked, normalized gene expression matrix, we can obtain: , in, For the first In the nth sample The expression value of each gene, For frequency index, .
[0057] Further performing sine and cosine periodic transformations on the angle value yields: ; .
[0058] By concatenating the above data, a preliminary gene expression embedding representation is obtained: .
[0059] When the gene expression value is set to a preset mask identifier value m, and m = -10, a learnable mask embedding vector can be introduced. The gene expression embedding representation at the corresponding position is replaced, and the mask indicator function can be defined as follows: .
[0060] Through the above steps, the final gene expression embedding representation can be obtained: .
[0061] Thus, the embedding matrix can be obtained: .
[0062] By performing dimensionality transformation on the gene-level protein sequence embedding feature matrix through the gene embedding sublayer, the gene-level embedding feature matrix is obtained.
[0063] Specifically, the gene-level protein sequence embedding feature matrix can be: , in, The total number of feature vectors embedded in gene-level protein sequences. For the embedded dimension.
[0064] To match the gene-level protein sequence embedding features with the unified feature space dimension within the model, a projection network consisting of two layers of linear mapping and nonlinear transformation is used to transform the dimension of the gene-level protein sequence embedding feature matrix, namely: , , in, For intermediate mapping features, It is a gene-level embedding feature. For non-linear activation functions, in this embodiment, the ReLU activation function can be used. For the first The gene-level protein sequence embedding feature vector of each gene. , This is the weight matrix. , This is a bias term.
[0065] Based on the above steps, the gene-level embedding feature matrix can be obtained: .
[0066] Furthermore, to adapt to batch sample input, the gene-level embedding feature matrix can be expanded in the sample dimension. Assuming the batch sample size is B, the expanded gene-level embedding feature matrix can be: .
[0067] Through the above design, the gene embedding sublayer can introduce the inherent sequence-level properties of genes into the unified feature space of the model and participate in the subsequent modeling process as static prior information, thereby enhancing the model's ability to express gene functional characteristics and potential regulatory relationships.
[0068] By using the sample context embedding sublayer, sample context embedding features are generated based on the masked, standardized gene expression matrix.
[0069] Specifically, the normalized gene expression matrix after masking can be: , in, For batch size, This represents the number of genes.
[0070] The process involves a first linear mapping, a nonlinear activation transformation, and a second linear mapping, namely: , , in, For intermediate mapping features, For the first The sample context embedding vector of each sample is used to characterize the potential state of the overall transcriptome expression of that sample. For non-linear activation functions, in this embodiment, the GELU activation function can be used. , This is the weight matrix. , This is a bias term.
[0071] To enable sample context to participate in gene-level modeling, the sample context embedding vector can be extended to gene-dimensional alignment. The sample context embedding features can be: .
[0072] Extending the sample context embedding features by expanding them along the gene dimension yields the expanded features: .
[0073] Thus, all genes within the same sample share the same context embedding vector, with replication and expansion occurring only at the gene dimension.
[0074] Based on the standardized gene expression matrix after masking, a nonlinear mapping is performed on the overall gene expression distribution of a single sample to generate a sample-level context representation, which is then extended in the gene dimension so that all genes within the same sample share the context embedding vector.
[0075] The sample context embedding sublayer is used to characterize the overall transcriptional state at the sample level, providing a global representation of the sample from the perspective of whole-gene expression distribution. This enables the model to perceive the overall expression pattern within a single sample. The sample context embedding sublayer uses a multilayer perceptron structure to perform nonlinear mapping on the sample-level expression vector, generating sample context embedding features. Through this sample context embedding feature construction method, the overall expression distribution of the sample is compressed into a low-dimensional latent representation, enabling the model to perceive the overall state of the sample during gene-level modeling. This improves the model's ability to express global changes at the sample level and provides prior global contextual knowledge for subsequent processing.
[0076] The batch embedding sublayer generates a batch embedding tensor based on batch identification information.
[0077] Specifically, assuming the number of batches is M, and each sample corresponds to a discrete batch identifier, that is: , in, For the first Batch number of each sample For batch size, .
[0078] Furthermore, define the batch embedding matrix: , Among them, the m-th row For the learnable batch embedding tensor corresponding to the m-th batch, Parameters can be updated during model training.
[0079] Through the above steps, discrete batch numbers can be mapped to continuous vector representations: .
[0080] The batch embedding tensor can be obtained: .
[0081] Since the model is modeled in the form of gene-level sequences, the sample-level batch embeddings are extended to the gene dimension. All genes within the same sample share the same batch embedding vector. The extended batch embedding tensor can be: .
[0082] The batch identification information corresponding to the samples is embedded to construct a representation to explicitly model the systematic biases introduced by different experimental batches, sequencing platforms or data sources. By introducing a learnable batch embedding tensor, the model can automatically learn and correct the impact of batch effects on gene expression patterns.
[0083] Through the feature fusion layer, the representation embedding matrix is processed. Gene-level embedding feature matrix Sample context embedding features and batch embedding tensors The fusion process is performed to obtain a unified fusion feature representation.
[0084] Specifically, the fusion process can be performed using an additive approach, and the fusion features can be: .
[0085] Furthermore, before modeling the fused features to obtain a general feature representation, layer normalization and layer discarding can be performed on the fused features to obtain the final fused features.
[0086] Specifically, performing layer normalization and layer discarding on the fused features can improve numerical stability and convergence speed. The processing can be achieved using the following formula: , , in, For layer normalization, This is a discard layer.
[0087] Furthermore, regarding Perform two-layer feedforward mapping: , in, , This is the weight matrix. The activation function is non-linear; in this embodiment, it can be the ReLU activation function.
[0088] The final fusion characteristics can be obtained: .
[0089] Based on this, modeling the fused features to obtain a general feature representation can include: modeling the final fused features to obtain a general feature representation.
[0090] In some embodiments of this application, obtaining the predicted expression value of the masked gene based on general feature representation and mask position index information may include: S11. Perform layer normalization on the general feature representation to obtain the gene-level embedding matrix.
[0091] Specifically, the general feature representation output of the gene structure and global interaction modeling layer. Perform layer normalization to obtain .
[0092] For the The nth sample, the th The embedding representation of a gene can be defined as: .
[0093] This allows us to obtain the gene-level embedding matrix: .
[0094] To obtain an overall representation of the sample, global pooling is performed on the gene dimension, for example, using global max pooling: .
[0095] The sample-level embedding matrix can be obtained: .
[0096] in, This sample-level embedding matrix can characterize the potential representation of the overall transcriptional state of a sample.
[0097] S12. Based on the gene-level embedding matrix, extract the corresponding hidden representation according to the mask position index information.
[0098] Specifically, based on the mask position index information, which can be a mask position index matrix, the hidden representation of the corresponding masked gene position is extracted from the gene-level embedding matrix.
[0099] Let the mask position index matrix be: , in, The number of masked genes for each sample.
[0100] Using the mask position index matrix, the hidden representation corresponding to the masked gene position is extracted from the gene-level embedding matrix: , in, For the first The first sample The location index of the masked gene.
[0101] The final result is: .
[0102] S13. Based on the hidden representation, a fully connected network is used for regression prediction to obtain the predicted expression value of the masked gene.
[0103] Specifically, a subset of masked features can be constructed using the hidden representation of the masked gene location, and then input into a mask prediction head for regression mapping. The mask prediction head comprises a first fully connected layer, a non-linear activation layer, and a second fully connected layer connected in sequence, which output the predicted expression value of the corresponding masked gene.
[0104] Regression prediction is performed using a two-layer fully connected network: , , in, , This is the weight matrix. , This is a bias term.
[0105] Finally, the predicted expression values of the masked genes were obtained: .
[0106] To enable the model to learn the intrinsic dependencies between expression values and thus obtain high-quality gene embedding representations, the mask regression loss can be defined as the mean squared error. Based on this, the mask regression loss can be: , in, This represents the true expression value of the masked gene. The predicted expression value of the masked gene.
[0107] In some embodiments of this application, the structured training data may include: a masked, standardized gene expression matrix, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information. The data format of the structured training data is adapted to the input requirements of the transcriptome model to be trained. (Reference) Figure 4 As shown, Figure 4 This application provides a flowchart for obtaining structured training data. Based on this, step S102, obtaining structured training data and masked gene true expression value labels based on gene expression matrix and sample batch source information, may include: S21. Based on the preset gene annotation information, the gene expression matrix is standardized to obtain standardized gene expression data.
[0108] Specifically, standardized expression data serves as the basis for subsequent gene sequence alignment and feature construction processing.
[0109] S22. Based on a preset gene sequence list, the standardized gene expression data are reordered to obtain a standardized gene expression matrix.
[0110] Specifically, based on a pre-defined gene sequence list, standardized gene expression data are reordered, and zeros are padded at the positions corresponding to missing genes to obtain a standardized gene expression matrix with consistent gene order. The pre-defined gene sequence list is constructed based on an authoritative gene annotation database and serves as the sole standard for aligning gene expression data during the training of the multidimensional embedding fusion transcriptome basic model. By constructing a unified gene sequence representation framework, transcriptome data from different sources and under different experimental conditions can be aligned into structurally consistent gene expression representations.
[0111] S23. Based on the preset gene sequence list, construct the corresponding gene-level protein sequence embedding feature matrix.
[0112] S24. Based on a preset masking strategy, mask some gene expression values in the standardized gene expression matrix to obtain the masked standardized gene expression matrix, masking position index information, and the true expression value label of the masked gene.
[0113] Specifically, the preset masking strategy can be as follows: From the standardized gene expression matrix, a portion of gene expression values are randomly selected as masking targets according to a preset ratio. These selected gene expression values are then masked, making them invisible in the model input or replaced with preset placeholder values. The location index of the masked genes and their actual expression values are recorded. The masked standardized gene expression matrix is used as input features, the actual expression values of the masked genes are used as prediction targets, and the masked location index information is used as auxiliary input information. This allows the model to utilize the remaining unmasked gene expression features as contextual information to reconstruct and predict the masked gene expression values.
[0114] The preset ratio is any ratio within a preset range of the total number of genes. Masking can be implemented in various ways, such as: replacing the selected gene expression value with zero; replacing the selected gene expression value with a preset random noise value; replacing the selected gene expression value with a mask identifier value used to identify missing states; keeping the gene expression value unchanged but marking it as a state to be predicted in the model, etc.
[0115] S25. Encode the source information of the sample batch to obtain batch identification information.
[0116] Specifically, batch origin information is used to characterize the experimental batch, sequencing batch, or data source batch of the sample. The sample batch origin information is encoded, mapping different batches to different discrete batch identifier values. These batch identifier values are then further converted into continuous vector representations to obtain batch identification information. This information, along with the standardized gene expression matrix, mask position index information, and gene-level protein sequence embedding feature matrix, constitutes structured training data for subsequent model input. For samples that cannot be matched to a known batch origin, a pre-defined unknown batch identifier value is uniformly assigned as batch identification information.
[0117] The processing flow of this embodiment can, to a certain extent, ensure that transcriptome data from different sources and batches have a unified data structure and comparability at the input level, providing a data foundation for the model to learn a stable and generalizable transcriptome representation.
[0118] In some embodiments of this application, the preset gene annotation information may include: a unique gene identifier and corresponding gene length information. Based on this, S21, the gene expression matrix is standardized based on the preset gene annotation information to obtain standardized gene expression data, which may include: S31. Based on the unique gene identifier, determine the gene length information corresponding to each gene in the gene expression matrix.
[0119] S32. For each gene, based on the corresponding gene length information, convert the original gene expression value into the number of reads per million transcripts.
[0120] Specifically, based on gene length information, the original gene expression values are normalized. This can include converting the original gene expression values into Transcripts Per Million (TPM) expression values based on gene length information to eliminate the impact of differences in sequencing depth and gene length.
[0121] S33. Transform the TPM expression values to obtain standardized gene expression data.
[0122] Specifically, the TPM expression value is transformed by log2(TPM+1) to compress the dynamic range of the expression value and enhance the numerical stability, thereby obtaining standardized transcriptome expression data at the gene level.
[0123] In some embodiments of this application, step S23, constructing a corresponding gene-level protein sequence embedding feature matrix based on a preset gene sequence list, may include: S41. Based on a preset gene sequence list, determine the protein sequence information corresponding to each gene.
[0124] Specifically, the protein sequence information corresponding to each gene can be determined based on a preset gene sequence list.
[0125] S42. Input the protein sequence information into the pre-trained biological sequence representation model to obtain the corresponding gene-level protein sequence embedding feature vector.
[0126] Specifically, the biological sequence representation model is trained using protein sequence information as training samples and the gene-level protein sequence embedding feature vectors corresponding to the protein sequence information as training labels.
[0127] S43. Based on the gene-level protein sequence embedding feature vector, obtain the gene-level protein sequence embedding feature matrix, and construct the gene-level protein sequence embedding feature matrix that corresponds one-to-one with the standardized gene expression matrix.
[0128] Specifically, the gene-level protein sequence embedding feature matrix is associated with the standardized gene expression matrix to construct a gene-level protein sequence embedding feature matrix that corresponds one-to-one with the standardized gene expression matrix. For genes lacking protein sequence information, zero vectors or preset default vectors are used to fill in the missing information to ensure that the feature dimensions are fixed.
[0129] In some embodiments of this application, a multi-GPU distributed data parallel training method can be adopted to support high-dimensional feature training on large-scale datasets.
[0130] Specifically, before training begins, process groups are initialized via a distributed communication module. Each GPU corresponds to an independent training process, and each process processes only a subset of the training data. To ensure data consistency across different GPUs, the training set is partitioned using a distributed sampler. At the start of each training cycle, the sampling random seed is reset, ensuring a different data order in each cycle and improving model generalization ability. After backpropagation, each process automatically synchronizes parameters through an efficient communication mechanism, ensuring consistent model parameters across all GPUs.
[0131] In some embodiments of this application, since the model dimension is large and the number of samples in a single batch is limited by the GPU memory, a gradient accumulation mechanism can be introduced.
[0132] Specifically, during actual training, after each mini-batch's forward computation and backpropagation, only the gradients are accumulated without immediately updating the model parameters. Once a preset number of steps are reached, a unified parameter update operation is performed. This approach achieves equivalent large-batch training results without increasing GPU memory usage, improves model convergence stability, and reduces training instability caused by gradient fluctuations.
[0133] In some embodiments of this application, an automatic mixed precision training mechanism can be used to improve training efficiency and reduce memory usage.
[0134] Specifically, during the forward propagation phase, the model is executed in low-precision floating-point format, thereby reducing memory usage, improving matrix multiplication speed, and increasing overall training throughput. When using half-precision floating-point format, a gradient scaling mechanism is introduced to avoid numerical underflow. When using bfloat16 format, good numerical stability can be maintained without gradient scaling. This strategy significantly improves training efficiency while ensuring that training accuracy is largely unaffected.
[0135] After each training cycle, the model's performance can be evaluated on the validation set, and the validation loss can be calculated. If the validation loss of the current training cycle is better than the historical best result, the current model parameters are saved as the best model; at the same time, periodic checkpoint files are saved at preset intervals so that training can be resumed in case of abnormal interruption.
[0136] Through the above training method, the model of this invention can support high-dimensional feature training at the scale of tens of thousands of genes, efficiently scale in a multi-GPU environment, improve memory utilization and computational efficiency, ensure the stability of the training process, and obtain embedding representations with good generalization ability. The training method described in this embodiment is applicable to large-scale biological transcriptome data and can also be extended to other high-dimensional structured biological data scenarios.
[0137] In some embodiments of this application, the transcriptome basic model trained based on the above embodiments can be used to process pre-set analysis tasks. The specific processing procedure is as follows: S51. Provide raw transcriptome data and batch source information for biological samples.
[0138] S52. Process the raw transcriptome data to obtain the gene expression matrix.
[0139] S53. Based on the gene expression matrix and sample batch source information, structured training data is obtained.
[0140] S54. Input the structured training data into the transcriptome basic model to obtain a general feature representation.
[0141] Specifically, during the training of the transcriptome basic model, a masked regression loss can be constructed based on the predicted expression values and the true expression values of the masked genes, using a loss function. This masked regression loss is then used to update the model parameters. The loss function can include the mean squared error loss function, the mean absolute error loss function, the cross-entropy loss, or a combination thereof.
[0142] S55. Based on the pre-defined analysis task, the prediction results corresponding to the pre-defined analysis task are obtained by using general feature representation.
[0143] Specifically, through the aforementioned structural design, the transcriptome base model can comprehensively model gene expression information, gene intrinsic attribute information, overall sample context information, and batch difference information within a unified embedding representation space. This enhances the model's representation and generalization capabilities for complex transcriptome data. Simultaneously, it supports self-supervised expression reconstruction and supervised classification tasks, and possesses flexible task extension capabilities, thereby improving the model's versatility, transferability, and practical application value. Furthermore, it provides stable and universal feature representations for various downstream analysis tasks. It should be noted that this step can be performed independently of the transcriptome base model, or it can be embedded as an output head for another downstream supervised task within the output layer of the transcriptome base model.
[0144] When the model is applied to downstream supervised tasks, other downstream supervised task output heads can be selectively enabled according to pre-defined analysis task requirements. For example, when a classification task is enabled, the sample-level embedding represents the classification prediction result obtained through a classification mapping network. Furthermore, classification loss can be calculated based on the true class labels to supervise and optimize the model parameters, enabling the model to learn sample-level discriminative features.
[0145] In some embodiments of this application, S54, inputting structured training data into the transcriptome basic model to obtain a general feature representation may include: Structured training data is obtained through the input layer.
[0146] Specifically, structured training data may include: a masked standardized gene expression matrix, mask position index information, gene-level protein sequence embedding features, and batch identification information.
[0147] By using a multi-dimensional fusion embedding layer, the standardized gene expression matrix after masking, gene-level protein sequence embedding features, and batch identification information are processed to obtain fused features.
[0148] Specifically, the masked standardized gene expression matrix, gene-level protein sequence embedding features, and batch identification information are fused to obtain fused features. By combining these multi-dimensional fused features within the model, the model can comprehensively consider inherent gene properties, the overall expression state of the sample, and systematic differences introduced by different experimental batches or data sources while learning gene expression patterns. This improves the model's robustness and generalization ability under complex transcriptome data conditions.
[0149] By modeling the fusion features through gene structure and global interaction modeling layers, a general feature representation is obtained.
[0150] Specifically, by integrating the input gene structure and global interaction modeling layer with the feature input, structural dependency modeling and global interaction modeling are performed to obtain a general feature representation. This representation can not only be used to predict the expression values of masked genes, but also for prediction of other analysis tasks. This avoids the need for repeated feature engineering design or model training for different tasks, and effectively improves the efficiency, consistency and scalability of transcriptome data analysis.
[0151] Through the output layer, based on the general feature representation and mask position index information, the predicted expression value of the masked gene is obtained.
[0152] Specifically, task-level mapping is performed on the general feature representations output by the gene structure and global interaction modeling layer to generate predicted expression values of the masked genes.
[0153] Through the above structural design, the transcriptome basic model can comprehensively model gene expression information, gene intrinsic attribute information, overall sample context information, and batch difference information under a unified fusion embedding representation space. This enhances the model's ability to represent and generalize complex transcriptome data. At the same time, it supports self-supervised expression reconstruction tasks and supervised classification tasks, achieving high-quality, generalized feature representation modeling of transcriptome data. It also has flexible task expansion capabilities, thereby improving the model's versatility, transferability, and practical application value, and providing stable and reusable general feature representations for various downstream analysis tasks.
[0154] In some embodiments of this application, the internal network architecture of the gene structure and global interaction modeling layer can be selected in various ways; one such architecture is described below. The gene structure and global interaction modeling layer may include a static graph structure modeling sublayer and a multi-layer global latent attention sublayer. This sublayer can perform structural dependency modeling and global interaction modeling on the fusion features output by the multi-dimensional fusion embedding layer, thereby enhancing the model's ability to characterize inter-gene regulatory relationships and global expression patterns. The final output is a structurally enhanced gene-level representation, i.e., a general feature representation. The feature modeling module includes a static graph structure modeling sublayer and a multi-layer global latent attention sublayer. These two types of structures work together to achieve joint modeling of structural relationships and global dependencies.
[0155] Based on this, by modeling the fused features through gene structure and global interaction modeling layers, a general feature representation is obtained, which may include: By modeling sub-layers using a static graph structure, a normalized adjacency matrix is obtained based on a pre-constructed gene relationship graph. The normalized adjacency matrix is then used to propagate the fused features after linear mapping through graph structure propagation. Combined with residual connections and normalization processing, an initial representation with enhanced structure is obtained.
[0156] Specifically, based on a pre-constructed gene relationship graph, graph structure propagation modeling is performed on the regulation or correlation between genes to obtain a normalized adjacency matrix. The gene relationship graph can be: , in, A set of gene nodes, Let be the set of edges connecting genes. The weights are the edge weights corresponding to the weighted adjacency matrix W.
[0157] To ensure numerical stability, the edge weights can be processed using absolute values, and the node degree and lower bound can be calculated and truncated. , , , in, For genes With genes The strength of the correlation between them This is a preset non-zero positive number used to perform minimum value constraint processing on node degree.
[0158] By following the steps above, a symmetric normalized adjacency matrix can be constructed: .
[0159] The normalized adjacency matrix is finally obtained: .
[0160] Fusion features By performing a linear mapping, we can obtain the fused features after linear mapping processing:
[0161] in, This is the weight matrix. This is a bias term.
[0162] Using normalized adjacency matrix Fusion features after linear mapping Propagate the graph structure: .
[0163] After performing residual connections and layer normalization, we can obtain the initial representation after structural enhancement: .
[0164] By using multiple global latent attention sublayers, interactive modeling is performed on the initial representation after structural enhancement to obtain a general feature representation.
[0165] Specifically, the initial representation after structural enhancement is processed to obtain the attention output, assuming the first... The initial representation of the layered structure after enhancement is as follows: .
[0166] For the The initial representation after the layer's structural enhancement is projected onto the query to obtain query features: , in, This is the weight matrix.
[0167] Rearranging the query into a multi-head form yields the rearranged query characteristics: , in, M is the batch size, and M is the number of attention heads. For the number of genes, The feature dimension corresponding to each attention head, d represents the total dimension of the model's hidden features.
[0168] After completing query projection and multi-head rearrangement, to improve training stability and avoid feature scales that are too large or too small, the query features can be further normalized using Root Mean Square Normalization (RMSNorm). RMSNorm normalization does not rely on mean centering; it scales the features based solely on their root mean square magnitude. This allows for scale standardization while preserving feature orientation information, thereby enhancing numerical stability and saving computational resources.
[0169] Query characteristics after rearrangement For each feature vector Calculate its root mean square value: , in, For feature vectors The component values in the k-th dimension To prevent division by zero constant.
[0170] Perform normalization and learnable scaling: , in, The k-th dimension feature after normalization These are learnable parameters.
[0171] Thus, the normalized Query feature representation is obtained: .
[0172] Subsequently, a low-rank compression strategy is applied to the key features and value features to reduce the feature dimension while preserving the inter-gene association information, thereby improving the model's computational efficiency. This yields: , in, This is a low-rank latent representation matrix. This is the weight matrix.
[0173] By performing an upprojection reconstruction using the shared low-rank space, we can obtain: , in, The key feature matrix, The eigenvalue matrix, This is the weight matrix.
[0174] The above steps enable key features and value features to share the same low-rank latent representation, thus reducing computational complexity while maintaining expressive power. Furthermore, scaling the dot product attention calculation: , Then, by connecting the linear mapping with the residual, the attention output is obtained: , in, This is the weight matrix.
[0175] Furthermore, the attention output can be nonlinearly enhanced to obtain a general feature representation.
[0176] Specifically, normalize the attention output: , The normalized attention output is nonlinearly transformed by a feedforward neural network module to enhance the model's expressive power.
[0177] right Performing a linear dimension upscaling transformation, mapping the feature dimension d to 4d through a linear layer, yields an intermediate representation: , in, , .
[0178] Dividing the intermediate feature into two sub-feature parts along the last dimension yields: , in, , .
[0179] The Swish activation function is applied to a portion of the features, and this signal is used as a gate signal to modulate another portion of the features, thus forming the Swiglu gated output: , , in, For the activation function, in this embodiment it can be the Sigmoid activation function. This is element-wise multiplication.
[0180] The gated features are dimensionality-reduced using a linear layer to restore the feature dimensions to their original dimensions, and then a general feature representation is obtained through residual connections. : , , Wherein, parameter matrix , .
[0181] Through the above gating mechanism, the model can adaptively control the information flow according to the input features, thereby enhancing important features and suppressing redundant information, improving feature expression ability, improving the model's ability to model complex transcriptome expression patterns, and maintaining high computational efficiency and training stability.
[0182] Multi-head attention interaction modeling is performed on the initial representation after structural enhancement. By query mapping, generating keys and values from low-rank latent representations, scaling dot product attention computation, residual connections and feedforward nonlinear enhancement operations, global dependencies across genes are captured layer by layer.
[0183] The gene structure and global interaction modeling layer in this embodiment, through the combination of a static graph structure modeling sublayer and a multi-layer global potential attention sublayer, can display the structural relationship information between introduced genes, capture global dependency features across genes, reduce computational complexity while maintaining high expression capacity, and improve the stability and generalization performance of the transcriptome basic model on large-scale transcriptome data.
[0184] Furthermore, the final general feature representation can be obtained by stacking the results of static graph structure modeling and multi-layer global latent attention modeling in multiple layers.
[0185] In some embodiments of this application, the process of pre-constructing a gene relationship map may include: S61. Determine a unified set of gene nodes based on a pre-defined gene sequence list.
[0186] Specifically, the process involves acquiring a gene set and its corresponding prior biological relationships. These prior biological relationships include at least one or more of the following: gene regulatory relationships, protein-protein interaction relationships, pathway co-existence relationships, or co-expression correlations. Using each gene in the gene set as a graph node, a node set is constructed and rearranged according to a predefined gene sequence list. The resulting unified gene node set can be: , Where N is the size of the uniform gene set, and its order is strictly consistent with the arrangement order of the gene expression matrix and gene embedding features.
[0187] S62. Traverse any two genes, calculate the correlation between any gene pair, and obtain the original correlation set of gene pairs.
[0188] Specifically, based on gene co-expression networks, the Pearson correlation coefficient between any pair of genes can be calculated to obtain the original set of correlations: , in, For genes With genes The strength of the correlation between them.
[0189] S63. Construct a weighted edge set based on the original set of gene pair correlations.
[0190] Specifically, an edge set is constructed based on the biological relationships between genes. When a preset relationship exists between two genes, a connection edge is established between the corresponding nodes, and edge weights are assigned to the connection edges. The edge weights are determined based on the correlation strength, interaction confidence, regulatory strength, or statistical correlation coefficient between genes.
[0191] To improve the stability and sparsity of the graph structure, correlations are filtered and processed to remove self-loops: For the retained gene pairs, construct a weighted edge set: , in, .
[0192] If the gene pair does not meet the preset threshold condition, such as the threshold being τ, then the corresponding edge is not constructed, thus obtaining a sparse graph structure.
[0193] S64. Construct a weighted adjacency matrix based on a unified set of gene nodes and a set of weighted edges.
[0194] Specifically, the edge weights are standardized by taking the absolute value of the edge weights, calculating the degree value of each node and performing symmetric normalization to obtain a normalized adjacency structure representation. The normalized adjacency structure representation is stored as static graph structure parameters for modeling gene-level features for structure propagation and relationship enhancement during model training and inference.
[0195] Based on the unified set of gene nodes and the set of weighted edges mentioned above, construct a weighted adjacency matrix: .
[0196] The elements of a weighted adjacency matrix are defined as follows: .
[0197] To ensure the preservation of node information, a unit self-loop term can be added to the adjacency matrix, resulting in: , in, It is an identity matrix.
[0198] S65. Using a unified set of gene nodes, a set of weighted edges, and a weighted adjacency matrix, a gene relationship graph is obtained.
[0199] Specifically, using a unified set of gene nodes Weighted edge set and weighted adjacency matrix To obtain a gene relationship map .
[0200] The gene relationship graph is constructed under a unified gene indexing system and can be aligned node-by-node with gene expression matrices and gene-protein value sequence embedding features, serving as a fixed topological foundation for static graph structure modeling. The pre-constructed gene relationship graph built through the above steps transforms external biological prior knowledge into a computable graph structure, enabling the model to utilize known relationships between genes when modeling transcriptome data, thereby enhancing structural expression capabilities and biological consistency.
[0201] In some embodiments of this application, when the number of genes is greater than a preset value, the complexity is... To reduce computational complexity, a block-based latent attention mechanism can be introduced. Based on this, the initial representation after structural enhancement is processed to obtain the attention output, which may include: S71. Divide the gene dimensions according to the preset block size.
[0202] Specifically, let the block size be... Then the gene dimension can be divided into: , in, For the number of genes, The number of blocks after partitioning, if Not divisible If so, then fill in.
[0203] S72. For each block, process the initial representation after structural enhancement to obtain the corresponding attention output.
[0204] Specifically, the j-th block can be represented as: , Global latent attention interaction modeling is performed independently within each block to obtain the attention output corresponding to each block: , S73. At the gene dimension, the attention outputs corresponding to each segment are spliced together to obtain the final attention output.
[0205] Specifically, by splicing the data at the gene level, the final attention output can be obtained: , Through this embodiment, the complexity can be reduced from Reduced to ,in, This structure significantly reduces computational costs while maintaining the ability to model local dependencies, enabling the model to run stably at a scale of approximately 20,000 genes. When N≤C, no block processing is required.
[0206] In some embodiments of this application, the pre-defined analysis task can be a classification task. Based on this, S55, based on the pre-defined analysis task, using general feature representation, the prediction result corresponding to the pre-defined analysis task can be obtained, which may include: S81. At the gene dimension, global average pooling is performed on the general feature representation to obtain the sample-level hidden representation.
[0207] Specifically, the normalized general feature representation can be globally averaged and pooled along the gene dimension: , We obtain the sample-level hidden representation: , in, This is the sample-level hidden representation corresponding to the b-th sample. Let be the normalized hidden representation vector corresponding to the i-th gene in the b-th sample.
[0208] S82. Using sample-level hidden representations, the classification prediction results are obtained through a feedforward network mapping.
[0209] Specifically, the sample-level hidden representation is used to map the input sample classification header to categories: , , Obtain the classification prediction output: , in, For activation function, , This is the weight matrix. , For bias terms, This represents the number of categories.
[0210] The sample classification head can include a first fully connected layer, a non-linear activation layer, and a second fully connected layer connected in sequence, which are used to output the predicted probability distribution of each category and obtain the classification prediction result.
[0211] Furthermore, when the pre-defined analysis task can be a classification task, in order to optimize the model in a targeted manner, a classification loss can be constructed based on the cross-entropy between the classification prediction result and the true classification result label, and the model parameters can be optimized and updated using the classification loss; alternatively, the mask regression loss and the classification loss can be weighted or directly summed to obtain the total model loss, which can be used to jointly optimize the model parameters.
[0212] For classification tasks, the classification loss can be constructed using the following formula based on the cross-entropy between the classification prediction result and the true classification result label: , in, To categorize the true results, This represents the classification prediction results.
[0213] In multi-task training scenarios, the following formula can be used to calculate the mask regression loss. With classification loss By performing weighted summation or direct summation, the total model loss can be obtained. : .
[0214] In the model pre-training stage, this embodiment uses the masked expression value prediction task as the self-supervised training objective. The model parameters are optimized using the masked regression loss function, enabling the model to learn the intrinsic relationships between gene expressions, thereby obtaining a general feature representation with universal representation capabilities. During this stage, other downstream supervised task output heads, such as classification output heads, can be constructed and integrated into the model structure, but they do not participate in loss calculation; their corresponding classification loss weights are set to zero or disabled. When other downstream supervised task output heads, such as classification output heads, are pre-integrated into the model architecture as predefined structural components of the basic large model, the model possesses the ability to directly support downstream classification tasks. In the basic large model pre-training stage, the classification module does not participate in loss calculation; it is only enabled and participates in training in specific downstream classification application tasks, thereby achieving seamless transfer from self-supervised pre-training to downstream supervised tasks.
[0215] The model in this embodiment supports the following three training modes in practical applications: First, only the mask expression value prediction task is enabled. In this mode, the model only calculates the mask regression loss for self-supervised pre-training to learn a general transcriptome representation. Second, only the sample-level classification task is enabled. In this mode, the model only calculates the classification loss for supervised learning scenarios, such as disease classification or sample type identification tasks. Third, both the mask expression value prediction task and the sample-level classification task are enabled. In this mode, the model adopts a multi-task joint training approach, optimizing both the mask regression loss and the classification loss simultaneously, so that the model has both expression reconstruction ability and classification discrimination ability, thereby further improving the generalization and discrimination ability of the embedded representation.
[0216] The following describes the multidimensional embedding fusion transcriptome basic model training device provided in the embodiments of this application. The multidimensional embedding fusion transcriptome basic model training device described below can be referred to in correspondence with the multidimensional embedding fusion transcriptome basic model training method described above.
[0217] Figure 5 This application provides a schematic diagram of a multidimensional embedding fusion transcriptome basic model training device. The multidimensional embedding fusion transcriptome basic model training device may include: Training data provision module 10 is used to provide raw transcriptome data and sample batch source information for biological samples; The raw data processing module 20 is used to process the raw transcriptome data to obtain the gene expression matrix; 30 sets of structured training data were acquired to obtain structured training data and labels of the true expression values of the masked genes based on the gene expression matrix and sample batch source information. The model training module 40 is used to input structured training data into the transcriptome basic model to obtain the predicted expression values of the masked genes. The model update module 50 is used to update the parameters of the transcriptome base model with the training objective of making the predicted expression value of the masked gene approach the true expression value of the masked gene.
[0218] As can be seen from the above embodiments, the multidimensional embedding fusion transcriptome basic model training device provided in this application includes: a training data providing module 10, used to provide raw transcriptome data and sample batch source information of biological samples; a raw data processing module 20, used to process the raw transcriptome data to obtain a gene expression matrix; a structured training data acquisition module 30, used to obtain structured training data and masked gene true expression value labels based on the gene expression matrix and sample batch source information; a model training module 40, used to input the structured training data into the transcriptome basic model to obtain the predicted expression values of the masked genes; and a model updating module 50, used to update the parameters of the transcriptome basic model with the training objective of the predicted expression values of the masked genes approaching the true expression value labels of the masked genes. This application processes raw transcriptome data to obtain structured training data, including a standardized gene expression matrix after masking, mask position index information, gene-level protein sequence embedding feature matrix, and batch identification information. During the training of the transcriptome basic model, the structured training data is processed to obtain fusion features, and based on the fusion features, the predicted expression values of the masked genes are obtained. This achieves unified fusion of multi-source features of transcriptome data, thereby constructing a robust transcriptome basic model with excellent generalization ability.
[0219] Furthermore, the transcriptome basic model can learn the correlation features between gene expression in transcriptome data through a multidimensional embedding mechanism, enhance the ability to model and fill in missing gene expression values, and generate transcriptome feature embeddings with universal expression capabilities, thereby providing unified and reusable feature input support for various downstream bioinformatics analysis tasks.
[0220] Optionally, the basic transcriptome model to be trained includes: a cascaded input layer, a multi-dimensional fusion embedding layer, a gene structure and global interaction modeling layer, and an output layer; The model training module 40 executes the process of inputting structured training data into the transcriptome base model to obtain the predicted expression values of the masked genes, which may include: Structured training data is obtained through the input layer; Through the multi-dimensional fusion embedding layer, fusion features are obtained by utilizing the masked standardized gene expression matrix, gene-level protein sequence embedding feature matrix, and batch identification information in the structured training data; The fusion features are modeled using the gene structure and global interaction modeling layer to obtain a general feature representation; The predicted expression value of the masked gene is obtained through the output layer based on the general feature representation and the mask position index information. Optionally, the structured training data further includes: a masked standardized gene expression matrix, a gene-level protein sequence embedding feature matrix, and batch identification information. The multi-dimensional fusion embedding layer includes: an expression embedding sublayer, a gene embedding sublayer, a sample context embedding sublayer, a batch embedding sublayer, and a feature fusion layer. The model training module 40 executes the process of obtaining fused features through the multi-dimensional fusion embedding layer using the masked standardized gene expression matrix, gene-level protein sequence embedding feature matrix, and batch identification information in the structured training data. This process may include: By using the expression embedding sublayer, the expression embedding matrix is obtained based on the masked normalized gene expression matrix; By transforming the dimension of the gene-level protein sequence embedding feature matrix through the gene embedding sublayer, the gene-level embedding feature matrix is obtained. By using the sample context embedding sublayer, sample context embedding features are generated based on the masked, standardized gene expression matrix. A batch embedding tensor is generated based on batch identification information through a batch embedding sublayer. The expression embedding matrix, the gene-level embedding feature matrix, the sample context embedding feature, and the batch embedding tensor are fused through the feature fusion layer to obtain fused features.
[0221] Optionally, the model training module 40 may perform the process of obtaining the predicted expression value of the masked gene based on the general feature representation and the mask position index information, which may include: The general feature representation is subjected to layer normalization to obtain a gene-level embedding matrix; Based on the gene-level embedding matrix, the corresponding hidden representation is extracted according to the mask position index information; Based on the hidden representation, regression prediction is performed using a fully connected network to obtain the predicted expression values of the masked genes.
[0222] Optionally, the structured training data includes: a standardized gene expression matrix after masking, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information. The structured training data acquisition process 30, based on the gene expression matrix and the sample batch source information, to obtain the structured training data and the true expression value labels of the masked genes, may include: Based on the preset gene annotation information, the gene expression matrix is standardized to obtain standardized gene expression data. Based on a preset gene sequence list, the standardized gene expression data is reordered to obtain a standardized gene expression matrix; Based on a pre-defined list of gene sequences, a corresponding gene-level protein sequence embedding feature matrix is constructed. Based on a preset masking strategy, the expression values of some genes in the standardized gene expression matrix are masked to obtain the masked standardized gene expression matrix, masking position index information, and labels of the true expression values of the masked genes. The source information of the sample batch is encoded to obtain batch identification information.
[0223] Optionally, the preset gene annotation information includes: a unique gene identifier and corresponding gene length information. The process of acquiring structured training data 30 by performing standardized processing on the gene expression matrix based on the preset gene annotation information to obtain standardized gene expression data may include: Based on the unique gene identifier, the gene length information corresponding to each gene in the gene expression matrix is determined; For each gene, based on the corresponding gene length information, the original gene expression value is converted into the number of reads per million transcripts. The expression value of reads per million transcripts was transformed to obtain standardized gene expression data.
[0224] Optionally, the process of acquiring structured training data 30, which involves constructing a corresponding gene-level protein sequence embedding feature matrix based on a preset gene sequence list, may include: Based on a pre-defined gene sequence list, determine the protein sequence information corresponding to each gene; Protein sequence information is input into a pre-trained biological sequence representation model to obtain the corresponding gene-level protein sequence embedding feature vector. The biological sequence representation model is trained using protein sequence information as training samples and the gene-level protein sequence embedding feature matrix vector corresponding to the protein sequence information as training labels. Based on the gene-level protein sequence embedding feature vector, the gene-level protein sequence embedding feature matrix is obtained, and a gene-level protein sequence embedding feature matrix corresponding one-to-one with the standardized gene expression matrix is constructed.
[0225] This application also provides a training device for a multidimensional embedded fusion transcriptome basic model. Figure 6 The hardware structure block diagram of the training device for the multidimensional embedding fusion transcriptome basic model is shown, with reference to Figure 6 The hardware structure of a training device for a multidimensional embedded fusion transcriptome basic model may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4. In this embodiment of the application, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4; Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. Memory 3 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device; The memory stores a program, and the processor can call the program stored in the memory. The program is used to implement the various processing steps in the aforementioned multidimensional embedding fusion transcriptome basic model training method.
[0226] This application embodiment also provides a storage medium that can store a program suitable for processor execution, the program being used to: implement each processing flow in the aforementioned multidimensional embedding fusion transcriptome basic model training method.
[0227] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0228] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined with each other, and the same or similar parts can be referred to each other.
[0229] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for training a multidimensional embedding fusion transcriptome basic model, characterized in that, include: Provide raw transcriptome data and batch origin information for biological samples; The raw transcriptome data is processed to obtain a gene expression matrix; Based on the gene expression matrix and the sample batch source information, structured training data and the true expression value labels of the masked genes are obtained. The structured training data includes: a standardized gene expression matrix after masking, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information. Structured training data is input into the transcriptome base model to be trained to obtain the predicted expression values of the masked genes. The transcriptome base model to be trained is configured to process the structured training data to obtain fusion features and, based on the fusion features, obtain the predicted expression values of the masked genes. The training objective is to make the predicted expression value of the masked gene approach the actual expression value of the masked gene. The parameters of the transcriptome base model to be trained are updated until convergence.
2. The method according to claim 1, characterized in that, The basic transcriptome model to be trained includes: a cascaded input layer, a multi-dimensional fusion embedding layer, a gene structure and global interaction modeling layer, and an output layer. The training process of the transcriptome base model to be trained includes: Structured training data is obtained through the input layer; Through the multi-dimensional fusion embedding layer, fusion features are obtained by utilizing the masked standardized gene expression matrix, gene-level protein sequence embedding feature matrix, and batch identification information in the structured training data; The fusion features are modeled using the gene structure and global interaction modeling layer to obtain a general feature representation; The predicted expression value of the masked gene is obtained through the output layer based on the general feature representation and the mask position index information. The parameters of the transcriptome base model are updated until convergence, with the training objective being that the predicted expression value of the masked gene approaches the actual expression value of the masked gene.
3. The method according to claim 2, characterized in that, The structured training data further includes: a masked standardized gene expression matrix, a gene-level protein sequence embedding feature matrix, and batch identification information. The multi-dimensional fusion embedding layer includes: an expression embedding sublayer, a gene embedding sublayer, a sample context embedding sublayer, a batch embedding sublayer, and a feature fusion layer. Through the multi-dimensional fusion embedding layer, using the masked standardized gene expression matrix, the gene-level protein sequence embedding feature matrix, and the batch identification information in the structured training data, fused features are obtained, including: By using the expression embedding sublayer, the expression embedding matrix is obtained based on the masked normalized gene expression matrix; By transforming the dimension of the gene-level protein sequence embedding feature matrix through the gene embedding sublayer, the gene-level embedding feature matrix is obtained. By using the sample context embedding sublayer, sample context embedding features are generated based on the masked, standardized gene expression matrix. A batch embedding tensor is generated based on batch identification information through a batch embedding sublayer. The expression embedding matrix, the gene-level embedding feature matrix, the sample context embedding feature, and the batch embedding tensor are fused through the feature fusion layer to obtain fused features.
4. The method according to claim 2, characterized in that, The process of obtaining the predicted expression value of the masked gene based on the general feature representation and the mask position index information includes: The general feature representation is subjected to layer normalization to obtain a gene-level embedding matrix; Based on the gene-level embedding matrix, the corresponding hidden representation is extracted according to the mask position index information; Based on the hidden representation, regression prediction is performed using a fully connected network to obtain the predicted expression values of the masked genes.
5. The method according to claim 1, characterized in that, The structured training data includes: a standardized gene expression matrix after masking, mask position index information, a gene-level protein sequence embedding feature matrix, and batch identification information. Based on the gene expression matrix and the sample batch source information, the structured training data and the true expression value labels of the masked genes are obtained, including: Based on the preset gene annotation information, the gene expression matrix is standardized to obtain standardized gene expression data. Based on a preset gene sequence list, the standardized gene expression data is reordered to obtain a standardized gene expression matrix; Based on a pre-defined list of gene sequences, a corresponding gene-level protein sequence embedding feature matrix is constructed. Based on a preset masking strategy, the expression values of some genes in the standardized gene expression matrix are masked to obtain the masked standardized gene expression matrix, masking position index information, and labels of the true expression values of the masked genes. The source information of the sample batch is encoded to obtain batch identification information.
6. The method according to claim 5, characterized in that, The preset gene annotation information includes: a unique gene identifier and corresponding gene length information. Based on the preset gene annotation information, the gene expression matrix is standardized to obtain standardized gene expression data, including: Based on the unique gene identifier, the gene length information corresponding to each gene in the gene expression matrix is determined; For each gene, based on the corresponding gene length information, the original gene expression value is converted into the number of reads per million transcripts. The expression value of reads per million transcripts was transformed to obtain standardized gene expression data.
7. The method according to claim 5, characterized in that, The construction of a corresponding gene-level protein sequence embedding feature matrix based on a preset gene sequence list includes: Based on a pre-defined gene sequence list, determine the protein sequence information corresponding to each gene; Protein sequence information is input into a pre-trained biological sequence representation model to obtain the corresponding gene-level protein sequence embedding feature vector. The biological sequence representation model is trained using protein sequence information as training samples and the gene-level protein sequence embedding feature vector corresponding to the protein sequence information as training labels. Based on the gene-level protein sequence embedding feature vector, the gene-level protein sequence embedding feature matrix is obtained, and a gene-level protein sequence embedding feature matrix corresponding one-to-one with the standardized gene expression matrix is constructed.
8. A training device for a multidimensional embedded fusion transcriptome basic model, characterized in that, include: The training data provision module provides raw transcriptome data and batch source information for biological samples. The raw data processing module is used to process the raw transcriptome data to obtain a gene expression matrix; The structured training data acquisition module is used to obtain structured training data and the true expression value labels of the masked genes based on the gene expression matrix and the sample batch source information. The model training module is used to input structured training data into the transcriptome basic model to obtain the predicted expression values of the masked genes. The transcriptome basic model to be trained is configured to process the structured training data to obtain fusion features and, based on the fusion features, obtain the predicted expression values of the masked genes. The model update module is used to update the parameters of the transcriptome base model until convergence, with the training objective being that the predicted expression value of the masked gene approaches the actual expression value label of the masked gene.
9. A training device for a multidimensional embedded fusion transcriptome basic model, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement each step of the training method for the multidimensional embedding fusion transcriptome basic model as described in any one of claims 1-7.
10. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the training method for the multidimensional embedding fusion transcriptome basic model as described in any one of claims 1-7.