Multi-omics sequence data integration analysis method
Through the preprocessing and fine-tuning of multiomic sequence data, the problem of difficult interaction between DNA, RNA and protein in multiomic data integration is solved, and efficient multiomic data integration and widely applicable biological analysis are achieved.
Patent Information
- Application Number
- CN202510387915.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
AI Technical Summary
Existing multiomics methods are difficult to effectively integrate the deep interaction between DNA, RNA and proteins, lack biological explanatory and stability, and it is difficult to fully reveal the complex relationship between genes, metabolism and phenotypes.
By obtaining multi-omic sequence data samples, preprocessing according to the inverse transcription and inverse translation laws of the central law, training data is constructed, and the network architecture of the general characterization model is constructed for pre-training. The integrated model is used for fine-tuning downstream tasks, and the hidden spatial characterization of the complete DNA sequence is output.
The efficient integration of multiomics data is achieved, which can effectively capture the interaction and regulatory relationship between DNA, RNA and proteins, reduce computing resource consumption, maintain model generalization capabilities, and demonstrate broad applicability in different downstream tasks.
Smart Images

Figure CN120279998A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics computing, and particularly relates to a method for integrated analysis of multi-omics sequence data. Background Art
[0002] In the field of molecular biology, DNA, RNA, and proteins are key macromolecules in life activities, and they interact with each other through the central dogma of biology. Current research has gradually applied deep learning techniques to analyze these macromolecular data, but usually only focuses on single-omics data (such as DNA sequence analysis or protein structure prediction). These methods are difficult to effectively integrate different types of biological macromolecules and fail to fully capture the complex relationships between them. Although existing multi-omics methods integrate different types of data to a certain extent, they do not adequately consider the deep interaction relationships between DNA, RNA, and proteins, and lack biological interpretability and stability.
[0003] Single-omics data models often have difficulty comprehensively revealing the complex relationships among genes, metabolism, and phenotypes. Specifically, DNA, RNA, and proteins each have unique roles in function, and there are significant functional differences between their coding and non-coding regions. DNA sequences not only contain coding regions (i.e., genes), but also regulatory elements and non-coding regions, which affect the expression and function of the final protein by regulating the transcription level and splicing pattern of RNA. In addition, as the role of non-coding RNA is gradually analyzed, RNA not only acts as a carrier for information transfer in the process of the central dogma, but also plays an important role in translation and regulation. In summary, current multi-omics data analysis of DNA, RNA, and proteins is difficult to effectively integrate various types of omics data and is not conducive to capturing the interaction and regulatory relationships between DNA, RNA, and proteins. Summary of the Invention
[0004] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a method for integrated analysis of multi-omics sequence data, which realizes the effective integration of multi-omics data through unified coding.
[0005] The purpose of the present invention can be achieved by the following technical solutions: A method for integrated analysis of multi-omics sequence data, comprising the following steps:
[0006] S1. Obtain multi-omics sequence data samples, and preprocess the multi-omics sequence data samples according to the reverse transcription and reverse translation rules of the central dogma to construct training data;
[0007] S2. Construct the network architecture of a general representation model, and pre-train it using the training data to obtain an integrated model;
[0008] S3. According to the reverse transcription and reverse translation rules of the central dogma, preprocess the multi-omics sequence data to be processed, then input it into the integration model, fine-tune the downstream tasks for the integration model, and output the hidden space representation of the corresponding complete DNA sequence.
[0009] Further, the step S1 specifically performs unified encoding processing on DNA, RNA, and amino acid sequences based on the reverse translation and reverse transcription rules of the central dogma to obtain training data uniformly represented at the nucleotide level.
[0010] Further, the step S1 includes the following steps:
[0011] S11. In the construction of training data, first align the DNA, RNA, and amino acids, perform local or global alignment on the DNA, RNA, and amino acid sequences respectively to find matching segments, and define a scoring matrix M(i,j) to quantify the matching situation between the DNA sequence and the RNA or amino acid sequence.
[0012] S12. Find the maximum scoring path in the scoring matrix through dynamic programming to determine the optimal alignment method of DNA and RNA or amino acid sequences, and define the maximum score as max(M(i,j)) to represent the cumulative score on the best alignment path.
[0013] S13. Use the reverse transcription and reverse translation rules in the central dogma to convert the RNA sequence into a DNA sequence and the amino acid sequence into a DNA code to achieve unified expression of data at the nucleotide level, obtain a nucleotide sequence, and uniformly represent the smallest constituent unit x in the data as x ∈ {A, T, G, C}, where A, T, G, and C are the four bases respectively.
[0014] Further, in the step S11, for DNA and RNA sequences, specifically perform base alignment according to the base pairing principle (such as A-U and G-C).
[0015] For amino acid sequences, specifically perform alignment based on the correspondence between codons and amino acids.
[0016] Further, the network architecture of the general representation model in the step S2 includes a tokenizer and an encoder. The tokenizer is used to encode local and global information in the nucleotide sequence.
[0017] The encoder is used to process DNA fragments in the coding region and non-coding region, and model the coding region and non-coding region respectively.
[0018] Furthermore, the working process of the tokenizer is as follows: a convolutional operation with a sliding window step size of 3 is selected, which is consistent with the three-base structure of codons. This operation enables every 3 bases to be encoded as a whole. Assume the DNA sequence is S DNA =(x1, x2, …, x n ). Define that each sliding window contains 3 bases, and the convolutional output is:
[0019]
[0020] where w j is the convolutional kernel weight. The convolution with a step size of 3 enables the 3-mer features in the coding region to be extracted with each token as a unit, and these tokens are the sequence representations input to the subsequent model.
[0021] Furthermore, the encoder includes an amino acid translator for the coding region and a context decoder for the non-coding region that are processed in parallel in the backbone network. The amino acid translator is used to convert 3-mer tokens into the corresponding amino acids of the protein;
[0022] The context decoder is used to randomly mask some tokens in the input sequence and predict the masked tokens;
[0023] The backbone network is used to fuse the output results of the amino acid translator and the context decoder and output the complete DNA sequence representation.
[0024] Furthermore, the loss function of the amino acid translator is:
[0025]
[0026] where X = (x1, x2, …, x n ) is the input DNA sequence, is the amino acid sequence generated by translation, Y = (y1, y2, …, y m ) is the true amino acid sequence, m is the length of the amino acid sequence, |A| is the size of the amino acid set, is the indicator function, which is 1 when the true amino acid y i is equal to the j-th amino acid and 0 otherwise, is the probability that the model predicts the i-th amino acid as the j-th amino acid;
[0027] The calculation formula of the amino acid translator is:
[0028] H coding = θ backbone (T coding )
[0029] where T codingThe token sequence representing the coding region, H coding The sequence feature after encoding, θ backbone are the backbone network parameters.
[0030] Furthermore, the context decoder is modeled through a Masked Language Model (MLM), and the corresponding loss function is:
[0031]
[0032] where, T noncoding is the token sequence of the non-coding region, and M is the index of the masked token.
[0033] Furthermore, the calculation formula of the backbone network is:
[0034] H = Concat(H coding , H noncoding )
[0035] where, H coding is the coding region feature obtained through the protein translation task, and H noncoding is the non-coding region feature obtained through the masked language model;
[0036] The loss function of the backbone network combines the loss L translation of the protein translation task and the masked language model loss L MLM of the non-coding region, and is defined as follows:
[0037] L = L translation + λ · L MLM
[0038] where, λ is the weight hyperparameter used to balance the two parts of the loss.
[0039] Furthermore, the specific process of fine-tuning the integrated model for the downstream task in step S3 is:
[0040] S31. During the fine-tuning process, keep the pre-trained backbone network parameters θ backbone frozen, and directly use the general features learned in the pre-training stage for the downstream task. The output of the backbone network is expressed as:
[0041] H frozen = f(x; θ backbone )
[0042] where, x is the input DNA sequence, and H frozen is the feature representation generated by the backbone network;
[0043] S32. For each specific downstream task, add task-specific parameter layers and only fine-tune the parameters θ of these layers. tuning , during this process, the backbone network θ backbone remains frozen, and by updating θ tuning to adapt to the downstream task requirements;
[0044] S33. In the fine-tuning of the downstream task, the goal is to minimize the task-specific loss function L downstream , and this loss function only acts on the parameters θ of the fine-tuning layer tuning rather than the backbone network. For the protein structure prediction task, the loss function is defined based on the difference between the prediction and the actual structure; for the phenotype prediction task, the loss function is the cross-entropy loss or the mean squared error loss;
[0045] L downstream = E (x,y)~D [l(g(H frozen ; θ tuning ), y)]
[0046] where D represents the dataset of the downstream task, l represents the loss function, y is the target output, and in each fine-tuning iteration, only the gradient of θ tuning is calculated and the parameter update is performed:
[0047]
[0048] where η is the learning rate;
[0049] S34. Introduce knowledge distillation in the fine-tuning. Use the existing large protein model as the teacher model to pre-generate high-level feature representations of proteins from biological sequences, and use the integrated model as the student model to predict the corresponding protein structures by learning the sequence representations of the teacher model;
[0050] The loss function of knowledge distillation consists of two parts. One part is the task-specific loss L downstream of the fine-tuning layer, and the other part is the knowledge distillation loss L distill , that is, the degree of imitation of the output of the teacher model by the student model. Define the total loss function as:
[0051] L total = L downstream + μL distill
[0052] where μ is the trade-off coefficient used to adjust the ratio of the fine-tuning loss to the distillation loss;
[0053] The knowledge distillation loss L distill is defined by the cross-entropy between the output P student of the student model and the output P teacher of the teacher model:
[0054]
[0055] By minimizing L distill , the performance of the student model is gradually approximated to that of the teacher model.
[0056] Compared with the prior art, the present invention has the following advantages:
[0057] The present invention obtains multiple omics sequence data samples, preprocesses the multiple omics sequence data samples according to the reverse transcription and reverse translation rules of the central dogma, constructs training data; then constructs the network architecture of a general representation model, pre-trains it using the training data to obtain an integrated model; then preprocesses the omics sequence data to be processed according to the reverse transcription and reverse translation rules of the central dogma, and then inputs it into the integrated model, fine-tunes the integrated model for downstream tasks, and outputs the hidden space representation of the corresponding complete DNA sequence. Thus, a unified representation at the nucleotide level is obtained based on the rules of the central dogma, and an integrated model is constructed. Through unified encoding, the efficient integration of multiple omics data is achieved, and the interaction and regulatory relationships between DNA, RNA, and proteins can be effectively captured.
[0058] The present invention converts DNA sequences, RNA sequences, and amino acid sequences into a unified nucleotide sequence according to the reverse transcription and reverse translation rules of the central dogma, converts RNA sequences into DNA sequences, and converts amino acid sequences into DNA codes, effectively realizing the unified expression of data.
[0059] The present invention designs a tokenizer and an encoder in the integrated model for the unified representation at the nucleotide level. On the one hand, it learns the long-range dependencies and global features in coding and non-coding gene sequences. On the other hand, in the coding region, a 3-mer (3-nucleotide) coding method based on codon degeneracy is adopted, and in the non-coding region, a unified modeling is carried out through a long sequence encoder that combines efficient long convolution and self-attention. Through pre-training, a general representation model that can effectively capture long-range interactions in the genome can be established.
[0060] During the downstream task fine-tuning process of the present invention, the pre-trained computational model is fine-tuned by using parameter-efficient fine-tuning and knowledge distillation. Specifically, the backbone network of the model in the pre-training stage is frozen, and only a small number of parameters are adjusted and knowledge is distilled, such as the adaptation layer or the task-specific head network. This method effectively reduces the number of parameters to be optimized during the fine-tuning process, maintains the generalization ability of the pre-trained model, and at the same time greatly reduces the consumption of computing resources, and can show wide applicability in different downstream tasks such as protein structure prediction, gene regulation, and phenotype prediction, and can maintain high efficiency even under limited computing resources. Description of the Drawings
[0061] Figure 1 It is a schematic diagram of the method flow of the present invention;
[0062] Figure 2 It is a process diagram for preprocessing multi-group biological data;
[0063] Figure 3 It is a process diagram for the inference of the integrated model;
[0064] Figure 4 It is a process diagram for the training of the tokenizer in the integrated model;
[0065] Figure 5 It is a process diagram for the pre-training, downstream task fine-tuning and knowledge distillation of the integrated model. Specific embodiments
[0066] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0067] Embodiment
[0068] As Figure 1 shown, a multi-omics sequence data integration and analysis method includes the following steps:
[0069] S1. Obtain multi-omics sequence data samples, and preprocess the multi-omics sequence data samples according to the reverse transcription and reverse translation rules of the central dogma to construct training data;
[0070] S2. Construct the network architecture of the general representation model, and pre-train it with the training data to obtain an integrated model;
[0071] S3. Preprocess the multi-omics sequence data to be processed according to the reverse transcription and reverse translation rules of the central dogma, then input it into the integrated model, perform downstream task fine-tuning on the integrated model, and output the hidden space representation of the corresponding complete DNA sequence.
[0072] This embodiment applies the above solution, and mainly has the following content:
[0073] First, in the process of constructing training data, based on the reverse translation and reverse transcription rules in the central dogma, preprocess DNA, RNA, amino acid and protein multi-species data to achieve unified representation at the nucleotide level, that is, convert DNA sequences, RNA sequences and amino acid sequences into unified nucleotide sequences.
[0074] II. During the pre-training process, use the training data for pre-training to learn the long-range dependencies and global features in coding and non-coding gene sequences; for the coding region, adopt a 3-mer (3-nucleotide) coding method based on codon degeneracy, and for the non-coding region, perform unified modeling through a long sequence encoder that combines efficient long convolution and self-attention. The goal of the pre-training stage is to establish a general representation model, i.e., an integration model, that can effectively capture long-range interactions in the genome.
[0075] On the one hand, based on the unified nucleotide sequence, utilize codon degeneracy to learn a dynamic tokenizer for vectorized encoding; on the other hand, use an encoding network that combines self-attention and long convolution operators to model the coding region and non-coding region respectively. Among them, for the coding region modeling, use the translation and distillation alignment tasks to model the central dogma, and for the non-coding region modeling, adopt the context masking task to learn context information.
[0076] III. During the fine-tuning of downstream tasks, use parameter-efficient fine-tuning and knowledge distillation methods to fine-tune the pre-trained computational model. The specific method is to freeze the backbone network of the model in the pre-training stage and only adjust a small number of parameters, such as the adaptation layer or task-specific head network.
[0077] In the first part of the content, as Figure 2 shown, the main process is as follows:
[0078] Step 1.1: In the construction of training data, first align the DNA, RNA, and amino acid sequences to find the corresponding matching regions. In this embodiment, align the DNA and RNA sequences according to the base pairing rules (such as A-U and G-C); for the amino acid sequence, perform the alignment based on the correspondence between codons and amino acids. Define a scoring matrix M(i,j) to quantify the matching degree between the DNA sequence and the RNA or amino acid sequence.
[0079] Step 1.2: Use the dynamic programming method to find the path with the highest score in the scoring matrix M(i,j) to determine the optimal alignment method between the DNA and RNA (or amino acid) sequences. The maximum score path is denoted as max(M(i,j)), which represents the cumulative score on the best alignment path.
[0080] Step 1.3: According to the reverse transcription and reverse translation rules of the central dogma, convert the RNA sequence into a DNA sequence and the amino acid sequence into a DNA encoding to achieve unified expression at the nucleotide level. All data is finally represented as x ∈ {A, T, G, C}, thus realizing the unified encoding of DNA, RNA, and amino acid sequences.
[0081] In the second part, as Figure 3 and Figure 4As shown in the figure, the present embodiment performs training and inference of a hybrid long sequence model, including the following steps:
[0082] Step 2.1: Using the nucleotide sequence obtained after preprocessing as input, segment the DNA sequence and construct a long convolutional tokenizer for encoding local and global information in the DNA sequence. To more effectively capture the features in the coding region, a convolutional operation with a sliding window step size of 3 is selected, which is consistent with the codon structure. This ensures that every 3 bases are encoded as a whole. Assuming the DNA sequence is S DNA =(d1, d2, …, d n ), it is defined that each sliding window contains 3 bases, and the convolutional output is:
[0083]
[0084] where w j is the convolutional kernel weight. The convolution with a step size of 3 enables the 3-mer features in the coding region to be extracted with each token as a unit. These tokens are the sequence representations input to the subsequent model.
[0085] Step 2.2: On the basis of Step 2.1, construct a backbone network (θ backbone ) that combines long convolution and attention to process DNA fragments in the coding region and non-coding region. An amino acid translation task is designed for the coding region, while only the pre-training method of the masked language model is used in the non-coding region to predict randomly masked tokens.
[0086] (a) Protein translation task in the coding region: To simulate the protein translation task, convert 3-mer tokens into the corresponding amino acids of the protein. This process is similar to a sequence-to-sequence translation model and is defined as:
[0087] H coding =θ backbone (T coding )
[0088] Specifically, given the DNA sequence X = (x1, x2, …, x n ), the amino acid sequence generated by the model is represented as while the true amino acid sequence is Y = (y1, y2, …, y m ). Where m is the length of the amino acid sequence.
[0089]
[0090] where |A| is the size of the amino acid set (20 standard amino acids). is the indicator function. When the true amino acid y iIt is 1 when it is equal to the j-th amino acid, and 0 otherwise. is the probability that the model predicts the i-th amino acid as the j-th amino acid.
[0091] where T coding represents the token sequence of the coding region, and H coding represents the encoded sequence features. To achieve the protein translation task, the model predicts these 3-mer encoded features so that each token corresponds to an amino acid.
[0092] (b) Masked Language Model for Non-coding Region: The coding region and the non-coding region are modeled by a random masked language model, that is, some tokens in the input sequence are randomly masked, and then the model is required to predict the masked tokens. Suppose the token sequence of the non-coding region is T noncoding , and the index of the masked token is M, then its loss function is defined as:
[0093]
[0094] (c) Overall Backbone Architecture: The coding region and the non-coding region are processed in parallel in the backbone (skeleton network) with a hybrid attention mechanism and a long convolution module to obtain a unified sequence representation:
[0095] H = Concat(H coding , H noncoding )
[0096] where H coding is the coding region feature obtained through the protein translation task, and H noncoding is the non-coding region feature obtained through the MLM (Masked Language Model) task. The two are fused in the backbone to form a complete DNA sequence representation.
[0097] The final loss function combines the loss L translation of the protein translation task and the masked language model loss L MLM of the non-coding region, and is defined as follows:
[0098] L = L translation + λ · L MLM
[0099] where λ is the weight hyperparameter used to balance the two parts of the loss. Through the training of this hybrid task, the model can simultaneously capture the translation information of the coding region and the sequence features of the non-coding region, so as to better adapt to the modeling requirements of multi-omics data in downstream tasks.
[0100] Finally, as Figure 5As shown, in this embodiment, the integrated model for fine-tuning and distillation for downstream tasks includes the following steps:
[0101] Step 3.1, during the fine-tuning process, keep the parameters θ of the pre-trained backbone network backbone frozen, so that the general features learned in the pre-training stage can be directly used for downstream tasks without having to be re-trained in the fine-tuning stage, saving a large amount of computing resources. The output of the backbone network is expressed as:
[0102] H frozen = f(x; θ backbone )
[0103] where x is the input DNA sequence and H frozen is the feature representation generated by the backbone network.
[0104] Step 3.2, for each specific downstream task (such as protein structure prediction or phenotype prediction), add task-specific parameter layers (such as an adaptation layer or a linear projection layer), and only fine-tune the parameters θ tuning of these layers. During this process, the backbone network θ backbone remains frozen, and the model updates θ tuning to adapt to the requirements of the downstream task.
[0105] Step 3.3, in the fine-tuning of the downstream task, the goal is to minimize the task-specific loss function L downstream :
[0106] L downstream = E (x,y)~D [l(g(H frozen ; θ tuning ), y)]
[0107] where D represents the dataset of the downstream task, l represents the loss function, and y is the target output.
[0108] This loss function only acts on the parameters θ tuning of the fine-tuning layer rather than the backbone network. For the protein structure prediction task, the loss function can be defined based on the difference between the prediction and the actual structure. For the phenotype prediction task, the loss function can be the cross-entropy loss or the mean squared error loss.
[0109] In each fine-tuning iteration, only calculate the gradient for θ tuning and update the parameters:
[0110]
[0111] where η is the learning rate. Since the backbone network θ backbone is frozen, the update only affects the parameters of the fine-tuning layer, significantly reducing the computational cost while maintaining the stability of the pre-trained features.
[0112] Step 3.4. Introduce knowledge distillation in fine-tuning. In this embodiment, the ESM protein large model serves as the teacher model, which pre-generates high-level feature representations of proteins from biological sequences. The integration model serves as the student model to learn the sequence representations of the teacher model to predict their corresponding protein structures. The loss function of knowledge distillation consists of two parts. One part is the task-specific loss L downstream of the fine-tuning layer, and the other part is the knowledge distillation loss L distill , that is, the degree to which the student model imitates the output of the teacher model. Define the total loss function as:
[0113] L total = L downstream + μL distill
[0114] where μ is a trade-off coefficient used to adjust the ratio of the fine-tuning loss to the distillation loss.
[0115] The knowledge distillation loss L distill can be defined by the cross-entropy between the output P student of the student model and the output P teacher of the teacher model:
[0116]
[0117] By minimizing L distill , the student model gradually approaches the performance of the teacher model, so as to achieve high-quality prediction results even under limited resources.
[0118] In summary, in view of the deficiencies of existing multi-omics methods in insufficient consideration of the deep interaction relationships among DNA, RNA, and proteins and lack of biological interpretability, this solution proposes an efficient integration solution for multi-omics data through unified encoding. First, a unified representation at the nucleotide level is obtained based on the central dogma rule, and then a mixed convolutional and attention structure is used to process the coding region and non-coding region respectively based on this representation to construct a general representation model that can effectively capture long-range interactions in the genome. And this model uses a parameter-efficient fine-tuning strategy. By freezing the pre-trained backbone network, only a small number of parameter adjustments and knowledge distillation are performed for downstream tasks to meet the requirements of different biological tasks. This solution has significant advantages in aspects such as multi-omics data integration modeling and sequence long-range interaction capture, and can provide a more comprehensive and efficient computational tool for biological research and applications.
Claims
1. A method for integrated analysis of multi-omics sequence data, characterized in that, It includes the following steps: S1. Obtain multiple omics sequence data samples, preprocess the multiple omics sequence data samples according to the reverse transcription and reverse translation rules of the central dogma, and construct training data; S2. Construct the network architecture of the general representation model, and perform pre-training using the training data to obtain an integrated model; S3. Preprocess the multi-omics sequence data to be processed according to the reverse transcription and reverse translation rules of the central dogma, then input it into the integrated model, fine-tune the downstream tasks for the integrated model, and output the hidden space representation of the corresponding complete DNA sequence.
2. The multi-omics sequence data integration and analysis method according to claim 1, wherein The specific step S1 is to perform unified encoding processing on DNA, RNA, and amino acid sequences based on the reverse translation and reverse transcription rules of the central dogma to obtain training data that is uniformly represented at the nucleotide level; The step S1 includes the following steps: S11. In the construction of training data, first align DNA, RNA, and amino acids, perform local or global alignment on DNA, RNA, and amino acid sequences respectively to find matching fragments, and define a scoring matrix M(i,j) to quantify the matching situation between the DNA sequence and the RNA or amino acid sequence; S12. Find the maximum scoring path in the scoring matrix through dynamic programming to determine the optimal DNA and RNA or amino acid sequence alignment method, and define the maximum score as max(M(i,j)), which is used to represent the cumulative score on the best alignment path; S13. Use the reverse transcription and reverse translation rules in the central dogma to convert the RNA sequence into a DNA sequence and the amino acid sequence into a DNA code to achieve unified expression of data at the nucleotide level, obtain a nucleotide sequence, and uniformly represent the smallest constituent unit x in the data as x∈{A,T,G,C}, where A, T, G, and C are the four bases respectively.
3. The multi-omics sequence data integration and analysis method according to claim 2, characterized in that In the step S11, for DNA and RNA sequences, specifically, base alignment is performed according to the base pairing principle (such as A-U and G-C); For amino acid sequences, specifically, alignment is performed based on the correspondence between codons and amino acids.
4. A multi-omics sequence data integration and analysis method according to claim 2, characterized in that In the step S2, the network architecture of the general representation model includes a tokenizer and an encoder, and the tokenizer is used to encode local and global information in the nucleotide sequence; The encoder is used to process DNA fragments in the coding region and non-coding region, and model the coding region and non-coding region respectively.
5. A multi-omics sequence data integration and analysis method according to claim 4, characterized in that The working process of the tokenizer is as follows: Select a convolution operation with a sliding window step size of 3, which is consistent with the three-base structure of codons. This operation enables every 3 bases to be encoded as a whole. Assume the DNA sequence is S DNA =(x1, x2, …, x n ). Define that each sliding window contains 3 bases, and the convolution output is: where, w j is the convolutional kernel weight. The convolution with a stride of 3 enables the 3-mer features in the encoding region to be extracted with each token as a unit, and these tokens are the sequence representations input to the subsequent model.
6. The multi-omics sequence data integration and analysis method according to claim 5, wherein, The encoder includes a coding region amino acid translator and a non-coding region context decoder that are processed in parallel in the backbone network. The amino acid translator is used to convert 3-mer tokens into the corresponding amino acids of the protein; The context decoder is used to randomly mask some tokens in the input sequence and predict the masked tokens; The backbone network is used to fuse the output results of the amino acid translator and the context decoder and output the complete DNA sequence representation.
7. A multi-omics sequence data integration and analysis method according to claim 6, characterized in that, The loss function of the amino acid translator is: where X = (x1, x2, …, x n ) is the input DNA sequence, is the amino acid sequence generated by translation, Y = (y1, y2, …, y m ) is the true amino acid sequence, m is the length of the amino acid sequence, |A| is the size of the amino acid set, is the indicator function, which is 1 when the true amino acid y i equals the j-th amino acid and 0 otherwise, is the probability that the model predicts the i-th amino acid as the j-th amino acid; The calculation formula of the amino acid translator is: H coding = θ backbone (T coding ) Among them, T coding represents the token sequence in the coding region, H coding represents the sequence features after coding, and θ backbone is the backbone network parameter.
8. A multi-omics sequence data integration and analysis method according to claim 7, characterized in that The context decoder is modeled through the masked language model MLM, and the corresponding loss function is: Among them, T noncoding is the token sequence of the non-coding region, and M is the index of the masked token.
9. The multi-omics sequence data integration and analysis method according to claim 8, wherein The calculation formula of the backbone network is: H = Concat(H coding , H noncoding ) Among them, H coding is the coding region feature obtained through the protein translation task, and H noncoding is the non-coding region feature obtained through the masked language model; The loss function of the backbone network combines the loss L of the protein translation task translation and the masked language model loss L of the non-coding region MLM , and is defined as follows: L = L translation + λ·L MLM Among them, λ is a weight hyperparameter used to balance the losses of the two parts.
10. A multi-omics sequence data integration and analysis method according to claim 9, wherein The specific process of fine-tuning the integrated model for downstream tasks in step S3 is as follows: S31. During the fine-tuning process, keep the parameters θ of the pre-trained backbone network backbone frozen, and directly use the general features learned in the pre-training stage for downstream tasks. The output of the backbone network is expressed as: H frozen = f(x; θ backbone ) where x is the input DNA sequence, and H frozen is the feature representation generated by the backbone network; S32. For each specific downstream task, add task-specific parameter layers and only fine-tune the parameters θ of these layers. tuning , during this process, the backbone network θ backbone remains frozen, and by updating θ tuning to adapt to the requirements of the downstream task; S33. In the fine-tuning of downstream tasks, the goal is to minimize the task-specific loss function L downstream , which only acts on the parameters θ of the fine-tuning layer tuning rather than the backbone network. For the protein structure prediction task, the loss function is defined based on the difference between the prediction and the actual structure; for the phenotype prediction task, the loss function is the cross-entropy loss or the mean squared error loss L downstream = E (x,y)~D [l(g(H frozen ; θ tuning ), y)] Among them, D represents the dataset of the downstream task, l represents the loss function, y is the target output. In each fine-tuning iteration, only calculate the gradient of θ tuning and update the parameters: where η is the learning rate; S34. Introduce knowledge distillation in the fine-tuning process. Use the existing large protein model as the teacher model to pre-generate high-level feature representations of proteins from biological sequences. Use the integrated model as the student model, and let the student model learn the sequence representations of the teacher model to predict the corresponding protein structures; The loss function of knowledge distillation consists of two parts. One part is the task-specific loss $L$ of the fine-tuning layer downstream , and the other part is the knowledge distillation loss $L$ distill , that is, the degree to which the student model imitates the output of the teacher model. The total loss function is defined as: L total = L downstream + μL distill where μ is the trade-off coefficient used to adjust the ratio of the fine-tuning loss to the distillation loss; Knowledge distillation loss L distill Defined by the cross entropy between the output P of the student model student and the output P of the teacher model teacher as follows: By minimizing L distill , the student model is gradually made to approximate the performance of the teacher model.