Multi-omics sequence data integration analysis method based on artificial intelligence

Through the multi-omic sequence data integration analysis method based on artificial intelligence, the problem of the intrinsic correlation in the existing technology cannot be handled, and the accurate relationship revealing between genes, metabolism, protein structure and phenotype and the integrated analysis of multi-omic data is achieved, which is suitable for multi-omic and cell-level biological prediction tasks.

CN120280003APending Publication Date: 2025-07-08WESTLAKE UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510387917.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to accurately reveal the complex relationship between genes, metabolism, protein structure and phenotype, and cannot effectively deal with the complex intrinsic correlations of multiple biological macromolecules across multiomics.

Method used

Using the multi-omic sequence data integration analysis method based on artificial intelligence, we use multi-omic sequence data samples to preprocess, construct training data and test data, construct a network architecture of the multi-omic computation model, and use training and test fine-tuning to obtain a Uni-Life calculation model to achieve prediction of nucleotide sequence, gene labeling information and protein structure.

Benefits of technology

It can accurately reveal the complex relationship between genes, metabolism, protein structure and phenotype, realize single-omic functional prediction and integrated analysis of multi-group biological data interaction, is compatible with multi-omic biological data modeling, and uniformly model multi-omic data into three levels of nucleotide, gene and function, predict the functional and structural information of single-omics, and model the complex interactive relationship of multi-omics data in specific scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120280003A_ABST
    Figure CN120280003A_ABST
Patent Text Reader

Abstract

The invention relates to an artificial intelligence-based multi-omics sequence data integration analysis method, which comprises the following steps of: acquiring multi-omics data samples, preprocessing the multi-omics data samples, and constructing training data and test data; constructing a network architecture of a multi-omics calculation model, and performing training and test fine tuning on the multi-omics calculation model by using the training data and the test data to obtain a Unii-Life calculation model; the method comprises the following steps: preprocessing to-be-processed multi-omics data, then inputting the to-be-processed multi-omics data into a Unii-Life calculation model, and outputting to obtain a nucleotide sequence, gene labeling information, a protein structure and species category information. Compared with the prior art, the method has the advantages that the multi-omics data is uniformly modeled into three levels of nucleotide, gene and function, so that the function, structure and other information of the single omics can be predicted, the complex relationship among the gene, metabolism, protein structure and phenotype can be accurately revealed, and the integrated analysis of single omics function prediction and multi-group physical data interaction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computational biology, and particularly to an integrated analysis method for multi-omics sequence data based on artificial intelligence. Background Art

[0002] In molecular biology, omics mainly includes genomics, proteomics, metabolomics, transcriptomics, lipidomics, immunomics, glycomics, and RNA-omics, etc. And multi-omics refers to jointly explaining scientific problems from the levels of species, genes, and metabolites through two or more omics research methods to better understand the disease lesion process and the metabolic pathways of substances in the body.

[0003] The goal of molecular biology is to study and analyze biological macromolecules such as DNA, RNA, amino acids, and proteins that affect life processes. In recent years, artificial intelligence technology has gradually been applied to model the complex internal laws of macromolecules. Existing deep learning technologies usually focus on single-omics data or the mutual relationship between two omics, which enables them to accurately model some relatively definite biological macromolecule conversion relationships. A successful case is the protein structure prediction model represented by the AlphaFold series. They are based on the basic law that the amino acid sequence determines the protein spatial structure, and train a protein calculation model that can accurately predict the protein structure from a large-scale labeled database. However, existing models and biological research cannot effectively process multiple biological macromolecules across multiple omics, and their complex internal correlations cannot be accurately captured by a single model or experiment, resulting in the difficulty of existing single-omics calculation models in accurately revealing the complex relationships among genes, metabolism, protein structure, and phenotype. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide an integrated analysis method for multi-omics sequence data based on artificial intelligence, which can accurately reveal the complex relationships among genes, metabolism, protein structure, and phenotype, and realize the integrated analysis of single-omics function prediction and multi-omics biological data interaction.

[0005] The purpose of the present invention can be achieved through the following technical solutions: An integrated analysis method for multi-omics sequence data based on artificial intelligence, comprising the following steps:

[0006] S1. Obtain multi-omics data samples, preprocess the multi-omics biological data samples, and construct training data and test data;

[0007] S2. Construct the network architecture of the multi-omics calculation model, and use the training data and test data to train and test and fine-tune the multi-omics calculation model to obtain the Uni-Life calculation model;

[0008] S3. Preprocess the multi-omics data to be processed, then input it into the Uni-Life computational model, and output the nucleotide sequence, gene annotation information, protein structure, and species category information.

[0009] Furthermore, the preprocessing of the multi-omics biological data samples in step S1 is specifically to convert the RNA sequence and the amino acid sequence into a unified DNA sequence according to the reverse transcription and reverse translation rules of the central dogma;

[0010] Then, based on the unified DNA sequence, the coding genes, non-coding RNA regions, regulatory elements and gene variation information are hierarchically matched and annotated to construct gene annotation data and generate training labels.

[0011] Furthermore, constructing training data and test data in step S1 includes the following steps:

[0012] S11. When constructing training data, DNA, RNA and amino acid sequences are matched according to gene indexes. The nucleotide sequences of 2000 bp before the transcription start site and after the terminator are intercepted with the gene fragment as the center. The transcription and translation rules of the central dogma are used to obtain the RNA and amino acid sequences corresponding to the DNA. The minimum unit of data encoding is unified into {A, T, G, C, N}, where ATGC and N represent 4 bases and placeholders, respectively.

[0013] S12. When integrating the labels of the training data, query the cis-regulatory elements (such as promoters, enhancers and terminators), non-coding RNAs (such as miRNAs and IncRNAs, etc.) and gene features (such as exons and introns) in the NCBI or UCSC database according to the gene index, align all the annotations to the nucleotide sequence to obtain an annotation sequence, and at the same time, search the NCBI database for the species category corresponding to the nucleotide sequence, use the annotation sequence and species category as supervised learning labels, and finally obtain each training data in the form of a 4-tuple: {nucleotide, gene annotation, amino acid, species category};

[0014] S13. When constructing test data, given any DNA, RNA, or amino acid sequence arbitrarily, first align the data with the DNA sequence as the template. It is necessary to perform local or global alignment on the RNA and amino acid sequences respectively to locate the matching fragments. By defining a scoring matrix M(i,j) to quantify the matching situation between the DNA sequence and the RNA or amino acid sequence, use the dynamic programming algorithm implemented by the sequence matching tool to find the maximum scoring path in the scoring matrix, so as to determine the optimal DNA and RNA (or amino acid) sequence alignment method. Then, using the reverse transcription and reverse translation rules of the central dogma, convert the RNA sequence into a DNA sequence, and convert the amino acid sequence into complementary DNA (cDNA) encoding according to species preferences to achieve the unified expression of multi-omics data.

[0015] Further, the network architecture of the multi-omics calculation model in step S2 includes a first layer, a second layer, and a third layer. The first layer is used to perform unified encoding and context modeling on the multi-omics sequences expressed by the DNA sequence, and reconstruct the gene coding sequence into the original DNA sequence S DNA ;

[0016] The second layer is based on gene-level embedding, models the long-range interaction relationship between the coding region and non-coding region sequence representations, and predicts the gene annotation information;

[0017] The third layer is used to model the central dogma protein translation and folding processes of the coding sequence, and predict the species category information of the amino acid sequence, protein structure, and nucleotide sequence.

[0018] Further, the first layer includes a dynamic tokenizer, a long sequence encoder, and a decoder. The dynamic tokenizer includes a sequence modeling network and a dynamic tokenization module. The sequence modeling network adopts a perceptron architecture combining long convolution and cross-attention to encode important local gene features and global information in the nucleotide sequence; the dynamic tokenization module divides the sequence fragments through a feature merging algorithm and merges them into the smallest unit;

[0019] The long sequence encoder adopts an encoder network ε combining hybrid long convolution and self-attention ψ , which is used to model the context information of the multi-omics sequence and obtain the gene-level representation sequence;

[0020] The decoder is used to reconstruct the gene-level representation sequence into the original DNA sequence S DNA .

[0021] Furthermore, the second layer includes a gene embedding encoder and a gene annotation decoder. The gene embedding encoder is used to model the long-range interaction relationship between coding region and non-coding region sequence representations, and the gene annotation decoder is used to predict the cis-regulatory elements and gene features of each functional segment on the nucleotide.

[0022] Furthermore, the third layer includes a protein structure decoder and a species classifier. The protein structure decoder is used to extract a representation sequence corresponding to the length of the amino acid sequence from the gene annotation information, and align the latent space features to an existing large protein model through knowledge distillation to predict the protein structure;

[0023] The species classifier introduces a global classification term to extract the global classification features of gene-level representations, and performs species classification through a non-linear classifier to output the species category probability.

[0024] Furthermore, the training of the multi-omics computational model in step S2 includes four tasks: DNA sequence and gene embedding sequence reconstruction task, gene annotation task, protein structure decoding task, and nucleotide sequence species classification task.

[0025] Furthermore, the specific process of the DNA sequence and gene embedding sequence reconstruction task is as follows: for the dynamic tokenizer and the decoder, referring to the training method of the BERT masked language model, randomly mask some bases in the input sequence with a probability of 15% to obtain the sequence M i where \(M_i\) represents whether the \(i\)-th site is masked, \(B(.)\) represents the Bernoulli distribution, and then the decoder is required to predict the masked bases. The loss function is defined as:

[0026]

[0027] where \(p(.)\) represents the probability predicted at the \(i\)-th masked site;

[0028] Similarly, a random masked sequence of gene representations can be input to the encoder and the gene-level representation context is reconstructed, and its prediction target is the dynamic vocabulary of each DNA sequence The form of the loss function is the same as above;

[0029] The specific process of the gene annotation task is as follows: The gene annotation decoder consists of multiple cross-attention modules and query words. Based on the gene-level representation sequence \(Z\) modeled by the gene embedding encoder gene , for each gene annotation category \(g\), the functional segment is segmented and decoded into where \(A_t\) t ∈{0,1} represents whether the \(t\)-th site of the DNA sequence belongs to the annotation category There are a total of G annotation categories, and the decoding calculation process is expressed as:

[0030] A c = D gene (CrossAttention(Q c , P DNA ))

[0031] Among them, the cross-attention module CrossAttention(.) takes Z gene as the key-value pair, and the learnable annotation category query word Q c as the query key. Through the segmentation prediction head D gene and the built-in Sigmoid function, a binary mask sequence output is realized. After calculating G times, the complete gene annotation result of the DNA sequence is obtained Combined with the gene annotation label Y c , the supervised segmentation loss function is defined as:

[0032]

[0033] Among them, A c,i is the mask of the i-th token in the DNA sequence output by the functional fragment segmentation decoding, and Y c,i is the label of the i-th token in the DNA sequence for the gene annotation label;

[0034] The specific process of the protein structure decoding task is as follows: Based on the exon fragments of gene annotation, on the one hand, the corresponding amino acid sequence is obtained during the construction of training data On the other hand, the gene-level representation is intercepted by the feature fusion module to obtain a representation sequence of the same length as the amino acid The protein structure decoder implemented by using the cross-attention module and the structure prediction head D protein will predict the amino acid sequence and its corresponding protein 3D structure according to Z AA This process is expressed as: The process is expressed as:

[0035] H protein = D protein (CrossAttention(Z eA , Z gene ))

[0036] Among them, the amino acid representation sequence Z AA is used as the query key for cross-attention, and the gene-level context representation Z gene is used as the key-value. Similarly, after the cross-attention module, the amino acid prediction head D AA is used to realize the decoding of the amino acid representation sequence to the amino acid sequence to constrain the amino acid representation sequence extracted by the model;

[0037] Using the pre-trained protein model h δ :S AA →H protein As the teacher model and using the protein structure decoder as the student model, knowledge distillation is performed on the protein structure decoder, and the sequence-level knowledge distillation loss L dist Align the output of the student model with the output of the teacher model, and use cross-entropy to define:

[0038]

[0039] where H teacher and H student respectively represent the protein structures predicted by the teacher and student models. At the same time, the translation loss is used to optimize the translation task of the amino acid sequence:

[0040]

[0041] where, is an indicator function, which is 1 when the true amino acid S i is equal to the j-th amino acid, and 0 otherwise, is the probability that the i-th site of the model-predicted Z AA sequence is the j-th amino acid. The distillation and translation losses complement each other to model the translation and folding processes from genes to amino acids and proteins;

[0042] The specific process of the nucleotide sequence species classification task is as follows: A global classification term is introduced into the encoder to extract the global classification features of gene-level representations. The classification term aggregates the species category information through the long convolution and self-attention mechanisms of the encoder, and passes through the non-linear classifier D cls for species classification, and the output species category probability is p(P cls ), and the cross-entropy classification loss L cls is used for training:

[0043] L cls =-Y S logp(P cls )

[0044] where, is the species-level category annotation;

[0045] The training process uses logarithmic weighting for the loss function to achieve multi-task balance, ensuring that multi-tasks with different difficulties have similar convergence speeds:

[0046] L total =log(L MLM )+log(L gene) + log(L trans ) + log(L dist ) + log(L cls )

[0047] where L total is the total loss function.

[0048] Furthermore, the specific process of testing and fine-tuning in step S2 is as follows:

[0049] First, in the fine-tuning stage, keep the parameters of the tokenizer, encoder, and decoder in the pre-trained computational model frozen to extract the general gene-level representation sequence Z gene , and introduce a low-rank fine-tuning branch and a downstream task decoder D down to the encoder. The parameters of the low-rank branch are ε LoRA ;

[0050] Then, considering two different implementation methods in specific scenarios, in the first case, for single-omics downstream tasks not covered in the pre-training stage and downstream tasks of splicing multiple omics data, convert the test data into nucleotide form and splice it together and input it into the computational model, X down = {S DNA , …}, and use the loss function L down of the downstream task and the label Y down to fine-tune the low-rank branch and the multi-layer perceptron (MLP) decoder D finetune ;

[0051] In the second case, considering the downstream task of the interaction between quantitative sequencing data and multi-omics sequence data, regard sequencing and quantitative prediction as results and multi-omics data as conditions, and introduce a cross-attention mechanism into the decoder D down . Use the feature encoding of the sequencing data X seq as the query key, and use the representation obtained by the multi-omics sequence X down through the encoder as the key value to achieve precise quantitative modeling in specific downstream tasks.

[0052] Compared with the prior art, the present invention has the following advantages:

[0053] The present invention first obtains multi-omics data samples, and preprocesses the multi-omics biological data samples to construct training data and test data; then constructs the network architecture of the multi-omics computational model, and uses the training data and test data to train and fine-tune the multi-omics computational model to obtain the Uni-Life computational model, which is used to reconstruct nucleotide sequences, predict gene annotation information, predict protein structures, and species category information. Thus, it can process biological data compatible with multi-omics, uniformly model multi-omics data into three levels of nucleotides, genes, and functions, and can not only predict information such as the functions and structures of single omics, but also model the complex interaction relationships of multi-omics data in specific scenarios.

[0054] The present invention converts DNA, RNA, and amino acids into a unified nucleotide representation based on the central dogma. Specifically, according to the reverse transcription and reverse translation rules of the central dogma, the RNA sequence and amino acid sequence are converted into a unified DNA sequence, and then based on the unified DNA sequence, information such as coding genes, non-coding RNA regions, regulatory elements, and gene variations is hierarchically matched and annotated to construct gene annotation data and generate training labels, which can achieve the unified expression of multi-omics data and ensure the reliability of the training data and test data.

[0055] The present invention implements a dynamic tokenizer on a large-scale DNA sequence to extract gene-level embedded codes, and realizes gene-level characterization modeling through an encoder that combines long convolutional and global attention modules, constructing the complex relationships between gene fragments and regulatory elements in coding and non-coding regions; on this basis, the central dogma modeling of nucleotides, genes, and functions at three levels is respectively realized, and the decoding and prediction of each layer of information are realized by corresponding decoders. Among them, in the first layer, the nucleotide decoder reconstructs the gene-level characterization into a nucleotide sequence, in the second layer, the gene annotation decoder predicts fine-grained gene annotation information, including the precise positions and functions of coding regions, non-coding regions, and regulatory elements; in the third layer of function decoding, the protein structure decoder predicts the protein structure of the coding sequence, and the species classifier predicts the species category of the nucleotide sequence. Thus, the modeling and prediction of biological information are respectively realized at the nucleotide, gene, and function levels, and the integrated analysis of single-omics function prediction and multi-omics biological data interaction is realized.

[0056] The present invention provides a parameter-efficient fine-tuning strategy for solving the requirements of multi-omics tasks in specific scenarios in biology, and performs parameter-efficient fine-tuning for multi-omics downstream tasks containing multiple types of macromolecular data; for quantitative sequencing tasks, an interface for fusing cell-level quantitative sequencing data features is provided to fuse the gene-level characterization sequences modeled by Uni-Life with the quantitative features of specific tasks, which can efficiently and accurately solve cell-level biological prediction tasks. Thus, cross-omics tasks that take into account the characteristics of specific application scenarios are realized, which can not only effectively maintain the generality of the pre-trained model and the efficiency of fine-tuning, but also efficiently model multi-omics or cell environment features to accurately complete multi-omics tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 is a schematic flow chart of the method of the present invention;

[0058] Figure 2 is a flow chart of multi-omics biological data preprocessing and integration;

[0059] Figure 3 is a schematic diagram of hierarchical gene annotation integration;

[0060] Figure 4 is a flow chart of single-omics task inference and multi-omics task fine-tuning of the Uni-Life computational model;

[0061] Figure 5 is a flow chart of pre-training and multi-omics task training of the Uni-Life computational model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0063] Embodiment

[0064] As the only rule that qualitatively describes the conversion rules between key macromolecules such as DNA, RNA, amino acids, and proteins in a general sense, the central dogma of biology has not been accurately modeled by existing bio-computing models. Among them, DNA, as the biological genetic material at the beginning of the central dogma, contains the core information that determines the expression and structure of biological macromolecules such as RNA, amino acids, and proteins. Therefore, whether it is possible to construct the processes of the central dogma such as transcription, translation, and folding starting from DNA, and to realize the prediction of important information such as gene annotation, regulatory element recognition, and protein structure for a given DNA sequence will play an important role in molecular biology and cross-omics biological research. Specifically, the central dogma can be modeled separately in the coding region and non-coding region of DNA, because the coding region contains the complete rules of transcription, translation, and protein folding of the central dogma, while the non-coding region usually only plays a regulatory role in the coding region. Taking the genes in the coding region as units, the DNA sequence can be divided into multiple gene sequences according to the transcription start site. Each gene contains gene features such as exons and introns. The exon sequence will become an amino acid sequence through template transcription, splicing, and translation, and fold into a protein with a spatial structure. At the same time, the cis-regulatory elements contained in the non-coding region will regulate the processes of the central dogma such as the transcription level and splicing mode of genes, and ultimately affect the expression and function of proteins. Therefore, this requires that the bio-computing model can adopt a unified and efficient coding method to accurately capture the conversion and regulatory relationships between DNA, RNA, amino acids, and proteins according to the central dogma. Therefore, this solution proposes an artificial intelligence-based integrated analysis method for multi-omics sequence data, such as Figure 1 shown, including the following steps:

[0065] S1. Obtain multi-omics data samples, preprocess the multi-omics biological data samples, and construct training data and test data;

[0066] S2. Construct the network architecture of the multi-omics computing model, and use the training data and test data to train and test and fine-tune the multi-omics computing model to obtain the Uni-Life computing model;

[0067] S3. Preprocess the multi-omics data to be processed, and then input it into the Uni-Life computing model to output nucleotide sequences, gene annotation information, protein structures, and species category information.

[0068] This embodiment applies the above solution, which is mainly divided into the following parts:

[0069] 1. After reading the multi-omics data to be processed, convert the RNA sequence and amino acid sequence into a unified DNA sequence according to the reverse transcription and reverse translation rules of the central dogma;

[0070] Hierarchically match and annotate information including coding genes, non-coding RNA regions, regulatory elements, and gene mutations based on a unified DNA sequence, construct gene annotation data, and generate training labels;

[0071] Second, implement a dynamic tokenizer on large-scale DNA sequences to extract gene-level embedded encodings, and build a gene-level representation model through an encoder that combines long convolutional and global attention modules to construct the complex relationships between gene fragments and regulatory elements in coding and non-coding regions;

[0072] On this basis, implement the central dogma modeling at three levels: nucleotide, gene, and function, and each level is decoded hierarchically by a specific decoder. First, the nucleotide decoder reconstructs the gene-level representation into a nucleotide sequence, then the gene annotation decoder predicts fine-grained gene annotation information, then the protein structure decoder predicts the protein structure of the coding sequence, and finally the species classifier predicts the species category of the nucleotide sequence;

[0073] Third, combine the quantitative sequencing data in downstream tasks and design an efficient fine-tuning technique to generalize the computational model to cell-level downstream tasks.

[0074] In the first part of the content, mainly through gene indexing, information such as regulatory elements, non-coding RNA regions, gene function annotations, and species categories is aligned to the DNA sequence of the same gene, and the nucleotide sequence containing upstream and downstream cis-regulatory elements, transcription start sites, and gene fragments is intercepted as the unified input form of a data sample. The amino acid sequence corresponding to the coding sequence is generated using the translation rules of the central dogma. When constructing test data, this solution converts DNA, RNA, and amino acid sequences into nucleotide forms according to the reverse transcription and reverse translation rules of the central dogma as input data.

[0075] Specifically, as Figure 1 shown, considering two cases of training and testing, the multi-omics data preprocessing includes the following steps:

[0076] Step (1): When constructing training data, match DNA, RNA, and amino acid sequences according to gene indexing, intercept the nucleotide sequences within a range of 2000 bp before the transcription start site and after the terminator with the gene fragment as the center, and use the transcription and translation rules of the central dogma to obtain the RNA and amino acid sequences corresponding to the DNA. The smallest unit of data encoding is unified as {A, T, G, C, N}, where ATGC and N represent 4 bases and placeholders respectively. In this embodiment, for DNA and RNA sequences, base substitution is performed according to the base pairing principle (such as A-U and G-C); for proteins, pairing is performed according to the codon degeneracy of 64 codons to 22 amino acids.

[0077] Step (2): When integrating the labels of the training data, as Figure 3 shown, consider but not limited to various regulatory elements in the non-coding region of the biological database, such as methylation modification, cis-regulatory elements, and transcription start sites, etc., and various gene features in the coding region, such as exons, introns, untranslated regions, start codons, and stop codons, etc. At the same time, consider the amino acid sequence corresponding to the coding sequence of the exon, and the species type corresponding to the nucleotide sequence. This solution queries cis-regulatory elements (such as promoters, enhancers, and terminators), non-coding RNAs (such as miRNAs and IncRNAs, etc.), and gene features (such as exons and introns) in the NCBI or UCSC database according to the gene index, and aligns all annotations to the nucleotide sequence to obtain an annotation sequence. At the same time, search for the species category corresponding to the nucleotide sequence in the NCBI database, and use the annotation sequence and the species category as the supervised learning labels. Each piece of training data finally obtained is in the form of a 4-tuple, {nucleotide, gene annotation, amino acid, species category}.

[0078] Step (3): When constructing the test data, given any DNA, RNA, or amino acid sequence, first perform data alignment with the DNA sequence as the template. It is necessary to perform local or global alignment on the RNA and amino acid sequences respectively to locate the matching fragments. A scoring matrix M(i,j) can be defined to quantify the matching situation between the DNA sequence and the RNA or amino acid sequence, and the dynamic programming algorithm implemented by the sequence matching tool is used to find the maximum scoring path in the scoring matrix, so as to determine the optimal DNA and RNA (or amino acid) sequence alignment method. Then, using the reverse transcription and reverse translation rules of the central dogma, convert the RNA sequence into a DNA sequence, and convert the amino acid sequence into complementary DNA (cDNA) encoding according to the species preference to achieve the unified expression of multi-omics data.

[0079] In the second part of the content, first use a dynamic tokenizer to adaptively segment and embed the long nucleotide sequence with low information density into a gene-level embedding sequence, and use a long sequence encoding model that combines efficient self-attention and gated long convolutional modules to model the context information on the gene encoding. The tokenizer and the decoder use a perception module that combines local pooling and cross-attention. This solution performs self-supervised pre-training on the tokenizer and the decoder, and the gene encoder and the gene decoder respectively using the mask reconstruction task of the nucleotide and gene-level characterization sequences.

[0080] Secondly, based on gene-level embeddings, the encoder models the long-range interaction relationships between coding and non-coding region sequence representations, so as to predict the cis-regulatory elements and gene features of each functional segment on nucleotides through a gene annotation decoder. The decoded gene annotation results can use the tags included in the biological database as the tags for gene annotation, so as to accurately model the central dogma transcription process of coding sequences and non-coding regions. In particular, the functions and regulatory roles of non-coding RNAs will also be modeled in this process, reflecting their important roles in gene expression regulation.

[0081] Furthermore, in order to model functional hierarchical features, this solution designs a protein structure decoder, which can extract the representation sequence corresponding to the length of the amino acid sequence according to the coding region annotation predicted by the gene annotation decoder, and let the protein decoder align the latent space features to the existing large protein model through knowledge distillation to predict the protein structure. At the same time, global classification tokens and species category classifiers are used in gene-level representations to predict the species types of nucleotide sequences.

[0082] In the third part of the content, in order to facilitate the solution of downstream tasks of multi-omics or single-cell, this solution further introduces a cross-attention interface on the basis of the parameter-efficient fine-tuning method of the computational model. Using the main modality of nucleotide input as the query of cross-attention and the latent space encoding of another multi-omics modality or cell environment sequencing result as the key-value of cross-attention, the latter information can be fused into the main modality representation, so as to achieve cross-omics tasks that take into account the characteristics of specific application scenarios. This method can not only effectively maintain the generality of the pre-trained model and the efficiency of fine-tuning, but also efficiently model multi-omics or cell environment features to accurately complete multi-omics tasks.

[0083] Figure 4 It shows the complete inference process of the network architecture, multi-omics information prediction and downstream task fine-tuning of the Uni-Life computational model in this solution, mainly including three aspects of design.

[0084] Design (1): This solution designs a dynamic tokenizer and an efficient long-sequence encoder for the input DNA sequence, and obtains gene-level representations through tokenization, merging, and encoding. The dynamic tokenizer includes a sequence modeling network and a dynamic tokenization module. Among them, the sequence modeling uses a perceiver architecture that combines long convolution and cross-attention to encode important local gene features and global information in the nucleotide sequence; the dynamic tokenization module divides the sequence segments through a feature merging algorithm (tokenmerge) and merges them into the smallest units. Since the coding region sequence has a minimum unit of 3-base combinations as codons, and the regulatory elements in the non-coding region have no obvious rules, this solution expects the dynamic tokenizer to adaptively learn the words required for coding and non-coding region sequences from large-scale nucleotide sequence data. Given a DNA sequence with a length of N Train a dynamic tokenizer T φ :S DNA →P DNA Embed, dynamically partition, and downsample the DNA sequence to obtain a chunked gene-level representation sequence of length L (usually L < N). Among them, the dynamic tokenization module processes the DNA sequence S DNA Calculate the information entropy to construct the vocabulary of each DNA sequence

[0085]

[0086] Among them, H(x i ) is the information entropy at the x i site, and p e (.) is the information entropy of the DNA sequence at a certain site relative to the previous site (next byte entropy). By setting the threshold θ g of the information gain, it can be screened whether each position needs to be split into a new fragment, that is, H(x t ) - H(x t-1 ) > θ g Pool the embeddings belonging to the same fragment using mean pooling to obtain the gene-level embedding sequence P DNA Then, by constructing an encoder network ε ψ that combines long convolutional and self-attention, the context information of the multi-omics sequence can be modeled to obtain the gene-level representation sequence The long convolution, combined with the gating mechanism, can model the medium-range long context information of the long sequence representation within linear complexity. Combined with the self-attention mechanism that explicitly models global information, it can realize the processing of local, medium-range, and long-range information of the unified representation of the multi-omics sequence.

[0087] Design (2): This solution models multi-omics information at three levels to accurately model the central dogma of biology. The first level is to uniformly encode and contextually model the multi-omics sequence expressed by the DNA sequence, as described in the above Design (1), and use a de-tokenizer to reconstruct the gene-encoded sequence into the original DNA sequence S DNA. The second level is to model the long-range interaction relationship of coding region and non-coding region sequence representations based on gene-level embeddings and predict gene annotation information. This solution designs a gene annotation decoder that can predict the cis-regulatory elements and gene features of each functional segment on nucleotides. where A t ∈ {0, 1} indicates whether the t-th site of the DNA sequence belongs to the annotation category (there are a total of G annotation categories). The third level is to model the function of the coding sequence. This solution designs a protein structure decoder to model the protein translation and folding process of the central dogma of the coding sequence. According to the coding region sequence representation predict the corresponding amino acid sequence and protein structure At the same time, predict the species category information of the nucleotide sequence through a species classifier

[0088] Design (3): For multi-omics interactions and cell-level downstream tasks, this solution designs a parameter-efficient fine-tuning strategy and gives specific designs in two cases. First, in the fine-tuning stage, keep the parameters of the tokenizer, encoder, and decoder in the pre-trained computational model frozen to extract the general gene-level representation sequence Z gene , and introduce a low-rank fine-tuning branch (LoRA fine-tune branch) and a downstream task decoder D down for the encoder. The parameters of the low-rank branch are ε LoRA . Then, this solution considers two different implementation methods in specific scenarios. In the first case, for single-omics downstream tasks not covered in the pre-training stage and downstream tasks of splicing multiple omics data, such as genome function prediction other than general gene annotation, protein design and gRNA design tasks of the CRISPR-Cas9 gene editing system, RNA-protein interaction tasks, etc., the characteristic is to input multiple omics data for in-context learning. In this case, the preprocessing method in the first part can cover possible omics data inputs. The omics data can be converted into nucleotide form and packed and input into the computational model, X down = {S DNA ,…}, and use the loss function L down of the downstream task and the label Y downto finely tune the low-rank branch and the multi-layer perceptron (MLP) decoder D finetune , and its learning objective is expressed as:

[0089]

[0090] In the second case, considering the downstream tasks of the interaction between quantitative sequencing data and multi-omics sequence data, sequencing and quantitative prediction are usually regarded as results and multi-omics data as conditions. For example, predicting various RNA alternative splicing sites in transcriptomics, modeling cell states in single-cell sequencing, cell fate transition, etc. For this reason, a cross-attention mechanism can be introduced into the decoder D down , using the feature encoding of the sequencing data X seq as the query key, and the representation obtained by the multi-omics sequence X down through the encoder as the key value, which can achieve precise quantitative modeling in specific downstream tasks.

[0091] Figure 5 Figure [ID] shows the pre-training process of the three levels of the Uni-Life computational model, mainly including the following four tasks:

[0092] (a) DNA sequence and gene embedding sequence reconstruction task: In order to enable the tokenizer and the encoder to better model the nucleotide sequence with low information density into a gene-level representation with high information density, this solution combines the decoder of the encoder and the detokenizer of the tokenizer to implement a self-supervised masked context modeling task. For the tokenizer and the detokenizer, referring to the training method of the BERT masked language model (masked language modeling, MLM), some bases in the input sequence are randomly masked (masked) with a probability of 15% to obtain the sequence M i where \(M_i\) represents whether the \(i\)-th site is masked, and \(B(.)\) represents the Bernoulli distribution. Then, the detokenizer model is required to predict the masked bases, and the loss function can be defined as:

[0093]

[0094] where \(p(.)\) represents the probability predicted for the \(i\)-th masked site. Similarly, for the encoder, a randomly masked sequence of gene representations can be input and the gene-level representation context can be reconstructed, and its prediction target is the dynamic vocabulary of each DNA sequence so the form of the loss function is the same as above.

[0095] (b) Gene annotation task: To accurately model gene-level information in the transcription process of the central dogma, this solution implements a gene annotation decoder for segmenting functional fragments of DNA sequences. It consists of multiple cross-attention modules and query tokens, and based on the gene-level representation sequence Z gene , for each gene annotation category g, it decodes the segmentation of functional fragments into where A j ∈{0,1} indicates whether the j-th site of the DNA sequence belongs to the annotation category There are a total of G annotation categories. The decoding calculation process is expressed as:

[0096] A c = D gene (CrossAttention(Q c , P DNA ))

[0097] where the cross-attention module CrossAttention(.) uses Z gene as the key-value pair, uses the learnable annotation category query token Q c as the query key, and through the segmentation prediction head D gene and the built-in Sigmoid function, it realizes the output of a binary mask sequence. By calculating G times, the complete gene annotation result of the DNA sequence can be obtained Combined with the gene annotation label Y c , the supervised segmentation loss function is defined as:

[0098]

[0099] (c) Protein structure decoding task: To further model the central dogma law at the functional level, this solution implements a protein structure decoder to predict the amino acid sequence and protein structure of the coding sequence. Based on the exon fragments of gene annotation, on the one hand, the corresponding amino acid sequence (using 20 standard amino acid codings) can be obtained during the construction of training data. On the other hand, the gene-level representation can be intercepted by the feature fusion module (token merge) to obtain a representation sequence of the same length as the amino acid The protein structure decoder implemented by using the cross-attention module and the structure prediction head D protein will predict the amino acid sequence and its corresponding protein 3D structure according to Z AA This process can be expressed as: This process can be expressed as:

[0100] H protein = D protein(CrossAttention(Z AA ,Z gene ))

[0101] Among them, the amino acid characterization sequence Z AA is used as the query key of cross-attention, and the gene-level context characterization Z ene is used as the key value. Similarly, after the cross-attention module, an amino acid prediction head D AA is used to implement the decoding of the amino acid characterization sequence into an amino acid sequence , which can constrain the amino acid characterization sequence extracted by the model. Since the large model technology for realizing the translation from amino acid sequence to protein structure has matured, a pre-trained protein model h δ :S AA →H protein can be obtained as the teacher model (such as ESM-2). Using the designed protein structure decoder as the student model, knowledge distillation is performed on the protein structure decoder. The sequence-level knowledge distillation loss L dist can align the output of the student model and the output of the teacher model, and cross-entropy can be used to define it:

[0102]

[0103] Among them, H teacher and H student represent the protein structures predicted by the teacher and student models respectively. At the same time, the translation task of the amino acid sequence can be optimized using the translation loss:

[0104]

[0105] Among them, is the indicator function, which is 1 when the true amino acid S i is equal to the j-th amino acid, and 0 otherwise. is the probability that the i-th site of the model-predicted Z AA sequence is the j-th amino acid. The distillation and translation losses complement each other in modeling the translation and folding processes from genes to amino acids and proteins.

[0106] (d) Nucleotide sequence species classification: For the functional level, phenotypic information can be further modeled, that is, a species classifier is used to predict the species type of the nucleotide sequence. In this scheme, a global classification token is introduced into the encoder to extract the global classification features of the gene-level characterization, and species classification is performed through a non-linear classifier D cls , and the output species category probability is p(P cls ). Combining the species-level category annotation using the cross-entropy classification loss Lcls as follows:

[0107] L cls = -Y S logp(P cls )

[0108] (e) Pre-training stage: Combining the above 4 training tasks, in the pre-training stage, the loss function is logarithmically weighted to achieve multi-task balance, ensuring that multi-tasks with different difficulties have similar convergence speeds. The total loss L total can be defined as follows:

[0109] L total = log(L MLM ) + log(L gene ) + log(L trans ) + log(L dist ) + log(L cls )

[0110] Through this pre-training of mixed tasks, the Uni-Life computational model can simultaneously model information at the nucleotide, gene, and function levels. It can not only achieve multi-omics sequence interaction to model general multi-omics representations, but also complete tasks such as single-omics gene annotation and protein structure prediction.

[0111] In summary, under the guidance of the central dogma and gene annotation (information such as protein-coding regions, non-coding RNA regions and functions, regulatory elements, and variant sites), this solution uses unified nucleotide coding to achieve integrated input of multi-omics data, and realizes the modeling and prediction of biological information at the nucleotide, gene, and function levels respectively, achieving integrated analysis of single-omics function prediction and multi-omics biological data interaction. It can not only directly predict single-omics biological information, but also solve multi-omics downstream tasks and cell-level analysis tasks through a computationally efficient parameter fine-tuning strategy, showing wide applicability in biological applications such as gene annotation, species classification, protein structure prediction, multi-omics interaction, and single-cell function prediction. The Uni-Life computational model proposed in this solution has significant advantages in aspects such as multi-omics data modeling and long-range interaction capture, and can provide a more comprehensive and efficient computational tool for multi-level biological research and applications.

Claims

1. A method for integrated analysis of multi-omics sequence data based on artificial intelligence, characterized in that, It includes the following steps: S1. Obtain multi-omics data samples, preprocess the multi-omics biological data samples, and construct training data and test data; S2. Construct the network architecture of the multi-omics computational model, and use the training data and test data to train and fine-tune the multi-omics computational model to obtain the Uni-Life computational model; S3. Preprocess the multi-omics data to be processed, and then input it into the Uni-Life computational model to output nucleotide sequences, gene annotation information, protein structures, and species category information.

2. The multi-omics sequence data integration and analysis method based on artificial intelligence according to claim 1, wherein, In step S1, the preprocessing of the multi-omics biological data samples specifically converts the RNA sequence and amino acid sequence into a unified DNA sequence according to the reverse transcription and reverse translation rules of the central dogma; Then, according to the unified DNA sequence, hierarchically match and annotate the coding gene, non-coding RNA region, regulatory element, and gene variation information, construct gene annotation data and generate training labels.

3. The method for integrated analysis of multi-omics sequence data based on artificial intelligence according to claim 2, wherein The construction of training data and test data in step S1 includes the following steps: S11. When constructing training data, according to the gene index, match the DNA, RNA, and amino acid sequences, intercept the nucleotide sequences in the ranges of 2000 bp before the transcription start site and after the terminator with the gene fragment as the center, and use the transcription and translation rules of the central dogma to obtain the RNA and amino acid sequences corresponding to the DNA, and unify the smallest unit of data coding into {A, T, G, C, N}, where ATGC and N represent 4 bases and placeholders respectively; S12. When integrating the labels of the training data, according to the gene index, query the cis-regulatory elements, non-coding RNAs, and gene features in the NCBI or UCSC database, align all the annotations to the nucleotide sequence to obtain an annotation sequence. At the same time, search for the species category corresponding to the nucleotide sequence in the NCBI database, and use the annotation sequence and species category as supervised learning labels. The final obtained each piece of training data is in the form of a 4-tuple: {nucleotide, gene annotation, amino acid, species category}; S13. When constructing test data, arbitrarily given DNA, RNA, or amino acid sequences, first align the data with the DNA sequence as the template. It is necessary to perform local or global alignment on the RNA and amino acid sequences respectively to locate the matching fragments. Define a scoring matrix M(i, j) to quantify the matching situation between the DNA sequence and the RNA or amino acid sequence, and use the dynamic programming algorithm implemented by the sequence matching tool to find the maximum scoring path in the scoring matrix, so as to determine the optimal DNA and RNA or amino acid sequence alignment method. Then, use the reverse transcription and reverse translation rules of the central dogma to convert the RNA sequence into a DNA sequence, and convert the amino acid sequence into a complementary DNA code according to the species preference to achieve the unified expression of multi-omics data.

4. An integrated analysis method for multi-omics sequence data based on artificial intelligence according to claim 1, characterized in that, The network architecture of the multi-omics computational model in step S2 includes a first level, a second level, and a third level. The first level is used to uniformly encode and contextually model the multi-omics sequences expressed by the DNA sequences, and reconstruct the gene coding sequences into the original DNA sequence S DNA ; The second level is based on gene-level embedding, models the long-range interaction relationship between the coding region and non-coding region sequence representations, and predicts the gene annotation information; The third level is used to model the central dogma protein translation and folding processes of coding sequences, and to predict the species category information of amino acid sequences, protein structures, and nucleotide sequences.

5. The multi-omics sequence data integration and analysis method based on artificial intelligence according to claim 4, characterized in that, The first level includes a dynamic tokenizer, a long sequence encoder, and a decoder. The dynamic tokenizer includes a sequence modeling network and a dynamic tokenization module. The sequence modeling network adopts a perceptron architecture combining long convolution and cross-attention, and is used to encode important local gene features and global information in the nucleotide sequence. The dynamic tokenization module divides sequence segments through a feature merging algorithm and merges them into the smallest units. The long sequence encoder adopts an encoder network ε that combines long convolution and self-attention ψ , which is used to model the context information of multi-omics sequences and obtain a gene-level characterization sequence; The word decoder is used to reconstruct the gene-level characterization sequence into the original DNA sequence S DNA .

6. The multi-omics sequence data integration and analysis method based on artificial intelligence according to claim 5, characterized in that The second level includes a gene embedding encoder and a gene annotation decoder. The gene embedding encoder is used to model the long-range interaction relationship between coding region and non-coding region sequence representations, and the gene annotation decoder is used to predict the cis-regulatory elements and gene features of each functional segment on the nucleotide.

7. An integrated analysis method for multi-omics sequence data based on artificial intelligence according to claim 6, characterized in that, The third level includes a protein structure decoder and a species classifier. The protein structure decoder is used to extract a representative sequence corresponding to the length of the amino acid sequence from the gene annotation information, and aligns the latent space features to an existing large protein model through knowledge distillation to predict the protein structure. The species classifier introduces global classification terms to extract global classification features of gene-level representations, and performs species classification through a non-linear classifier to output species category probabilities.

8. The method for integrated analysis of multi-omics sequence data based on artificial intelligence according to claim 7, characterized in that, In step S2, training the multi-omics computational model includes four tasks: DNA sequence and gene embedding sequence reconstruction tasks, gene annotation tasks, protein structure decoding tasks, and nucleotide sequence species classification tasks.

9. The method for integrated analysis of multi-omics sequence data based on artificial intelligence according to claim 8, wherein The specific process of the DNA sequence and gene embedding sequence reconstruction task is as follows: For the dynamic tokenizer and detokenizer, referring to the training method of the BERT masked language model, randomly mask some bases in the input sequence with a probability of 15% to obtain the sequence M i where $M_{i}$ represents whether the $i$-th site is masked, and $B(.)$ represents the Bernoulli distribution. Then, the detokenizer is required to predict the masked bases, and the loss function is defined as: Among them, p(.) represents the probability predicted by the i-th masked site. Similarly, a randomly masked sequence of gene representations can be input to the encoder and the gene-level representation context can be reconstructed, with the prediction target being the dynamic vocabulary for each DNA sequence The form of the loss function is the same as above; The specific process of the gene annotation task is as follows: The gene annotation decoder consists of multiple cross-attention modules and query words, and based on the gene-level representation sequence Z modeled by the gene embedding encoder gene , the functional segment segmentation of each gene annotation category g is decoded into where A t ∈ {0, 1} indicates whether the t-th site of the DNA sequence belongs to the annotation category There are a total of G annotation categories, and the decoding calculation process is expressed as: A c = D gene (CrossAttention(Q c , P DNA )) Among them, the cross-attention module CrossAttention(.) uses Z gene as key-value pairs, and uses the learnable labeled category query word Q c as the query key. Through the segmentation prediction head D gene and the built-in Sigmoid function, a binary mask sequence output is realized. After calculating G times, the complete gene annotation result of the DNA sequence is obtained Combined with the gene annotation label Y c , the supervised segmentation loss function is defined as: Among them, A c,i is the mask of the i-th token in the DNA sequence output by the functional fragment segmentation decoding, and Y c,i is the label of the i-th token in the DNA sequence for the gene annotation label; The specific process of the protein structure decoding task is as follows: Based on the exon fragments annotated by genes, on the one hand, the corresponding amino acid sequence is obtained during the construction of training data. On the other hand, the gene-level representation is intercepted by the feature fusion module to obtain a representation sequence of the same length as the amino acid. The cross-attention module and the structure prediction head D are used. protein The implemented protein structure decoder will be based on Z. AA Predict the amino acid sequence and its corresponding protein 3D structure. This process is expressed as: H protein = D protgin (CrossAttention(Z AA , z gene )) Among them, the amino acid characterization sequence Z AA serves as the query key for cross-attention, and the gene-level context characterization Z gene serves as the key-value. Similarly, after the cross-attention module, an amino acid prediction head D AA is used to realize the decoding from the amino acid characterization sequence to the amino acid sequence for constraining the amino acid characterization sequence extracted by the model; Using the pre-trained protein model h δ :S AA →H protein As the teacher model, using the protein structure decoder as the student model, knowledge distillation is performed on the protein structure decoder, and the sequence-level knowledge distillation loss L dist Align the output of the student model with the output of the teacher model, using the cross-entropy definition: where H teacher and H student represent the protein structures predicted by the teacher and student models, respectively. Meanwhile, the translation loss is used to optimize the translation task of the amino acid sequence: Among them, is an indicator function, which is 1 when the true amino acid S i is equal to the j-th amino acid and 0 otherwise. is the probability that the i-th site of the model-predicted Z AA sequence is the j-th amino acid, and the distillation and translation losses complement each other to model the translation and folding processes from genes to amino acids and proteins; The specific process of the nucleotide sequence species classification task is as follows: Introduce a global classification term in the encoder to extract the global classification features of gene-level representations. The classification term aggregates species category information through the long convolution and self-attention mechanisms of the encoder, and passes through the non-linear classifier D cls for species classification, and the output species category probability is p(P cls ). The cross-entropy classification loss L cls is used for training: L cls = -Y S logp(P cls ) Among them, is the species-level category annotation; In the training process, the loss function is logarithmically weighted to achieve multi-task balance, ensuring that multi-tasks with different difficulties have similar convergence speeds. L total = log(L MLM ) + log(L gene ) + log(L trans ) + log(L dist ) + log(L cls ) where, L total is the total loss function.

10. The method for integrated analysis of multi-omics sequence data based on artificial intelligence according to claim 9, wherein, The specific process of testing and fine-tuning in step S2 is as follows: First, during the fine-tuning stage, keep the parameters of the tokenizer, encoder, and decoder in the pre-trained computational model frozen to extract the general gene-level characterization sequence Z gene , and introduce a low-rank fine-tuning branch and a downstream task decoder D to the encoder down , where the parameters of the low-rank branch are ε LoRA ; Then, considering two different implementation methods in a specific scenario, in the first case, for single-omics downstream tasks not covered in the pre-training stage and downstream tasks of splicing multiple omics data, the test data is converted into nucleotide form and spliced together and input into the computational model, X down ={S DNA ,…}, and the loss function L down of the downstream task and the label Y down are used to fine-tune the low-rank branch and the multi-layer perceptron (MLP) decoder D finetune ; In the second case, considering the downstream tasks of the interaction between quantitative sequencing data and multi-omics sequence data, sequencing and quantitative prediction are regarded as results while multi-omics data are regarded as conditions. In decoder D down a cross-attention mechanism is introduced, using the feature encoding of sequencing data X seq as the query key, and using the representation obtained from multi-omics sequence X down through the encoder as the key value to achieve precise quantitative modeling in specific downstream tasks.

Citation Information

Cited By

  • Insect multi-tissue expression quantity prediction method and system based on multi-modal feature fusion

    CN122090926A