DNA sequence alignment method based on deep learning
Through the deep learning-based DNA sequence alignment method, the memory usage and alignment speed of minimap2 during ultra-long DNA sequencing fragments is solved, and efficient and accurate sequence alignment is achieved, especially in the process of high-density structural variations and complex repeat sequences.
Patent Information
- Application Number
- CN202510358349.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
When faced with ultra-long DNA sequencing fragments, existing sequence alignment tools such as minimap2 have problems such as a significant increase in memory footprint, a significant decrease in alignment speed and insufficient alignment accuracy. Especially when dealing with high-density structural mutations, complex repeat sequences and long insertion deletions, it is difficult to fully and accurately reflect the true genomic mutation characteristics.
Using a deep learning-based DNA sequence alignment method, the DNA sequencing sequence is converted into word vectors through BPE encoding, and the feature vector is extracted using the pre-trained model DNABert2 and dimensionality reduction processing is performed to construct a deep learning classification model, predict the chromosome category to which the DNA sequencing fragment belongs, and use the minimap2 tool for high-precision sequence alignment.
It significantly reduces the range of candidate regions, improves the speed and accuracy of alignment, reduces memory usage, and can better cope with high-density structural variations and complex repeat sequences, fully reflects the characteristics of real genomic mutations, and improves the efficiency and accuracy of alignment.
Smart Images

Figure CN120299519A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a DNA sequence alignment method based on deep learning. Background Art
[0002] DNA sequence alignment plays an important role in genomic research, and its applications cover aspects such as identification of sequence differences, improvement of gene annotation accuracy, and analysis of genetic mechanisms. By comparing the genomic sequences of different individuals or species, researchers can effectively reveal the similarities and differences between sequences, accurately locate mutation and genetic polymorphism sites, and thus deeply understand the genetic basis of diseases and phenotypes. In addition, sequence alignment can also identify conserved and variable regions in the genome, simplify the process of identifying gene functions and regulatory elements, and promote the research on gene expression and epigenetic regulation mechanisms. In the fields of gene editing, mutation analysis, and functional genomics, high-quality sequence alignment significantly improves the accuracy of gene manipulation and functional research. At the same time, through sequence alignment analysis of special genetic markers such as mitochondrial DNA and sex chromosomes, the genetic transmission pattern of human populations can be effectively restored, revealing the historical trajectory of human evolution and migration. Therefore, DNA sequence alignment is of crucial significance for precision medicine, genetics, genomics, and evolutionary biology research.
[0003] Currently, mainstream sequence alignment tools such as minimap2 can achieve fast and accurate alignment of fragments from hundreds of base pairs to tens of thousands of base pairs, and are widely used in genomic research. minimap2 adopts a seed index and a chaining algorithm, and performs outstandingly in terms of alignment efficiency. However, with the continuous development of sequencing technology, the length of sequencing fragments has gradually increased, and some are even close to one million bases. The emergence of such ultra-long fragments poses a severe challenge to minimap2, specifically manifested as a significant increase in memory occupancy and a significant decrease in alignment speed. In addition, when facing problems such as high-density structural variations, complex repetitive sequences, and long insertions and deletions, minimap2 has the defect of insufficient accuracy and is difficult to comprehensively and accurately reflect the true genomic variation characteristics. Therefore, it is urgent to develop or optimize a new generation of sequence alignment tools to meet the requirements of high-precision and high-efficiency alignment analysis of ultra-long DNA sequencing fragments. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems that existing mainstream sequence alignment tools such as minimap2 have a significant increase in memory occupancy and a significant decrease in alignment speed when facing ultra-long DNA sequencing fragments, and have insufficient alignment accuracy and are difficult to comprehensively and accurately reflect the true genomic variation characteristics when dealing with problems such as high-density structural variations, complex repetitive sequences, and long insertions and deletions, and to propose a DNA sequence alignment method based on deep learning.
[0005] The specific process of a DNA sequence alignment method based on deep learning is as follows:
[0006] Step 1: Convert a DNA sequencing sequence of one million in length from string form into word vectors based on BPE encoding;
[0007] Step 2: Input the word vectors into the pre-trained model DNABert2, and the pre-trained model DNABert2 outputs feature vectors;
[0008] Perform dimensionality reduction on the feature vectors by average pooling to obtain the dimensionality-reduced features;
[0009] Step 3: Construct a deep learning classification model;
[0010] Train the deep learning classification model based on the training set to obtain a trained deep learning classification model;
[0011] Input the dimensionality-reduced features in Step 2 into the trained deep learning classification model, and the trained deep learning classification model predicts the chromosome category to which the sequence corresponding to the dimensionality-reduced features in Step 2 belongs;
[0012] Record the category results, and the DNA sequences corresponding to each category are stored in 1 fasta format file;
[0013] Step 4: Align the sequences with the recorded categories in each file with the DNA sequences of the corresponding chromosomes in the reference sequence, and save the alignment results;
[0014] The reference sequence is the human genome sequence.
[0015] The beneficial effects of the present invention are as follows:
[0016] The present invention proposes a DNA sequence alignment method based on deep learning. This method aims at ultra-long DNA sequencing fragments of one million in length, predicts the chromosome category to which they belong through a deep learning model, and thus replaces the genomic sequence with the DNA sequence of the corresponding chromosome as the candidate alignment region during the alignment process. This significantly reduces the range of the alignment candidate region, improves the alignment speed, and effectively reduces the memory occupancy, solving the problems of excessive memory occupancy and slow alignment speed existing in traditional alignment tools (such as minimap2) when processing DNA sequencing fragments. Through this method, the alignment tool can use ultra-long DNA sequencing fragments for alignment, can better handle problems such as high-density structural variations, complex repetitive sequences, and long insertions and deletions, comprehensively and accurately reflect the variation characteristics of the real genome, and greatly improve the alignment accuracy and efficiency. Description of the Drawings
[0017] Figure 1 is the flow chart of the present invention;
[0018] Figure 2 is the structural diagram of the deep learning classification model of the present invention, L = 5. Specific Embodiments
[0019] Specific Embodiment 1: The specific process of a DNA sequence alignment method based on deep learning in this embodiment is as follows:
[0020] First, encode the DNA sequencing fragments through the BPE method and input them into the pre-trained model DNABert2 to obtain sequence feature representations; subsequently, perform dimensionality reduction on the features through average pooling operation to obtain the overall sequence representation; finally, use a deep learning classifier to realize the prediction of the chromosome category of the sequencing fragments.
[0021] According to the predicted chromosome category, select the corresponding chromosome sequence as the reference sequence for the DNA sequencing fragments, and use the minimap2 tool to complete the high-precision sequence alignment analysis.
[0022] Step 1: Convert a DNA sequencing sequence of one million in length from string form to word vectors based on BPE encoding;
[0023] Step 2: Input the word vectors into the pre-trained model DNABert2, and the pre-trained model DNABert2 outputs feature vectors;
[0024] Perform dimensionality reduction on the feature vectors by using average pooling to obtain high-quality dimensionality-reduced features;
[0025] Step 3: Construct a deep learning classification model ( Figure 2 );
[0026] Train the deep learning classification model based on the training set to obtain a trained deep learning classification model;
[0027] Input the dimensionality-reduced features in Step 2 into the trained deep learning classification model, and the trained deep learning classification model predicts the chromosome category to which the sequence corresponding to the dimensionality-reduced features in Step 2 belongs;
[0028] Record the category results, and the DNA sequences corresponding to each category are stored in 1 fasta format file;
[0029] Step 4: Align the sequences with the recorded categories in each file with the corresponding chromosome DNA sequences in the reference sequence, and save the alignment results;
[0030] The reference sequence is the human genome sequence.
[0031] Specific Embodiment 2: The difference between this embodiment and Specific Embodiment 1 is that in Step 1, a DNA sequencing sequence of one million in length is converted from string form to word vectors based on BPE encoding; the specific process is as follows:
[0032] Step 1: Use the set of human genome sequences D = {S1, S2, …, S i , …, S n} as the corpus (the set of all DNA sequences), where i = 1, 2, …, n;
[0033] For example, S1 is the DNA sequence of chromosome 1;
[0034] Step 2: Consider each base (1 base is 1 character) in the i-th sequence Si in the sequence set D as an independent Token; i
[0035] Step 3: Calculate the occurrence frequency of each Token pair (x, y) in the i-th sequence Si in the corpus; i
[0036] Each Token pair (x, y) is composed of adjacent Tokens, and there is no overlap between Token pairs (x, y);
[0037] Step 4: Repeat Step 3 until the occurrence frequency of each Token pair (x, y) in all n sequences in the sequence set D is obtained;
[0038] Step 5: Add up the occurrence frequencies corresponding to the same Token pair (x, y) in the n sequences to obtain the total occurrence frequency of the same Token pair (x, y) in the n sequences; expressed as:
[0039]
[0040] where count((x, y), Si i ) represents the occurrence frequency corresponding to the same Token pair (x, y) in the i-th sequence Si i ;
[0041] Step 6: Combine the Token pairs (x, y) corresponding to the maximum total occurrence frequency in each sequence to obtain 1 new Token, denoted as z = (x, y), and update the vocabulary;
[0042] Step 7: Repeat Steps 3 to 6 M times to obtain all the final Tokens, number all the final Tokens (1, 2, 3, 4, etc.), and the numbers are non-repeating;
[0043] All the final Tokens and the corresponding number for each Token are used as the final vocabulary;
[0044] For example, the initial vocabulary is {"A": 0, "T": 1, "C": 2, "G": 3}
[0045] After one update, it becomes {"A":0,"T":1,"C":2,"G":3,"AT":4}
[0046] M is 4088;
[0047] In the finally obtained vocabulary, it contains both single-base Tokens and merged Tokens that can better represent common subsequences, so as to achieve a more compact and efficient vector representation when segmenting and encoding new sequences.
[0048] Step 1-8: To accelerate the prediction speed of the subsequent model, and since the DNA sequence itself contains rich sequence feature information, for a DNA sequencing fragment with a length of one million, a DNA sequence S′ with a length of ten thousand is selected backward from the head end of the DNA sequencing fragment, a DNA sequence S″ with a length of ten thousand is selected backward from the middle end of the DNA sequencing fragment, and a DNA sequence S″′ with a length of ten thousand is selected forward from the tail end of the DNA sequencing fragment;
[0049] Based on the final vocabulary, after BPE encoding the DNA sequence S′ with a length of ten thousand selected backward from the head end of the DNA sequencing fragment, a string of Token sequences is obtained:
[0050] {x′1,x′2,…,x′ j ,…,x′ m} = BPE(S′)
[0051] Based on the final vocabulary, after BPE encoding the DNA sequence S″ with a length of ten thousand selected backward from the middle end of the DNA sequencing fragment, a string of Token sequences is obtained:
[0052] {x″1,x″2,…,x″ j ,…,x″ m} = BPE(S″)
[0053] Based on the final vocabulary, after BPE encoding the DNA sequence S″′ with a length of ten thousand selected forward from the tail end of the DNA sequencing fragment, a string of Token sequences is obtained:
[0054] {x″′1,x″′2,…,x″′ j ,…,x m} = BPE(S″′)
[0055] Among them,
[0056] m represents the number of Tokens after encoding;
[0057] {x′1,x′2,…,x′ j ,…,x′ m} represents a sequence of tokens obtained by BPE encoding a DNA sequence of length ten thousand selected from the head end of a DNA sequencing fragment backward; x′1 represents the first token obtained after BPE encoding the first token in the DNA sequence of length ten thousand; x′2 represents the second token obtained after BPE encoding the second token in the DNA sequence of length ten thousand; x′ j represents the j-th token obtained after BPE encoding the j-th token in the DNA sequence of length ten thousand; x′ m represents the m-th token obtained after BPE encoding the m-th token in the DNA sequence of length ten thousand;
[0058] {x″1, x″2, …, x″ j , …, x″ m} represents a sequence of tokens obtained by BPE encoding a DNA sequence of length ten thousand selected from the middle end of a DNA sequencing fragment backward; x″1 represents the first token obtained after BPE encoding the first token in the DNA sequence of length ten thousand; x″2 represents the second token obtained after BPE encoding the second token in the DNA sequence of length ten thousand; x″ j represents the j-th token obtained after BPE encoding the j-th token in the DNA sequence of length ten thousand; x″ m represents the m-th token obtained after BPE encoding the m-th token in the DNA sequence of length ten thousand;
[0059] {x″′1, x″′2, …, x″′ j , …, x″′ m} represents a sequence of tokens obtained by BPE encoding a DNA sequence of length ten thousand selected from the tail end of a DNA sequencing fragment forward; x″′1 represents the first token obtained after BPE encoding the first token in the DNA sequence of length ten thousand; x″′2 represents the second token obtained after BPE encoding the second token in the DNA sequence of length ten thousand; x″′ j represents the j-th token obtained after BPE encoding the j-th token in the DNA sequence of length ten thousand; x″′ m represents the m-th token obtained after BPE encoding the m-th token in the DNA sequence of length ten thousand;
[0060] Step 19. Map the sequence of tokens obtained after BPE encoding to the corresponding word vector embedding space to obtain the word vectors of the token sequence: It is represented as:
[0061] {e′1, e′2, …, e′ j , …, e′ m} = Embed{x′1, x′2, …, x′ j , …, x′ m}
[0062] {e″1, e″2, …, e″ j , …, e″ m} = Embed{x″1, x″2, …, x″ j , …, x″ m}
[0063] {e″′1, e″′2, …, e″′ j , …, e″′ m} = Embed{x″′1, x″′2, …, x″′ j , …, x″′ m}
[0064] Among them,
[0065] {e′1, e′2, …, e′ j , …, e′ m} represents the word vectors corresponding to each Token obtained by mapping {x′1, x′2, …, x′ j , …, x′ m} to the corresponding word vector embedding space; e′1 represents the word vector corresponding to the 1st Token obtained by mapping x′1 to the corresponding word vector embedding space; e′2 represents the word vector corresponding to the 2nd Token obtained by mapping x′2 to the corresponding word vector embedding space; e′ j represents the word vector corresponding to the jth Token obtained by mapping x′ j to the corresponding word vector embedding space; e′ m represents the word vector corresponding to the mth Token obtained by mapping x′ m to the corresponding word vector embedding space;
[0066] {e″1, e″2, …, e″ j , …, e″ m} represents the word vectors corresponding to each Token obtained by mapping {x″1, x″2, …, x″ j , …, x′ m} to the corresponding word vector embedding space; e″1 represents the word vector corresponding to the 1st Token obtained by mapping x″1 to the corresponding word vector embedding space; e″2 represents the word vector corresponding to the 2nd Token obtained by mapping x″2 to the corresponding word vector embedding space; e″j Indicates mapping x″ j to the corresponding word vector embedding space, and obtaining the word vector corresponding to the j-th Token; e″ m Indicates mapping x″ m to the corresponding word vector embedding space, and obtaining the word vector corresponding to the m-th Token;
[0067] {e″′1, e″′2, …, e″′ j , …, e″′ m} indicates mapping {x″′1, x″′2, …, x″′ j , …, x″′ m} to the corresponding word vector embedding space, and obtaining the word vector corresponding to each Token; e″′1 indicates mapping x″′1 to the corresponding word vector embedding space, and obtaining the word vector corresponding to the 1st Token; e″′2 indicates mapping x″′2 to the corresponding word vector embedding space, and obtaining the word vector corresponding to the 2nd Token; e″′ j Indicates mapping x″′ j to the corresponding word vector embedding space, and obtaining the word vector corresponding to the j-th Token; e″′ m Indicates mapping x″′ m to the corresponding word vector embedding space, and obtaining the word vector corresponding to the m-th Token;
[0068] Embed represents word embedding;
[0069] The embedding vector e output via Embed(·) i usually contains information such as the context relationship and position of the sequence.
[0070] Other steps and parameters are the same as those in the first specific implementation manner.
[0071] Specific implementation manner three: The difference between this implementation manner and the first or second specific implementation manner is that after encoding the DNA sequence S′ with a length of ten thousand selected from the head end of the DNA sequencing fragment backward based on the final vocabulary, a string of Token sequences is obtained:
[0072] {x′1, x′2, …, x′ j , …, x′ m} = BPE(S′)
[0073] The specific process is as follows:
[0074] 1), Regarding 1 base in the DNA sequence S′ with a length of ten thousand selected from the head end of the DNA sequencing fragment backward as 1 character;
[0075] 2), Let \(i = 1\) and \(j = 1\); \(i\) can be understood as the starting position of the search, and \(j\) is the length of the token to be searched.
[0076] 3), Starting from the \(i\)-th character in the sequence \(S'\) and moving backward, search for the adjacent \(j + 1\) characters including the \(i\)-th character in the final vocabulary.
[0077] If the adjacent \(j + 1\) characters can be found in the final vocabulary, determine whether \(i + j - 1=10000\) (termination signal because the sequence length is only ten thousand). If it is satisfied, execute 5); if not, execute 4).
[0078] If the adjacent \(j + 1\) characters cannot be found in the final vocabulary, execute 5).
[0079] The purpose of this step is to find tokens with longer lengths as much as possible. Because in addition to converting the string into numbers, the meaning of encoding can also make the length of the encoded sequence much smaller than the original sequence, improving the subsequent calculation efficiency.
[0080] 4). Let \(j = j + 1\) and continue to execute 3).
[0081] 5). Starting from the \(i\)-th character and moving backward, regard the adjacent \(j\) characters including the \(i\)-th character as a token, and encode it according to the content of the final vocabulary (the corresponding number is whatever number it is), and determine whether \(i + j - 1 = 10000\) (termination signal because the sequence length is only ten thousand). If it is satisfied, end and obtain a string of Token sequences \(\{x'_1,x'_2,\cdots,x' m \}\);
[0082] If not, let \(i = i + j\), \(j = 1\), and execute 3).
[0083] Other steps and parameters are the same as those in the first or second specific implementation manner.
[0084] Specific implementation manner four: The difference between this implementation manner and one of the first to third specific implementation manners is that after the DNA sequence \(S''\) with a length of ten thousand selected from the middle end of the DNA sequencing fragment is encoded by BPE based on the final vocabulary, a string of Token sequences is obtained:
[0085] \(\{x''_1,x''_2,\cdots,x'' j ,\cdots,x'' m \}=BPE(S'')
[0086] The specific process is as follows:
[0087] 1), Regard 1 base in the DNA sequence \(S''\) with a length of ten thousand selected from the middle end of the DNA sequencing fragment as 1 character.
[0088] 2), let \(i = 1\) and \(j = 1\); \(i\) can be understood as the starting position for searching, and \(j\) is the length of the token to be searched for.
[0089] 3), Starting from the \(i\)-th character in the sequence \(S''\) and moving backward, search for the adjacent \(j + 1\) characters including the \(i\)-th character in the final vocabulary.
[0090] If the adjacent \(j + 1\) characters can be found in the final vocabulary, determine whether \(i + j - 1 = 10000\) (termination signal because the sequence length is only ten thousand). If it is satisfied, execute 5); if not, execute 4).
[0091] If the adjacent \(j + 1\) characters cannot be found in the final vocabulary, execute 5).
[0092] The approach in this step is to try to find tokens with longer lengths as much as possible. Because in addition to converting the string to numbers, the meaning of encoding can also make the length of the encoded sequence much smaller than the original sequence, improving the subsequent calculation efficiency.
[0093] 4). Let \(j = j + 1\), and continue to execute 3).
[0094] 5). Starting from the \(i\)-th character and moving backward, regard the adjacent \(j\) characters including the \(i\)-th character as a token, and encode it according to the content of the final vocabulary (the corresponding number is whatever number it is), and determine whether \(i + j - 1 = 10000\) (termination signal because the sequence length is only ten thousand). If it is satisfied, end and obtain a string of Token sequences \(\{x''_1, x''_2, \ldots, x''\) m \};
[0095] If it is not satisfied, let \(i = i + j\), \(j = 1\), and execute 3).
[0096] Other steps and parameters are the same as one of the specific embodiments one to three.
[0097] Specific embodiment five: The difference between this embodiment and one of the specific embodiments one to four is that after the DNA sequence \(S'''\) with a length of ten thousand selected from the end of the DNA sequencing fragment is encoded by BPE based on the final vocabulary, a string of Token sequences is obtained:
[0098] \(\{x'''_1, x'''_2, \ldots, x'''\) j , \ldots, x'''\) m \} = BPE(S''')
[0099] The specific process is as follows:
[0100] 1), Regard 1 base in the DNA sequence \(S'''\) with a length of ten thousand selected from the end of the DNA sequencing fragment as 1 character.
[0101] 2), Let \(i = 1\) and \(j = 1\); \(i\) can be understood as the starting position of the search, and \(j\) is the length of the token to be searched.
[0102] 3), Starting from the \(i\)-th character in the sequence \(S'''\), search for the adjacent \(j + 1\) characters including the \(i\)-th character in the final vocabulary.
[0103] If the adjacent \(j + 1\) characters can be found in the final vocabulary, determine whether \(i + j - 1 = 10000\) (termination signal because the sequence length is only ten thousand). If satisfied, execute 5); if not satisfied, execute 4).
[0104] If the adjacent \(j + 1\) characters cannot be found in the final vocabulary, execute 5).
[0105] The approach in this step is to find tokens with longer lengths as much as possible. Because in addition to converting the string to numbers, the meaning of encoding can also make the length of the encoded sequence much smaller than the original sequence, improving the subsequent calculation efficiency.
[0106] 4). Let \(j = j + 1\) and continue to execute 3).
[0107] 5). Starting from the \(i\)-th character, regard the adjacent \(j\) characters including the \(i\)-th character as a token, and encode it according to the content of the final vocabulary (the corresponding number is whatever number it is), and determine whether \(i + j - 1 = 10000\) (termination signal because the sequence length is only ten thousand). If satisfied, end and obtain a string of Token sequences \(\{x'''_1, x'''_2, \ldots, x''' \) m \(\}\);
[0108] If not satisfied, let \(i = i + j\), \(j = 1\), and execute 3).
[0109] Other steps and parameters are the same as those in any one of the specific embodiments one to four.
[0110] Specific Embodiment Six: The difference between this embodiment and any one of the specific embodiments one to five is that in step two, the word vector is input into the pre-trained model DNABert2, and the pre-trained model DNABert2 outputs a feature vector.
[0111] Adopt average pooling to perform dimensionality reduction processing on the feature vector to obtain high-quality reduced-dimensional features.
[0112] The specific process is as follows:
[0113] Step Two One: Input the word vector into the pre-trained model DNABert2, and the pre-trained model DNABert2 outputs a feature vector; the specific process is as follows:
[0114] 1), Input the word vectors {e′1, e′2, …, e′ j , …, e′ m} corresponding to each Token into the pre-trained model DNABert2. The pre-trained model DNABert2 outputs the feature vectors {h′1, h′2, …, h′ j , …, h′ m} corresponding to each Token; It is expressed as:
[0115] {h′1, h′2, …, h′ j , …, h′ m} = DNABert2({e′1, e′2, …, e′ j , …, e′ m})
[0116] 2), Input the word vectors {e″1, e″2, …, e″ j , …, e″ m} corresponding to each Token into the pre-trained model DNABert2. The pre-trained model DNABert2 outputs the feature vectors {h″1, h″2, …, h″ j , …, h″ m} corresponding to each Token; It is expressed as:
[0117] {h″1, h″2, …, h″ j , …, h″ m} = DNABert2({e″1, e″2, …, e″ j , …, e″ m})
[0118] 3), Input the word vectors {e″′1, e″′2, …, e″′ j , …, e″′ m} corresponding to each Token into the pre-trained model DNABert2. The pre-trained model DNABert2 outputs the feature vectors {h″′1, h″′2, …, h″′ j , …, h″′ m} corresponding to each Token; It is expressed as:
[0119] {h″′1, h″′2, …, h″′ j , …, h″′ m} = DNABert2({e″′1, e″′2, …, e″′ j , …, e″′ m})
[0120] Among them,
[0121] $h'_1$ represents the feature vector output by the pre-trained model DNABert2 when the word vector $e'_1$ is input into it; $h'_2$ represents the feature vector output by the pre-trained model DNABert2 when the word vector $e'_2$ is input into it; $h'$ j represents the word vector $e'$ j input into the pre-trained model DNABert2, and the feature vector output by the pre-trained model DNABert2; $h'$ m represents the word vector $e'$ m input into the pre-trained model DNABert2, and the feature vector output by the pre-trained model DNABert2;
[0122] $h''_1$ represents the feature vector output by the pre-trained model DNABert2 when the word vector $e''_1$ is input into it; $h''_2$ represents the feature vector output by the pre-trained model DNABert2 when the word vector $e''_2$ is input into it; $h''$ j represents the word vector $e''$ j input into the pre-trained model DNABert2, and the feature vector output by the pre-trained model DNABert2; $h''$ m represents the word vector $e''$ m input into the pre-trained model DNABert2, and the feature vector output by the pre-trained model DNABert2;
[0123] $h'''_1$ represents the feature vector output by the pre-trained model DNABert2 when the word vector $e'''_1$ is input into it; $h'''_2$ represents the feature vector output by the pre-trained model DNABert2 when the word vector $e'''_2$ is input into it; $h'''$ j represents the word vector $e'''$ j input into the pre-trained model DNABert2, and the feature vector output by the pre-trained model DNABert2; $h'''$ m represents the word vector $e'''$ m input into the pre-trained model DNABert2, and the feature vector output by the pre-trained model DNABert2;
[0124] Step Two: Use average pooling to reduce the dimension of the feature vector to obtain high-quality reduced-dimensional features; the specific process is as follows:
[0125] For the next chromosome classification prediction task, it is necessary to reduce the dimension of the output word vector; use the average pooling method, and the specific approach is to take the mean of all feature vectors to obtain the overall feature vector of the selected region sequence;
[0126] 1), Use average pooling to perform dimensionality reduction on the feature vectors {h′1, h′2, …, h′ j , …, h′ m} to obtain high-quality dimension-reduced features O′; expressed as:
[0127]
[0128] 2), Use average pooling to perform dimensionality reduction on the feature vectors {h″1, h″2, …, h″ j , …, h″ m} to obtain high-quality dimension-reduced features O″; expressed as:
[0129]
[0130] 3), Use average pooling to perform dimensionality reduction on the feature vectors {h″′1, h″′2, …, h″′ j , …, h″′ m} to obtain high-quality dimension-reduced features O″′; expressed as:
[0131]
[0132] Other steps and parameters are the same as those in any one of the first to fifth specific embodiments.
[0133] Specific Embodiment Seven: The difference between this embodiment and any one of the first to sixth specific embodiments is that in step three, a deep learning classification model ( Figure 2 ) is constructed;
[0134] Based on the training set, the deep learning classification model is trained to obtain a trained deep learning classification model;
[0135] Input the dimension-reduced features in step two into the trained deep learning classification model, and the trained deep learning classification model predicts the chromosome category to which the sequence corresponding to the dimension-reduced features in step two belongs;
[0136] Record the category results, and the DNA sequences corresponding to each category are stored in one fasta format file;
[0137] The specific process is as follows:
[0138] Step Three One, construct a deep learning classification model ( Figure 2 ); the specific process is as follows:
[0139] The deep learning classification model sequentially includes:
[0140] Input layer, first module, second module, third module, fourth module, fifth module, twenty - first fully - connected layer Linear, eleventh Dropout layer, sixteenth LeakyReLU activation function layer, twenty - second fully - connected layer Linear, twelfth Dropout layer, seventeenth LeakyReLU activation function layer, twenty - third fully - connected layer Linear, output layer;
[0141] The first module includes:
[0142] First BN layer, first fully - connected layer Linear, first LeakyReLU activation function, first Dropout layer, second fully - connected layer Linear, second LeakyReLU activation function, second Dropout layer, third fully - connected layer Linear, fourth fully - connected layer Linear, third LeakyReLU activation function;
[0143] The second module includes:
[0144] Second BN layer, fifth fully - connected layer Linear, fourth LeakyReLU activation function, third Dropout layer, sixth fully - connected layer Linear, fifth LeakyReLU activation function, fourth Dropout layer, seventh fully - connected layer Linear, eighth fully - connected layer Linear, sixth LeakyReLU activation function;
[0145] The third module includes:
[0146] Third BN layer, ninth fully - connected layer Linear, seventh LeakyReLU activation function, fifth Dropout layer, tenth fully - connected layer Linear, eighth LeakyReLU activation function, sixth Dropout layer, eleventh fully - connected layer Linear, twelfth fully - connected layer Linear, ninth LeakyReLU activation function;
[0147] The fourth module includes:
[0148] Fourth BN layer, thirteenth fully - connected layer Linear, tenth LeakyReLU activation function, seventh Dropout layer, fourteenth fully - connected layer Linear, eleventh LeakyReLU activation function, eighth Dropout layer, fifteenth fully - connected layer Linear, sixteenth fully - connected layer Linear, twelfth LeakyReLU activation function;
[0149] The fifth module includes:
[0150] The fifth BN layer, the seventeenth fully connected layer Linear, the thirteenth LeakyReLU activation function, the ninth Dropout layer, the eighteenth fully connected layer Linear, the fourteenth LeakyReLU activation function, the tenth Dropout layer, the nineteenth fully connected layer Linear, the twentieth fully connected layer Linear, the fifteenth LeakyReLU activation function;
[0151] Step 3-2: Train the deep learning classification model based on the training set to obtain a trained deep learning classification model;
[0152] Step 3-3: Input the features after dimensionality reduction in Step 2 into the trained deep learning classification model, and the trained deep learning classification model predicts the chromosome category to which the sequence corresponding to the features after dimensionality reduction in Step 2 belongs;
[0153] Step 3-4: Record the category results, and the DNA sequences corresponding to each category are stored in 1 fasta format file.
[0154] After completing sequence encoding and using the DNABert2 model to extract the feature vectors of the sequences, this method constructs a feedforward neural network (DNAMatch) based on the deep residual structure for chromosome category prediction. The backbone network consists of multiple consecutive residual modules, and each residual module is composed of three fully connected layers, with the following characteristics:
[0155] Layer-by-layer feature refinement: The input and output dimensions of each residual module are the same. Inside the module, a linear transformation that first expands to 3072 dimensions and then returns to 768 dimensions is adopted to strengthen the non-linear expression of the features, and along with the LeakyReLU non-linear activation function, the feature expression ability is further increased.
[0156] Residual connection: Prevent the problem of gradient disappearance when the number of network layers deepens and improve the feature propagation ability.
[0157] Generally speaking, the model first gradually refines the features through multiple residual modules. The modules are all composed of linear transformation, LeakyReLU activation function and Dropout layer, which enhances the generalization ability of the network and reduces the problem of gradient disappearance. Finally, the prediction result is obtained through multi-layer linear dimensionality reduction operation, realizing the accurate classification of the chromosome to which the DNA fragment belongs.
[0158] Other steps and parameters are the same as those in any one of the specific embodiments 1 to 6.
[0159] Specific embodiment 8: The difference between this embodiment and any one of the specific embodiments 1 to 7 is that in Step 3-2, the deep learning classification model is trained based on the training set to obtain a trained deep learning classification model; the specific process is as follows:
[0160] 1), Use the simlord tool to simulate and generate DNA sequencing fragments of length 10,000 from the human reference genome for model training;
[0161] 2), After BPE encoding the simulated DNA sequencing fragments of length 10,000 based on the final vocabulary, obtain a sequence;
[0162] 3), Map the sequence obtained after BPE encoding to the corresponding word vector embedding space to obtain the word vectors of the sequence;
[0163] 4), Input the word vectors of the sequence obtained in 3) into the pre-trained model DNABert2, and the pre-trained model DNABert2 outputs feature vectors;
[0164] 5), Use average pooling to perform dimensionality reduction on the feature vectors output by the pre-trained model DNABert2 to obtain high-quality dimensionality-reduced features as the training set;
[0165] 6), Use the dimensionality-reduced features (with labels) as the input of the deep learning classification model, and the chromosome category to which the DNA sequence corresponding to the dimensionality-reduced features belongs as the output of the deep learning classification model, and train the deep learning classification model to obtain a trained deep learning classification model; The specific process is as follows:
[0166] 61), Input the dimensionality-reduced features into the first module through the input layer, and the first module outputs feature B'; The specific process is as follows:
[0167] The dimensionality-reduced features are successively input into the first BN layer, the first fully connected layer Linear, the first LeakyReLU activation function, the first Dropout layer, the second fully connected layer Linear, the second LeakyReLU activation function, the second Dropout layer, and the third fully connected layer Linear through the input layer. The third fully connected layer Linear outputs feature A'. After adding the feature A' output by the third fully connected layer Linear and the dimensionality-reduced features element by element, they are successively input into the fourth fully connected layer Linear and the third LeakyReLU activation function. The third LeakyReLU activation function outputs feature B', and feature B' is used as the output feature of the first module;
[0168] 62), Input the output feature B' of the first module into the second module, and the second module outputs feature D'; The specific process is as follows:
[0169] The output feature B' of the first module is sequentially input into the second BN layer, the fifth fully connected layer Linear, the fourth LeakyReLU activation function, the third Dropout layer, the sixth fully connected layer Linear, the fifth LeakyReLU activation function, the fourth Dropout layer, and the seventh fully connected layer Linear. The seventh fully connected layer Linear outputs the feature C'. After adding the output feature C' of the seventh fully connected layer Linear and the output feature B' of the first module element-wise, they are sequentially input into the eighth fully connected layer Linear and the sixth LeakyReLU activation function. The sixth LeakyReLU activation function outputs the feature D', and the feature D' is used as the output feature of the second module;
[0170] 63), Input the output feature D' of the second module into the third module, and the third module outputs the feature F'; The specific process is as follows:
[0171] The output feature D' of the second module is sequentially input into the third BN layer, the ninth fully connected layer Linear, the seventh LeakyReLU activation function, the fifth Dropout layer, the tenth fully connected layer Linear, the eighth LeakyReLU activation function, the sixth Dropout layer, and the eleventh fully connected layer Linear. The eleventh fully connected layer Linear outputs the feature E'. After adding the output feature E' of the eleventh fully connected layer Linear and the output feature D' of the second module element-wise, they are sequentially input into the twelfth fully connected layer Linear and the ninth LeakyReLU activation function. The ninth LeakyReLU activation function outputs the feature F', and the feature F' is used as the output feature of the third module;
[0172] 64), Input the output feature F' of the third module into the fourth module, and the fourth module outputs the feature H'; The specific process is as follows:
[0173] The output feature F' of the third module is sequentially input into the fourth BN layer, the thirteenth fully connected layer Linear, the tenth LeakyReLU activation function, the seventh Dropout layer, the fourteenth fully connected layer Linear, the eleventh LeakyReLU activation function, the eighth Dropout layer, and the fifteenth fully connected layer Linear. The fifteenth fully connected layer Linear outputs the feature G'. After adding the output feature G' of the fifteenth fully connected layer Linear and the output feature F' of the third module element-wise, they are sequentially input into the sixteenth fully connected layer Linear and the twelfth LeakyReLU activation function. The twelfth LeakyReLU activation function outputs the feature H', and the feature H' is used as the output feature of the fourth module;
[0174] 65), Input the output feature H' of the fourth module into the fifth module, and the fifth module outputs the feature J'; The specific process is as follows:
[0175] The output feature H' of the fourth module is successively input into the fifth BN layer, the seventeenth fully connected layer Linear, the thirteenth LeakyReLU activation function, the ninth Dropout layer, the eighteenth fully connected layer Linear, the fourteenth LeakyReLU activation function, the tenth Dropout layer, and the nineteenth fully connected layer Linear. The nineteenth fully connected layer Linear outputs the feature I'. After the feature I' output by the nineteenth fully connected layer Linear and the output feature H' of the fourth module are added element-wise, they are successively input into the twentieth fully connected layer Linear and the fifteenth LeakyReLU activation function. The fifteenth LeakyReLU activation function outputs the feature J', and the feature J' is used as the output feature of the fifth module;
[0176] 66) The output feature J' of the fifth module is successively input into the twenty-first fully connected layer Linear, the eleventh Dropout layer, the sixteenth LeakyReLU activation function layer, the twenty-second fully connected layer Linear, the twelfth Dropout layer, the seventeenth LeakyReLU activation function layer, and the twenty-third fully connected layer Linear. The twenty-third fully connected layer Linear outputs the chromosome category to which the DNA sequence corresponding to the predicted dimensionality-reduced feature belongs, and the category result is output through the output layer;
[0177] The deep learning classification model is trained to obtain a trained deep learning classification model.
[0178] Other steps and parameters are the same as those in any one of the specific embodiments one to seven.
[0179] Specific Embodiment Nine: The difference between this embodiment and any one of the specific embodiments one to eight is that in step 33, the dimensionality-reduced feature obtained in step 2 is input into the trained deep learning classification model, and the trained deep learning classification model predicts the chromosome category (24 categories) to which the sequence corresponding to the dimensionality-reduced feature in step 2 belongs; the specific process is as follows:
[0180] 1) The dimensionality-reduced feature obtained in step 2 is input into the first module through the input layer, and the first module outputs the feature B''; the specific process is as follows:
[0181] The features after dimensionality reduction are input into the first BN layer, the first fully connected layer Linear, the first LeakyReLU activation function, the first Dropout layer, the second fully connected layer Linear, the second LeakyReLU activation function, the second Dropout layer, and the third fully connected layer Linear in sequence through the input layer. The third fully connected layer Linear outputs feature A″. After adding the output feature A″ of the third fully connected layer Linear and the features after dimensionality reduction element by element, they are input into the fourth fully connected layer Linear and the third LeakyReLU activation function in sequence. The third LeakyReLU activation function outputs feature B″, and feature B″ is used as the output feature of the first module;
[0182] 2) Input the output feature B″ of the first module into the second module, and the second module outputs feature D″; The specific process is as follows:
[0183] The output feature B″ of the first module is input into the second BN layer, the fifth fully connected layer Linear, the fourth LeakyReLU activation function, the third Dropout layer, the sixth fully connected layer Linear, the fifth LeakyReLU activation function, the fourth Dropout layer, and the seventh fully connected layer Linear in sequence. The seventh fully connected layer Linear outputs feature C″. After adding the output feature C″ of the seventh fully connected layer Linear and the output feature B″ of the first module element by element, they are input into the eighth fully connected layer Linear and the sixth LeakyReLU activation function in sequence. The sixth LeakyReLU activation function outputs feature D″, and feature D″ is used as the output feature of the second module;
[0184] 3) Input the output feature D″ of the second module into the third module, and the third module outputs feature F″; The specific process is as follows:
[0185] The output feature D″ of the second module is input into the third BN layer, the ninth fully connected layer Linear, the seventh LeakyReLU activation function, the fifth Dropout layer, the tenth fully connected layer Linear, the eighth LeakyReLU activation function, the sixth Dropout layer, and the eleventh fully connected layer Linear in sequence. The eleventh fully connected layer Linear outputs feature E″. After adding the output feature E″ of the eleventh fully connected layer Linear and the output feature D″ of the second module element by element, they are input into the twelfth fully connected layer Linear and the ninth LeakyReLU activation function in sequence. The ninth LeakyReLU activation function outputs feature F″, and feature F″ is used as the output feature of the third module;
[0186] 4) Input the output feature F″ of the third module into the fourth module, and the fourth module outputs feature H″; The specific process is as follows:
[0187] The output feature F″ of the third module is sequentially input into the fourth BN layer, the thirteenth fully connected layer Linear, the tenth LeakyReLU activation function, the seventh Dropout layer, the fourteenth fully connected layer Linear, the eleventh LeakyReLU activation function, the eighth Dropout layer, and the fifteenth fully connected layer Linear. The fifteenth fully connected layer Linear outputs the feature G″. After adding the output feature G″ of the fifteenth fully connected layer Linear and the output feature F″ of the third module element-wise, they are sequentially input into the sixteenth fully connected layer Linear and the twelfth LeakyReLU activation function. The twelfth LeakyReLU activation function outputs the feature H″, and the feature H″ is used as the output feature of the fourth module;
[0188] 5), Input the output feature H″ of the fourth module into the fifth module, and the fifth module outputs the feature J″; The specific process is as follows:
[0189] The output feature H″ of the fourth module is sequentially input into the fifth BN layer, the seventeenth fully connected layer Linear, the thirteenth LeakyReLU activation function, the ninth Dropout layer, the eighteenth fully connected layer Linear, the fourteenth LeakyReLU activation function, the tenth Dropout layer, and the nineteenth fully connected layer Linear. The nineteenth fully connected layer Linear outputs the feature I″. After adding the output feature I″ of the nineteenth fully connected layer Linear and the output feature H″ of the fourth module element-wise, they are sequentially input into the twentieth fully connected layer Linear and the fifteenth LeakyReLU activation function. The fifteenth LeakyReLU activation function outputs the feature J″, and the feature J″ is used as the output feature of the fifth module;
[0190] 6), Input the output feature J″ of the fifth module into the twenty-first fully connected layer Linear, the eleventh Dropout layer, the sixteenth LeakyReLU activation function layer, the twenty-second fully connected layer Linear, the twelfth Dropout layer, the seventeenth LeakyReLU activation function layer, and the twenty-third fully connected layer Linear. The twenty-third fully connected layer Linear outputs the chromosome category to which the DNA sequence corresponding to the dimension-reduced feature in step two of the prediction belongs, and the category result is output through the output layer.
[0191] Other steps and parameters are the same as those in any one of the first to eighth specific embodiments.
[0192] Specific embodiment ten: The difference between this embodiment and any one of the first to ninth specific embodiments is that in step four, the sequence of the record category in each file is compared with the DNA sequence of the corresponding chromosome in the reference sequence, and the comparison result is saved; The reference sequence is the human genome sequence; The specific process is as follows:
[0193] Use the minimap2 tool to align the DNA sequences stored in each file with the corresponding chromosomal DNA sequences in the reference sequence to obtain the alignment results of the DNA sequences stored in each file;
[0194] The alignment results include: the alignment position of the sequence, the matching length, the base mismatch or insertion / deletion situation, the alignment score, and the alignment quality;
[0195] Save the alignment results as a sam format file;
[0196] The DNA sequences stored in each file are the DNA sequences of the known chromosomal categories predicted in step three three.
[0197] After using the DNABert2 model and the deep learning classification model (DNAMatch) to predict the chromosomal categories of DNA sequencing fragments, further use the minimap2 tool to perform accurate sequence alignment on the predicted sequences and save the alignment results. Specifically, first, according to the chromosomal classification prediction results of the aforementioned DNA sequencing fragments, sort and classify different DNA sequencing fragments according to the predicted chromosomal categories and save them as corresponding fasta format files. Subsequently, for each prediction label, select the corresponding reference chromosomal sequence from the complete reference genome as the object for subsequent alignment. In addition, for DNA sequencing fragments that are difficult to predict, alignment is still performed across the entire genome.
[0198] Then, in the sequence alignment stage, call minimap2 to perform sequence alignment on each classified fasta file respectively, and detailed alignment results can be obtained. The minimap2 output results are saved in the standardized SAM format, which is convenient for compatible calling and data analysis by various bioinformatics analysis software. The result file usually contains the following key information: the alignment position of the sequence, the matching length, the base mismatch or insertion / deletion situation, the alignment score, and the alignment quality, etc.
[0199] Other steps and parameters are the same as those in any one of the specific embodiments one to nine.
[0200] The present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A DNA sequence alignment method based on deep learning, characterized in that: The specific process of the method is as follows: Step 1: Convert a DNA sequencing sequence of one million in length from a string form into word vectors based on BPE encoding; Step 2: Input the word vectors into the pre-trained model DNABert2, and the pre-trained model DNABert2 outputs feature vectors; Perform dimensionality reduction processing on the feature vectors by using average pooling to obtain the dimensionality-reduced features; Step 3: Construct a deep learning classification model; Train the deep learning classification model based on the training set to obtain a trained deep learning classification model; Input the dimensionality-reduced features in Step 2 into the trained deep learning classification model, and the trained deep learning classification model predicts the chromosome category to which the sequence corresponding to the dimensionality-reduced features in Step 2 belongs; Record the category results, and the DNA sequences corresponding to each category are stored in one fasta format file; Step 4: Align the sequences recording the categories in each file with the DNA sequences of the corresponding chromosomes in the reference sequence, and save the alignment results; The reference sequence is the human genome sequence.
2. The DNA sequence alignment method based on deep learning according to claim 1, wherein: In Step 1, the DNA sequencing sequence of one million in length is converted from a string form into word vectors based on BPE encoding; the specific process is as follows: Step 1: Use the human genome sequence set D = {S1, S2, …, S i , …, S n} as the corpus, where i = 1, 2, …, n; Steps 1 and 2: Treat each base in the $i$-th sequence $S$ in the sequence set $D$ as an independent Token; i Step 13. Calculate the occurrence frequency of each Token pair (x, y) in the i-th sequence S in the corpus i ; Each Token pair (x, y) is composed of adjacent Tokens, and there is no overlap between Token pairs (x, y); Step 14: Repeat Step 13 until the occurrence frequency of each Token pair (x, y) in all n sequences in the sequence set D is obtained; Step 15: Add the occurrence frequencies corresponding to the same Token pair (x, y) in the n sequences to obtain the total occurrence frequency of the same Token pair (x, y) in the n sequences; expressed as: where count((x,y),S i ) represents the occurrence frequency corresponding to the same Token pair (x,y) in the i-th sequence S i ; Step 16: Merge the Token pair (x, y) corresponding to the maximum total occurrence frequency in each sequence to obtain a new Token, and the new Token is denoted as z = (x, y), and update the vocabulary; Step 17: Repeat Steps 13 to 16 M times to obtain all the final Tokens, number all the final Tokens, and the numbers are not repeated; All the final Tokens and the number corresponding to each Token are used as the final vocabulary; M is 4088; Step 18: For a DNA sequencing fragment of one million in length, select a DNA sequence S′ of ten thousand in length from the head end of the DNA sequencing fragment backward, select a DNA sequence S″ of ten thousand in length from the middle end of the DNA sequencing fragment backward, and select a DNA sequence S″′ of ten thousand in length from the tail end of the DNA sequencing fragment forward; Based on the final vocabulary, after BPE encoding the DNA sequence S′ of ten thousand in length selected from the head end of the DNA sequencing fragment backward, a string of Token sequences is obtained: {x′1,x′2,…,x′ j ,…,x′ m} = BPE(S′) Based on the final vocabulary, after BPE encoding the DNA sequence S″ of ten thousand in length selected from the middle end of the DNA sequencing fragment backward, a string of Token sequences is obtained: {x″1,x″2,…,x″ j ,…,x″ m} = BPE(S″) Based on the final vocabulary, after BPE encoding the DNA sequence S″′ of ten thousand in length selected from the tail end of the DNA sequencing fragment forward, a string of Token sequences is obtained: {x‴1, x‴2, …, x‴ j , …, x‴ m} = BPE(S‴) Wherein, m represents the number of Tokens after encoding; {x′1, x′2, …, x′ j , …, x′ m} represents a sequence of Tokens obtained by BPE encoding a DNA sequence of length 10,000 selected from the head of a DNA sequencing fragment; x′1 represents the first Token obtained after BPE encoding the first Token in the DNA sequence of length 10,000; x′2 represents the second Token obtained after BPE encoding the second Token in the DNA sequence of length 10,000; x′ j represents the j-th Token obtained after BPE encoding the j-th Token in the DNA sequence of length 10,000; x′ m represents the m-th Token obtained after BPE encoding the m-th Token in the DNA sequence of length 10,000; {x″1, x″2, …, x″ j , …, x″ m} represents a sequence of tokens obtained by BPE encoding a DNA sequence of length ten thousand selected from the middle end to the back of a DNA sequencing fragment; x″1 represents the first token obtained after BPE encoding the first token in the DNA sequence of length ten thousand; x″2 represents the second token obtained after BPE encoding the second token in the DNA sequence of length ten thousand; x″ j represents the j-th token obtained after BPE encoding the j-th token in the DNA sequence of length ten thousand; x″ m represents the m-th token obtained after BPE encoding the m-th token in the DNA sequence of length ten thousand; {x″′1, x″′2, …, x″′ j , …, x″′ m} represents a sequence of Tokens obtained by BPE encoding a DNA sequence of length ten thousand selected from the end of a DNA sequencing fragment towards the front; x″′1 represents the first Token obtained after BPE encoding the first Token in the DNA sequence of length ten thousand; x″′2 represents the second Token obtained after BPE encoding the second Token in the DNA sequence of length ten thousand; x″′ j represents the j-th Token obtained after BPE encoding the j-th Token in the DNA sequence of length ten thousand; x″′ m represents the m-th Token obtained after BPE encoding the m-th Token in the DNA sequence of length ten thousand; Step 19. Map a string of Token sequences obtained after BPE encoding to the corresponding word vector embedding space to obtain the word vectors of the Token sequences, which is expressed as: {e′1,e′2,…,e′ j ,…,e′ m} = Embed{x′1,x′2,…,x′ j ,…,x′ m} {e″1,e″2,…,e″ j ,…,e″ m} = Embed{x″1,x″2,…,x″ j ,…,x″ m} {e″′1, e″′2, …, e″′ j , …, e″′ m} = Embed{x″′1, x″′2, …, x″′ j , …, x″′ m} where, {e′1, e′2, …, e′ j , …, e′ m} represents mapping {x′1, x′2, …, x′ j , …, x′ m} to the corresponding word vector embedding space to obtain the word vectors corresponding to each Token; e′1 represents the word vector corresponding to the first Token obtained by mapping x′1 to the corresponding word vector embedding space; e′2 represents the word vector corresponding to the second Token obtained by mapping x′2 to the corresponding word vector embedding space; e′ j represents the word vector corresponding to the j-th Token obtained by mapping x′ j to the corresponding word vector embedding space; e′ m represents the word vector corresponding to the m-th Token obtained by mapping x′ m to the corresponding word vector embedding space; {e″1, e″2, …, e″ j , …, e″ m} represents mapping {x″1, x″2, …, x″ j , …, x″ m} to the corresponding word vector embedding space to obtain the word vectors corresponding to each Token; e″1 represents the word vector corresponding to the 1st Token obtained by mapping x″1 to the corresponding word vector embedding space; e″2 represents the word vector corresponding to the 2nd Token obtained by mapping x″2 to the corresponding word vector embedding space; e″ j represents the word vector corresponding to the jth Token obtained by mapping x″ j to the corresponding word vector embedding space; e″ m represents the word vector corresponding to the mth Token obtained by mapping x″ m to the corresponding word vector embedding space; {e″′1,e″′2,…,e″′ j ,…,e″′ m} represents mapping {x″′1,x″′2,…,x″′ j ,…,x″′ m} to the corresponding word vector embedding space, and the word vectors corresponding to each Token obtained; e″′1 represents the word vector corresponding to the first Token obtained by mapping x″′1 to the corresponding word vector embedding space; e″′2 represents the word vector corresponding to the second Token obtained by mapping x″′2 to the corresponding word vector embedding space; e″′ j represents the word vector corresponding to the j-th Token obtained by mapping x″′ j to the corresponding word vector embedding space; x″′ m represents the word vector corresponding to the m-th Token obtained by mapping x″′ m to the corresponding word vector embedding space; Embed represents word embedding.
3. The method for DNA sequence alignment based on deep learning according to claim 2, characterized in that: After BPE encoding the DNA sequence S' with a length of 10,000 selected from the head end of the DNA sequencing fragment backward based on the final vocabulary, a string of Token sequences is obtained: {x′1,x′2,…,x′ j ,…,x′ m} = BPE(S′) The specific process is as follows: 1). Regard 1 base in the DNA sequence S' with a length of 10,000 selected from the head end of the DNA sequencing fragment backward as 1 character; 2). Let i = 1 and j = 1; 3). Starting from the i-th character in the sequence S', search for the adjacent j + 1 characters including the i-th character in the final vocabulary; If the adjacent j + 1 characters can be found in the final vocabulary, judge whether i + j - 1 = 10000 is satisfied. If it is satisfied, execute 5); if not, execute 4); If the adjacent j + 1 characters cannot be found in the final vocabulary, execute 5); 4). Let j = j + 1 and continue to execute 3); 5). Starting from the $i$-th character and moving backward, consider adjacent $j$ characters including the $i$-th character as a token, and encode it according to the content of the final vocabulary. Determine whether $i + j - 1 = 10000$ is satisfied. If it is satisfied, end the process to obtain a sequence of tokens $\{x_1', x_2', \ldots, x' m}\; If not, let i = i + j, j = 1, and execute 3).
4. A DNA sequence alignment method based on deep learning according to claim 3, characterized in that: After BPE encoding the DNA sequence S'' with a length of 10,000 selected from the middle end of the DNA sequencing fragment backward based on the final vocabulary, a string of Token sequences is obtained: {x″1,x″2,…,x″ j ,…,x″ m} = BPE(S″) The specific process is as follows: 1). Regard 1 base in the DNA sequence S'' with a length of 10,000 selected from the middle end of the DNA sequencing fragment backward as 1 character; 2). Let i = 1 and j = 1; 3). Starting from the i-th character in the sequence S'', search for the adjacent j + 1 characters including the i-th character in the final vocabulary; If the adjacent j + 1 characters can be found in the final vocabulary, judge whether i + j - 1 = 10000 is satisfied. If it is satisfied, execute 5); if not, execute 4); If the adjacent j + 1 characters cannot be found in the final vocabulary, execute 5); 4). Let j = j + 1 and continue to execute 3); 5). Starting from the i-th character and moving backward, consider adjacent j characters including the i-th character as a token, and encode it according to the content of the final vocabulary. Determine whether i + j - 1 = 10000 is satisfied. If it is satisfied, end to obtain a sequence of tokens {x″1, x″2, …, x″ m}; If not, let i = i + j, j = 1, and execute 3).
5. The DNA sequence alignment method based on deep learning according to claim 4, characterized in that: After BPE encoding the DNA sequence S''' with a length of 10,000 selected from the tail end of the DNA sequencing fragment forward based on the final vocabulary, a string of Token sequences is obtained: {x″′1, x″′2, …, x″′ j , …, x″′ m} = BPE(S″′) The specific process is as follows: 1). Regard 1 base in the DNA sequence S''' with a length of 10,000 selected from the tail end of the DNA sequencing fragment forward as 1 character; 2). Let i = 1 and j = 1; 3). Starting from the i-th character in the sequence S''', search for the adjacent j + 1 characters including the i-th character in the final vocabulary; If the adjacent j + 1 characters can be found in the final vocabulary, judge whether i + j - 1 = 10000 is satisfied. If it is satisfied, execute 5); if not, execute 4); If the adjacent j + 1 characters cannot be found in the final vocabulary, execute 5); 4). Let j = j + 1 and continue to execute 3); 5). Starting from the $i$-th character and moving backward, consider $j$ adjacent characters including the $i$-th character as a token, and encode it according to the content of the final vocabulary. Determine whether $i + j - 1 = 10000$ is satisfied. If it is satisfied, end the process to obtain a sequence of tokens $\{x'''_1, x'''_2, \ldots, x'''_{ m}\}$; m} If not, let i = i + j, j = 1, and execute 3).
6. The DNA sequence alignment method based on deep learning according to claim 5, wherein: In Step 2, input the word vectors into the pre-trained model DNABert2, and the pre-trained model DNABert2 outputs feature vectors; The feature vector is dimensionally reduced by average pooling to obtain the feature after dimensional reduction; The specific process is as follows: Step 2-1: Input the word vector into the pre-trained model DNABert2, and the pre-trained model DNABert2 outputs the feature vector; the specific process is as follows: 1), Input the word vectors {e′1, e′2, …, e′ j , …, e′ m} corresponding to each Token into the pre-trained model DNABert2, and the pre-trained model DNABert2 outputs the feature vectors {h′1, h′2, …, h′ j , …, h′ m} corresponding to each Token; It is expressed as: {h′1,h′2,…,h′ j ,…,h′ m} = DNABert2({e′1,e′2,…,e′ j ,…,e′ m}) 2), input the word vectors {e″1, e″2, …, e″ j , …, e″ m} corresponding to each Token into the pre-trained model DNABert2, and the pre-trained model DNABert2 outputs the feature vectors {h″1, h″2, …, h″ j , …, h″ m} corresponding to each Token; It is expressed as: {h″1, h″2, …, h″ j , …, h″ m} = DNABert2({e″1, e″2, …, e″ j , …, e″ m}) 3), Input the word vectors {e″′1, e″′2, …, e″′ j , …, e″′ m} corresponding to each Token into the pre-trained model DNABert2, and the pre-trained model DNABert2 outputs the feature vectors {h″′1, h″′2, …, h″′ j , …, h″′ m} corresponding to each Token; It is expressed as: {h″′1, h″′2, …, h″′ j , …, h″′ m} = DNABert2({e″′1, e″′2, …, e″′ j , …, e″′ m}) Among them, $h'_1$ represents the feature vector output by the pre-trained model DNABert2 when the word vector $e'_1$ is input into it; $h'_2$ represents the feature vector output by the pre-trained model DNABert2 when the word vector $e'_2$ is input into it; $h'$ j represents the word vector $e'$ j input into the pre-trained model DNABert2, and the feature vector output by the pre-trained model DNABert2; $h'$ m represents the word vector $e'$ m input into the pre-trained model DNABert2, and the feature vector output by the pre-trained model DNABert2; $h''_1$ represents the feature vector output by the pre-trained model DNABert2 when the word vector $e''_1$ is input into it; $h''_2$ represents the feature vector output by the pre-trained model DNABert2 when the word vector $e''_2$ is input into it; $h''$ j represents the word vector $e''$ j input into the pre-trained model DNABert2, and the feature vector output by the pre-trained model DNABert2; $h''$ m represents the word vector $e''$ m input into the pre-trained model DNABert2, and the feature vector output by the pre-trained model DNABert2; $h'''_1$ represents the feature vector output by the pre-trained model DNABert2 when the word vector $e'''_1$ is input into the pre-trained model DNABert2; $h'''_2$ represents the feature vector output by the pre-trained model DNABert2 when the word vector $e'''_2$ is input into the pre-trained model DNABert2; $h'''$ j represents the word vector $e'''$ j input into the pre-trained model DNABert2, and the feature vector output by the pre-trained model DNABert2; $h'''$ m represents the word vector $e'''$ m input into the pre-trained model DNABert2, and the feature vector output by the pre-trained model DNABert2; Step 2-2: The feature vector is dimensionally reduced by average pooling to obtain the feature after dimensional reduction; The specific process is as follows: 1) Use average pooling to perform dimensionality reduction on the feature vectors {h′1, h′2, …, h′ m} to obtain the feature O′ after dimensionality reduction, which is expressed as: 2) Use average pooling to perform dimensionality reduction on the feature vectors {h″1, h″2, …, h″ m}, and obtain the feature O" after dimensionality reduction, which is expressed as: 3), average pooling is used to perform dimensionality reduction on the feature vectors {h″′1, h″′2, …, h″′ m}, and the feature O″′ after dimensionality reduction is obtained, which is expressed as:
7. The DNA sequence alignment method based on deep learning according to claim 6, wherein: In step 3, a deep learning classification model is constructed; The deep learning classification model is trained based on the training set to obtain a trained deep learning classification model; The feature after dimensional reduction in step 2 is input into the trained deep learning classification model, and the trained deep learning classification model predicts the chromosome category to which the sequence corresponding to the feature after dimensional reduction in step 2 belongs; Record the category results, and the DNA sequences corresponding to each category are stored in one fasta format file; The specific process is as follows: Step 3-1: Construct a deep learning classification model; The specific process is as follows: The deep learning classification model successively includes: Input layer, first module, second module, third module, fourth module, fifth module, twenty-first fully connected layer Linear, eleventh Dropout layer, sixteenth LeakyReLU activation function layer, twenty-second fully connected layer Linear, twelfth Dropout layer, seventeenth LeakyReLU activation function layer, twenty-third fully connected layer Linear, output layer; The first module includes: First BN layer, first fully connected layer Linear, first LeakyReLU activation function, first Dropout layer, second fully connected layer Linear, second LeakyReLU activation function, second Dropout layer, third fully connected layer Linear, fourth fully connected layer Linear, third LeakyReLU activation function; The second module includes: Second BN layer, fifth fully connected layer Linear, fourth LeakyReLU activation function, third Dropout layer, sixth fully connected layer Linear, fifth LeakyReLU activation function, fourth Dropout layer, seventh fully connected layer Linear, eighth fully connected layer Linear, sixth LeakyReLU activation function; The third module includes: Third BN layer, ninth fully connected layer Linear, seventh LeakyReLU activation function, fifth Dropout layer, tenth fully connected layer Linear, eighth LeakyReLU activation function, sixth Dropout layer, eleventh fully connected layer Linear, twelfth fully connected layer Linear, ninth LeakyReLU activation function; The fourth module includes: The fourth BN layer, the thirteenth fully connected layer Linear, the tenth LeakyReLU activation function, the seventh Dropout layer, the fourteenth fully connected layer Linear, the eleventh LeakyReLU activation function, the eighth Dropout layer, the fifteenth fully connected layer Linear, the sixteenth fully connected layer Linear, the twelfth LeakyReLU activation function; The fifth module includes: The fifth BN layer, the seventeenth fully connected layer Linear, the thirteenth LeakyReLU activation function, the ninth Dropout layer, the eighteenth fully connected layer Linear, the fourteenth LeakyReLU activation function, the tenth Dropout layer, the nineteenth fully connected layer Linear, the twentieth fully connected layer Linear, the fifteenth LeakyReLU activation function; Step 32: Train the deep learning classification model based on the training set to obtain a trained deep learning classification model; Step 33: Input the features after dimensionality reduction in Step 2 into the trained deep learning classification model, and the trained deep learning classification model predicts the chromosome category to which the sequence corresponding to the features after dimensionality reduction in Step 2 belongs; Step 34: Record the category results, and the DNA sequences corresponding to each category are stored in one fasta format file.
8. The method for DNA sequence alignment based on deep learning according to claim 7, wherein: In Step 32, the deep learning classification model is trained based on the training set to obtain a trained deep learning classification model; the specific process is as follows: 1) Use the simlord tool to simulate and generate DNA sequencing fragments with a length of 10,000 for the human reference genome; 2) Based on the final vocabulary, the simulated DNA sequencing fragments with a length of 10,000 are encoded by BPE to obtain a string of sequences; 3) Map the string of sequences obtained after BPE encoding to the corresponding word vector embedding space to obtain the word vectors of the sequences; 4) Input the word vectors of the sequences obtained in 3) into the pre-trained model DNABert2, and the pre-trained model DNABert2 outputs feature vectors; 5) Use average pooling to perform dimensionality reduction processing on the feature vectors output by the pre-trained model DNABert2 to obtain the features after dimensionality reduction as the training set; 6) Use the features after dimensionality reduction as the input of the deep learning classification model, and the chromosome category to which the DNA sequence corresponding to the features after dimensionality reduction belongs as the output of the deep learning classification model to train the deep learning classification model to obtain a trained deep learning classification model; the specific process is as follows: 61) Input the features after dimensionality reduction into the first module through the input layer, and the first module outputs feature B′; the specific process is as follows: The features after dimensionality reduction are sequentially input into the first BN layer, the first fully-connected layer Linear, the first LeakyReLU activation function, the first Dropout layer, the second fully-connected layer Linear, the second LeakyReLU activation function, the second Dropout layer, and the third fully-connected layer Linear through the input layer. The third fully-connected layer Linear outputs the feature A'. After adding the output feature A' of the third fully-connected layer Linear and the features after dimensionality reduction element by element, they are sequentially input into the fourth fully-connected layer Linear and the third LeakyReLU activation function. The third LeakyReLU activation function outputs the feature B'. The feature B' is used as the output feature of the first module; 62), Input the output feature B' of the first module into the second module, and the second module outputs the feature D'; The specific process is as follows: The output feature B' of the first module is sequentially input into the second BN layer, the fifth fully-connected layer Linear, the fourth LeakyReLU activation function, the third Dropout layer, the sixth fully-connected layer Linear, the fifth LeakyReLU activation function, the fourth Dropout layer, and the seventh fully-connected layer Linear. The seventh fully-connected layer Linear outputs the feature C'. After adding the output feature C' of the seventh fully-connected layer Linear and the output feature B' of the first module element by element, they are sequentially input into the eighth fully-connected layer Linear and the sixth LeakyReLU activation function. The sixth LeakyReLU activation function outputs the feature D'. The feature D' is used as the output feature of the second module; 63), Input the output feature D' of the second module into the third module, and the third module outputs the feature F'; The specific process is as follows: The output feature D' of the second module is sequentially input into the third BN layer, the ninth fully-connected layer Linear, the seventh LeakyReLU activation function, the fifth Dropout layer, the tenth fully-connected layer Linear, the eighth LeakyReLU activation function, the sixth Dropout layer, and the eleventh fully-connected layer Linear. The eleventh fully-connected layer Linear outputs the feature E'. After adding the output feature E' of the eleventh fully-connected layer Linear and the output feature D' of the second module element by element, they are sequentially input into the twelfth fully-connected layer Linear and the ninth LeakyReLU activation function. The ninth LeakyReLU activation function outputs the feature F'. The feature F' is used as the output feature of the third module; 64), Input the output feature F' of the third module into the fourth module, and the fourth module outputs the feature H'; The specific process is as follows: The output feature F' of the third module is sequentially input into the fourth BN layer, the thirteenth fully connected layer Linear, the tenth LeakyReLU activation function, the seventh Dropout layer, the fourteenth fully connected layer Linear, the eleventh LeakyReLU activation function, the eighth Dropout layer, and the fifteenth fully connected layer Linear. The fifteenth fully connected layer Linear outputs the feature G'. The feature G' output by the fifteenth fully connected layer Linear and the feature F' output by the third module are added element-wise and then sequentially input into the sixteenth fully connected layer Linear and the twelfth LeakyReLU activation function. The twelfth LeakyReLU activation function outputs the feature H', and the feature H' is used as the output feature of the fourth module; 65), The output feature H' of the fourth module is input into the fifth module, and the fifth module outputs the feature J'; The specific process is as follows: The output feature H' of the fourth module is sequentially input into the fifth BN layer, the seventeenth fully connected layer Linear, the thirteenth LeakyReLU activation function, the ninth Dropout layer, the eighteenth fully connected layer Linear, the fourteenth LeakyReLU activation function, the tenth Dropout layer, and the nineteenth fully connected layer Linear. The nineteenth fully connected layer Linear outputs the feature I'. The feature I' output by the nineteenth fully connected layer Linear and the feature H' output by the fourth module are added element-wise and then sequentially input into the twentieth fully connected layer Linear and the fifteenth LeakyReLU activation function. The fifteenth LeakyReLU activation function outputs the feature J', and the feature J' is used as the output feature of the fifth module; 66), The output feature J' of the fifth module is sequentially input into the twenty-first fully connected layer Linear, the eleventh Dropout layer, the sixteenth LeakyReLU activation function layer, the twenty-second fully connected layer Linear, the twelfth Dropout layer, the seventeenth LeakyReLU activation function layer, and the twenty-third fully connected layer Linear. The twenty-third fully connected layer Linear outputs the chromosome category to which the DNA sequence corresponding to the predicted dimensionality-reduced feature belongs, and the category result is output through the output layer; The deep learning classification model is trained to obtain a trained deep learning classification model.
9. A DNA sequence alignment method based on deep learning according to claim 8, characterized in that: In step 33, the dimensionality-reduced feature in step 2 is input into the trained deep learning classification model, and the trained deep learning classification model predicts the chromosome category to which the sequence corresponding to the dimensionality-reduced feature in step 2 belongs; The specific process is as follows: 1), The dimensionality-reduced feature in step 2 is input into the first module through the input layer, and the first module outputs the feature B''; The specific process is as follows: The features after dimensionality reduction are sequentially input into the first BN layer, the first fully connected layer Linear, the first LeakyReLU activation function, the first Dropout layer, the second fully connected layer Linear, the second LeakyReLU activation function, the second Dropout layer, and the third fully connected layer Linear through the input layer. The third fully connected layer Linear outputs the feature A″. After adding the feature A″ output by the third fully connected layer Linear and the features after dimensionality reduction element by element, they are sequentially input into the fourth fully connected layer Linear and the third LeakyReLU activation function. The third LeakyReLU activation function outputs the feature B″, and the feature B″ is used as the output feature of the first module; 2), Input the output feature B″ of the first module into the second module, and the second module outputs the feature D″; The specific process is as follows: The output feature B″ of the first module is sequentially input into the second BN layer, the fifth fully connected layer Linear, the fourth LeakyReLU activation function, the third Dropout layer, the sixth fully connected layer Linear, the fifth LeakyReLU activation function, the fourth Dropout layer, and the seventh fully connected layer Linear. The seventh fully connected layer Linear outputs the feature C″. After adding the feature C″ output by the seventh fully connected layer Linear and the output feature B″ of the first module element by element, they are sequentially input into the eighth fully connected layer Linear and the sixth LeakyReLU activation function. The sixth LeakyReLU activation function outputs the feature D″, and the feature D″ is used as the output feature of the second module; 3), Input the output feature D″ of the second module into the third module, and the third module outputs the feature F″; The specific process is as follows: The output feature D″ of the second module is sequentially input into the third BN layer, the ninth fully connected layer Linear, the seventh LeakyReLU activation function, the fifth Dropout layer, the tenth fully connected layer Linear, the eighth LeakyReLU activation function, the sixth Dropout layer, and the eleventh fully connected layer Linear. The eleventh fully connected layer Linear outputs the feature E″. After adding the feature E″ output by the eleventh fully connected layer Linear and the output feature D″ of the second module element by element, they are sequentially input into the twelfth fully connected layer Linear and the ninth LeakyReLU activation function. The ninth LeakyReLU activation function outputs the feature F″, and the feature F″ is used as the output feature of the third module; 4), Input the output feature F″ of the third module into the fourth module, and the fourth module outputs the feature H″; The specific process is as follows: The output feature F″ of the third module is sequentially input into the fourth BN layer, the thirteenth fully connected layer Linear, the tenth LeakyReLU activation function, the seventh Dropout layer, the fourteenth fully connected layer Linear, the eleventh LeakyReLU activation function, the eighth Dropout layer, and the fifteenth fully connected layer Linear. The fifteenth fully connected layer Linear outputs the feature G″. The feature G″ output by the fifteenth fully connected layer Linear and the feature F″ output by the third module are added element-wise and then sequentially input into the sixteenth fully connected layer Linear and the twelfth LeakyReLU activation function. The twelfth LeakyReLU activation function outputs the feature H″, and the feature H″ is used as the output feature of the fourth module; 5), Input the output feature H″ of the fourth module into the fifth module, and the fifth module outputs the feature J″; The specific process is as follows: The output feature H″ of the fourth module is sequentially input into the fifth BN layer, the seventeenth fully connected layer Linear, the thirteenth LeakyReLU activation function, the ninth Dropout layer, the eighteenth fully connected layer Linear, the fourteenth LeakyReLU activation function, the tenth Dropout layer, and the nineteenth fully connected layer Linear. The nineteenth fully connected layer Linear outputs the feature I″. The feature I″ output by the nineteenth fully connected layer Linear and the feature H″ output by the fourth module are added element-wise and then sequentially input into the twentieth fully connected layer Linear and the fifteenth LeakyReLU activation function. The fifteenth LeakyReLU activation function outputs the feature J″, and the feature J" is used as the output feature of the fifth module; 6), Input the output feature J" of the fifth module into the twenty-first fully connected layer Linear, the eleventh Dropout layer, the sixteenth LeakyReLU activation function layer, the twenty-second fully connected layer Linear, the twelfth Dropout layer, the seventeenth LeakyReLU activation function layer, and the twenty-third fully connected layer Linear. The twenty-third fully connected layer Linear outputs the chromosome category to which the DNA sequence corresponding to the predicted dimensionality-reduced feature in step two belongs, and the category result is output through the output layer.
10. A DNA sequence alignment method based on deep learning according to claim 9, characterized in that: In step four, the sequence of the record category in each file is compared with the DNA sequence of the corresponding chromosome in the reference sequence, and the comparison result is saved; The reference sequence is the human genome sequence; The specific process is as follows: Use the minimap2 tool to compare the DNA sequence stored in each file with the DNA sequence of the corresponding chromosome in the reference sequence to obtain the comparison result of the DNA sequence stored in each file; The comparison result includes: the alignment position of the sequence, the matching length, the base mismatch or insertion / deletion situation, the alignment score, and the alignment quality; Save the comparison result as a sam format file; The DNA sequence stored in each file is the DNA sequence of the known chromosome category predicted in step three three.