DNA sequence processing model pre-training and sequence processing method and related product

By using preset fusion structure feature extraction and encoding of DNA forward and reverse strands in the DNA sequence processing model, the problem that existing models fail to make full use of the properties and structural characteristics of DNA sequences is solved, achieving more full feature expression and higher accuracy of analysis tasks.

CN120015111APending Publication Date: 2025-05-16BIOMAP (BEIJING) INTELLIGENCE TECH LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510091317.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing deep learning models fail to fully utilize the sequence properties and structural characteristics of DNA when processing DNA sequences, resulting in insufficient feature expression.

Method used

By extracting and encoding the DNA forward strand and reverse strand embedding vectors based on the preset fusion structure feature extraction module and sequence coding module, the fusion forward strand and reverse strand output feature sequences, generate a fusion forward strand feature sequence, and input the sequence feature decoder to adjust the model parameters.

Benefits of technology

The DNA sequence processing model utilizes DNA-specific sequence properties and structural characteristics, enhances the diversity and accuracy of feature expression, and improves the accuracy of downstream DNA sequence analysis tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015111A_ABST
    Figure CN120015111A_ABST
Patent Text Reader

Abstract

The invention provides a DNA sequence processing model pre-training and sequence processing method and related products. According to a specific embodiment of the pre-training method for the DNA sequence processing model, feature coding is performed on a first forward chain embedded vector sequence and a first reverse chain embedded vector sequence, so that reverse complementation between double-helix nucleotide chains of DNA is reflected; extracting structural features by using a preset fusion structural feature extraction module to reflect the structural features of the DNA sequence; the first forward chain output feature sequence and the first reverse chain output feature sequence are fused to obtain the first fused forward chain feature sequence, so that feature expression can be performed on the reverse complementary property between the double-helix nucleotide chains of the DNA on the feature level, and the diversity of DNA sequence feature expression information is enriched; and subsequently, the specific DNA sequence analysis task is finely adjusted based on the DNA sequence processing model, so that the accuracy of the downstream DNA sequence analysis task can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the technical field of bioinformatics analysis, and specifically to a DNA sequence processing model pre-training and sequence processing method and related products. Background Art

[0002] With the rapid development of genomics and bioinformatics, the analysis and interpretation of DNA sequences have had a profound impact on fields such as biomedicine, genetics, and biotechnology. Traditional DNA sequence analysis methods are usually based on simple statistical models (such as hidden Markov models, statistical characteristics analysis of coding regions, principal component analysis, and Fisher discriminant, etc.), which make it difficult to effectively extract deep information from sequences. In recent years, the introduction of deep learning technology has provided new possibilities for improving the accuracy and efficiency of DNA analysis.

[0003] However, existing deep learning models often fail to fully utilize the unique sequence properties and structural characteristics of DNA when processing DNA sequences, resulting in insufficient feature expression. Summary of the invention

[0004] The embodiments of the present disclosure propose a DNA sequence processing model pre-training and encoding method, device, electronic device, storage medium and computer program product.

[0005] In a first aspect, an embodiment of the present disclosure provides a DNA sequence processing model pre-training method, the method comprising:

[0006] Based on a preset fusion structural feature extraction module, structural feature extraction is performed on the first forward chain embedded vector sequence and the first reverse chain embedded vector sequence respectively to obtain a first forward chain structural feature sequence and a first reverse chain structural feature sequence, wherein the first forward chain embedded vector sequence and the first reverse chain embedded vector sequence are respectively vector sequences obtained by embedding a first DNA forward chain sequence and a first DNA reverse chain sequence that is reverse complementary to the first DNA forward chain sequence;

[0007] Based on a preset sequence encoding module, feature encoding is performed on the first forward chain structure feature sequence and the first reverse chain structure feature sequence respectively to obtain a first forward chain output feature sequence and a first reverse chain output feature sequence;

[0008] Fusing the first forward chain output characteristic sequence and the first reverse chain output characteristic sequence to obtain a first fused forward chain characteristic sequence;

[0009] Inputting the first fused forward chain feature sequence into a preset sequence feature decoder to obtain a first decoded forward chain sequence;

[0010] The model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder are adjusted based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence.

[0011] In some optional implementations, before respectively performing structural feature extraction on the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence based on the preset fusion structural feature extraction module, the method further includes:

[0012] Obtain the original forward strand sequence of the sample DNA;

[0013] Generating a first DNA forward strand sequence based on the original forward strand sequence, wherein the first DNA forward strand sequence includes original words that are the same as the original forward strand sequence and words to be predicted that are different from the original forward strand sequence;

[0014] generating a first DNA reverse strand sequence which is reversely complementary to the first DNA forward strand sequence;

[0015] Based on a preset word element embedding representation module, the first DNA forward chain sequence and the first DNA reverse chain sequence are respectively embedded and represented to obtain a first forward chain embedding vector sequence and a first reverse chain embedding vector sequence.

[0016] In some optional embodiments, generating a first DNA forward strand sequence based on the original forward strand sequence comprises:

[0017] Generating a first DNA forward strand sequence identical to the original forward strand sequence;

[0018] Randomly selecting some nucleotide identifiers in the first DNA forward chain sequence as words to be predicted;

[0019] Replacing a first preset proportion of the to-be-predicted words in the first DNA forward strand sequence with random words;

[0020] The to-be-predicted words in the first DNA forward chain sequence that are less than or equal to a second preset ratio are replaced with preset mask words, wherein the sum of the first preset ratio and the second preset ratio is less than or equal to 100%.

[0021] In some optional embodiments, replacing the to-be-predicted words in the first DNA forward strand sequence that are less than or equal to a second preset ratio with preset mask words includes:

[0022] Randomly reducing the second preset ratio according to a preset small probability to obtain a mask ratio;

[0023] The to-be-predicted word of the mask ratio in the first DNA forward chain sequence is replaced with the preset mask word.

[0024] In some optional embodiments, the preset fusion structure feature extraction module includes at least one preset DNA structure feature extraction submodule; and

[0025] The method of extracting structural features from the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence based on the preset fusion structural feature extraction module to obtain the first forward chain structural feature sequence and the first reverse chain structural feature sequence includes:

[0026] Using each of the preset DNA structure feature extraction submodules, respectively extract structure features from the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence to obtain a forward chain DNA structure feature sequence and a reverse chain DNA structure feature sequence corresponding to the preset DNA structure feature extraction submodule;

[0027] A first forward strand structural feature sequence and a first reverse strand structural feature sequence are generated based on the forward strand DNA structural feature sequence and the reverse strand DNA structural feature sequence corresponding to each of the preset DNA structural feature extraction submodules, respectively.

[0028] In some optional embodiments, the at least one preset DNA structure feature extraction submodule includes at least one of the following: a first DNA structure feature map extraction submodule, a second DNA structure feature map extraction submodule and a third DNA structure feature map extraction submodule, wherein the first DNA structure feature map extraction submodule includes a first convolution layer composed of at least one 1*3 convolution kernel, the second DNA structure feature map extraction submodule includes a second convolution layer composed of at least one 1*5 convolution kernel, and the third DNA structure feature map extraction submodule includes a third convolution layer composed of at least one 1*7 convolution kernel, wherein the first convolution layer is used to extract codon structure and / or DNA minor groove structure, the second convolution layer is used to extract DNA minor groove structure and / or DNA major groove structure, and the third convolution layer is used to extract DNA major groove structure.

[0029] In some optional embodiments, the preset sequence encoding module includes a sliding window attention layer and a feedforward neural network layer connected in sequence, the sliding window attention layer is used to use at least two attention windows of different sizes to perform attention weight extraction on the first forward chain structure feature sequence and the first reverse chain structure feature sequence, respectively, to obtain the in-window forward chain attention feature map and the in-window reverse chain attention feature map corresponding to the corresponding attention window, and the in-window forward chain attention feature map and the in-window reverse chain attention feature map extracted based on each attention window, respectively, to generate a forward chain attention feature map and a reverse chain attention feature map, and the feedforward neural network layer is used to process based on the forward chain attention feature map and the reverse chain attention feature map, respectively, to obtain the first forward chain output feature sequence and the first reverse chain output feature sequence.

[0030] In some optional embodiments, the fusing the first forward chain output characteristic sequence and the first reverse chain output characteristic sequence to obtain a first fused forward chain characteristic sequence comprises:

[0031] Reverse processing is performed on the first reverse chain output feature sequence to obtain a reversed feature sequence of the first reverse chain;

[0032] Based on a preset complementary linear transformation module, a linear transformation is performed on the reversed feature sequence of the first reverse chain to obtain a reversed complementary feature sequence of the first reverse chain;

[0033] The first forward chain output characteristic sequence and the first reverse chain reverse complement characteristic sequence are fused to obtain the first fused forward chain characteristic sequence.

[0034] In some optional embodiments, adjusting the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module, and the preset sequence feature decoder based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence includes:

[0035] Based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence, the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module, the preset complementary linear transformation module and the preset sequence feature decoder are adjusted.

[0036] In some optional embodiments, adjusting the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module, and the preset sequence feature decoder based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence includes:

[0037] Based on the difference between the first decoded forward chain sequence and the part corresponding to the word to be predicted in the original forward chain sequence corresponding to the first DNA forward chain sequence, the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder are adjusted.

[0038] In a second aspect, an embodiment of the present disclosure provides a DNA sequence processing method, the method comprising:

[0039] Based on a preset fusion structural feature extraction module, structural feature extraction is performed on the second forward chain embedded vector sequence and the second reverse chain embedded vector sequence respectively to obtain a second forward chain structural feature sequence and a second reverse chain structural feature sequence, wherein the second forward chain embedded vector sequence and the second reverse chain embedded vector sequence are respectively vector sequences obtained by embedding representation based on a second DNA forward chain sequence and a second DNA reverse chain sequence that is reverse complementary to the second DNA forward chain sequence;

[0040] Based on a preset sequence encoding module, feature encoding is performed on the second forward chain structure feature sequence and the second reverse chain structure feature sequence respectively to obtain a second forward chain output feature sequence and a second reverse chain output feature sequence, wherein the preset fusion structure feature extraction module and the preset sequence encoding module are pre-trained by the method described in any implementation manner of the first aspect;

[0041] The second forward chain output characteristic sequence and the second reverse chain output characteristic sequence are fused to obtain a second fused forward chain characteristic sequence.

[0042] In some optional implementations, before respectively extracting structural features from the second forward chain embedding vector sequence and the second reverse chain embedding vector sequence based on the preset fusion structural feature extraction module, the method further includes:

[0043] Obtaining the forward strand sequence of the DNA to be encoded;

[0044] Determining a reverse strand sequence of the DNA to be encoded that is reverse complementary to the forward strand sequence of the DNA to be encoded;

[0045] Based on a preset word element embedding representation module, the DNA forward chain sequence to be encoded and the DNA reverse chain sequence to be encoded are respectively embedded and represented to obtain a second forward chain embedding vector sequence and a second reverse chain embedding vector sequence.

[0046] In some optional embodiments, the method further comprises:

[0047] Obtaining a DNA sequence analysis result label of the target DNA sequence analysis task for the forward strand sequence of the DNA to be encoded;

[0048] Inputting the second fusion forward strand characteristic sequence into the target DNA sequence analysis task decoder to obtain a DNA sequence analysis result;

[0049] The model parameters of the target DNA sequence analysis task decoder are adjusted based on the difference between the DNA sequence analysis result and the DNA sequence analysis result label.

[0050] In some optional embodiments, the fusing the second forward chain output characteristic sequence and the second reverse chain output characteristic sequence to obtain a second fused forward chain characteristic sequence comprises:

[0051] The second forward chain output characteristic sequence and the second reverse chain output characteristic sequence are spliced ​​to obtain the second fused forward chain characteristic sequence; or, the sum of the second forward chain output characteristic sequence and the reverse complement characteristic sequence of the second reverse chain is determined as the second fused forward chain characteristic sequence.

[0052] In a third aspect, an embodiment of the present disclosure provides a DNA sequence processing model pre-training device, the device comprising:

[0053] A first structural feature fusion module is configured to extract structural features from the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence based on a preset fusion structural feature extraction module, respectively, to obtain a first forward chain structural feature sequence and a first reverse chain structural feature sequence, wherein the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence are respectively vector sequences obtained by embedding representation based on a first DNA forward chain sequence and a first DNA reverse chain sequence that is reverse complementary to the first DNA forward chain sequence;

[0054] A first sequence encoding module is configured to perform feature encoding on the first forward chain structure feature sequence and the first reverse chain structure feature sequence based on a preset sequence encoding module, respectively, to obtain a first forward chain output feature sequence and a first reverse chain output feature sequence;

[0055] A first bidirectional sequence fusion module is configured to fuse the first forward chain output feature sequence and the first reverse chain output feature sequence to obtain a first fused forward chain feature sequence;

[0056] A first sequence decoding module is configured to input the first fused forward chain feature sequence into a preset sequence feature decoder to obtain a first decoded forward chain sequence;

[0057] The first model training module is configured to adjust the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence.

[0058] In some optional embodiments, the DNA sequence processing model pre-training device further includes a first bidirectional sequence generation module, which is configured to: before extracting structural features from the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence respectively based on a preset fusion structural feature extraction module:

[0059] Obtain the original forward strand sequence of the sample DNA;

[0060] Generating a first DNA forward strand sequence based on the original forward strand sequence, wherein the first DNA forward strand sequence includes original words that are the same as the original forward strand sequence and words to be predicted that are different from the original forward strand sequence;

[0061] generating a first DNA reverse strand sequence which is reversely complementary to the first DNA forward strand sequence;

[0062] Based on a preset word element embedding representation module, the first DNA forward chain sequence and the first DNA reverse chain sequence are respectively embedded and represented to obtain a first forward chain embedding vector sequence and a first reverse chain embedding vector sequence.

[0063] In some optional embodiments, generating a first DNA forward strand sequence based on the original forward strand sequence comprises:

[0064] Generating a first DNA forward strand sequence identical to the original forward strand sequence;

[0065] Randomly selecting some nucleotide identifiers in the first DNA forward chain sequence as words to be predicted;

[0066] Replacing a first preset proportion of the to-be-predicted words in the first DNA forward strand sequence with random words;

[0067] The to-be-predicted words in the first DNA forward chain sequence that are less than or equal to a second preset ratio are replaced with preset mask words, wherein the sum of the first preset ratio and the second preset ratio is less than or equal to 100%.

[0068] In some optional embodiments, replacing the to-be-predicted words in the first DNA forward strand sequence that are less than or equal to a second preset ratio with preset mask words includes:

[0069] Randomly reducing the second preset ratio according to a preset small probability to obtain a mask ratio;

[0070] The to-be-predicted word of the mask ratio in the first DNA forward chain sequence is replaced with the preset mask word.

[0071] In some optional embodiments, the preset fusion structure feature extraction module includes at least one preset DNA structure feature extraction submodule; and

[0072] The first structural feature fusion module includes:

[0073] The substructure feature extraction unit is configured to use each of the preset DNA structure feature extraction submodules to respectively perform structure feature extraction on the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence to obtain a forward chain DNA structure feature sequence and a reverse chain DNA structure feature sequence corresponding to the preset DNA structure feature extraction submodule;

[0074] The substructure feature fusion unit is configured to generate a first forward chain structure feature sequence and a first reverse chain structure feature sequence based on the forward chain DNA structure feature sequence and the reverse chain DNA structure feature sequence corresponding to each of the preset DNA structure feature extraction submodules.

[0075] In some optional embodiments, the at least one preset DNA structure feature extraction submodule includes at least one of the following: a first DNA structure feature map extraction submodule, a second DNA structure feature map extraction submodule and a third DNA structure feature map extraction submodule, wherein the first DNA structure feature map extraction submodule includes a first convolution layer composed of at least one 1*3 convolution kernel, the second DNA structure feature map extraction submodule includes a second convolution layer composed of at least one 1*5 convolution kernel, and the third DNA structure feature map extraction submodule includes a third convolution layer composed of at least one 1*7 convolution kernel, wherein the first convolution layer is used to extract codon structure and / or DNA minor groove structure, the second convolution layer is used to extract DNA minor groove structure and / or DNA major groove structure, and the third convolution layer is used to extract DNA major groove structure.

[0076] In some optional embodiments, the preset sequence encoding module includes a sliding window attention layer and a feedforward neural network layer connected in sequence, the sliding window attention layer is used to use at least two attention windows of different sizes to perform attention weight extraction on the first forward chain structure feature sequence and the first reverse chain structure feature sequence, respectively, to obtain the in-window forward chain attention feature map and the in-window reverse chain attention feature map corresponding to the corresponding attention window, and the in-window forward chain attention feature map and the in-window reverse chain attention feature map extracted based on each attention window, respectively, to generate a forward chain attention feature map and a reverse chain attention feature map, and the feedforward neural network layer is used to process based on the forward chain attention feature map and the reverse chain attention feature map, respectively, to obtain the first forward chain output feature sequence and the first reverse chain output feature sequence.

[0077] In some optional embodiments, the first bidirectional sequence fusion module includes:

[0078] A reverse unit, configured to perform reverse processing on the first reverse chain output feature sequence to obtain a reversed feature sequence of the first reverse chain;

[0079] A complementary unit is configured to perform a linear transformation on the first reverse chain reverse feature sequence based on a preset complementary linear transformation module to obtain a first reverse chain reverse complementary feature sequence;

[0080] The bidirectional fusion unit is configured to fuse the first forward chain output characteristic sequence and the first reverse chain reverse complement characteristic sequence to obtain the first fused forward chain characteristic sequence.

[0081] In some optional implementations, the first model training module is further configured to:

[0082] Based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence, the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module, the preset complementary linear transformation module and the preset sequence feature decoder are adjusted.

[0083] In some optional implementations, the first model training module is further configured to:

[0084] Based on the difference between the first decoded forward chain sequence and the part corresponding to the word to be predicted in the original forward chain sequence corresponding to the first DNA forward chain sequence, the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder are adjusted.

[0085] In a fourth aspect, an embodiment of the present disclosure provides a DNA sequence encoding device, the device comprising:

[0086] A second structural feature fusion module is configured to extract structural features from the second forward chain embedding vector sequence and the second reverse chain embedding vector sequence based on the preset fusion structural feature extraction module, respectively, to obtain a second forward chain structural feature sequence and a second reverse chain structural feature sequence, wherein the second forward chain embedding vector sequence and the second reverse chain embedding vector sequence are respectively vector sequences obtained by embedding representation based on a second DNA forward chain sequence and a second DNA reverse chain sequence that is reverse complementary to the second DNA forward chain sequence;

[0087] A second sequence encoding module is configured to perform feature encoding on the second forward chain structure feature sequence and the second reverse chain structure feature sequence respectively based on a preset sequence encoding module to obtain a second forward chain output feature sequence and a second reverse chain output feature sequence, wherein the preset fusion structure feature extraction module and the preset sequence encoding module are pre-trained by the method described in any implementation of the first aspect;

[0088] The second bidirectional sequence fusion module is configured to fuse the second forward chain output feature sequence and the second reverse chain output feature sequence to obtain a second fused forward chain feature sequence.

[0089] In some optional embodiments, the DNA sequence encoding device further includes a second bidirectional sequence generation module, which is configured to: before the preset fusion structure feature extraction module extracts structure features from the second forward chain embedding vector sequence and the second reverse chain embedding vector sequence respectively:

[0090] Obtaining the forward strand sequence of the DNA to be encoded;

[0091] Determining a reverse strand sequence of the DNA to be encoded that is reverse complementary to the forward strand sequence of the DNA to be encoded;

[0092] Based on a preset word element embedding representation module, the DNA forward chain sequence to be encoded and the DNA reverse chain sequence to be encoded are respectively embedded and represented to obtain a second forward chain embedding vector sequence and a second reverse chain embedding vector sequence.

[0093] In some optional embodiments, the DNA sequence encoding device further comprises:

[0094] A label acquisition module is configured to obtain a DNA sequence analysis result label of the forward strand of the DNA to be encoded for the target DNA sequence analysis task;

[0095] A sequence analysis module is configured to input the second fusion forward strand characteristic sequence into a target DNA sequence analysis task decoder to obtain a DNA sequence analysis result;

[0096] The second model training module is configured to adjust the model parameters of the target DNA sequence analysis task decoder based on the difference between the DNA sequence analysis result and the DNA sequence analysis result label.

[0097] In some optional embodiments, the second bidirectional sequence fusion module is further configured to:

[0098] The second forward chain output characteristic sequence and the second reverse chain output characteristic sequence are spliced ​​to obtain the second fused forward chain characteristic sequence; or, the sum of the second forward chain output characteristic sequence and the reverse complement characteristic sequence of the second reverse chain is determined as the second fused forward chain characteristic sequence.

[0099] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect and / or the second aspect.

[0100] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by one or more processors, implements the method described in any implementation manner in the first aspect and / or the second aspect.

[0101] In a seventh aspect, an embodiment of the present disclosure provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the method described in any implementation manner in the first aspect and / or the second aspect.

[0102] In order to solve the problem that the existing deep learning models often fail to fully utilize the unique sequence properties and structural characteristics of DNA when processing DNA sequences, resulting in insufficient feature expression, the embodiments of the present invention provide a DNA sequence processing model pre-training and encoding method, device, electronic device, storage medium and computer program product, which extracts structural features of the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence based on a preset fusion structural feature extraction module, respectively, to obtain a first forward chain structural feature sequence and a first reverse chain structural feature sequence, wherein the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence are respectively based on the first DNA forward chain sequence and the first DNA sequence that is reverse complementary to the first DNA forward chain sequence. The reverse chain sequence is embedded to represent the vector sequence obtained; based on the preset sequence encoding module, the first forward chain structure feature sequence and the first reverse chain structure feature sequence are respectively feature encoded to obtain the first forward chain output feature sequence and the first reverse chain output feature sequence; then, based on the first forward chain output feature sequence and the first reverse chain output feature sequence, the first fused forward chain feature sequence is determined; then the first fused forward chain feature sequence is input into the preset sequence feature decoder to obtain the first decoded forward chain sequence; finally, based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence, the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder are adjusted. The above method can achieve technical effects including but not limited to the following:

[0103] First, by using a parameter-sharing preset fusion structure feature extraction module and a preset sequence encoding module to respectively perform feature encoding on the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence, the reverse complementary nature between the double helix nucleotide chains of DNA is reflected, making the extraction of DNA sequence features more comprehensive and improving the performance of the DNA sequence processing model;

[0104] Second, by extracting structural features using a preset fusion structural feature extraction module, various structural features of DNA sequences can be fused;

[0105] Third, by fusing the first forward chain output characteristic sequence and the first reverse chain output characteristic sequence to obtain the first fused forward chain characteristic sequence, the reverse complementary properties between the double-helix nucleotide chains of DNA can be characteristically expressed at the characteristic level, enriching the diversity of DNA sequence characteristic expression information.

[0106] Fourth, the DNA sequence processing model obtained by pre-training based on the characteristics of the DNA sequence can improve the accuracy of downstream DNA sequence analysis tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0107] Other features, objects and advantages of the present disclosure will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings. The drawings are only for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. In the drawings:

[0108] Figure 1 is an exemplary system architecture diagram in which an embodiment of the present disclosure may be applied;

[0109] Figure 2A is a flowchart of an embodiment of a DNA sequence processing model pre-training method according to the present disclosure;

[0110] Figure 2B is a decomposed flow chart of one embodiment of step 202' according to the present disclosure;

[0111] Figure 2C is a decomposed flow chart of one embodiment of step 203 according to the present disclosure;

[0112] Figure 2D is a decomposed flow chart of one embodiment of step 201 according to the present disclosure;

[0113] Figure 3 is a specific example of the DNA sequence processing model pre-training method according to the present disclosure;

[0114] Figure 4 is a flow chart of an embodiment of a DNA sequence processing method according to the present disclosure;

[0115] Figure 5 A schematic diagram of the structure of an embodiment of a DNA sequence processing model pre-training device according to the present disclosure;

[0116] Figure 6 A schematic structural diagram of an embodiment of a DNA sequence encoding device according to the present disclosure;

[0117] Figure 7 A schematic diagram of the structure of a computer system of an electronic device suitable for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION

[0118] The present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It is understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.

[0119] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0120] Figure 1 An exemplary system architecture 100 is shown to which embodiments of the DNA sequence processing model pre-training and DNA sequence processing methods, apparatuses, electronic devices, and storage media disclosed herein can be applied.

[0121] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0122] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, 103, such as DNA sequence processing model pre-training applications, DNA sequence encoding applications, etc.

[0123] Terminal devices 101, 102, 103 can be hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with information input devices (e.g., keyboard, mouse, touch screen, microphone, camera, etc.) and information output devices (e.g., display screen, speaker, etc.), including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III, Moving Picture Experts Compression Standard Audio Layer 3), MP4 (Moving Picture Experts Group Audio Layer IV, Moving Picture Experts Compression Standard Audio Layer 4) players, laptop computers and desktop computers, etc. When terminal devices 101, 102, 103 are software, they can be installed in the terminal devices listed above. It can be implemented as multiple software or software modules (for example, used to provide model pre-training services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0124] In some cases, the DNA sequence processing model pre-training and DNA sequence processing method provided by the present disclosure may be performed by the terminal devices 101, 102, 103, and accordingly, the DNA sequence processing model pre-training and DNA sequence encoding apparatus may be provided in the terminal devices 101, 102, 103. In this case, the system architecture 100 may also not include the server 105.

[0125] In some cases, the DNA sequence processing model pre-training and DNA sequence processing method provided by the present disclosure can be jointly performed by the terminal devices 101, 102, 103 and the server 105. For example, the step of "extracting structural features of the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence respectively based on the preset fusion structural feature extraction module" can be performed by the terminal devices 101, 102, 103, and the steps of "inputting the first fusion forward chain feature sequence into the preset sequence feature decoder to obtain the first decoded forward chain sequence" can be performed by the server 105. The present disclosure does not limit this. Accordingly, the model training device based on sequence data can also be respectively set in the terminal devices 101, 102, 103 and the server 105.

[0126] In some cases, the sequence data-based model training and sequence data encoding methods provided in the present disclosure can be executed by the server 105. Accordingly, the DNA sequence processing model pre-training and DNA sequence encoding device can also be set in the server 105. In this case, the system architecture 100 may not include the terminal devices 101, 102, and 103.

[0127] It should be noted that the server 105 can be hardware or software. When the server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server 105 is software, it can be implemented as multiple software or software modules (for example, for providing distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0128] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0129] Continue to refer Figure 2A , which shows a process 200 of an embodiment of a DNA sequence processing model pre-training method according to the present disclosure, the DNA sequence processing model pre-training method comprises the following steps:

[0130] Step 201: Based on a preset fusion structural feature extraction module, structural feature extraction is performed on the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence respectively to obtain a first forward chain structural feature sequence and a first reverse chain structural feature sequence.

[0131] Here, the first forward strand embedded vector sequence and the first reverse strand embedded vector sequence are respectively vector sequences obtained by embedding and representing the first DNA forward strand sequence and the first DNA reverse strand sequence which is reverse complementary to the first DNA forward strand sequence.

[0132] The first forward chain embedding vector sequence is composed of the embedding vectors corresponding to each word (i.e., token) in the first DNA forward chain sequence, arranged in the order of the corresponding word in the first DNA forward chain sequence. Specifically, it can be obtained by dividing the first DNA forward chain sequence into units of single nucleotide identifiers. Correspondingly, the first reverse chain embedding vector sequence is composed of the embedding vectors corresponding to each word in the first DNA reverse chain sequence, arranged in the order of the corresponding word in the first DNA reverse chain sequence. Specifically, it can be obtained by dividing the first DNA reverse chain sequence into units of single nucleotide identifiers.

[0133] Here, the preset fusion structure feature extraction module can be various feature extraction models for extracting DNA structure features. Specifically, the model structure and model parameters of the preset fusion structure feature extraction module can be formulated by technicians according to the structural characteristics of the DNA sequence.

[0134] Step 202: Based on a preset sequence encoding module, feature encoding is performed on the first forward chain structure feature sequence and the first reverse chain structure feature sequence respectively to obtain a first forward chain output feature sequence and a first reverse chain output feature sequence.

[0135] Here, various models suitable for encoding feature vector sequences can be used as the preset sequence encoding module. For example, the preset sequence encoding module can be an encoder in a Transformer, and so on.

[0136] Specifically, the first forward chain structural feature sequence and the first reverse chain structural feature sequence can be respectively input into a preset sequence encoding module to obtain a first forward chain output feature sequence and a first reverse chain output feature sequence.

[0137] The first forward strand output feature sequence may be composed of output features corresponding to each word (i.e., token) in the first DNA forward strand sequence arranged in the order of the corresponding word in the first DNA forward strand sequence. The first forward strand output feature sequence may also be understood as a feature representation obtained after the first DNA forward strand sequence is feature represented by a preset fusion structure feature extraction module and a preset sequence encoding module.

[0138] The first reverse strand output feature sequence may be composed of output features corresponding to each word (i.e., token) in the first DNA reverse strand sequence, arranged in the order of the corresponding word in the first DNA reverse strand sequence. The first reverse strand output feature sequence may also be understood as a feature representation obtained after the first DNA reverse strand sequence is feature represented by a preset fusion structure feature extraction module and a preset sequence encoding module.

[0139] Step 203: fuse the first forward chain output feature sequence and the first reverse chain output feature sequence to obtain a first fused forward chain feature sequence.

[0140] Since the forward chain needs to be decoded and predicted later, and the features expressed by the first reverse chain output feature sequence are different from those expressed by the first forward chain output feature sequence, they cannot be directly fused. It is necessary to map the first reverse chain output feature sequence to the first forward chain output feature sequence, and then fuse the mapped feature sequence with the first forward chain output feature sequence to obtain the first fused forward chain feature sequence.

[0141] In practice, according to the reverse complementary characteristics of DNA sequences, mapping the first reverse strand output feature sequence to the first forward strand output feature sequence requires two operations, one is reverse and the other is complementary. The reverse operation is easy to implement, but it is impossible to directly complement each other at the representation level (i.e., the forward strand output feature and the reverse strand output feature).

[0142] To solve the above complementary problem, in some optional implementations, step 203 may include: Figure 2B The following steps 2031 to 2033 are shown:

[0143] Step 2031, reverse processing is performed on the first reverse chain output feature sequence to obtain the first reverse chain reverse feature sequence.

[0144] Step 2032: Based on a preset complementary linear transformation module, a linear transformation is performed on the reversed feature sequence of the first reverse chain to obtain a reverse complementary feature sequence of the first reverse chain.

[0145] Here, the preset complementary linear transformation module may be various linear transformation models, which are used to perform a linear transformation on the reverse feature sequence of the first reverse chain to obtain the reverse complementary feature sequence of the first reverse chain.

[0146] As an example, the feature sequence after the first reverse chain is reversed is a matrix M with R rows and C columns. 1 The preset complementary linear transformation module can be a C*C matrix M L , accordingly, the first reverse strand is reverse complemented with the characteristic sequence M 2It is also a matrix with R rows and C columns. Among them, the matrix M L The values ​​of the elements in are learnable and adjustable.

[0147] That is to say, a parameterized learnable preset complementary linear transformation module is used here, and the feature sequence after the first reverse chain is reversed is transformed to the output feature sequence of the first forward chain, that is, the complementary operation is converted into a learnable preset complementary linear transformation module, so that the preset complementary linear transformation module can learn how to complement each other during the process of parameter adjustment (i.e. pre-training), thereby solving the problem of complementarity.

[0148] It can be understood that, accordingly, step 205, adjusting the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence, can be performed as follows:

[0149] Based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence, the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module, the preset complementary linear transformation module and the preset sequence feature decoder are adjusted.

[0150] That is, in step 205, the preset complementary linear transformation module that performs the complementary operation is also optimized.

[0151] Step 2033, fusing the first forward chain output characteristic sequence and the first reverse chain reverse complement characteristic sequence to obtain a first fused forward chain characteristic sequence.

[0152] Here, various methods may be used to fuse the first forward chain output characteristic sequence and the first reverse chain reverse complement characteristic sequence to obtain a first fused forward chain characteristic sequence.

[0153] For example, the sum of the first forward chain output characteristic sequence and the first reverse chain reverse complement characteristic sequence can be determined as the first fusion forward chain characteristic sequence.

[0154] Alternatively, the first forward chain output characteristic sequence and the first reverse chain reverse complement characteristic sequence may be spliced ​​to obtain the first fused forward chain characteristic sequence.

[0155] Through the above steps 2031 to 2033, it is possible to map the first reverse chain output feature sequence to the first forward chain output feature sequence, and then fuse the mapped feature sequence with the first forward chain output feature sequence to obtain a first fused forward chain feature sequence, and then the features in the first fused forward chain feature sequence are expressed in the same direction (for example, the forward chain direction).

[0156] Step 204: input the first fused forward chain feature sequence into a preset sequence feature decoder to obtain a first decoded forward chain sequence.

[0157] Here, the preset sequence feature decoder is used to characterize the correspondence between the feature sequence and the DNA sequence. The preset sequence feature decoder may include a linear transformation layer and / or a nonlinear transformation layer. For example, the preset sequence feature decoder may include a fully connected network and an activation function layer connected in sequence. The first fused forward chain feature sequence is input into the fully connected network to obtain N values ​​corresponding to each word in the first DNA forward chain sequence, and then the above N values ​​are normalized to obtain N probability values, which correspond to the probability values ​​belonging to the N nucleotide identifiers, respectively, and the nucleotide identifier with the largest probability value corresponding to each word is used as the decoding nucleotide identifier corresponding to the corresponding word in the first DNA forward chain sequence. Finally, the decoding forward chain sequence can be generated using the decoding nucleotide identifier corresponding to each word in the first DNA forward chain sequence. That is, the decoding forward chain sequence is formed by sequentially arranging the decoding nucleotide identifiers, and the sequence length of the decoding forward chain sequence can be the same as the sequence length of the first DNA forward chain sequence.

[0158] Step 205, adjusting model parameters of a preset fusion structure feature extraction module, a preset sequence encoding module and a preset sequence feature decoder based on the difference between the first decoded forward strand sequence and the original forward strand sequence corresponding to the first DNA forward strand sequence.

[0159] Here, various implementation methods can be used to adjust the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder with the optimization goal of minimizing the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence.

[0160] In some optional implementations, before step 201, the DNA sequence processing model pre-training method may include the following steps 201' to 204':

[0161] Step 201 ′, obtaining the original forward strand sequence of the sample DNA.

[0162] Here, obtaining the original forward chain sequence of the sample DNA may include the following methods:

[0163] It is obtained through various sequencing technologies, or through gene libraries (such as genomic libraries or cDNA libraries), or through reverse transcription, artificial synthesis, genomic databases, etc.

[0164] The original forward strand sequence of the sample DNA refers to a sequence directly obtained through the above-mentioned various methods, which is formed by nucleotide identifiers being arranged in order as words.

[0165] Step 202', generating a first DNA forward strand sequence based on the original forward strand sequence.

[0166] Here, various implementation methods can be used to generate a first DNA forward chain sequence based on the original forward chain sequence. The generated first DNA forward chain sequence includes original words that are the same as the original forward chain sequence and words to be predicted that are different from the original forward chain sequence. Specifically, the first DNA forward chain sequence can be the same length as the original forward chain sequence, but some partial fragments are different from the original forward chain, the words in the different fragments are words to be predicted, and the words in the same fragments can be original words, and specifically, the original words can be nucleotide identifiers.

[0167] It should be noted that, here, the so-called word to be predicted may be a nucleotide identifier or other word that is not a nucleotide identifier.

[0168] After step 202', the word to be predicted in the first DNA forward chain sequence is different from the corresponding word in the original forward chain sequence. Accordingly, in the subsequent step 205, the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder are adjusted based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence, or the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder are adjusted based on the difference between the first decoded forward chain sequence and the part corresponding to the word to be predicted in the original forward chain sequence.

[0169] As an example, step 202' may be performed as follows:

[0170] First, a first DNA forward strand sequence is generated that is identical to the original forward strand sequence.

[0171] Then, a third preset ratio of nucleotide identifiers is randomly selected in the first DNA forward chain sequence as word units to be masked.

[0172] Next, some nucleotide identifiers in the first DNA forward chain sequence are randomly selected as words to be masked.

[0173] Here, various methods may be used to randomly select some nucleotide identifiers in the first DNA forward chain sequence as the word elements to be masked.

[0174] For example, a third preset ratio of nucleotide identifiers can be randomly selected in the first DNA forward chain sequence as word elements to be masked, wherein the third preset ratio is greater than 0 and less than 100%, optionally, the third preset ratio is greater than 0 and less than 40%, preferably, the third preset ratio is greater than 0 and less than 20%, and more preferably, the third preset ratio is 15%.

[0175] Finally, the to-be-masked word in the first DNA forward chain sequence is replaced with the preset mask word.

[0176] Here, the preset mask token is another identifier different from the nucleotide identifier, e.g. <mask>Word element.

[0177] In practice, the preset mask word unit generally does not appear in the fine-tuning process of the downstream DNA sequence analysis task decoder. Therefore, the optional method of randomly selecting the third preset ratio of nucleotide identifiers to replace the preset mask word unit will lead to inconsistency between the pre-training and the downstream DNA sequence analysis task decoder fine-tuning. In order to avoid the above inconsistency problem, step 202' may optionally include the following: Figure 2C The following steps 2021' to 2024' are shown:

[0178] Step 2021', generating a first DNA forward strand sequence that is identical to the original forward strand sequence.

[0179] Step 2022', randomly select some nucleotide identifiers in the first DNA forward chain sequence as the word to be predicted.

[0180] Here, various methods may be used to randomly select some nucleotide identifiers in the first DNA forward strand sequence as the word to be predicted. For example, a third preset ratio of nucleotide identifiers may be randomly selected in the first DNA forward strand sequence as the word to be predicted, wherein the third preset ratio is greater than 0 and less than 100%, optionally, the third preset ratio is greater than 0 and less than 40%, preferably, the third preset ratio is greater than 0 and less than 20%, and more preferably, the third preset ratio is 15%.

[0181] Step 2023', replacing a first preset ratio of the to-be-predicted words in the first DNA forward strand sequence with random words.

[0182] The first preset ratio is greater than 0 and less than 100%. Optionally, the first preset ratio is greater than 0 and less than 40%. Preferably, the first preset ratio is greater than 0 and less than 20%. More preferably, the first preset ratio is 10%.

[0183] Here, the random word element may be a nucleotide identifier randomly selected from the nucleotide identifiers. For example, the random word element identifier may be any one of the adenine identifier "A", the thymine identifier "T", the guanine identifier "G", and the cytosine identifier "C".

[0184] Step 2024', replace the to-be-predicted words in the first DNA forward strand sequence that are less than or equal to a second preset ratio with preset mask words.

[0185] Here, the sum of the first preset ratio and the second preset ratio is less than or equal to 100%. As an example, the first preset ratio is 10% and the second preset ratio is 80%.

[0186] It is understandable that when the sum of the first preset ratio and the second preset ratio is equal to 100%, all the words to be predicted in the first DNA forward chain sequence are replaced by other words that are not nucleotide markers, that is, they are replaced by preset random words or preset mask words.

[0187] When the sum of the first preset ratio and the second preset ratio is less than 100%, there are still (1-first preset ratio-second preset ratio) to-be-predicted words in the first DNA forward chain sequence that have not been replaced, that is, (1-first preset ratio-second preset ratio) of the to-be-predicted words in the first DNA forward chain sequence are still original nucleotide identifiers.

[0188] Through the optional implementation of the above steps 2021' to 2024', by replacing some nucleotide identifiers in the first DNA forward chain sequence with random words instead of uniformly replacing them with preset mask words, the inconsistency problem between pre-training and downstream DNA sequence analysis task decoder fine-tuning can be avoided.

[0189] In order to further maintain the consistency of the subsequent pre-training process of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder of the DNA sequence with the fine-tuning process of the downstream DNA sequence analysis task decoder, optionally, step 2024' can also be performed as follows:

[0190] First, the second preset ratio is randomly reduced according to a preset small probability to obtain a mask ratio.

[0191] That is, the probability of reducing the second preset ratio is a preset small probability, and a value is randomly determined between 0 and the second preset ratio as the mask ratio.

[0192] The preset small probability is a relatively small probability value between 0 and 1. As an example, the preset small probability may be 0.02.

[0193] Then, the to-be-predicted word units of the mask ratio in the first DNA forward strand sequence are replaced with the preset mask word units.

[0194] Adopting the optional implementation of the above step 2024' can further promote the consistency of the pre-training process of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder with the fine-tuning process of the downstream DNA sequence analysis task decoder.

[0195] Random words and preset mask words are added to the first DNA forward chain sequence obtained by the optional implementation method of steps 2021' to 2024', and optionally some words keep the original nucleotide identifiers unchanged, so as to be more consistent with the fine-tuning of the downstream DNA sequence analysis task decoder, thereby improving the efficiency of fine-tuning the downstream DNA sequence analysis task decoder and improving the task completion accuracy of the final downstream DNA sequence analysis task decoder.

[0196] Step 203', generating a first DNA reverse strand sequence that is reverse complementary to the first DNA forward strand sequence.

[0197] Here, a first DNA reverse strand sequence that is reverse complementary to the first DNA forward strand sequence can be generated according to the reverse complementary characteristics of the DNA forward strand sequence and the reverse strand sequence.

[0198] For example, this can be done as follows:

[0199] First, the first DNA forward chain sequence is reversed to obtain the reversed first DNA forward chain sequence.

[0200] Then, the original word (ie, nucleotide identifier) ​​in the first DNA forward chain sequence after reverse reaction is replaced with a complementary nucleotide identifier to obtain the first DNA reverse chain sequence.

[0201] For example, the adenine marker "A" in the first DNA forward chain sequence after reverse transcription is replaced by the thymine marker "T", the thymine marker "T" is replaced by the adenine marker "A", the guanine marker "G" is replaced by the cytosine marker "C", and the cytosine marker "C" is replaced by the guanine marker "G".

[0202] It is understandable that, in the complementary process, if the optional method of the above step 2023' is adopted, some of the words to be predicted in the first DNA forward chain sequence are replaced with random words. Since the random words are also nucleotide identifiers, "A" is replaced with "T", "T" is replaced with "A", "G" is replaced with "C", and "C" is replaced with "G". The preset mask words are not replaced, and the preset mask words remain unchanged.

[0203] Step 204', based on a preset word unit embedding representation module, embedding representation is performed on the first DNA forward strand sequence and the first DNA reverse strand sequence respectively to obtain a first forward strand embedding vector sequence and a first reverse strand embedding vector sequence.

[0204] Here, the preset word unit embedding representation module can be various currently known or future developed models for converting word units into vectors. As an example, the preset word unit embedding representation module can include but is not limited to the following models: One-Hot Encoding, Word2Vec, GloVe (Global Vectors for Word Representation), FastText, BERT (Bidirectional Encoder Representations from Transformers), etc.

[0205] Adopting the above steps 201' to 204', by generating a first DNA forward chain sequence that is partially identical to the original forward chain sequence of the sample DNA and partially different from the sequence, there are also partial vectors in the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence based on the above first DNA forward chain sequence that reflect the original nucleotide identifiers in the original forward chain, and partial vectors are vectors of other word units, which provide a basis for the subsequent prediction of the difference between the decoded forward chain and the original forward chain and the corresponding parts of the above other word units, so as to facilitate the pre-training of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder. Accordingly, step 205 can be performed as follows: based on the difference between the first decoded forward chain sequence and the part corresponding to the word unit to be predicted in the original forward chain sequence corresponding to the first DNA forward chain sequence, the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder are adjusted.

[0206] In practice, DNA sequences have a variety of structural features. Therefore, different structural feature extraction models can be designed based on the various structural features of the DNA sequences to adapt to the specific structural features of the DNA sequences and improve the compatibility between the DNA sequence processing model and the nature of the DNA sequences. In some optional embodiments, the preset fusion structural feature extraction module recorded in step 201 may include at least one preset DNA structural feature extraction submodule. Accordingly, step 201 may include the following: Figure 2D The following steps 2011 and 2012 are shown:

[0207] Step 2011, using each preset DNA structure feature extraction submodule, respectively extract structure features from the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence to obtain the forward chain DNA structure feature sequence and the reverse chain DNA structure feature sequence corresponding to the preset DNA structure feature extraction submodule.

[0208] For example, suppose there are K preset DNA structure feature extraction submodules M 1 , …, M i , …, M K Then, using the preset DNA structure feature extraction submodule M i The structure feature extraction of the first forward strand embedding vector sequence SSV and the first reverse strand embedding vector sequence ASV can be performed respectively to obtain the DNA structure feature extraction submodule M i The corresponding forward strand DNA structural characteristic sequence SSF i and reverse strand DNA structural characteristic sequence ASF i .

[0209] Step 2012: Generate a first forward strand structural feature sequence and a first reverse strand structural feature sequence based on the forward strand DNA structural feature sequence and the reverse strand DNA structural feature sequence corresponding to each preset DNA structural feature extraction submodule.

[0210] Continuing with the above example, we can integrate K preset DNA structure feature extraction submodules M 1 , …, M i , …, M K The corresponding forward strand DNA structural characteristic sequence SSF 1 , …, SSF i , …, SSF K Generate the first forward chain structural feature sequence FSSF, integrating K preset DNA structural feature extraction submodules M 1 , …, M i , …, M K The corresponding reverse strand DNA structural characteristic sequence ASF 1 , …, ASF i , …, ASF K Generate the first reverse chain structural feature sequence FASF.

[0211] Here, various methods can be used to generate the first forward strand structural feature sequence and the first reverse strand structural feature sequence. For example, the forward strand DNA structural feature sequence SSF corresponding to each preset DNA structural feature extraction submodule can be i The sum of the reverse strand DNA structure feature sequences ASF corresponding to each preset DNA structure feature extraction submodule is determined as the first forward strand structure feature sequence. i The sum is determined as the first reverse chain structural characteristic sequence. It can be specifically expressed by the following formula:

[0212]

[0213] For another example, the first forward strand embedding vector sequence SSV and the forward strand DNA structure feature sequence SSF corresponding to each preset DNA structure feature extraction submodule may be embedded in the vector sequence SSV. i The sum is determined as the first forward chain structural feature sequence. That is, the first forward chain embedding vector sequence and the forward chain DNA structural feature sequence under different DNA structural feature extraction models are fused by using residual linking.

[0214] Alternatively, the first reverse strand embedding vector sequence ASV and the reverse strand DNA structure feature sequence ASF corresponding to each preset DNA structure feature extraction submodule may be embedded in the vector sequence ASV. i The sum is determined as the first reverse chain structural characteristic sequence. It can be specifically expressed by the following formula:

[0215]

[0216] That is, by adopting the residual connection method to fuse the first forward chain embedding vector sequence SSV and the forward chain DNA structure feature sequence SSF under different DNA structure feature extraction models i , so that the first forward chain structural feature sequence FSSF better represents the DNA forward chain sequence, and the first reverse chain embedding vector sequence ASV and the reverse chain DNA structural feature sequence ASF under different DNA structural feature extraction models are fused by residual linking i , the first reverse strand structural characteristic sequence FASF better characterizes the DNA reverse strand sequence.

[0217] According to biological knowledge and technology, DNA sequences have the following characteristics:

[0218] Each DNA codon encodes a specific amino acid and each DNA codon consists of three nucleotides.

[0219] The minor groove of DNA contains 3-5 nucleotides.

[0220] The major groove of DNA contains 5-7 nucleotides.

[0221] Based on the above structural characteristics of the above-mentioned DNA sequence, in some optional embodiments, at least one preset DNA structure feature extraction submodule may include at least one of the following: a first DNA structure feature map extraction submodule, a second DNA structure feature map extraction submodule and a third DNA structure feature map extraction submodule.

[0222] The first DNA structure feature map extraction submodule includes a first convolution layer composed of at least one 1*3 convolution kernel, the second DNA structure feature map extraction submodule includes a second convolution layer composed of at least one 1*5 convolution kernel, and the third DNA structure feature map extraction submodule includes a third convolution layer composed of at least one 1*7 convolution kernel.

[0223] Among them, the first convolutional layer is used to extract the codon structure and / or the DNA minor groove structure, the second convolutional layer is used to extract the DNA minor groove structure and / or the DNA major groove structure, and the third convolutional layer is used to extract the DNA major groove structure.

[0224] Assume that the dimension of the vector in the first forward chain embedding vector sequence SSV is D and the sequence length is L.

[0225] Here, each convolution kernel of size 1*3 in the first convolution layer can be used to perform sliding convolution along the vector sequence extension direction of the first forward chain embedding vector sequence SSV with a 1*3 sliding window, and finally obtain the forward chain DNA structure feature sequence CSSF corresponding to the first convolution layer. 1 .

[0226] Using each convolution kernel of size 1*5 in the second convolution layer, a sliding convolution is performed along the vector sequence extension direction of the first forward chain embedding vector sequence SSV with a 1*5 sliding window, and finally the forward chain DNA structure feature sequence CSSF corresponding to the second convolution layer is obtained. 2 .

[0227] Using each convolution kernel of size 1*7 in the third convolution layer, a sliding convolution is performed along the vector sequence extension direction of the first forward chain embedding vector sequence SSV with a 1*7 sliding window, and finally the forward chain DNA structure feature sequence CSSF corresponding to the third convolution layer is obtained. 3 .

[0228] Optionally, the forward strand DNA structural feature sequence CSSF can be directly fused here 1 , CSSF 2 and CSSF 3 Generate the first forward chain structure feature sequence FSSF. For example, you can 1 , CSSF 2 and CSSF 3 For another example, the residual connection method can be used to embed the first forward chain into the vector sequence SSV, CSSF 1 , CSSF 2 and CSSF 3 Add them together to get FSSF.

[0229] Optionally, the first DNA structure feature extraction submodule may further include a first nonlinear activation function layer connected to the first convolution layer. The second DNA structure feature extraction submodule may further include a second nonlinear activation function layer connected to the second convolution layer. The third DNA structure feature extraction submodule may further include a third nonlinear activation function layer connected to the third convolution layer.

[0230] As an example, the first non-linear activation function layer, the second non-linear activation function layer, and the third non-linear activation function layer may be GELU activation functions (Gaussian Error Linear Unit).

[0231] In this way, the forward chain DNA structure characteristic sequences CSSF corresponding to the first convolutional layer, the second convolutional layer and the third convolutional layer can be obtained. 1 , CSSF 2 and CSSF 3 After that, CSSF 1 , CSSF 2 and CSSF 3 Input the first nonlinear activation function layer, the second nonlinear activation function layer and the third nonlinear activation function layer, and obtain the forward chain DNA structure feature sequence SSF corresponding to the first DNA structure feature extraction submodule, the second DNA structure feature extraction submodule and the third DNA structure feature extraction submodule respectively. 1 、SSF 2 and SSF 3 , and then the forward strand DNA structural characteristic sequence SSF can be directly fused 1 , SSF 2 and SSF 3 Generate the first forward chain structure feature sequence FSSF, or use the residual connection method to embed the first forward chain into the vector sequence SSV, SSF 1 , SSF 2 and SSF 3 Add them together to get FSSF.

[0232] By adopting a similar method, the reverse strand DNA structure feature sequence ASF corresponding to the first DNA structure feature extraction submodule, the second DNA structure feature extraction submodule and the third DNA structure feature extraction submodule can be obtained. 1 、ASF 2 and ASF 3 Then, the possible fusion method mentioned above is used to fuse the reverse strand DNA structure characteristic sequence ASF 1 , ASF 2 and ASF 3 Generate the first forward chain structural feature sequence ASSF.

[0233] use Figure 2D The first forward chain structural feature sequence and the first reverse chain structural feature sequence generated in step 2011 and step 2012 incorporate multiple preset DNA structural feature extraction submodules designed for the properties and structural characteristics of DNA sequences, which are more suitable for the specific scenarios of DNA sequences.

[0234] It should be noted that the above steps 201 to 205 only show the process of pre-training the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder based on a pair of first forward chain embedding vector sequence and the first reverse chain embedding vector sequence. In practice, the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder can be pre-trained multiple times. Each pre-training can use at least one pair of first forward chain embedding vector sequence and first reverse chain embedding vector sequence to execute steps 201 to 204, and in step 205, based on each pair of first forward chain embedding vector sequence and first reverse chain embedding vector sequence, the difference between the corresponding first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence is calculated, and the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder are adjusted based on the sum of the calculated differences.

[0235] According to biological knowledge, the length of an RNA sequence is about 1000 BP (Base Pair), and the length of a protein sequence is the length of RNA divided by three, which is about several hundred BP. A significant difference between DNA sequences and RNA sequences and protein sequences is that they are generally very long. For example, a DNA sequence may have millions of BPs, but at the same time, some DNA sequences are not particularly long. Therefore, the coding model suitable for RNA sequences or protein sequences is often not suitable for DNA sequences. Since DNA sequences are particularly long, there will be a problem of large amount of calculation when encoding DNA sequences. In addition, the length change of DNA sequences will have problems of long-range dependence and local dependence.

[0236] Based on the consideration of the above technical issues, in some optional implementations, the preset sequence encoding module described in step 202 may include a sliding window attention layer and a feedforward neural network layer connected in sequence.

[0237] Here, the sliding window attention layer is used to use at least two attention windows of different sizes to extract the attention weights of the first forward chain structure feature sequence and the first reverse chain structure feature sequence obtained in step 201, respectively, to obtain the in-window forward chain attention feature map and the in-window reverse chain attention feature map corresponding to the corresponding attention window, and to generate the forward chain attention feature map and the reverse chain attention feature map based on the in-window forward chain attention feature map and the in-window reverse chain attention feature map extracted from each attention window. That is, here, by using attention windows of different sizes, different attention heads are used for different length regions of the DNA sequence at different levels to capture long-range dependencies and local features. Specifically, by using the self-attention mechanism, the local region and the long sequence region in the DNA sequence are focused to enhance the learning ability of relevant information. In other words, in order to solve the long attention problem, that is, the sequence is too long, which will make it difficult to train later, or to make the model compatible with the features of short sequences and long sequences at the same time, different attention window sizes are used to make the model compatible with short sequences and long sequences. Optionally, there can be a full attention window among at least two attention windows, and the full attention window is used to extract the attention feature map of the entire DNA sequence. In this way, the model can also learn all the attention features of the DNA sequence, further enhancing the feature compatibility of the model.

[0238] The feedforward neural network layer is used to process the forward chain attention feature map and the reverse chain attention feature map to obtain the first forward chain output feature sequence and the first reverse chain output feature sequence respectively.

[0239] The method of the above steps 201 to 205 is used for training to obtain a preset fusion structure feature extraction module, a preset sequence encoding module and a preset sequence feature decoder, wherein the preset fusion structure feature extraction module and the preset sequence encoding module can be used as a DNA sequence encoding pre-training model for encoding DNA sequence data to obtain a feature sequence. In the above DNA sequence encoding pre-training model, the reverse complementary nature between the double helix nucleotide chains of DNA is reflected by feature encoding the forward chain embedding vector sequence and the reverse chain embedding vector sequence; in addition, by extracting structural features using the preset fusion structure feature extraction module, the structural features of the DNA sequence can be reflected; and by fusing the forward chain output feature sequence and the reverse chain output feature sequence to obtain a fusion forward chain feature sequence, the reverse complementary nature between the double helix nucleotide chains of DNA can be feature expressed at the feature level, enriching the diversity of DNA sequence feature expression information. Furthermore, based on the above DNA sequence encoding pre-training model, decoders of other DNA sequence analysis tasks can be spliced ​​to perform downstream DNA sequence analysis tasks. Since the above DNA sequence encoding pre-training model has been specially pre-trained for the characteristics of the DNA sequence, the accuracy of the downstream DNA sequence analysis task can be improved.

[0240] Specifically, please refer to Figure 3 , Figure 3 A specific example of a method for pre-training a model based on a DNA sequence processing is shown. Figure 3 As shown, first, the original forward strand sequence of the sample DNA is randomly selected and the word to be predicted is replaced (for example, refer to the relevant records about step 202' above) to obtain a first DNA forward strand sequence.

[0241] Then, the first DNA forward strand sequence is reversely complemented to obtain the first DNA reverse strand sequence.

[0242] Next, based on a preset word unit embedding representation module, the first DNA forward chain sequence and the first DNA reverse chain sequence are respectively embedded and represented to obtain a first forward chain embedding vector sequence and a first reverse chain embedding vector sequence.

[0243] Afterwards, based on the preset fusion structural feature extraction module, structural feature extraction is performed on the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence respectively to obtain the first forward chain structural feature sequence and the first reverse chain structural feature sequence.

[0244] Here, the preset fusion structure feature extraction module includes three parallel DNA structure feature extraction submodules, namely: a first convolution layer and a first nonlinear activation function layer connected in sequence, a second convolution layer and a second nonlinear activation function layer connected in sequence, and a third convolution layer and a third nonlinear activation function layer connected in sequence. The preset fusion structure feature extraction module also includes a residual connection layer, that is, the data input to the preset fusion structure feature extraction module, the data output by the first nonlinear activation function layer, the data output by the second nonlinear activation function layer, and the data output by the third nonlinear activation function layer are added to obtain the data output by the preset fusion structure feature extraction module.

[0245] Next, based on a preset sequence encoding module, feature encoding is performed on the first forward chain structure feature sequence and the first reverse chain structure feature sequence respectively to obtain a first forward chain output feature sequence and a first reverse chain output feature sequence.

[0246] Here, the preset sequence encoding module includes a sliding window attention layer and a feedforward neural network layer connected in sequence. The sliding window attention layer is as follows: Figure 3 As shown, three sliding attention windows of different sizes and one full attention window may be included. The three sliding attention windows are a window of 1*128, a window of 1*512, and a window of 1*2048, respectively, and the full attention window is a window of 1*8192.

[0247] Among them, the 1*128 window is used to extract the forward chain attention feature map or the reverse chain attention feature map in the window with a sequence length less than or equal to 128, the 1*512 window is used to extract the forward chain attention feature map or the reverse chain attention feature map in the window with a sequence length less than or equal to 512, and the 1*2048 window is used to extract the forward chain attention feature map or the reverse chain attention feature map in the window with a sequence length less than or equal to 2048.

[0248] Here, it is assumed that the longest sequence length supported by the preset sequence encoding module is 8192, so the 1*8192 full attention window is used to extract the overall forward chain attention feature map or the overall reverse chain attention feature map.

[0249] By fusing the forward chain attention feature maps within the window with a length less than or equal to 128, a length less than or equal to 512, and a length less than or equal to 2048 and the overall forward chain attention feature map, the first forward chain output feature sequence can be obtained.

[0250] By fusing the in-window reverse chain attention feature maps with a length less than or equal to 128, a length less than or equal to 512, and a length less than or equal to 2048 and the overall reverse chain attention feature map, a first reverse chain output feature sequence can be obtained.

[0251] Then, the first forward chain output feature sequence and the first reverse chain output feature sequence are fused to obtain a first fused forward chain feature sequence.

[0252] The first fused forward chain feature sequence is then input into a preset sequence feature decoder to obtain a first decoded forward chain sequence.

[0253] Finally, based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence, the model parameters of the preset word element embedding representation module, the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder are adjusted.

[0254] The DNA sequence processing model pre-training provided by the above-mentioned embodiment of the present disclosure is to extract structural features of the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence respectively based on a preset fusion structural feature extraction module to obtain a first forward chain structural feature sequence and a first reverse chain structural feature sequence, wherein the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence are respectively vector sequences obtained by embedding representation based on the first DNA forward chain sequence and the first DNA reverse chain sequence that is reverse complementary to the first DNA forward chain sequence; then, based on the preset sequence encoding module, feature encoding is performed on the first forward chain structural feature sequence and the first reverse chain structural feature sequence respectively to obtain a first forward chain output feature sequence and a first reverse chain output feature sequence; then, based on the first forward chain output feature sequence and the first reverse chain output feature sequence, a first fused forward chain feature sequence is determined; then, the first fused forward chain feature sequence is input into a preset sequence feature decoder to obtain a first decoded forward chain sequence; finally, based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence, the model parameters of the preset fusion structural feature extraction module, the preset sequence encoding module and the preset sequence feature decoder are adjusted. The above method can achieve the following technical effects including but not limited to:

[0255] First, by encoding the first forward strand embedding vector sequence and the first reverse strand embedding vector sequence, the reverse complementary nature between the double helix nucleotide chains of DNA is reflected;

[0256] Second, by extracting structural features using a preset fusion structural feature extraction module, the structural features of the DNA sequence can be reflected;

[0257] Third, by fusing the first forward chain output characteristic sequence and the first reverse chain output characteristic sequence to obtain the first fused forward chain characteristic sequence, the reverse complementary properties between the double-helix nucleotide chains of DNA can be characteristically expressed at the characteristic level, enriching the diversity of DNA sequence characteristic expression information.

[0258] Reference below Figure 4 , which shows a process 400 of an embodiment of a DNA sequence processing method according to the present disclosure. The DNA sequence processing method comprises the following steps:

[0259] Step 401: Based on a preset fusion structural feature extraction module, structural feature extraction is performed on the second forward chain embedding vector sequence and the second reverse chain embedding vector sequence respectively to obtain a second forward chain structural feature sequence and a second reverse chain structural feature sequence.

[0260] Here, the second forward strand embedded vector sequence and the second reverse strand embedded vector sequence are respectively vector sequences obtained by embedding and representing the second DNA forward strand sequence and the second DNA reverse strand sequence which is reverse complementary to the second DNA forward strand sequence.

[0261] The second forward chain embedding vector sequence is composed of the embedding vectors corresponding to each word (i.e., token) in the second DNA forward chain sequence, arranged in the order of the corresponding word in the second DNA forward chain sequence. Specifically, it can be obtained by cutting the second DNA forward chain sequence in units of single nucleotide identifiers. Correspondingly, the second reverse chain embedding vector sequence is composed of the embedding vectors corresponding to each word in the second DNA reverse chain sequence, arranged in the order of the corresponding word in the second DNA reverse chain sequence. Specifically, it can be obtained by cutting the second DNA reverse chain sequence in units of single nucleotide identifiers.

[0262] Here, you can use Figure 2A In the illustrated embodiment, the same or similar method as that described in the embodiment illustrated in step 201 and its optional implementation manners executes step 401, which will not be described in detail herein.

[0263] In some optional implementations, before step 401, the method flow 400 may further include the following steps 401' to 403':

[0264] Step 401', obtaining the forward strand sequence of the DNA to be encoded.

[0265] Here, the forward strand sequence of the DNA to be encoded can be obtained through various sequencing techniques, or through a gene library (eg, a genomic library or a cDNA library), or through reverse transcription, artificial synthesis, genomic databases, and the like.

[0266] Step 402', determining the reverse strand sequence of the DNA to be encoded that is reverse complementary to the forward strand sequence of the DNA to be encoded.

[0267] First, the forward chain sequence of the DNA to be encoded is reversed to obtain the reversed forward chain sequence of the DNA to be encoded.

[0268] Then, the original word element (ie, nucleotide identifier) ​​in the forward chain sequence of the DNA to be encoded after reverse conversion is replaced with a complementary nucleotide identifier to obtain the reverse chain sequence of the DNA to be encoded.

[0269] Step 403', based on a preset word unit embedding representation module, embedding representation is performed on the DNA forward strand sequence to be encoded and the DNA reverse strand sequence to be encoded respectively, to obtain a second forward strand embedding vector sequence and a second reverse strand embedding vector sequence.

[0270] Here, the preset word unit embedding representation module can be any model that is currently known or will be developed in the future to convert word units into vectors. Figure 2A In the optional implementation of the DNA sequence processing model pre-training method shown, the word unit embedding representation module obtained after the model parameter part of the preset word unit embedding representation module is adjusted.

[0271] Step 402: Based on a preset sequence encoding module, feature encoding is performed on the second forward chain structure feature sequence and the second reverse chain structure feature sequence respectively to obtain a second forward chain output feature sequence and a second reverse chain output feature sequence.

[0272] Here, you can use Figure 2A In the illustrated embodiment, the same or similar method as that described in the embodiment shown in step 202 and its optional implementation manner is used to execute step 402, which will not be described in detail herein.

[0273] Step 403: fuse the second forward chain output feature sequence and the second reverse chain output feature sequence to obtain a second fused forward chain feature sequence.

[0274] Here, you can use Figure 2A In the illustrated embodiment, the same or similar method as that described in the embodiment shown in step 203 and its optional implementation manner is used to execute step 403, which will not be described in detail herein.

[0275] Optionally, step 403 may be performed as follows:

[0276] splicing the second forward chain output characteristic sequence and the second reverse chain output characteristic sequence to obtain a second fused forward chain characteristic sequence; or,

[0277] The sum of the second forward chain output characteristic sequence and the characteristic sequence after reverse complementation of the second reverse chain is determined as the second fusion forward chain characteristic sequence.

[0278] Here, the preset fusion structure feature extraction module and the preset sequence encoding module can be implemented as follows: Figure 2A The DNA sequence processing model pre-training method recorded in the illustrated embodiment and its optional implementation mode is pre-trained. The preset fusion structure feature extraction module and the preset sequence encoding module connected in sequence can be used as a DNA sequence encoding pre-training model for encoding DNA sequence data to obtain a feature sequence. In the above DNA sequence encoding pre-training model, the reverse complementary nature between the double helix nucleotide chains of DNA is reflected by feature encoding the forward chain embedding vector sequence and the reverse chain embedding vector sequence; in addition, by extracting structural features using the preset fusion structure feature extraction module, the structural features of the DNA sequence can be reflected; and by fusing the forward chain output feature sequence and the reverse chain output feature sequence to obtain a fused forward chain feature sequence, the reverse complementary nature between the double helix nucleotide chains of DNA can be feature expressed at the feature level, enriching the diversity of DNA sequence feature expression information. Furthermore, based on the above DNA sequence encoding pre-training model, decoders of other DNA sequence analysis tasks can be spliced ​​to perform downstream DNA sequence analysis tasks. Since the above DNA sequence encoding pre-training model has been specially pre-trained for the characteristics of the DNA sequence, the accuracy of the downstream DNA sequence analysis task can be improved.

[0279] In some optional implementations, the above DNA sequence processing method may further include the following steps 404 to 406:

[0280] Step 404, obtaining a DNA sequence analysis result label of the target DNA sequence analysis task for the forward strand sequence of the DNA to be encoded.

[0281] Here, the target DNA sequence analysis task may be various tasks for analyzing DNA sequences, such as gene function prediction, gene mutation impact assessment, and the like.

[0282] The DNA sequence analysis result label of the target DNA sequence analysis task for the forward strand sequence of the DNA to be encoded may be a confirmed correct analysis result of performing the target DNA sequence analysis task on the forward strand sequence of the DNA to be encoded.

[0283] Step 405, input the second fusion forward strand characteristic sequence into the target DNA sequence analysis task decoder to obtain a DNA sequence analysis result.

[0284] Here, the second fused forward chain characteristic sequence obtained by sequence encoding the DNA forward chain sequence to be encoded outputted in step 403 can be inputted into the target DNA sequence analysis task decoder to obtain the DNA sequence analysis result outputted by the target DNA sequence analysis task decoder.

[0285] Step 406: adjust the model parameters of the target DNA sequence analysis task decoder based on the difference between the DNA sequence analysis result and the DNA sequence analysis result label.

[0286] Here, various parameter optimization methods can be used to adjust the model parameters of the target DNA sequence analysis task decoder based on the difference between the DNA sequence analysis result output by the target DNA sequence analysis task decoder in step 405 and the DNA sequence analysis result label obtained in step 404.

[0287] The DNA sequence processing method provided by the above-mentioned embodiment of the present disclosure, steps 404 to 406 can be implemented on the basis of a DNA sequence encoding pre-training model that uses a preset fusion structure feature extraction module and a preset sequence encoding module as the basis for encoding DNA sequence data to obtain a feature sequence, thereby completing fine-tuning of the decoder of the target DNA sequence analysis task. Since the above-mentioned DNA sequence encoding pre-training model has been specially pre-trained according to the characteristics of the DNA sequence, it is only necessary to fine-tune the decoder of the downstream DNA sequence analysis task to improve the accuracy of the downstream DNA sequence analysis task.

[0288] Further references Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a DNA sequence processing model pre-training device. Figure 2A Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0289] like Figure 5 As shown, the DNA sequence processing model pre-training device 500 of this embodiment includes: a first structural feature fusion module 501, a first sequence encoding module 502, a first bidirectional sequence fusion module 503, a first sequence decoding module 504 and a first model training module 505. Among them, the first structural feature fusion module 501 is configured to extract structural features of the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence based on a preset fusion structural feature extraction module, respectively, to obtain a first forward chain structural feature sequence and a first reverse chain structural feature sequence, wherein the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence are respectively vector sequences obtained by embedding representation based on the first DNA forward chain sequence and the first DNA reverse chain sequence that is reverse complementary to the first DNA forward chain sequence; the first sequence encoding module 502 is configured to extract features of the first forward chain structural feature sequence and the first reverse chain structural feature sequence based on a preset sequence encoding module, respectively. encoding to obtain a first forward chain output feature sequence and a first reverse chain output feature sequence; a first bidirectional sequence fusion module 503 is configured to fuse the first forward chain output feature sequence and the first reverse chain output feature sequence to obtain a first fused forward chain feature sequence; a first sequence decoding module 504 is configured to input the first fused forward chain feature sequence into a preset sequence feature decoder to obtain a first decoded forward chain sequence; a first model training module 505 is configured to adjust the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence.

[0290] In this embodiment, the specific processing of the first structural feature fusion module 501, the first sequence encoding module 502, the first bidirectional sequence fusion module 503, the first sequence decoding module 504 and the first model training module 505 of the DNA sequence processing model pre-training device 500 and the technical effects thereof can be referred to respectively. Figure 2A The relevant descriptions of step 201, step 202, step 203, step 204 and step 205 in the corresponding embodiment are not repeated here.

[0291] In some optional embodiments, the DNA sequence processing model pre-training device 500 may also include a first bidirectional sequence generation module 501', which is configured to, before performing structural feature extraction on the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence based on a preset fusion structural feature extraction module: obtain the original forward chain sequence of the sample DNA; generate a first DNA forward chain sequence based on the original forward chain sequence, wherein the first DNA forward chain sequence includes original words that are the same as the original forward chain sequence and words to be predicted that are different from the original forward chain sequence; generate a first DNA reverse chain sequence that is reverse complementary to the first DNA forward chain sequence; and embed and represent the first DNA forward chain sequence and the first DNA reverse chain sequence based on a preset word embedding representation module to obtain a first forward chain embedding vector sequence and a first reverse chain embedding vector sequence.

[0292] In some optional embodiments, generating a first DNA forward strand sequence based on the original forward strand sequence may include:

[0293] Generating a first DNA forward strand sequence identical to the original forward strand sequence;

[0294] Randomly selecting some nucleotide identifiers in the first DNA forward chain sequence as words to be predicted;

[0295] Replacing a first preset proportion of the to-be-predicted words in the first DNA forward strand sequence with random words;

[0296] The to-be-predicted words in the first DNA forward chain sequence that are less than or equal to a second preset ratio are replaced with preset mask words, wherein the sum of the first preset ratio and the second preset ratio is less than or equal to 100%.

[0297] In some optional embodiments, replacing the to-be-predicted words that are less than or equal to the second preset ratio in the first DNA forward strand sequence with preset mask words may include:

[0298] Randomly reducing the second preset ratio according to a preset small probability to obtain a mask ratio;

[0299] The to-be-predicted word of the mask ratio in the first DNA forward chain sequence is replaced with the preset mask word.

[0300] In some optional embodiments, the preset fusion structure feature extraction module includes at least one preset DNA structure feature extraction submodule; and

[0301] The first structural feature fusion module 501 may include:

[0302] The substructure feature extraction unit 5011 is configured to use each of the preset DNA structure feature extraction submodules to respectively perform structure feature extraction on the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence to obtain a forward chain DNA structure feature sequence and a reverse chain DNA structure feature sequence corresponding to the preset DNA structure feature extraction submodule;

[0303] The substructure feature fusion unit 5012 is configured to generate a first forward strand structure feature sequence and a first reverse strand structure feature sequence based on the forward strand DNA structure feature sequence and the reverse strand DNA structure feature sequence corresponding to each of the preset DNA structure feature extraction submodules.

[0304] In some optional embodiments, the at least one preset DNA structure feature extraction submodule may include at least one of the following: a first DNA structure feature map extraction submodule, a second DNA structure feature map extraction submodule and a third DNA structure feature map extraction submodule, wherein the first DNA structure feature map extraction submodule includes a first convolution layer composed of at least one 1*3 convolution kernel, the second DNA structure feature map extraction submodule includes a second convolution layer composed of at least one 1*5 convolution kernel, and the third DNA structure feature map extraction submodule includes a third convolution layer composed of at least one 1*7 convolution kernel, wherein the first convolution layer is used to extract codon structure and / or DNA minor groove structure, the second convolution layer is used to extract DNA minor groove structure and / or DNA major groove structure, and the third convolution layer is used to extract DNA major groove structure.

[0305] In some optional embodiments, the preset sequence encoding module may include a sliding window attention layer and a feedforward neural network layer connected in sequence, the sliding window attention layer is used to use at least two attention windows of different sizes to perform attention weight extraction on the first forward chain structure feature sequence and the first reverse chain structure feature sequence, respectively, to obtain the in-window forward chain attention feature map and the in-window reverse chain attention feature map corresponding to the corresponding attention window, and the in-window forward chain attention feature map and the in-window reverse chain attention feature map extracted based on each attention window, respectively, to generate a forward chain attention feature map and a reverse chain attention feature map, and the feedforward neural network layer is used to process the forward chain attention feature map and the reverse chain attention feature map, respectively, to obtain the first forward chain output feature sequence and the first reverse chain output feature sequence.

[0306] In some optional implementations, the first bidirectional sequence fusion module 503 may include:

[0307] The reverse unit 5031 is configured to perform reverse processing on the first reverse chain output feature sequence to obtain a reversed feature sequence of the first reverse chain;

[0308] The complementary unit 5032 is configured to perform a linear transformation on the first reverse chain reversed feature sequence based on a preset complementary linear transformation module to obtain a first reverse chain reverse complementary feature sequence;

[0309] The bidirectional fusion unit 5033 is configured to fuse the first forward chain output characteristic sequence and the first reverse chain reverse complement characteristic sequence to obtain the first fused forward chain characteristic sequence.

[0310] In some optional implementations, the first model training module 505 may be further configured to:

[0311] Based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence, the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module, the preset complementary linear transformation module and the preset sequence feature decoder are adjusted.

[0312] In some optional implementations, the first model training module 505 may be further configured to:

[0313] Based on the difference between the first decoded forward chain sequence and the part corresponding to the word to be predicted in the original forward chain sequence corresponding to the first DNA forward chain sequence, the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder are adjusted.

[0314] It should be noted that the implementation details and technical effects of each module in the molecular identification sequence encoding module device provided in the embodiment of the present disclosure can be referred to the description of other embodiments in the present disclosure, and will not be repeated here.

[0315] Reference below Figure 6 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a DNA sequence encoding device. Figure 4 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0316] like Figure 6 As shown, the DNA sequence encoding device 600 of this embodiment includes: a second structural feature fusion module 601, a second sequence encoding module 602 and a second bidirectional sequence fusion module 603. Among them, the second structural feature fusion module 601 is configured to extract structural features of the second forward chain embedded vector sequence and the second reverse chain embedded vector sequence respectively based on the preset fusion structural feature extraction module to obtain the second forward chain structural feature sequence and the second reverse chain structural feature sequence, wherein the second forward chain embedded vector sequence and the second reverse chain embedded vector sequence are respectively vector sequences obtained by embedding representation based on the second DNA forward chain sequence and the second DNA reverse chain sequence that is reverse complementary to the second DNA forward chain sequence; the second sequence encoding module 602 is configured to perform feature encoding on the second forward chain structural feature sequence and the second reverse chain structural feature sequence respectively based on the preset sequence encoding module to obtain the second forward chain output feature sequence and the second reverse chain output feature sequence, wherein the preset fusion structural feature extraction module and the preset sequence encoding module are pre-trained by the method described in any implementation method of the first aspect; the second bidirectional sequence fusion module 603 is configured to fuse the second forward chain output feature sequence and the second reverse chain output feature sequence to obtain the second fused forward chain feature sequence.

[0317] In this embodiment, the specific processing of the second structural feature fusion module 601, the second sequence encoding module 602 and the second bidirectional sequence fusion module 603 of the DNA sequence encoding device 600 and the technical effects thereof can be referred to respectively. Figure 4 The relevant descriptions of step 401, step 402 and step 403 in the corresponding embodiment are not repeated here.

[0318] In some optional embodiments, the DNA sequence encoding device 600 may further include a second bidirectional sequence generating module 601', which is configured to: before extracting structural features from the second forward strand embedding vector sequence and the second reverse strand embedding vector sequence based on the preset fusion structural feature extraction module:

[0319] Obtaining the forward strand sequence of the DNA to be encoded;

[0320] Determining a reverse strand sequence of the DNA to be encoded that is reverse complementary to the forward strand sequence of the DNA to be encoded;

[0321] Based on a preset word element embedding representation module, the DNA forward chain sequence to be encoded and the DNA reverse chain sequence to be encoded are respectively embedded and represented to obtain a second forward chain embedding vector sequence and a second reverse chain embedding vector sequence.

[0322] In some optional embodiments, the DNA sequence encoding device 600 may further include:

[0323] The label acquisition module 604 is configured to acquire the DNA sequence analysis result label of the forward strand sequence of the DNA to be encoded for the target DNA sequence analysis task;

[0324] The sequence analysis module 605 is configured to input the second fusion forward strand characteristic sequence into a target DNA sequence analysis task decoder to obtain a DNA sequence analysis result;

[0325] The second model training module 606 is configured to adjust the model parameters of the target DNA sequence analysis task decoder based on the difference between the DNA sequence analysis result and the DNA sequence analysis result label.

[0326] In some optional implementations, the second bidirectional sequence fusion module 603 may be further configured as follows:

[0327] The second forward chain output characteristic sequence and the second reverse chain output characteristic sequence are spliced ​​to obtain the second fused forward chain characteristic sequence; or, the sum of the second forward chain output characteristic sequence and the reverse complement characteristic sequence of the second reverse chain is determined as the second fused forward chain characteristic sequence.

[0328] It should be noted that the implementation details and technical effects of each module in the molecular identification sequence encoding module device provided in the embodiment of the present disclosure can be referred to the description of other embodiments in the present disclosure, and will not be repeated here.

[0329] Reference below Figure 7 , which shows a schematic diagram of the structure of a computer system 700 suitable for implementing the electronic device of the present disclosure. Figure 7 The computer system 700 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0330] like Figure 7 As shown, the computer system 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 to a random access memory (RAM) 703. Various programs and data required for the operation of the computer system 700 are also stored in the RAM 703. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0331] Typically, the following devices may be connected to the I / O interface 705: input devices 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, etc.; output devices 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 708 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 709. The communication devices 709 may allow the computer system 700 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 7 The computer system 700 of the electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0332] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 709, or installed from a storage device 708, or installed from a ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0333] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0334] The computer-readable medium may be included in the electronic device, or may exist independently without being installed in the electronic device.

[0335] The computer readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device implements the following Figure 2A The DNA sequence processing model pre-training method shown in the embodiment and its optional implementation mode and / or Figure 4 The illustrated embodiment and its alternative implementations illustrate a method for processing DNA sequences.

[0336] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages, such as Python, Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0337] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0338] The modules involved in the embodiments described in the present disclosure may be implemented by software or hardware. The name of the module does not limit the module itself in some cases. For example, the first sequence decoding module may also be described as "a module for inputting the first fused forward chain feature sequence into a preset sequence feature decoder to obtain a first decoded forward chain sequence".

[0339] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.< / mask>

Claims

1. A DNA sequence processing model pre-training method, comprising: Based on a preset fusion structural feature extraction module, structural feature extraction is performed on the first forward chain embedded vector sequence and the first reverse chain embedded vector sequence respectively to obtain a first forward chain structural feature sequence and a first reverse chain structural feature sequence, wherein the first forward chain embedded vector sequence and the first reverse chain embedded vector sequence are respectively vector sequences obtained by embedding a first DNA forward chain sequence and a first DNA reverse chain sequence that is reverse complementary to the first DNA forward chain sequence; Based on a preset sequence encoding module, feature encoding is performed on the first forward chain structure feature sequence and the first reverse chain structure feature sequence respectively to obtain a first forward chain output feature sequence and a first reverse chain output feature sequence; Fusing the first forward chain output characteristic sequence and the first reverse chain output characteristic sequence to obtain a first fused forward chain characteristic sequence; Inputting the first fused forward chain feature sequence into a preset sequence feature decoder to obtain a first decoded forward chain sequence; The model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder are adjusted based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence.

2. The method according to claim 1, wherein: Before respectively performing structural feature extraction on the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence based on the preset fusion structural feature extraction module, the method further includes: Obtain the original forward strand sequence of the sample DNA; Generating a first DNA forward strand sequence based on the original forward strand sequence, wherein the first DNA forward strand sequence includes original words that are the same as the original forward strand sequence and words to be predicted that are different from the original forward strand sequence; generating a first DNA reverse strand sequence which is reverse complementary to the first DNA forward strand sequence; Based on a preset word element embedding representation module, the first DNA forward chain sequence and the first DNA reverse chain sequence are respectively embedded and represented to obtain a first forward chain embedding vector sequence and a first reverse chain embedding vector sequence.

3. The method according to claim 1, wherein: The preset fusion structure feature extraction module includes at least one preset DNA structure feature extraction submodule; as well as The method of extracting structural features from the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence based on the preset fusion structural feature extraction module to obtain the first forward chain structural feature sequence and the first reverse chain structural feature sequence includes: Using each of the preset DNA structure feature extraction submodules, respectively extract structure features from the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence to obtain a forward chain DNA structure feature sequence and a reverse chain DNA structure feature sequence corresponding to the preset DNA structure feature extraction submodule; A first forward strand structural feature sequence and a first reverse strand structural feature sequence are generated based on the forward strand DNA structural feature sequence and the reverse strand DNA structural feature sequence corresponding to each of the preset DNA structural feature extraction submodules, respectively.

4. The method according to claim 3, wherein: The at least one preset DNA structure feature extraction submodule includes at least one of the following: a first DNA structure feature map extraction submodule, a second DNA structure feature map extraction submodule and a third DNA structure feature map extraction submodule, wherein the first DNA structure feature map extraction submodule includes a first convolution layer composed of at least one 1*3 convolution kernel, the second DNA structure feature map extraction submodule includes a second convolution layer composed of at least one 1*5 convolution kernel, and the third DNA structure feature map extraction submodule includes a third convolution layer composed of at least one 1*7 convolution kernel, wherein the first convolution layer is used to extract codon structure and / or DNA minor groove structure, the second convolution layer is used to extract DNA minor groove structure and / or DNA major groove structure, and the third convolution layer is used to extract DNA major groove structure.

5. The method according to claim 1, wherein: The preset sequence encoding module includes a sliding window attention layer and a feedforward neural network layer connected in sequence. The sliding window attention layer is used to use at least two attention windows of different sizes to extract attention weights for the first forward chain structure feature sequence and the first reverse chain structure feature sequence respectively, and obtain the in-window forward chain attention feature map and the in-window reverse chain attention feature map corresponding to the corresponding attention window, and generate the forward chain attention feature map and the reverse chain attention feature map based on the in-window forward chain attention feature map and the in-window reverse chain attention feature map extracted from each attention window respectively. The feedforward neural network layer is used to process the forward chain attention feature map and the reverse chain attention feature map respectively to obtain the first forward chain output feature sequence and the first reverse chain output feature sequence.

6. The method according to claim 1, wherein: The fusing the first forward chain output characteristic sequence and the first reverse chain output characteristic sequence to obtain a first fused forward chain characteristic sequence comprises: Reverse processing is performed on the first reverse chain output feature sequence to obtain a reversed feature sequence of the first reverse chain; Based on a preset complementary linear transformation module, a linear transformation is performed on the reversed feature sequence of the first reverse chain to obtain a reversed complementary feature sequence of the first reverse chain; The first forward chain output characteristic sequence and the first reverse chain reverse complement characteristic sequence are fused to obtain the first fused forward chain characteristic sequence.

7. A DNA sequence processing method, comprising: Based on a preset fusion structural feature extraction module, structural feature extraction is performed on the second forward chain embedded vector sequence and the second reverse chain embedded vector sequence respectively to obtain a second forward chain structural feature sequence and a second reverse chain structural feature sequence, wherein the second forward chain embedded vector sequence and the second reverse chain embedded vector sequence are respectively vector sequences obtained by embedding representation based on a second DNA forward chain sequence and a second DNA reverse chain sequence that is reverse complementary to the second DNA forward chain sequence; Based on a preset sequence encoding module, feature encoding is performed on the second forward chain structure feature sequence and the second reverse chain structure feature sequence respectively to obtain a second forward chain output feature sequence and a second reverse chain output feature sequence, wherein the preset fusion structure feature extraction module and the preset sequence encoding module are pre-trained by the method according to any one of claims 1 to 6; The second forward chain output characteristic sequence and the second reverse chain output characteristic sequence are fused to obtain a second fused forward chain characteristic sequence.

8. The method according to claim 7, wherein: Before respectively extracting structural features from the second forward chain embedding vector sequence and the second reverse chain embedding vector sequence based on the preset fusion structural feature extraction module, the method further includes: Obtaining the forward strand sequence of the DNA to be encoded; Determining a reverse strand sequence of the DNA to be encoded that is reverse complementary to the forward strand sequence of the DNA to be encoded; Based on a preset word element embedding representation module, the DNA forward chain sequence to be encoded and the DNA reverse chain sequence to be encoded are respectively embedded and represented to obtain a second forward chain embedding vector sequence and a second reverse chain embedding vector sequence.

9. The method according to claim 8, wherein: The method further comprises: Obtaining a DNA sequence analysis result label of the target DNA sequence analysis task for the forward strand sequence of the DNA to be encoded; Inputting the second fusion forward strand characteristic sequence into the target DNA sequence analysis task decoder to obtain a DNA sequence analysis result; The model parameters of the target DNA sequence analysis task decoder are adjusted based on the difference between the DNA sequence analysis result and the DNA sequence analysis result label.

10. A DNA sequence processing model pre-training device, comprising: A first structural feature fusion module is configured to extract structural features from the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence based on a preset fusion structural feature extraction module, respectively, to obtain a first forward chain structural feature sequence and a first reverse chain structural feature sequence, wherein the first forward chain embedding vector sequence and the first reverse chain embedding vector sequence are respectively vector sequences obtained by embedding representation based on a first DNA forward chain sequence and a first DNA reverse chain sequence that is reverse complementary to the first DNA forward chain sequence; A first sequence encoding module is configured to perform feature encoding on the first forward chain structure feature sequence and the first reverse chain structure feature sequence based on a preset sequence encoding module, respectively, to obtain a first forward chain output feature sequence and a first reverse chain output feature sequence; A first bidirectional sequence fusion module is configured to fuse the first forward chain output feature sequence and the first reverse chain output feature sequence to obtain a first fused forward chain feature sequence; A first sequence decoding module is configured to input the first fused forward chain feature sequence into a preset sequence feature decoder to obtain a first decoded forward chain sequence; The first model training module is configured to adjust the model parameters of the preset fusion structure feature extraction module, the preset sequence encoding module and the preset sequence feature decoder based on the difference between the first decoded forward chain sequence and the original forward chain sequence corresponding to the first DNA forward chain sequence.

11. A DNA sequence encoding device, comprising: A second structural feature fusion module is configured to extract structural features from the second forward chain embedding vector sequence and the second reverse chain embedding vector sequence based on the preset fusion structural feature extraction module, respectively, to obtain a second forward chain structural feature sequence and a second reverse chain structural feature sequence, wherein the second forward chain embedding vector sequence and the second reverse chain embedding vector sequence are respectively vector sequences obtained by embedding representation based on a second DNA forward chain sequence and a second DNA reverse chain sequence that is reverse complementary to the second DNA forward chain sequence; A second sequence encoding module is configured to perform feature encoding on the second forward chain structure feature sequence and the second reverse chain structure feature sequence respectively based on a preset sequence encoding module to obtain a second forward chain output feature sequence and a second reverse chain output feature sequence, wherein the preset fusion structure feature extraction module and the preset sequence encoding module are pre-trained by the method according to any one of claims 1 to 6; The second bidirectional sequence fusion module is configured to fuse the second forward chain output feature sequence and the second reverse chain output feature sequence to obtain a second fused forward chain feature sequence.

12. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 6 and / or the method according to any one of claims 7 to 9.

13. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by one or more processors, the method according to any one of claims 1 to 6 and / or the method according to any one of claims 7 to 9 is implemented.

14. A computer program product, comprising a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the method according to any one of claims 1 to 6 and / or the method according to any one of claims 7 to 9 is implemented.

Citation Information

Cited By

  • Probiotic tolerance characteristic prediction method and device, electronic equipment and storage medium

    CN122117042A