Base recognition method, base recognition model training method, and electronic device
Through self-supervised pre-training and supervised optimization methods, the base recognition model is trained using label-free data, which solves the problems of high training cost and low prediction accuracy, and achieves more efficient base recognition.
Patent Information
- Application Number
- PCT/CN2024/073973
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-31
AI Technical Summary
In the prior art, the base recognition model based on deep learning is expensive to train and has low prediction accuracy, especially when the training data is insufficient, the prediction accuracy of the model cannot be guaranteed.
The self-supervised pre-training method combined with the supervised optimization method is used to train the base recognition model through label-free historical sequencing data, and the encoded feature vectors and context relationships of the sequencing data are extracted using the feature encoder and context network, and decoded using the decoder to reduce the dependence on labeled data.
It reduces the difficulty and cost of model training, improves the accuracy of base recognition, reduces the need for a large amount of labeled data, and improves the prediction accuracy of the model.
Smart Images

Figure CN2024073973_31072025_PF_FP_ABST
Abstract
Description
Base recognition method, base recognition model training method and electronic equipment Technical Field
[0001] The present application relates to the field of biomedicine technology, and specifically to a base recognition method, a base recognition model training method, and an electronic device. Background Art
[0002] Nanopore sequencing is a high-throughput sequencing method based on single-molecule current measurement. It uses nanopores made of proteins or solid-state materials to guide deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA) molecules one by one through the pore. The current is then measured as different bases in the DNA and / or RNA pass through the pore, generating the corresponding electrical signals as sequencing data. The base sequence of the DNA and / or RNA can then be inferred by analyzing the sequencing data.
[0003] In the related art, when training deep learning-based base recognition models, supervised training methods are typically used. For example, this requires a large amount of sequencing data, including real base sequences as labels, as training data, which is costly. Furthermore, supervised training methods have a long model training iteration cycle. The more training data, the longer the model training takes. When training data is insufficient, the model's prediction accuracy cannot be guaranteed. Therefore, the base recognition models in the related art suffer from high training costs and low prediction accuracy.
[0004] Summary of the Invention
[0005] In view of the above, it is necessary to propose a base recognition method, a base recognition model training method and an electronic device that can solve the problems of high training cost and low prediction accuracy of the base recognition model in related technologies.
[0006] An embodiment of the present application provides a base recognition method, which includes: inputting sequencing data into a pre-trained base recognition model; using a feature encoder of the base recognition model to feature encode the sequencing data to obtain multiple encoded feature vectors corresponding to the sequencing data; using a context network of the base recognition model to extract timing information and contextual relationships between the multiple encoded feature vectors to obtain multiple context feature vectors corresponding to the sequencing data; inputting the multiple context feature vectors into a preset decoder, and using the decoder to decode the multiple context feature vectors to obtain a base sequence corresponding to the sequencing data.
[0007] In one embodiment, the method further includes collecting the sequencing data, including: collecting electrical signal data obtained by nanopore sequencing of deoxyribonucleic acid DNA or ribonucleic acid RNA, the electrical signal data including electrical signal amplitude data that changes with time; and using the electrical signal data as the sequencing data.
[0008] In one embodiment, the base recognition model training method includes a self-supervised pre-training method and a supervised optimization method.
[0009] In one embodiment, the feature encoder includes a preset first number of network modules, wherein each network module includes: a one-dimensional convolution layer for extracting features from input data, wherein the input data includes the sequencing data; a normalization layer for normalizing the output data of the one-dimensional convolution layer; and an activation function unit for performing nonlinear mapping on the output data of the normalization layer.
[0010] In one embodiment, the context network includes a preset second number of Transformer network structures.
[0011] In one embodiment, the preset decoder includes a Connection Temporal Classification (CTC) decoder.
[0012] An embodiment of the present application provides a base recognition model training method, the method comprising: collecting sample data, the sample data comprising unlabeled historical sequencing data and labeled historical sequencing data; performing self-supervised pre-training on a preset model based on the unlabeled historical sequencing data to obtain a preset model that meets preset requirements as an initial model, wherein the preset model comprises a feature encoder, a quantization module, a mask module, and a context network; optimizing the initial model using the labeled historical sequencing data, and using the optimized initial model as the base recognition model.
[0013] In one embodiment, the collecting sample data includes: collecting historical electrical signal data obtained by nanopore sequencing of historical deoxyribonucleic acid DNA or historical ribonucleic acid RNA, the historical electrical signal data including electrical signal amplitude data that changes with time; using the historical electrical signal data as the unlabeled historical sequencing data; obtaining the true base sequence of the historical DNA or the historical RNA corresponding to the historical electrical signal data; and using the true base sequence as a label for the corresponding historical electrical signal data to obtain the labeled historical sequencing data.
[0014] In one embodiment, before performing self-supervised pre-training on a preset model based on the unlabeled historical sequencing data, the method further includes: preprocessing the unlabeled historical sequencing data, wherein the preprocessing includes one or more of data cleaning processing, data normalization processing, data interpolation processing, and data segmentation processing.
[0015] In one embodiment, the self-supervised pre-training of the preset model based on the unlabeled historical sequencing data includes: inputting the unlabeled historical sequencing data into the feature encoder, using the feature encoder to perform feature encoding on the data of multiple time steps in the historical sequencing data to obtain multiple encoded feature vectors; using the quantization module to perform a product quantization operation on the multiple encoded feature vectors to obtain multiple quantized feature vectors; using the masking module to mask the multiple encoded feature vectors to obtain multiple masked encoded feature vectors; using the context network to extract the timing information and context relationship between the multiple masked encoded feature vectors to obtain multiple context feature vectors; using each quantized feature vector as the learning target of the corresponding context feature vector, constructing a loss function based on the learning target, and if the loss function reaches a preset convergence condition, it is determined that the preset model meets the preset requirements.
[0016] In one embodiment, the using the quantization module to perform a product quantization operation on the multiple coding feature vectors to obtain multiple quantized feature vectors includes: determining multiple code table vectors corresponding to each coding feature vector in the multiple coding feature vectors, including: splitting each coding feature vector into multiple discrete sub-vectors on average, and using a preset quantization method to determine the code table vector corresponding to each discrete sub-vector from each code table in a preset multiple code tables to obtain the multiple code table vectors; wherein the quantization method includes a Gumbel-Softmax quantization method; and splicing the multiple code table vectors into a quantized feature vector corresponding to each coding feature vector.
[0017] In one embodiment, the masking module is used to mask the multiple encoding feature vectors to obtain multiple masked encoding feature vectors, including: determining the sequence length of the data sequence corresponding to each encoding feature vector in the multiple encoding feature vectors; determining the number of mask positions for masking the data sequence based on the sequence length, a preset mask probability value, and a preset mask length; randomly selecting multiple mask positions corresponding to the number of mask positions from the data sequence; at each mask position in the multiple mask positions, using a preset mask vector to mask multiple continuous data corresponding to the mask length in the data sequence to obtain a masked data sequence; wherein the dimension of the mask vector is equal to the mask length; and determining each masked encoding feature vector based on the masked data sequence.
[0018] In one embodiment, the method also includes determining multiple interference vectors, including: determining multiple mask positions when masking the multiple encoding feature vectors using the mask module; selecting multiple interference positions outside the multiple mask positions from the data sequence composed of the multiple encoding feature vectors; and determining an interference vector at each interference position among the multiple interference positions, wherein the dimension of the interference vector is equal to the dimension of each encoding feature vector.
[0019] In one embodiment, the loss function includes a first loss function, and the formula used by the first loss function includes:
[0020] Among them, L m Denotes the first loss function, q′∈Q t ={q t ,q′1,q′2,…,q′ i ,…,q′ K},q t represents the quantized feature vector at the tth time step, c t represents the context feature vector of the tth time step, sim() represents the preset similarity function, q′ i represents the i-th interferer vector, i is a positive integer in the range of [1, K], K represents the total number of interferer vectors; k represents the preset temperature parameter.
[0021] In one embodiment, the loss function includes a second loss function, and the formula used by the second loss function includes:
[0022] Among them, L d represents the second loss function, It represents the average probability that all encoded feature vectors are mapped to the vth code table vector in the gth code table, v represents the vth code table vector, the value of v is a positive integer in the range [1, V], V represents the total number of code table vectors in each code table vector; the value of g is a positive integer in the range [1, G], G represents the total number of code tables; -H() represents the entropy function.
[0023] In one embodiment, the loss function includes a weighted sum of a preset first loss function and a preset second loss function.
[0024] In one embodiment, optimizing the initial model using the labeled historical sequencing data includes: inputting the labeled historical sequencing data into the initial model; performing feature encoding on the labeled historical sequencing data using the feature encoder of the initial model to obtain a plurality of encoded feature vectors corresponding to the labeled historical sequencing data; extracting the timing information and contextual relationship between the plurality of encoded feature vectors using the context network of the initial model to obtain a plurality of context feature vectors corresponding to the labeled historical sequencing data; inputting the plurality of context feature vectors into a preset decoder, decoding the plurality of context feature vectors using the decoder to obtain a predicted base sequence corresponding to the labeled historical sequencing data; and determining a loss value between the predicted base sequence corresponding to the labeled historical sequencing data and the true base sequence using a preset classification loss function, and adjusting the model parameters of the initial model according to the loss value.
[0025] An embodiment of the present application provides a base recognition device, which includes: an input module for inputting sequencing data into a pre-trained base recognition model; a feature encoding module for using the feature encoder of the base recognition model to feature encode the sequencing data to obtain multiple encoded feature vectors corresponding to the sequencing data; a context learning module for using the context network of the base recognition model to extract the timing information and context relationship between the multiple encoded feature vectors to obtain multiple context feature vectors corresponding to the sequencing data; and a decoding module for inputting the multiple context feature vectors into a preset decoder, and using the decoder to decode the multiple context feature vectors to obtain a base sequence corresponding to the sequencing data.
[0026] An embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the processor is configured to implement the base recognition method or the base recognition model training method when executing a computer program stored in the memory.
[0027] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the base recognition method or the base recognition model training method is implemented.
[0028] In summary, the base recognition method described in the present application can obtain multiple encoding feature vectors of the sequencing data by inputting the sequencing data into a pre-trained base recognition model, using the feature encoder of the base recognition model, and using the context network of the base recognition model to extract the temporal information and contextual relationship between the multiple encoding feature vectors, thereby obtaining multiple context feature vectors with context-related features, and then using the decoder to decode the multiple context feature vectors, which can improve the accuracy of identifying the base sequence of the sequencing data. The base recognition model training method described in the present application realizes the algorithm development of the base recognition model through the pre-training + fine-tuning training method. Compared with the traditional supervised learning idea, it can use more nanopore sequencing data, can improve the convergence speed of the model, and can eliminate the need to obtain a large amount of labeled data after annotation, thereby reducing the training difficulty and training cost of the model, improving the prediction accuracy of the model, and thus improving the accuracy of base recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] FIG1 is a structural diagram of an electronic device provided in one embodiment of the present application.
[0030] FIG2 is a flow chart of a base recognition method provided in one embodiment of the present application.
[0031] FIG3 is an example diagram of the application reasoning process of the base recognition model provided in one embodiment of the present application.
[0032] FIG4 is an example diagram of a position encoder provided in an embodiment of the present application.
[0033] FIG5 is a flowchart of a base recognition model training method provided in an embodiment of the present application.
[0034] FIG6 is an example diagram of the architecture of the base recognition model training process provided in one embodiment of the present application.
[0035] FIG7 is a flowchart of a method for self-supervised pre-training of a preset model provided by an embodiment of the present application.
[0036] FIG8 is an example diagram of how the loss value corresponding to the loss function of the pre-training process provided in an embodiment of the present application changes with the training process.
[0037] FIG9 is a structural diagram of a base recognition device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0038] In order to more clearly understand the above-mentioned objectives, features and advantages of the present application, the present application is described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features therein can be combined with each other in the absence of conflict.
[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application pertains. The terms used herein in the specification of this application are for the purpose of describing embodiments in one embodiment only and are not intended to limit this application.
[0040] It should be noted that in this application, "at least one" means one or more, and "more than one" means two or more than two. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of this application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0041] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete manner. The following embodiments and features in the embodiments may be combined with each other unless there is a conflict.
[0042] In one embodiment, nanopore sequencing technology is a high-throughput sequencing method based on single-molecule current measurement. It can use nanopores composed of proteins or solid-state materials to guide deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA) molecules one by one through the pore. The current is measured when different bases in the DNA and / or RNA pass through the pore, and the electrical signal data corresponding to the DNA and / or RNA is obtained as sequencing data. By analyzing the sequencing data, the base sequence of the DNA and / or RNA can be inferred.
[0043] In the related art, when training deep learning-based base recognition models, supervised training methods are typically used. For example, this requires a large amount of sequencing data, including real base sequences as labels, as training data, which is costly. Furthermore, supervised training methods have a long model training iteration cycle. The more training data, the longer the model training takes. When training data is insufficient, the model's prediction accuracy cannot be guaranteed. Therefore, the base recognition models in the related art suffer from high training costs and low prediction accuracy.
[0044] To solve the above problems, the embodiment of the present application provides a base recognition method, which inputs sequencing data into a pre-trained base recognition model, uses the feature encoder of the base recognition model to obtain multiple encoding feature vectors of the sequencing data, and uses the context network of the base recognition model to extract the temporal information and contextual relationship between the multiple encoding feature vectors, thereby obtaining multiple context feature vectors with context-related features, and then uses the decoder to decode the multiple context feature vectors to obtain the base sequence of the sequencing data. Among them, the base recognition model is trained by the base recognition model training method provided in the embodiment of the present application, and the preset model is self-supervised pre-trained using unlabeled historical sequencing data, and the initial model obtained by pre-training is fine-tuned to obtain the base recognition model. It is possible to obtain a large amount of labeled data without the need to obtain labeled data, thereby reducing the training difficulty and training cost of the model, improving the prediction accuracy of the model, and thus improving the accuracy of base recognition.
[0045] Figure 1 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 10 can be a computer, server, mobile phone, tablet computer, laptop computer, cloud server, cloud computer, etc. The embodiment of the present application does not impose any restrictions on the specific type of electronic device.
[0046] As shown in Figure 1, the electronic device 10 may include a communication module 101, a memory 102, a processor 103, an input / output (I / O) interface 104, and a bus 105. The processor 103 is coupled to the communication interface 101, the memory 102, and the I / O interface 104 via the bus 105.
[0047] The communication module 101 may include a wired communication module and / or a wireless communication module. The wired communication module may provide one or more wired communication solutions such as universal serial bus (USB) and controller area network bus (CAN). The wireless communication module may provide one or more wireless communication solutions such as wireless fidelity (Wi-Fi), Bluetooth (BT), mobile communication network, frequency modulation (FM), near field communication technology (NFC), infrared technology (IR), etc.
[0048] The memory 102 may include one or more random access memories (RAMs) and one or more non-volatile memories (NVMs). The RAM can be directly read and written by the processor 103 and can be used to store executable programs (e.g., machine instructions) of the operating system or other running programs, as well as user and application data. The RAM may include static random-access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), etc.
[0049] The non-volatile memory can also store executable programs and user and application data, etc., and can be pre-loaded into the random access memory for direct reading and writing by the processor 110. The non-volatile memory can include disk storage devices and flash memory.
[0050] The memory 102 is configured to store one or more computer programs. The one or more computer programs are configured to be executed by the processor 103. The one or more computer programs include multiple instructions. When the multiple instructions are executed by the processor 103, the base calling method executed on the electronic device 10 can be implemented.
[0051] In other embodiments, the electronic device 10 further includes an external memory interface for connecting to an external memory to expand the storage capacity of the electronic device 10 .
[0052] The processor 103 may include one or more processing units. For example, the processor 103 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0053] The processor 103 provides computing and control capabilities. For example, the processor 103 is configured to execute a computer program stored in the memory 102 to implement the above-mentioned base recognition method.
[0054] The I / O interface 104 provides a channel for user input or output. For example, the I / O interface 104 can be used to connect various input and output devices, such as a mouse, keyboard, touch screen device, and display screen, allowing users to enter information or visualize information. Alternatively, the I / O interface 104 can also provide a data transmission channel with a nanopore sequencing device, allowing electronic devices to obtain sequencing data (e.g., electrical signals) from the nanopore sequencing device.
[0055] The bus 105 is at least used to provide a channel for mutual communication among the communication module 101 , the memory 102 , the processor 103 , and the I / O interface 104 in the electronic device 10 .
[0056] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 10. In other embodiments of the present application, the electronic device 10 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0057] FIG2 is a flow chart of a base recognition method according to an embodiment of the present application. The base recognition method is applied to an electronic device, such as the electronic device 10 in FIG1 , and specifically includes the following steps. The order of the steps in the flow chart may be changed, and some steps may be omitted, depending on different requirements.
[0058] S201, inputting sequencing data into a pre-trained base recognition model.
[0059] In one embodiment, the method further includes collecting the sequencing data, including: collecting electrical signal data obtained by nanopore sequencing of deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), the electrical signal data including electrical signal amplitude data that changes with time; and using the electrical signal data as the sequencing data.
[0060] In one embodiment, the electronic device can obtain electrical signal data obtained by sequencing DNA or RNA from a nanopore sequencing device, wherein the electrical signal data can represent electrical signal data obtained during the current sequencing process, or can be electrical signal data obtained after the sequencing process is completed. For example, the electronic device can collect the original electrical signal file (e.g., a file in fast5 format) obtained after the nanogene sequencer sequences DNA or RNA, and read the original electrical signal data of the DNA or RNA from the original electrical signal file as electrical signal data. The electrical signal data may include electrical signal amplitude data that changes over time. For example, as shown in FIG3 , the electrical signal data may be a raw waveform of a segment of an electrical signal, wherein each amplitude in the waveform represents the electrical signal amplitude at each time node, and thus the electrical signal data can be regarded as a one-dimensional data vector or sequence composed of multiple electrical signal amplitudes.
[0061] In one embodiment, based on the length of input data that the base recognition model can receive, the electrical signal data of the corresponding length can be input as sequencing data into the base recognition model, so that the base recognition model can recognize it, and there can be no overlapping data between different sequencing data. For example, if the length of input data that the base recognition model can receive is N, then the sequencing data can be a data sequence {x1, x2, ..., x n ,…,x N}, where each data can be represented as x n , n={1,2,...,N}, N represents the total length of the sequencing data, and the value of N can be set according to actual needs. This application does not impose any specific restrictions on this. For example, the value of N is 5000. The training method of the base recognition model can refer to the description of the embodiment shown in FIG5 .
[0062] S202 , performing feature encoding on the sequencing data using a feature encoder of the base recognition model to obtain a plurality of encoding feature vectors corresponding to the sequencing data.
[0063] In one embodiment, the process of feature encoding the sequencing data by the feature encoder can be understood as the process of capturing the hidden or underlying data representation of the sequencing data, so as to obtain the most relevant and meaningful latent representations (e.g., encoded feature vectors) in the sequencing data. For example, the dimension N*1 sequencing data {x1, x2, ..., x n ,…,x N} Input the feature encoder and get T m*1 dimensional encoding feature vectors {z1,z2,…,z t ,…,z T}, where the encoding feature vector z t It can be considered as a vector obtained by feature encoding the data at the tth time step in the sequencing data, where t = {1, 2, ..., T}, and T represents the total number of encoded feature vectors. In one example, the values of N, T, and m can be set according to actual needs and are not specifically limited in this application. For example, N can be a value such as 5000, T can be a value such as 800, 1000, 1250, and m can be a value such as 768.
[0064] In one embodiment, the feature encoder may include a preset first number of network modules, wherein each network module includes: a one-dimensional convolutional layer (1D Convolutional Neural Network, 1D-CNN\1D-conv), for extracting features from input data, wherein the input data includes the sequencing data, such as sequencing data at each time step; a normalization layer, for normalizing the output data of the one-dimensional convolution layer; and an activation function unit, for performing nonlinear mapping on the output data of the normalization layer. The preset first number can be set according to actual needs, and can also be determined through a parameter optimization process during model training. For example, according to the model training process, it can be determined that the first number is 6 (for example, as shown in FIG3 ), indicating that the performance of the model obtained by superimposing 6 of the above-mentioned network modules in the feature encoder is the best.
[0065] In one embodiment, since sequencing data is a one-dimensional data sequence, a 1D-CNN using a feature encoder from a trained base recognition model can extract features from the sequencing data. Specifically, the 1D-CNN can perform a convolution operation on the input data within a sliding window of a fixed size (e.g., sequencing data at a certain time step) using a trained convolution kernel, thereby achieving local perception and feature extraction of the input data within the window, and then mapping these features to the next layer.
[0066] In one embodiment, the normalization layer (Layer Normalization) can normalize the input data (e.g., the output features of a 1D-CNN) by calculating the mean and variance of the input data across its dimensions, normalizing the input data to a distribution with a mean of 0 and a variance of 1. This approach reduces the variance of the input data across feature dimensions and outputs a more stable feature representation with greater generalization capabilities, thereby improving the model's performance.
[0067] In one embodiment, the activation function unit can use a preset activation function (such as the Gelu (Gaussian Error Linear Unit) function) to perform element-by-element nonlinear mapping on the feature representation output by the normalization layer, thereby learning a more complex and nonlinear feature representation and improving the feature prediction and expression capabilities of the model. Among them, the Gelu function combines the characteristics of the Gaussian distribution and the nonlinear characteristics of the sigmoid function, and can achieve a smooth, continuous and differentiable nonlinear transformation. While keeping the positive part of the input unchanged, it compresses the negative part and maintains a high numerical stability, so that it can better handle the gradient disappearance and gradient explosion problems and improve model performance. In another example, other activation functions or approximate Gelu function formulas, such as the Fast Gelu function, can also be used to reduce computational complexity and improve the efficiency of nonlinear mapping.
[0068] In one embodiment, the feature encoder may also include other more or fewer network structures. For example, a pooling layer may be used in the feature encoder to reduce the dimension of the output of the convolutional layer to reduce the computational complexity of the model while improving the robustness and generalization ability of the model.
[0069] S203 , using the context network of the base recognition model to extract the temporal information and contextual relationship between the multiple encoding feature vectors, and obtain multiple contextual feature vectors corresponding to the sequencing data.
[0070] In one embodiment, since the encoding feature vector is obtained by feature encoding multiple time-series data in the sequencing data, the multiple encoding feature vectors have a time-series antecedent relationship. The process of the context network extracting the time sequence information and contextual relationship between the multiple encoding feature vectors can be understood as learning the relationship between any one of the multiple encoding feature vectors and several adjacent encoding feature vectors before its time sequence, or learning the relationship between any one of the encoding feature vectors and several adjacent encoding feature vectors after its time sequence. Therefore, although the dimension of the context feature vector corresponding to each time step output by the context network (for example, 768) is equal to the dimension of the encoding feature vector, the context feature vector (contextualized representation) of each time step has the implicit features of multiple time steps of the sequencing data.
[0071] For example, according to the time sequence, if any coding feature vector is the coding feature vector z corresponding to the t-th time step among the multiple coding feature vectors, t , the context network extracts the relationship between any encoding feature vector and several adjacent encoding feature vectors before it in time sequence, so the context network can learn the encoding feature vector z t and the encoded feature vector z before it t-1 and the encoded feature vector z t-2 The relationship between these two adjacent encoded feature vectors, where t represents an integer greater than or equal to 3.
[0072] In one embodiment, the context network (Context Network) includes a preset second number of Transformer network structures (transformer blocks). The Transformer network structure is a neural network architecture based on a self-attention mechanism (self-attention), which can learn the context information (such as the relationship between timing information and context) of each position (such as the position of the tth time step) in the input sequence (such as multiple encoded feature vectors) through the self-attention mechanism. The preset second number can be set according to actual needs, and can also be determined through the parameter optimization process in the model training process. For example, according to the model training process, it can be determined that the second number is 12 (such as shown in Figure 3), indicating that the performance of the model obtained by superimposing 12 Transformer network structures in the context network is the best.
[0073] For example, the dimension of each hidden layer unit in the 12-layer Transformer network used by the context network is 768 (= the dimension of each encoded feature vector), and the 1D-conv combined with the grouped convolution method is introduced to learn the context information (such as the temporal information and contextual relationship) of each position (such as the t-th position) in the input sequence (such as multiple encoded feature vectors). For example, as shown in Figure 4, the grouped convolution divides the input sequence into multiple (such as 16) groups, each group uses 1D-conv to perform a convolution operation independently, and uses the GELU (Gaussian Error Linear Unit) activation function to perform a nonlinear mapping transformation on the output of 1D-conv, and then splices the outputs of each group. Through the above method, the effect of the positional embedding device can be achieved, thereby capturing the contextual position information or sequential relationship in the input sequence.
[0074] S204: Input the multiple context feature vectors into a preset decoder, and use the decoder to decode the multiple context feature vectors to obtain a base sequence corresponding to the sequencing data.
[0075] In one embodiment, the preset decoder includes a Connectionist Temporal Classification decoder (CTC decoder). Since the task of the decoder of the base recognition model is to obtain the classification result (e.g., base sequence) of the base corresponding to the entire sequencing data based on multiple context feature vectors, and the dimension of the context feature vector may be different from the dimension of the base recognition result, a CTC decoder that can handle the situation where the input (e.g., multiple context feature vectors) and the label (e.g., base sequence) are not fully aligned can be used as the decoder of the base recognition model, thereby improving the recognition accuracy and robustness of the model.
[0076] For example, if the sequencing data is obtained by sequencing DNA, the dimension of the context feature vector for each time step of the sequencing data is 768. Decoding and classifying the multiple context feature vectors of the sequencing data can obtain a base sequence consisting of the following four bases: adenine (A), thymine (T), guanine (G), and cytosine (C). The base sequence can be regarded as a label for the sequencing data or multiple context feature vectors of the sequencing data. The dimension of the label may be different from the total dimension of the multiple context feature vectors (for example, the dimension of the label is much smaller than the total dimension of the multiple context feature vectors). In this case, the decoder required is a decoder that can handle decoding operations where the input and label are not fully aligned.
[0077] In one embodiment, when decoding and classifying multiple context feature vectors of sequencing data, the CTC decoder introduces a special blank marker (or placeholder) and uses a beam search algorithm to explore multiple candidate classification results at each time step. The multiple candidate classification results are aggregated based on the probability distribution and decoder parameters, and the classification result with the highest probability is selected to obtain the base sequence corresponding to the sequencing data. For example, if the sequencing data is obtained by sequencing DNA, the multiple candidate classification results include: adenine A, thymine T, guanine G, cytosine C, and placeholders, where the placeholders are used to solve the problem that the dimension of the input (e.g., multiple context feature vectors) is larger than the dimension of the label (e.g., the base sequence).
[0078] The base recognition method provided in the embodiment of the present application inputs sequencing data into a pre-trained base recognition model, uses the feature encoder of the base recognition model to obtain multiple encoding feature vectors of the sequencing data, and uses the context network of the base recognition model to extract the timing information and contextual relationship between the multiple encoding feature vectors, thereby obtaining multiple context feature vectors with context-related features, and then uses a decoder to decode the multiple context feature vectors, which can improve the accuracy of identifying the base sequence of the sequencing data.
[0079] The above embodiment introduces the application reasoning process of the base recognition model. Next, the base recognition model training method provided by the embodiment of the present application will be described. Referring to Figure 5, a flow chart of the base recognition model training method provided by one embodiment of the present application is shown. The base recognition model training method is applied to an electronic device, such as the electronic device 10 in Figure 1, and specifically includes the following steps. Depending on different needs, the order of the steps in the flow chart can be changed, and some steps can be omitted.
[0080] S301, collecting sample data.
[0081] In one embodiment, as shown in Figure 6, for example, an example diagram of the architecture of the base recognition model training process provided in an embodiment of the present application is provided. The training process of the base recognition model mainly includes two parts, wherein the first part is to use a large amount of (for example, greater than 100,000 hours, greater than 30 million seconds, etc.) of unlabeled historical sequencing data to perform self-supervised pretraining (Self-supervised Pretraining) on the preset model, and the second part is to use a small amount (for example, less than 10 hours) of labeled historical sequencing data to perform supervised optimization (or supervised fine-tune) on the initial model. Therefore, the training sample needs to include a set consisting of a large amount of unlabeled historical sequencing data (hereinafter referred to as the first set), and a set consisting of a small amount of labeled historical sequencing data (hereinafter referred to as the second set).
[0082] In one embodiment, the sample data includes unlabeled historical sequencing data and labeled historical sequencing data. Collecting the sample data includes: collecting historical electrical signal data obtained by nanopore sequencing historical deoxyribonucleic acid (DNA) or historical ribonucleic acid (RNA), the historical electrical signal data including time-varying electrical signal amplitude data; using the historical electrical signal data as the unlabeled historical sequencing data; obtaining the true base sequence of the historical DNA or RNA corresponding to the historical electrical signal data; and using the true base sequence as a label for the corresponding historical electrical signal data to obtain the labeled historical sequencing data.
[0083] In one embodiment, referring to the description in S201, the electronic device can collect the downloaded fast5 files of multiple (for example, more than 50) sequencing processes of the nanopore gene sequencer, and extract multiple historical electrical signal data from the fast5 files, wherein each historical electrical signal data corresponds to a historical DNA or historical RNA.
[0084] In one embodiment, when obtaining the true base sequence corresponding to the historical electrical signal data, a small amount (for example, 0.1%) of historical electrical signal data can be selected from multiple historical electrical signal data, and the true base sequence corresponding to the historical electrical signal data can be read from the fast5 file corresponding to the selected historical electrical signal data; the true base sequence obtained by the user by annotating the selected historical electrical signal data can also be received.
[0085] In one embodiment, since the fine-tuning process uses only a small amount of sample data, it is necessary to obtain the best possible model during self-supervised pre-training. Therefore, the unlabeled historical sequencing data can be pre-processed to improve the quality of the training samples, thereby improving the training efficiency and performance of the model. Before performing self-supervised pre-training on the preset model based on the unlabeled historical sequencing data, the method further includes: pre-processing the unlabeled historical sequencing data, wherein the pre-processing includes one or more of data cleaning, data normalization, data interpolation, and data segmentation.
[0086] Specifically, data cleaning processing can include filtering operations, removal of outliers or peaks, and other methods, which can remove bad data such as noise and outliers in the data (such as historical sequencing data) and improve data quality; data normalization processing can include zero-mean standardization method, which can normalize the data to a range of mean 0 and variance 1, eliminate the scale differences between different data, and facilitate comparison and calculation of subsequent processes; data interpolation processing can use linear interpolation, Lagrange difference and other methods to fill in missing data based on the changing trend of existing data points when there is missing or discontinuous data, thereby improving the smoothness of the data; data segmentation processing can adaptively segment longer data sequences into suitable lengths, facilitate the processing and analysis of subsequent processes, and improve the training efficiency of the model.
[0087] In one embodiment, the sample data may be cleaned from multiple dimensions such as noise, speed, signal uniformity, and accuracy, thereby improving the quality of the sample data.
[0088] S302 , performing self-supervised pre-training on a preset model based on unlabeled historical sequencing data, and obtaining a preset model that meets preset requirements as an initial model.
[0089] In one embodiment, the preset model may include a feature encoder, a quantization module, a mask module, and a context network. It is understood that the modules may be directly implemented in the form of algorithms. For example, the quantization module may be directly implemented using corresponding quantization operations, and the mask module may be directly implemented using corresponding masking methods. The preset model may also include other more or fewer network structures, and this application does not impose specific limitations on this.
[0090] In one embodiment, the self-supervised pre-training process can use the data in the first set to iteratively update the preset model at least once. In each update of the at least one iterative update, the data of the batch corresponding to each update in the first set can be used for updating (refer to the process shown in Figure 7). If the updated preset model does not meet the preset requirements, the data in the first set of the next batch is used for the next update until a preset model that meets the preset requirements is obtained as the initial model. The preset requirements may include but are not limited to a combination of one or more of the following requirements: the loss function corresponding to the preset model meets the preset convergence condition; the number of the at least one iterative update reaches a preset iteration threshold.
[0091] As shown in FIG7 , the method for self-supervised pre-training of a preset model includes the following process:
[0092] S401 , inputting unlabeled historical sequencing data into a feature encoder, and using the feature encoder to perform feature encoding on data of multiple time steps in the historical sequencing data to obtain multiple encoded feature vectors.
[0093] In one embodiment, before inputting the unlabeled historical sequencing data into the feature encoder, the unlabeled historical sequencing data may be normalized. For example, data of multiple time steps in the unlabeled historical sequencing data may be normalized to data with a mean of 0 and a variance of 1.
[0094] In one embodiment, a detailed introduction to the feature encoder can be referred to the description in S202. The use of the feature encoder in the model training process only requires replacing the sequencing data in S202 with unlabeled historical sequencing data, and no further description will be given.
[0095] In one embodiment, the feature encoder in S401 is obtained after the last update of the preset model; the multiple encoded feature vectors obtained in S401 (and the multiple encoded feature vectors involved in the subsequent pre-training process) are obtained by feature encoding the unlabeled historical sequencing data. For example, as shown in FIG6 , the unlabeled historical sequencing data {x1, x2, ..., x n ,…,x N} Input the feature encoder of the preset model obtained after the last update (including 6 1D-conv) to obtain the unlabeled historical sequencing data {x1, x2,…, x n ,…,x N} corresponding to the encoded feature vector {z1,z2,…,z t ,…,z T}, wherein n = {1, 2, ..., N}, N represents the total length of the unlabeled historical sequencing data, the value of N can be set according to actual needs, and this application does not impose a specific restriction on this. For example, the value of N is 5000; t = {1, 2, ..., T}, t represents the t-th time step, T represents the total number of encoded feature vectors, and the value of T can be set according to actual needs, and this application does not impose a specific restriction on this. For example, T can be 500, 1000, 1250, etc.
[0096] S402: Using a quantization module, perform a product quantization operation on the plurality of encoded feature vectors to obtain a plurality of quantized feature vectors.
[0097] In one embodiment, the quantization module is used to perform a product quantization operation on the multiple coded feature vectors to obtain multiple quantized feature vectors, including: determining multiple code table vectors corresponding to each of the multiple coded feature vectors, including: evenly splitting each coded feature vector into multiple discrete sub-vectors, using a preset quantization method to determine the code table vector (entry) corresponding to each discrete sub-vector from each code table in a preset multiple code tables (codebook), to obtain the multiple code table vectors (entries); wherein the quantization method includes a Gumbel-Softmax quantization method; splicing the multiple code table vectors into quantized feature vectors corresponding to each coded feature vector. The number of code tables can be set according to actual needs, and this application does not impose specific restrictions on this. For example, the number of code tables can be 3, 4, 5, 6, 8, etc.; the quantized feature vector can be used as a target for subsequent self-supervised learning (refer to the first loss function below).
[0098] In one embodiment, the process of product quantization operation can be understood as: converting the encoding feature vector z in the continuous space into t Split the discrete low-dimensional data into multiple subspaces (e.g., discrete subvectors), where each subspace corresponds to a learnable code table; quantize the discrete low-dimensional data in each subspace, and represent the discrete low-dimensional data in each subspace using a code table vector in the corresponding code table; concatenate all code table vectors corresponding to the discrete low-dimensional data in all subspaces to obtain the quantized feature vector q corresponding to each encoded feature vector t Through the product quantization operation, the encoded feature vector can be mapped from the high-dimensional feature expression space to the low-dimensional discrete space, thereby improving the robustness and representation ability of the feature.
[0099] In one embodiment, the codebook represents a randomly initialized matrix that can learn changes during model training and be optimized and updated according to the corresponding loss function (see the second loss function below). In one example, the matrix dimension of each codebook is 320*128, which means that each codebook contains 320 128*1 dimension entries. For any encoded feature vector z t When G = 4 codebooks are used, the encoded feature vector z t It will be mapped into a vector of G*128=4*128=512 dimensions.
[0100] In one embodiment, the Gumbel-Softmax quantization method is a reparameterization technique for approximating discrete values in the quantization process. By introducing noise that follows a Gumbel distribution into the code table vector of each code table, randomness is provided in the selection of the code table vector. The softmax function is applied to obtain the discrete distribution of the code table vector in each code table. The code table vector with the highest probability in each code table can be determined as the code table vector corresponding to each discrete sub-vector, so that random discrete quantization features can be used during model training while ensuring that gradients can be backpropagated.
[0101] S403: Mask the multiple coded feature vectors using a masking module to obtain multiple masked coded feature vectors.
[0102] In one embodiment, the masking module is used to mask the multiple encoding feature vectors to obtain multiple masked encoding feature vectors, including: determining the sequence length of the data sequence corresponding to each encoding feature vector in the multiple encoding feature vectors (for example, the dimension of each encoding feature vector, for example, 768); determining the number of mask positions for masking the data sequence based on the sequence length, a preset mask probability value (mask_prob, for example, 60%), and a preset mask length (mask_length, for example, 5); randomly selecting multiple mask positions corresponding to the number of mask positions from the data sequence; at each mask position in the multiple mask positions, using a preset mask vector to mask multiple continuous data corresponding to the mask length in the data sequence to obtain a masked data sequence; wherein the dimension of the mask vector (for example, 5) is equal to the mask length; and determining each masked encoding feature vector based on the masked data sequence.
[0103] In one embodiment, the masking process can be understood as a process of updating and replacing the encoded feature vector using a preset mask vector (mask vector). In one example, the masking process includes: 1) determining the position of the mask: for example, if the sequence length of the encoded feature vector is 1000, mask_prob is 60%, and mask_length is 5, then the number of data points (span) that need to be masked is 1000*60% / 5=120; 120 positions can be randomly selected from the 1000 data point positions of the encoded feature vector as the starting position of the mask span; wherein, overlapping masks may occur, for example, the first data point and the second data point are included in the 120 randomly selected positions, that is, the first to fifth data points and the second to sixth data points need to be masked, resulting in repeated masks; 2) assigning a randomly initialized mask vector to the mask position, which is a vector that can be learned and trained and is shared by all mask positions.
[0104] In one embodiment, the method also includes determining multiple interference vectors, including: determining multiple (for example, 100) mask positions when masking the multiple encoded feature vectors using the mask module; selecting multiple interference positions other than the multiple mask positions from the data sequence composed of the multiple encoded feature vectors; and determining an interference vector (or called distractors) at each of the multiple interference positions, wherein the dimension of the interference vector is equal to the dimension of each encoded feature vector.
[0105] In one embodiment, the data of the encoded feature vector at the mask position in the mask process is equivalent to the positive sample data, and the interference vector is equivalent to the negative sample data. Combined with the quantized feature vector q used as the training target in the above process, t In the subsequent training process, the model output can be compared with the above negative sample data and the quantized feature vector q t The distance between them is used to determine whether the model meets the preset requirements (refer to the first loss function below).
[0106] S404: Utilize a context network to extract temporal information and contextual relationships between multiple masked encoding feature vectors to obtain multiple contextual feature vectors.
[0107] In one embodiment, the introduction of the context network can refer to the description in S203. The use of the context network in the model training process only requires replacing the sequencing data in S203 with unlabeled historical sequencing data, and no further description will be given.
[0108] In one embodiment, the context network in S404 is obtained after the last update of the preset model; the multiple context feature vectors obtained in S404 (and the multiple context feature vectors involved in the subsequent pre-training process) are context feature vectors corresponding to the unlabeled historical sequencing data. For example, the unlabeled historical sequencing data {x1, x2, ..., x n ,…,x N}The corresponding multiple masked encoded feature vectors are input into the context network of the preset model obtained after the last update (including 12 transformer blocks) to obtain the unlabeled historical sequencing data {x1, x2,…, x n ,…,x N} corresponding context feature vector {c1,c2,…,c t ,…,c T}, wherein n = {1, 2, ..., N}, N represents the total length of the unlabeled historical sequencing data, the value of N can be set according to actual needs, and this application does not impose a specific restriction on this. For example, the value of N is 5000; t = {1, 2, ..., T}, t represents the t-th time step, T represents the total number of context feature vectors, and the value of T can be set according to actual needs, and this application does not impose a specific restriction on this. For example, T can be 500, 1000, 1250, etc.
[0109] S405 , taking each quantized feature vector as a learning target of the corresponding context feature vector, constructing a loss function based on the learning target, and determining that the preset model meets the preset requirements if the loss function reaches a preset convergence condition.
[0110] In one embodiment, the loss function includes a first loss function, which can be calculated based on the context feature vector c t With positive sample q t The distance between them and the context feature vector c t and negative sample q′ i The distance between them is used to construct a first loss function, and the formula used by the first loss function includes:
[0111] Among them, L m Denotes the first loss function, q′∈Q t ={q t ,q′1,q′2,…,q′ i ,…,q′ K},q t represents the quantized feature vector at the tth time step, c t represents the context feature vector of the tth time step, sim() represents the preset similarity function, q′ irepresents the i-th disruptor vector, i is a positive integer in the range [1, K], K represents the total number of disruptor vectors, and k represents the preset temperature parameter. t ,q′ i ,q t The dimensions are the same.
[0112] In one embodiment, sim() may use cosine similarity, for example, sim()=0.5×(1-cosine distance); K may be set to 100; k may be used to control the scaling of similarity and may be set to 0.1.
[0113] In one embodiment, the first loss function (Contrastive Loss) can be used to measure the ability of the context network to predict the future. The smaller the first loss value corresponding to the first loss function, the better the context feature vector c t With positive sample q t The smaller the distance between them and the context feature vector c t and negative sample q′ i The larger the distance between them, the better the ability of the context network to predict the future.
[0114] In one embodiment, the loss function includes a second loss function, which can be constructed based on the probability distribution of the code table vector in each code table. The formula used by the second loss function includes:
[0115] Among them, L d represents the second loss function, represents the average probability that all encoded feature vectors are mapped to the vth code table vector in the gth code table, v represents the vth code table vector, the value of v is a positive integer in the range [1, V], V represents the total number of code table vectors in each code table vector (for example, 320); the value of g is a positive integer in the range [1, G], G represents the total number of code tables (for example, 4); -H() represents the entropy function.
[0116] In one embodiment, The entropy function is used to measure the diversity of the probability distribution of the code table vectors in each code table. It can be used to optimize the probability distribution in the code table, making the probability of each code table vector being selected more uniform, thereby improving the diversity of the code table. Therefore, the second loss function (Diversity Loss) can be used to improve the expressive power of the quantized code table. For example, when G = 1 and V = 320, if all 1000 encoded feature vectors are mapped to the 10th vector of the code table, then the probability distribution of the selected code table vector is very concentrated. At this time, the corresponding entropy is the smallest and the value of the second loss function is the largest, which is the worst case. If the probability of 1000 encoded feature vectors being mapped to each code table vector in the code table is the same, then the probability distribution of the selected code table vector is very dispersed. At this time, the corresponding entropy is the largest and the value of the second loss function is the smallest, which is the optimal case.
[0117] In one embodiment, the loss function includes a weighted sum of a preset first loss function and a preset second loss function, for example, the loss function L=L m +αL d , where α represents the preset weight coefficient, which can be set to 0.1.
[0118] In one embodiment, the above-mentioned convergence condition can be set according to actual needs, and this application does not impose specific restrictions on this. For example, the convergence condition can be that the loss value calculated by L is less than 1.3. For example, as shown in Figure 8, an example diagram of the change of the loss value corresponding to the loss function of the pre-training process provided in an embodiment of the present application as the training process changes. Among them, the horizontal axis represents the number of samples used (trained sample num), the unit is 1 million, and the vertical axis represents the loss value (loss). It can be seen that 18 million seconds of unlabeled historical sequencing data are used to perform self-supervised pre-training on the preset model, and the loss value gradually converges to about 1.2.
[0119] S303: Optimize the initial model using the labeled historical sequencing data, and use the optimized initial model as the base recognition model.
[0120] In one embodiment, the process of optimizing the initial model using labeled historical sequencing data is similar to the commonly used labeled fine-tuning method. The labels of the labeled historical sequencing data can be used as the output targets of the initial model, thereby performing targeted training on the initial model, so that the initial model can adapt to base recognition tasks and improve model performance and generalization ability.
[0121] In one embodiment, optimizing the initial model using the labeled historical sequencing data includes: inputting the labeled historical sequencing data into the initial model; performing feature encoding on the labeled historical sequencing data using a feature encoder of the initial model to obtain a plurality of encoded feature vectors corresponding to the labeled historical sequencing data; extracting temporal information and contextual relationships between the plurality of encoded feature vectors using a context network of the initial model to obtain a plurality of context feature vectors corresponding to the labeled historical sequencing data; inputting the plurality of context feature vectors into a preset decoder, decoding the plurality of context feature vectors using the decoder to obtain a predicted base sequence corresponding to the labeled historical sequencing data; determining a loss value between the predicted base sequence corresponding to the labeled historical sequencing data and the true base sequence using a preset classification loss function, and adjusting the model parameters of the initial model (e.g., hyperparameters such as connection weights between neurons and parameters in the loss function) based on the loss value. When updating the model parameters of the initial model, the parameters in the feature encoder may be updated; or the parameters of the feature encoder may be fixed so that the parameters in the feature encoder are not updated, while the parameters other than the fixed parameters in the initial model are updated.
[0122] In one embodiment, the first part of the above optimization process can refer to the application reasoning process of the base recognition model (e.g., S201-S203). It is only necessary to replace the sequencing data with the labeled historical sequencing data, and no detailed description is given.
[0123] In one embodiment, as shown in FIG7 , for example, the decoder used in the fine-tuning process can be a linear layer added after the context network, and the linear layer can classify the output of the pre-trained model into C=5 categories, thereby realizing feature decoding. Among them, the linear layer usually includes a fully connected layer, and its output dimension is set to 5, corresponding to the number of categories of the base recognition task. For example, if the historical sequencing data is obtained by sequencing DNA, the five categories of the base recognition task include: adenine A, thymine T, guanine G, cytosine C, and placeholders, where the placeholders are used to solve the problem that the dimension of the input (such as multiple context feature vectors) is larger than the dimension of the label (such as the base sequence). In another example, the decoder can also be a CTC decoder.
[0124] In one embodiment, the classification loss function used during fine-tuning may be the CTC (Connectionist Temporal Classification) loss function (CTC-loss) corresponding to the CTC encoder. The CTC loss function is a loss function used for sequence classification tasks and can be used to handle problems with variable-length output sequences. For more information about CTC, please refer to S204. A larger loss value between the predicted base sequence corresponding to the labeled historical sequencing data calculated by the CTC loss function and the true base sequence indicates a lower prediction accuracy of the initial model, a greater number of hyperparameters in the initial model that require fine-tuning, or a larger adjustment range for certain hyperparameters.
[0125] The base recognition model training method provided in the embodiment of the present application creatively proposes and uses the idea and scheme of training a large pre-trained nanopore sequencing model (e.g., an initial model) in a nanopore sequencing scenario. The pre-trained large model fully considers the inherent characteristics of the nanopore sequencing signal. After training with a large amount of unlabeled nanopore sequencing signals, it can fully characterize the original electrical signal, and then can be simply fine-tuned for use in various downstream analyses of nanopore sequencing data. The base recognition model algorithm is developed through a pre-training + fine-tuning training method. Compared with traditional supervised learning ideas, more massive nanopore sequencing data can be used, which can improve the convergence speed of the model and eliminate the need to obtain a large amount of labeled data after annotation, thereby reducing the training difficulty and training cost of the model, improving the prediction accuracy of the model, and thus improving the accuracy of base recognition.
[0126] In other embodiments, other schemes can also be used as schemes for pre-training large models for base recognition tasks. For example, the base recognition model can be pre-trained based on the training scheme of the speech pre-training large model Whisper. Unlike models such as Wav2vec, the Whisper model uses a large amount of weakly labeled data, which further improves the accuracy of automatic speech recognition. At the same time, the Whisper model can directly perform multi-task learning without the need for fine-tuning for specific tasks (such as base recognition tasks). Therefore, as a base recognition task of sequencing signals that has similarities with the speech signal recognition task, the training method of the Whisper model can be used in the base recognition algorithm.
[0127] In one embodiment, the hardware environment and software development environment used in the model training process in the embodiment of the present application can be exemplified in Table 1 below:
[0128] Table 1
[0129] Using a graphics processing unit (GPU) to assist a central processing unit (CPU) in model training can effectively improve the training efficiency of the model.
[0130] FIG9 is a structural diagram of a base recognition device provided in one embodiment of the present application.
[0131] In some embodiments, the base call apparatus 70 may include multiple functional modules composed of computer program segments. The computer programs of the various program segments in the base call apparatus 70 may be stored in a memory of an electronic device and executed by at least one processor to perform base call functions (see FIG. 2 for details).
[0132] In this embodiment, the base call device 70 can be divided into multiple functional modules according to the functions it performs. These functional modules may include: an input module 701, a feature encoding module 702, a context learning module 703, and a decoding module 704. As referred to herein, a module refers to a series of computer program segments that can be executed by at least one processor and that can perform a fixed function, and is stored in a memory. In this embodiment, the functional implementation of each module in the base call device 70 can be found in the definition of the base call method above and will not be repeated here.
[0133] The input module 701 is used to input sequencing data into a pre-trained base recognition model.
[0134] The feature encoding module 702 is configured to perform feature encoding on the sequencing data using the feature encoder of the base recognition model to obtain a plurality of encoded feature vectors corresponding to the sequencing data.
[0135] The context learning module 703 is configured to extract the temporal information and contextual relationships between the multiple encoding feature vectors using the context network of the base recognition model to obtain multiple contextual feature vectors corresponding to the sequencing data.
[0136] The decoding module 704 is configured to input the multiple context feature vectors into a preset decoder, and use the decoder to decode the multiple context feature vectors to obtain a base sequence corresponding to the sequencing data.
[0137] In another embodiment, the present application also provides a base recognition model training device, which may include a plurality of functional modules consisting of computer program segments. The computer program of each program segment in the base recognition model training device may be stored in a memory of an electronic device and executed by at least one processor to perform (see FIG5 for details) the function of base recognition model training. Regarding the functional implementation of each module in the base recognition model training device, please refer to the above definition of the base recognition model training method, which will not be repeated here.
[0138] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions. The method implemented when the program instructions are executed can refer to the methods in the above-mentioned embodiments of the present application.
[0139] The computer-readable storage medium may be an internal memory of the electronic device described in the above embodiment, such as a hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, etc. equipped on the electronic device.
[0140] In some embodiments, the computer-readable storage medium may include a program storage area and a data storage area, wherein the program storage area may store an operating system, applications required for at least one function, etc.; the data storage area may store data created according to the use of the electronic device, etc.
[0141] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0142] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0143] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0144] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0145] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A base recognition method, characterized in that, The method includes: Inputting sequencing data into a pre-trained base recognition model; Performing feature encoding on the sequencing data by using a feature encoder of the base recognition model to obtain a plurality of encoded feature vectors corresponding to the sequencing data; Extracting temporal information and context relationships among the plurality of encoded feature vectors by using a context network of the base recognition model to obtain a plurality of context feature vectors corresponding to the sequencing data; Inputting the plurality of context feature vectors into a preset decoder, and decoding the plurality of context feature vectors by using the decoder to obtain a base sequence corresponding to the sequencing data.
2. The base recognition method according to claim 1, wherein The method further includes collecting the sequencing data, including: Collecting electrical signal data obtained by performing nanopore sequencing on deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), where the electrical signal data includes electrical signal amplitude data that changes over time; Using the electrical signal data as the sequencing data.
3. The base recognition method according to claim 1, characterized in that, The training method of the base recognition model includes a self-supervised pre-training method and a supervised optimization method.
4. The base recognition method according to claim 1, wherein The feature encoder includes a preset first number of network modules, where each network module includes: A one-dimensional convolutional layer for extracting features from input data, where the input data includes the sequencing data; A normalization layer for performing normalization processing on the output data of the one-dimensional convolutional layer; An activation function unit for performing non-linear mapping on the output data of the normalization layer.
5. The base recognition method according to claim 1, wherein The context network includes a preset second number of Transformer network structures.
6. The base recognition method according to claim 1, characterized in that, The preset decoder includes a connectionist temporal classification (CTC) decoder.
7. A method for training a base recognition model, characterized in that, The method includes: Collecting sample data, where the sample data includes unlabeled historical sequencing data and labeled historical sequencing data; Performing self-supervised pre-training on a preset model based on the unlabeled historical sequencing data to obtain a preset model that meets preset requirements as an initial model, where the preset model includes a feature encoder, a quantization module, a masking module, and a context network; Optimizing the initial model by using the labeled historical sequencing data, and using the optimized initial model as the base recognition model.
8. The base recognition model training method according to claim 7, wherein The collecting of the sample data includes: Collecting historical electrical signal data obtained by performing nanopore sequencing on historical DNA or historical RNA, where the historical electrical signal data includes electrical signal amplitude data that changes over time; Using the historical electrical signal data as the unlabeled historical sequencing data; Obtaining the true base sequence of the historical DNA or the historical RNA corresponding to the historical electrical signal data; Using the true base sequence as the label of the corresponding historical electrical signal data to obtain the labeled historical sequencing data.
9. The base recognition model training method according to claim 7, wherein Before performing self-supervised pre-training on a preset model based on the unlabeled historical sequencing data, the method further includes: preprocessing the unlabeled historical sequencing data, where the preprocessing includes one or more of data cleaning processing, data normalization processing, data interpolation processing, and data segmentation processing.
10. The base recognition model training method according to claim 7, wherein The performing self-supervised pre-training on a preset model based on the unlabeled historical sequencing data includes: Input the untagged historical sequencing data into the feature encoder, and use the feature encoder to perform feature encoding on the data of multiple time steps in the historical sequencing data to obtain multiple encoded feature vectors; Use the quantization module to perform product quantization operations on the multiple encoded feature vectors to obtain multiple quantized feature vectors; Use the masking module to mask the multiple encoded feature vectors to obtain multiple masked encoded feature vectors; Use the context network to extract the temporal information and context relationship between the multiple masked encoded feature vectors to obtain multiple context feature vectors; Take each quantized feature vector as the learning target of the corresponding context feature vector, construct a loss function based on the learning target, and if the loss function reaches the preset convergence condition, determine that the preset model meets the preset requirements.
11. The base recognition model training method according to claim 10, wherein The step of using the quantization module to perform product quantization operations on the multiple encoded feature vectors to obtain multiple quantized feature vectors includes: Determine multiple codebook vectors corresponding to each encoded feature vector in the multiple encoded feature vectors, including: evenly splitting each encoded feature vector into multiple discrete sub-vectors, and using a preset quantization method to determine the codebook vector corresponding to each discrete sub-vector from each of the preset multiple codebooks to obtain the multiple codebook vectors; wherein, the quantization method includes the Gumbel-Softmax quantization method; Concatenate the multiple codebook vectors into the quantized feature vector corresponding to each encoded feature vector.
12. The base recognition model training method according to claim 11, wherein The step of using the masking module to mask the multiple encoded feature vectors to obtain multiple masked encoded feature vectors includes: Determine the sequence length of the data sequence corresponding to each encoded feature vector in the multiple encoded feature vectors; According to the sequence length, a preset masking probability value, and a preset masking length, determine the number of masking positions for masking the data sequence; Randomly select multiple masking positions corresponding to the number of masking positions from the data sequence; At each masking position among the multiple masking positions, use a preset masking vector to mask multiple consecutive data corresponding to the masking length in the data sequence to obtain a masked data sequence; wherein, the dimension of the masking vector is equal to the masking length; Determine each masked encoded feature vector according to the masked data sequence.
13. The base recognition model training method according to claim 10, wherein The method further includes determining multiple interference vectors, including: Determine multiple masking positions when using the masking module to mask the multiple encoded feature vectors; Select multiple interference positions outside the multiple masking positions from the data sequence formed by the multiple encoded feature vectors; Determine an interference vector at each interference position among the multiple interference positions, wherein the dimension of the interference vector is equal to the dimension of each encoded feature vector.
14. The base recognition model training method according to claim 13, wherein The loss function includes a first loss function, and the formula used by the first loss function includes: Among them, L m represents the first loss function, q' ∈ Q t ={q t , q'1, q'2, …, q' i , …, q' K}, q t represents the quantization feature vector at the t-th time step, c t represents the context feature vector at the t-th time step, sim() represents a preset similarity function, q' i represents the i-th interferer vector, where the value of i is a positive integer in the range of [1, K], K represents the total number of interferer vectors; k represents a preset temperature parameter.
15. The base recognition model training method according to claim 11, characterized in that The loss function includes a second loss function, and the formula used by the second loss function includes: Among them, L d represents the second loss function, It represents the average probability that all encoded feature vectors are mapped to the v-th codebook vector in the g-th codebook. Here, v represents the v-th codebook vector, and the value of v is a positive integer within the range of [1, V], where V represents the total number of codebook vectors in each codebook; the value of g is a positive integer within the range of [1, G], and G represents the total number of codebooks; -H() represents the entropy function.
16. The base recognition model training method according to claim 10, wherein The loss function includes the weighted sum of a preset first loss function and a preset second loss function.
17. The base recognition model training method according to claim 8, wherein The optimizing the initial model by using the labeled historical sequencing data includes: Inputting the labeled historical sequencing data into the initial model; Performing feature encoding on the labeled historical sequencing data by using the feature encoder of the initial model to obtain a plurality of encoded feature vectors corresponding to the labeled historical sequencing data; Extracting the temporal information and context relationship among the plurality of encoded feature vectors by using the context network of the initial model to obtain a plurality of context feature vectors corresponding to the labeled historical sequencing data; Inputting the plurality of context feature vectors into a preset decoder, and decoding the plurality of context feature vectors by using the decoder to obtain a predicted base sequence corresponding to the labeled historical sequencing data; Using a preset classification loss function to determine the loss value between the predicted base sequence corresponding to the labeled historical sequencing data and the true base sequence, and adjusting the model parameters of the initial model according to the loss value.
18. An electronic device, characterized in that, The electronic device includes a processor and a memory. When the processor executes the computer program stored in the memory, it implements the base recognition method according to any one of claims 1 to 6, or implements the base recognition model training method according to any one of claims 7 to 17.
Citation Information
Patent Citations
Method for quickly identifying single-molecule nanopore sequencing bases based on deep network
CN112183486A
A self-supervised learning-based speech emotion recognition method and computer device
CN114937465A
Pre-training method and device for self-supervised emotion recognition model based on electroencephalogram signals
CN117171557A
Cited By
Nanopore sequencing method, terminal equipment and storage medium
CN121884951A