Base recognition method based on fluorescently labeled dNTP gene sequencing, sequencer and medium
By acquiring multi-channel fluorescence images in next-generation sequencing and combining them with cyclic timing information, the crosstalk problem between fluorescence signal acquisition units was solved, improving the accuracy of base identification and sequencing precision.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN SALUS BIOMED CO LTD
- Filing Date
- 2023-09-20
- Publication Date
- 2026-07-28
AI Technical Summary
In existing second-generation gene sequencing technologies, the accuracy of base recognition is limited by spatial crosstalk between fluorescence signal acquisition units and unknown brightness interference factors, especially in the case of high-density samples, which leads to a decrease in sequencing accuracy.
A base recognition method based on fluorescently labeled dNTPs is adopted. Multiple fluorescence images are acquired during gene sequencing to form multi-channel input data. Position codes are generated by combining cyclic temporal information. Convolutional networks are used for feature extraction and classification prediction to overcome spatial crosstalk and adapt to different density conditions.
It improves the accuracy of base identification, can more effectively handle crosstalk caused by various uncertainties, and improves the accuracy of sequencing results.
Smart Images

Figure CN117274614B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of gene sequencing technology, and in particular to a base identification method based on fluorescently labeled dNTP gene sequencing, a gene sequencer, and a computer-readable storage medium. Background Technology
[0002] Currently, gene sequencing technologies can be mainly divided into three generations. The first generation, the Sanger sequencing method, is a sequencing technology based on DNA synthesis reactions, also known as the SBS method or end-termination method. It was proposed by Sanger in 1975, and the first complete genome sequence of an organism was published in 1977. The second generation, represented by the Illumina platform, achieved high-throughput sequencing, representing a revolutionary advancement that made large-scale parallel sequencing a reality and greatly promoted the development of genomics in the life sciences. The third generation, Nanopore sequencing technology, is a new generation of single-molecule real-time sequencing technology. It primarily uses the changes in electrical signals caused by ssDNA or RNA template molecules passing through nanopores to infer base composition for real-time sequencing.
[0003] In second-generation sequencing (NGS) technology, fluorescence microscopy is used to capture the signals of fluorescent molecules in images. The base sequence is then obtained by decoding the fluorescence signals in these images. To distinguish between different bases, filters are used to capture images of the fluorescence intensity of the sequencing chip at different frequencies, thus obtaining the spectral characteristics of the fluorescent molecules' emission. Multiple images need to be captured in the same scene. These images are then located and registered, and point signals are extracted and analyzed for brightness information to obtain the base sequence. With the development of NGS technology, sequencers are now equipped with software for real-time processing of sequencing data. Different sequencing platforms use different optical systems and fluorescent dyes, resulting in variations in the spectral characteristics of fluorescent molecule emission. If the algorithm cannot obtain appropriate features or find suitable parameters to handle these different features, it may lead to significant errors in base classification, thus affecting sequencing quality.
[0004] Furthermore, next-generation sequencing technology utilizes the fact that different fluorescent molecules have different fluorescence emission wavelengths. When these fluorescent molecules are irradiated by a laser, they emit fluorescence signals of the corresponding wavelengths, such as... Figure 1 By selectively filtering out non-specific wavelengths of light after laser irradiation, a fluorescence signal of a specific wavelength can be obtained, such as... Figure 2 In DNA sequencing, four fluorescent labels are commonly used. These four fluorescent labels are simultaneously added to a cycle, and images of the fluorescence signals are captured using a camera. Since each fluorescent label corresponds to a specific wavelength, we can separate the fluorescence signals corresponding to different fluorescent labels from the image, thus obtaining the corresponding fluorescence images, such as... Figure 3 During this process, camera focus adjustments and sampling parameter settings can be made to ensure optimal quality of the obtained TIF grayscale image. However, in practical applications, the brightness of base clusters in fluorescence images is always affected by various factors, mainly including spatial crosstalk between base clusters within the image, crosstalk within channels, and crosstalk between cycles (phasing, prephasing). Known base identification techniques primarily normalize crosstalk and intensity, but the correction methods vary. One approach corrects fluorescence intensity values by using the crosstalk matrix and the phasing to prephasing ratio within each cycle to remove crosstalk noise, and then identifies bases using the intensity values of the four channels, such as... Figure 4 However, existing base recognition technologies can only correct for known brightness interference factors, such as brightness crosstalk between channels and phasing and prephasing caused by premature or delayed reactions between cycles. They cannot correct for brightness interference caused by other unknown biochemical or environmental factors, resulting in low recognition accuracy. When the sample density is higher, the base clusters are denser, and the brightness crosstalk between base clusters is more severe, which greatly reduces sequencing accuracy. Summary of the Invention
[0005] To address the existing technical problems, embodiments of this application provide a base identification method, gene sequencer, and computer-readable storage medium based on fluorescently labeled dNTP gene sequencing that can overcome spatial crosstalk between base signal acquisition units and adapt to different base signal acquisition unit densities, thereby effectively improving base identification accuracy.
[0006] To achieve the above objectives, the technical solution of this application embodiment is implemented as follows:
[0007] In a first aspect, embodiments of this application provide a base recognition method based on fluorescently labeled dNTP gene sequencing, comprising:
[0008] Multiple fluorescence images corresponding to sequencing signal responses of different base types for the sequencing chip are acquired in each cycle of gene sequencing to form multi-channel input image data for each cycle.
[0009] For each loop, the input of the base recognition model is formed based on the loop timing information and the multi-channel input image data. The position encoding layer of the base recognition model generates a position code based on the loop timing information. The convolutional network layer of the base recognition model performs feature extraction based on the multi-channel input image data. The position code and feature extraction results are combined to perform classification prediction and output the base recognition result corresponding to each channel input image data.
[0010] Secondly, embodiments of this application provide a gene sequencer, including a processor and a memory connected to the processor. The memory stores a computer program that can be executed by the processor. When the computer program is executed by the processor, it implements the base recognition method for gene sequencing based on fluorescently labeled dNTPs as described in any embodiment of this application.
[0011] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the base identification method based on fluorescently labeled dNTP gene sequencing as described in any embodiment of this application.
[0012] In the above embodiments, multiple fluorescence images corresponding to sequencing signal responses of different base types are collected in each cycle of the gene sequencing process to form multi-channel input image data. The image type data input of the base recognition model is a multi-channel input formed by multiple fluorescence images corresponding to different base categories in each cycle. Therefore, the prediction of base recognition results can maintain the relative magnitude relationship of the brightness values of the base signal acquisition units in multiple channels. It has strong adaptability to overcome spatial crosstalk between base signal acquisition units caused by various uncertain factors and to adapt to different base signal acquisition unit densities. It can learn richer feature representations, thereby effectively improving the accuracy of base recognition results. Secondly, the base recognition model... The input also includes the cycle timing information corresponding to each cycle. The cycle timing information is associated with the fluorescence images to be tested collected in the corresponding cycle. Since different cycles represent different reaction periods in the gene sequencing process, the quality of fluorescence images collected in different reaction periods under the same image acquisition conditions and the degree of spatial crosstalk between base signal acquisition units are significantly related to the number of reaction periods. The cycle timing information is generated into position codes through the position coding layer. The prediction of base recognition results takes into account the cycle timing information of the cycle to which the multi-channel input image data belongs. This allows the capture of image features related to the reaction period, which can more effectively solve the problem of inaccurate base recognition caused by crosstalk between cycles, and further improve the accuracy of base recognition.
[0013] In the above embodiments, the gene sequencer and computer-readable storage medium belong to the same concept as the corresponding embodiment of the base identification method based on fluorescently labeled dNTP gene sequencing, and thus have the same technical effect as the corresponding embodiment of the base identification method based on fluorescently labeled dNTP gene sequencing, which will not be repeated here. Attached Figure Description
[0014] Figure 1 This is a schematic diagram showing the distribution of fluorescence signal wavelengths of different fluorescent molecules in one embodiment;
[0015] Figure 2 This is a schematic diagram illustrating the principle of an imaging device acquiring a fluorescence image in one embodiment, wherein the imaging device selectively filters out light of non-specific wavelengths using a filter to obtain an image of a fluorescence signal of a specific wavelength.
[0016] Figure 3 This is a schematic diagram of four fluorescence images corresponding to the sequencing signal responses of the four base types ATCG in one embodiment, and a partially enlarged schematic diagram of one of the fluorescence images.
[0017] Figure 4 Here is a known base recognition flowchart from one embodiment;
[0018] Figure 5 This is a schematic diagram of a chip and a base signal acquisition unit on the chip in one embodiment;
[0019] Figure 6 This is a flowchart of a base recognition method based on fluorescently labeled dNTP gene sequencing in one embodiment;
[0020] Figure 7 This is a model architecture diagram of a base recognition model in one embodiment;
[0021] Figure 8 This is a schematic diagram of the residual block in one embodiment;
[0022] Figure 9 This is a schematic diagram of the logic for training a base recognition model in one embodiment;
[0023] Figure 10 for Figure 9 A schematic diagram illustrating the working principle of the medium base recognition model;
[0024] Figure 11 The flowchart is a specific example of a base recognition method based on fluorescently labeled dNTP gene sequencing;
[0025] Figure 12 This is a schematic diagram of the structure of a gene sequencer in one embodiment. Detailed Implementation
[0026] The technical solution of this application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0027] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0028] In the following description, the phrase "some embodiments" refers to a subset of all possible embodiments. It should be noted that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0029] In the following description, the terms "first, second, and third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, and third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0030] Next-generation sequencing (NGS) technology can sequence hundreds of thousands to millions of DNA molecules simultaneously. Known NGS sequencers generally record base information using optical signals, converting the light signals into base sequences. The base cluster positions generated by image processing and fluorescence localization techniques serve as references for template site positions in subsequent chip sequencing. Therefore, image processing and fluorescence localization techniques are directly related to the accuracy of the base sequence data. The base identification method based on fluorescently labeled dNTP gene sequencing provided in this application uses fluorescence images acquired from the sequencing chip as input data in fluorescently labeled dNTP gene sequencing, and is mainly applied to NGS technology. Fluorescent labeling is a measurement technique using optical signals, commonly used in industry for DNA sequencing, cell labeling, and drug research. The optical signal method used in NGS gene sequencing utilizes different wavelengths of fluorescent labeling for different bases. After filtering with a filter, successful ligation of specific bases excites light of a specific wavelength, which is then identified as the DNA base sequence to be tested. This technology, which generates images by collecting light signals and then converts them into base sequences, is the main principle of second-generation gene sequencing technology.
[0031] Second-generation sequencers, taking Illumina sequencers as an example, mainly include four stages in their sequencing process: sample preparation, cluster generation, sequencing, and data analysis.
[0032] Sample preparation, also known as library construction, refers to breaking down the basic DNA to be tested into a large number of DNA fragments, adding adapters to both ends of each DNA fragment, and including sequencing binding sites, indices (information identifying the origin of the DNA fragment), and specific sequences complementary to oligonucleotides on the sequencing chip (flowcell).
[0033] Cluster generation, also known as the process of seeding a library onto a flowcell and using bridged DNA amplification, involves the formation of a base cluster from a single DNA fragment.
[0034] Sequencing refers to sequencing reads for each base cluster on the flowcell. Fluorescently labeled dNTP sequencing primers are added during sequencing. One end of the dNTP chemical formula is linked to an azide group, which prevents polymerization during strand extension, ensuring that only one base is extended per cycle, corresponding to one sequencing read—that is, sequencing while synthesis occurs. In one cycle, each base cluster is identified by a fluorescently labeled dNTP. Specific fluorescent signals correspond to different base types in the sequencing signal response. Laser scanning identifies which base corresponds to each base cluster in the current cycle based on the emitted fluorescence color. In one cycle, tens of millions of base clusters are sequenced simultaneously in the flowcell. One fluorescent dot represents the fluorescence emitted by one base cluster, and one base cluster corresponds to one read in the FastQ database. During the sequencing phase, fluorescence images of the flowcell surface are captured using an infrared camera. These images undergo image processing and fluorescence spot location for base cluster detection. Based on the base cluster detection results from multiple fluorescence images corresponding to sequencing signal responses of different base types, templates are constructed to determine the locations of all base cluster template points (clusters) on the flowcell. Based on the templates, fluorescence intensity is extracted from the filtered images, then corrected. Finally, the maximum intensity at each base cluster template point location is used to identify the base, calculate a score, and output a FastQ base sequence file. Please refer to [link to documentation]. Figure 5 These are schematic diagrams of Flowcell (e.g.) Figure 5 (a) Fluorescence images of the corresponding parts of the flowcell taken in one cycle (e.g., in the middle (a)). Figure 5 (b) and a schematic diagram showing the sequencing results in the fastq file (e.g.) Figure 5 (c)
[0035] Data analysis involves analyzing millions of reads representing all DNA fragments. For each sample, the base sequences from the same library are clustered using unique indexes in adapters introduced during library construction. Reads are paired to generate continuous sequences, which are then compared with a reference genome for mutation identification.
[0036] It should be noted that the above uses Illumina sequencing technology as an example of massively parallel sequencing (MPS) to illustrate the sequencing process. By amplifying the DNA molecule to be tested using specific amplification techniques, base clusters are formed for each DNA fragment (single-stranded library molecule). The base cluster detection results are used to construct template points for the base clusters on the sequencing chip, so that subsequent operations such as base identification can be performed based on the template points of the base clusters, thereby improving the efficiency and accuracy of base identification. It is understood that the base identification method based on fluorescently labeled dNTP gene sequencing provided in this application is based on the location detection and base type identification of base clusters after amplification of single-stranded library molecules on the sequencing chip. Here, each base cluster refers to a base signal acquisition unit, and therefore it is not limited to the amplification technology used for the single-stranded library molecules. That is, the base identification method based on fluorescently labeled dNTP gene sequencing provided in this application can also be applied to the location detection and base type identification of base signal acquisition units on sequencing chips in other massively parallel sequencing technologies. For example, the base signal acquisition unit can refer to the base clusters obtained by bridge amplification technology in Illumina sequencing technology, as well as nanospheres obtained by rolling circle amplification (RCA) technology, etc. This application does not limit this.
[0037] Please see Figure 6 The base recognition method based on fluorescently labeled dNTP gene sequencing provided in one embodiment of this application includes the following steps:
[0038] S301: Acquire multiple fluorescence images of the sequencing chip corresponding to different base types and sequencing signal responses in each cycle of gene sequencing, forming multi-channel input image data for each cycle.
[0039] In each fluorescence image to be tested, each fluorescent spot corresponds one-to-one with a base signal acquisition unit of the corresponding base type. Base types typically refer to the four base types: A, C, G, and T. Since different base types correspond to the fluorescence signals of different fluorescently labeled dNTPs, theoretically there is no overlap between base signal acquisition units of different fluorescently labeled dNTPs. For each base type, the corresponding fluorescence image to be tested contains a base signal acquisition unit of the same base type located at the corresponding position in the sequencing chip.
[0040] However, the inventors discovered in their research that, in practical applications, the brightness of the base signal acquisition unit in the fluorescence image corresponding to different sequencing signal responses can be affected by a variety of factors.
[0041] For example, regarding the crosstalk problem, in the gene sequencing process, due to the overlap of wavelength distributions of fluorescent molecules with different fluorescently labeled dNTPs, multiple fluorescence images corresponding to sequencing signal responses of different base types will experience intensity crosstalk. For instance, if a fluorescent bright spot appears in the fluorescence image corresponding to base type A, the shadow of this bright spot may also appear in the fluorescence image corresponding to base type T, causing the fluorescence position in the fluorescence image corresponding to base type T to also have a certain brightness, which is called crosstalk. In addition, the sampling efficiency and filter filtration efficiency of fluorescence images of different base types will differ, resulting in the intensity distribution of fluorescence images of base types A and T not being at the same level. For example, the average brightness of fluorescence in image type A may be 100, while the average brightness of fluorescence in image type T may be 150. These image differences are caused by the image acquisition equipment, and these differences will affect the accuracy of subsequent base identification.
[0042] Regarding the issues of phasing and prephasing, in the gene sequencing pipeline, base addition and sequencing reactions occur in each sequencing cycle (one cycle). Each fluorophore contains many fluorescent molecules and copies, which react synchronously. That is, when the sequencing reaction of the current DNA base sequence reaches base type A, all copies of the fluorophore will react at the position of base type A and emit a fluorescent signal in channel A, which is ultimately manifested as the light intensity of the fluorophore. The corresponding fluorescence image represents the fluorescent bright spot formed by the base signal acquisition unit of base type A. However, due to incomplete fluorescence excision efficiency, fluorophores may be incompletely excised or not completely washed, resulting in incompletely excised fluorescence still having a certain light intensity in the same channel image of the next sequencing cycle, which is the fluorophore reaction lag effect (phasing). At the same time, premature reactions can also occur in the fluorophore. The fluorophore may exhibit partial fluorescence that should have reacted in the next sequencing cycle, but instead shows light intensity in the current sequencing cycle, which is the fluorophore reaction prephasing effect. The lag and advance of these reactions reflect the asynchrony and inconsistency of copy reactions in fluorophores, which are also the main reasons affecting sequencing length and base recognition error rate.
[0043] As analyzed above, the crosstalk effect of fluorescent spots can originate from the influence of different base types in different sequencing cycles, or from the influence of different base types and surrounding fluorescent sites within the same sequencing cycle. Existing base identification technologies, which correct fluorescence intensity values using the crosstalk matrix and phasing / prephasing ratio within each sequencing cycle, are only effective for known brightness interference factors and involve complex algorithms. Furthermore, current known machine learning strategies, such as Support Vector Machines (SVM), Convolutional Neural Networks (CNN), and Recurrent Neural Networks (RNN), all use each acquired fluorescence image as input. By extracting image features from the fluorescence images, they essentially use the extracted center brightness information as the basis for predicting the center position of the base arrowheads after extracting the brightness of the fluorescence images. This approach has limited effectiveness in addressing base identification errors caused by crosstalk, phasing, and prephasing problems.
[0044] In this embodiment, multiple fluorescence images corresponding to sequencing signal responses of different base types are acquired within each cycle of gene sequencing, targeting the sequencing chip. These multiple fluorescence images acquired within each cycle are superimposed to form a multi-channel input image data. Each fluorescence image includes the location information of a base signal acquisition unit of one base type. Based on the location information of the base signal acquisition units contained in each of the multiple fluorescence images, the location information of the complete multiple types of base signal acquisition units contained at the target site of the sequencing chip can be obtained. The target site can be a local location on the surface of the sequencing chip or the entire surface of the sequencing chip, and is usually related to the imaging area range that a single fluorescence image can encompass.
[0045] The fluorescence image to be tested refers to the raw fluorescence image captured on the surface of the sequencing chip during the sequencing phase of the sequencing process. In this embodiment, the A, C, G, and T bases correspond to the fluorescence signals of four different fluorescently labeled dNTPs, and theoretically, there is no overlap between the base signal acquisition units of the four different fluorescently labeled dNTPs. Acquiring multiple fluorescence images to be tested corresponding to the sequencing signal responses of different base types collected for the target site of the sequencing chip in each cycle means that within each cycle (reaction cycle) of gene sequencing, fluorescence images corresponding to the fluorescence signals of four different fluorescently labeled dNTPs are captured for the same target site of the sequencing chip. Utilizing the different brightness of the four bases A, C, G, and T under different wavelengths of light, corresponding fluorescence images (four raw fluorescence images) are collected for the same field of view (the same target site of the sequencing chip) where the four bases A, C, G, and T are excited and illuminated by the fluorescence signals of the four different fluorescently labeled dNTPs (four different environments). These are used as multiple fluorescence images to be tested corresponding to the sequencing signal responses of different base types in the corresponding cycle, i.e., one sequencing cycle.
[0046] Multiple fluorescence images corresponding to sequencing signal responses of different base types are grouped together and stacked along the channel dimension to form a multi-channel input image data. For example, four fluorescence images corresponding to sequencing signal responses of the four base types A, C, G, and T are stacked along the channel dimension to form a 4-channel input image data, whose dimension can be represented as (4, H, W), where H and W are the height and width of the fluorescence image to be tested.
[0047] S303, for each loop, the input of the base recognition model is formed based on the loop timing information and the multi-channel input image data. The position encoding layer of the base recognition model generates a position code based on the loop timing information. The convolutional network layer of the base recognition model performs feature extraction based on the multi-channel input image data. The position code and feature extraction results are combined to perform classification prediction and output the base recognition result corresponding to each channel input image data.
[0048] Cycle timing information refers to information representing the number of rounds in a base sequencing pipeline. For example, cycle timing information can be the time point within the total time period of the current cycle in a base sequencing pipeline, or the cycle number within the total number of cycles in the current cycle. It can be understood that one cycle corresponds to one reaction period, that is, one base type identification at the location of the base signal acquisition unit in the sequencing chip. Therefore, the total number of cycles in a gene sequencing pipeline usually corresponds to the length of the base sequence. For each cycle, the input to the base identification model includes the cycle timing information of the corresponding cycle and multi-channel input image data formed by multiple fluorescence images of the target cells corresponding to sequencing signal responses of different base types acquired within the corresponding cycle. The cycle timing information is encoded through a position encoding layer, and features are extracted from the multi-channel input image data through a convolutional network layer, so that the base type classification prediction results of the multi-channel input image data obtained from the fluorescence images acquired in each cycle are correlated with the cycle timing information.
[0049] The base recognition model outputs base recognition results corresponding to the input image data of each channel, which can be presented in different forms, such as multi-channel or single-channel output. Multi-channel output means that the output corresponds one-to-one with the image data of multiple input channels. For example, the recognition result of channel 1 is the recognition result of the base signal acquisition unit of type A base in the current loop, the recognition result of channel 2 is the recognition result of the base signal acquisition unit of type C base in the current loop, the recognition result of channel 3 is the recognition result of the base signal acquisition unit of type G base in the current loop, and the recognition result of channel 4 is the recognition result of the base signal acquisition unit of type T base in the current loop. Single-channel output refers to the output of a single-channel recognition result that includes the location of the base signal acquisition unit and its base type, based on the multi-channel recognition results corresponding to the image data of multiple input channels. For example, the recognition result of the A base type, C base type, G base type, and T base type corresponding to the recognition results obtained by processing the input image data of each channel is used to form a recognition result that simultaneously includes A, C, G, and T base signal acquisition units in the current loop.
[0050] Furthermore, the base recognition results can be presented in different forms. This is reflected in the fact that the form representing the recognition results of the base signal acquisition unit can be a data matrix identifying the base types of each base signal acquisition unit in the current loop, or an image identifying the base types of each base signal acquisition unit. Taking multi-channel output as an example, each output corresponds to the recognition result of a base signal acquisition unit of one base type. Channel 1 can be a coordinate data matrix representing the position information of the center of the base signal acquisition unit of base type A. Thus, the coordinate data matrix output by Channel 1 represents the recognition result of the base signal acquisition unit of base type A in the current loop. Similarly, the coordinate data matrix of Channel 2 corresponds to the recognition result of the base signal acquisition unit of base type C, the coordinate data matrix of Channel 3 corresponds to the recognition result of the base signal acquisition unit of base type G, and the coordinate data matrix of Channel 4 corresponds to the recognition result of the base signal acquisition unit of base type T. Taking a single-channel output as an example, based on the identification results of A, C, G, and T obtained from channels 1, 2, 3, and 4, a coordinate data matrix is formed at the corresponding position of the center of the base signal acquisition unit containing all base types in the current loop, with the base type label marked thereon. It should be noted that although the output of the base recognition model includes the coordinate data matrix of the center of the base signal acquisition unit, it expresses the identification of the base type at the center of different base signal acquisition units within the current loop, thus achieving base type identification.
[0051] The coordinate data matrix mentioned above can be in other forms that can characterize the base type at the center of each base signal acquisition unit, such as a probability data matrix representing whether a pixel is the center of a base signal acquisition unit of a certain base type. The probability value at the center of the base signal acquisition unit represents the probability that the base signal acquisition unit belongs to the A, C, G or T base type.
[0052] Other forms of characterizing the base type at the center of each base signal acquisition unit can also be in the form of images. For example, based on the coordinate data matrix and probability data matrix, the positions of the base signal acquisition units of A, C, G, and T base types can be obtained, and fluorescent images with base type labels marked at the positions of the centers of each base signal acquisition unit in the current cycle can be directly output.
[0053] Based on the various possible presentation forms of the base recognition results provided above, it can be seen that the base recognition model output corresponding to the base recognition results of each channel input image data is a base recognition result obtained after processing multiple fluorescence images to be tested in the current cycle by the base recognition model. It can identify the base type at the center of each base signal acquisition unit in the current cycle. It may not be limited to a specific form and is not restricted here.
[0054] In the above embodiments, multiple fluorescence images corresponding to sequencing signal responses of different base types are collected in each cycle of the gene sequencing process to form multi-channel input image data. The image type data input of the base recognition model is a multi-channel input formed by multiple fluorescence images corresponding to different base categories in each cycle. Therefore, the prediction of base recognition results can maintain the relative magnitude relationship of the brightness values of the base signal acquisition units in multiple channels. It has strong adaptability to overcome spatial crosstalk between base signal acquisition units caused by various uncertain factors and to adapt to different base signal acquisition unit densities. It can learn richer feature representations, thereby effectively improving the accuracy of base recognition results. Secondly, the base recognition model... The input also includes the cycle timing information corresponding to each cycle. The cycle timing information is associated with the fluorescence images to be tested collected in the corresponding cycle. Since different cycles represent different reaction periods in the gene sequencing process, the quality of fluorescence images collected in different reaction periods under the same image acquisition conditions and the degree of spatial crosstalk between base signal acquisition units are significantly related to the number of reaction periods. The cycle timing information is generated into position codes through the position coding layer. The prediction of base recognition results takes into account the cycle timing information of the cycle to which the multi-channel input image data belongs. This allows the capture of image features related to the reaction period, which can more effectively solve the problem of inaccurate base recognition caused by crosstalk between cycles, and further improve the accuracy of base recognition.
[0055] In some embodiments, the convolutional network layer is a fully convolutional network layer including a symmetric encoder and decoder; the convolutional network layer of the base recognition model performs feature extraction based on the multi-channel input image data, combines the position encoding and feature extraction results to perform classification prediction, and outputs the base recognition result corresponding to each channel of input image data, including:
[0056] The encoders of the fully convolutional network layer of the base recognition model sequentially perform feature extraction on the multi-channel input image and linearly superimpose the positional encoding to obtain the corresponding feature map.
[0057] The decoder and the encoder are connected by a skip connection. Each decoder receives a first feature map output by the encoder that is symmetrical to it, and receives a second feature map of the corresponding scale that is sequentially decoded and passed on during the decoding process. The corresponding base classification prediction information is obtained by splicing the first feature map and the second feature map based on the channel dimension and the position encoding.
[0058] Based on the base classification prediction information, the base recognition results corresponding to the input image data of each channel are output.
[0059] Please see Figure 7A symmetrical fully convolutional network layer with encoders and decoders refers to a Unet-like structure formed by the same number of encoders and decoders connected sequentially. The positional encoding layer encodes the cyclic temporal information, and the resulting positional encoding information is linearly superimposed onto each encoder. This associates the image feature extraction results of each encoder for multi-channel input image data with the cyclic temporal information of the current cycle. This significantly improves upon the limitations of conventional convolutional network models, which, due to locality or translation invariance, cannot handle image quality differences in fluorescence images acquired within different cycles and cannot capture image features related to cyclic temporal information. As a result, the base recognition model captures differentiated image features based on the cyclic temporal information of the corresponding cycle when obtaining base recognition results from multiple fluorescence images acquired in each cycle. This demonstrates higher sensitivity to fluorescence spot crosstalk caused by different reaction cycles in the gene sequencing process. In this design, a skip connection is used between the decoder and encoder. This means that each convolutional layer in the encoder is directly connected to the convolutional layer of its symmetrical counterpart, the decoder. Thus, the decoder's convolutional layers can directly receive the first feature map output from the encoder's convolutional layer. Simultaneously, based on their sequential connections within the fully convolutional network, they can also receive the second feature map of the corresponding scale, which is sequentially passed forward during the decoding process. The positional encoding layer linearly superimposes the positional encoded information obtained from encoding the cyclic temporal information into each decoder. This allows the base classification prediction result to be made by combining the extracted image features from the multi-channel input image data with the cyclic temporal information of the current loop. This not only effectively learns the complex feature representation of the input data but also accurately captures the temporal relationships between image features under different reaction cycles, thereby improving the base recognition model's ability to handle temporally related feature differences and enhancing the accuracy of base recognition results.
[0060] In some embodiments, the encoder includes a plurality of residual blocks and pooling layers, and the decoder includes a plurality of residual blocks and deconvolution layers;
[0061] Each residual block includes a first convolutional layer and a second convolutional layer for feature extraction of the multi-channel input image, and a linear layer located between the first convolutional layer and the second convolutional layer and superimposed on the feature channels between the first convolutional layer and the second convolutional layer; wherein, the linear layer is connected to the position encoding layer and is used to superimpose the position encoding generated by the position encoding layer based on the cyclic temporal information onto the image features extracted by the convolutional layer;
[0062] In each residual block, the linear layer is connected to the position encoding layer; the second convolutional layer of each residual block takes the feature map output by the first convolutional layer superimposed with the position encoding of the linear layer as input, and takes the union of its own extracted feature map and the feature map input by the first convolutional layer as output.
[0063] Please see Figure 8 A linear layer is added between the first and second convolutional layers of each residual block (ResBlock). Each linear layer is connected to the positional encoding layer and is responsible for encoding the positional encoding (the Cycle vector obtained by encoding the cyclic timing information) obtained by the positional encoding layer. The positional encoding is then linearly superimposed with the output of the first convolutional layer of the residual block and used as the input to the second convolutional layer. Please refer again. Figure 7 The encoder consists of three residual blocks and a max-pooling layer to extract rich and multi-scale feature information from multi-channel input image data. Each residual block provides an efficient way to learn complex representations of the input data, while the max-pooling layer retains the most salient features while reducing data dimensionality. The decoder consists of three residual blocks and a deconvolutional layer, designed to generate images corresponding to the target base classification. Furthermore, to preserve the details of the original image and reduce information loss due to convolution operations, skip connections are used between the residual blocks of the encoder and decoder. This allows the intermediate feature maps output by each residual block in the encoding process to be directly concatenated and fused with the feature maps of the corresponding scale in the decoding process based on the channel dimension.
[0064] Residual blocks, as crucial structural units in base recognition models, primarily mitigate the vanishing and exploding gradient problems during deep neural network training. Each ResBlock mainly consists of two convolutional layers, with a linear layer added between them specifically responsible for encoding cyclic temporal information (the Cycle vector output by the positional encoding layer). These two convolutional layers extract spatial features from the input data to reveal local structural features in multi-channel input image data. Generally, this is achieved by sliding a fixed-size filter (i.e., a convolutional kernel) across the input data and then performing a weighted summation operation on the regions covered by the filter, ultimately forming a new feature map.
[0065] The multi-channel input image data for each cycle originates from fluorescence images generated at different reaction cycles in the gene sequencing process. The quality of these fluorescence images is significantly correlated with the number of cycles. For example, as the reaction progresses to later stages (i.e., the number of cycles increases), the influence of biochemical interferences such as phasing gradually intensifies. However, conventional convolutional operations, due to their locality and translation invariance, cannot directly capture image features related to the number of cycles. In this embodiment, a cycle coding layer (i.e., a linear layer) is inserted between the two convolutional layers of each residual block. The main task of the cycle coding layer is to assign a unique code to the input data of each residual block, thereby enabling the base recognition model to more accurately capture the correspondence between the quality of the fluorescence images acquired in each cycle and the number of cycles. This design ensures computational efficiency while allowing the base recognition model to exhibit higher sensitivity to the number of cycles. In this way, the base recognition model can not only effectively learn the complex feature representation of multi-channel input image data, but also accurately capture the temporal relationship between the image features of fluorescence images collected in different cycles. It has good targeting for the identification of key difference information contained in the relative temporal relationship of different cycles. The base recognition model can better understand and process the differences caused by temporal sequence, thereby improving the base recognition model's ability to correct the brightness interference caused by unknown biochemical or environmental influences and thus improve the accuracy of base recognition.
[0066] In some embodiments, the position encoding layer generates position codes using sine and cosine functions; wherein, the position encoding layer generated based on the cyclic time series information by the base recognition model includes:
[0067] During a gene sequencing process, the cyclic timing information corresponding to the fluorescence images to be tested within each cycle is determined by accumulating the acquisition sequence of the fluorescence images to be tested. The cyclic timing information is used as the input to the position coding layer of the base recognition model, and the corresponding position code is generated by using sine and cosine functions.
[0068] The positional encoding layer is used to encode the cyclic temporal information of each loop, hence it can also be called the cycle-embedding layer. The principle of encoding the cyclic information of each loop is essentially the same as the positional encoding technique in deep learning. In deep learning models, the main goal of positional encoding is to encode sequential information into the model. For example, when processing sequential data (such as text or time series data), the position or order of input elements often has a significant impact on the model's output. However, many deep learning models, such as fully connected networks, convolutional neural networks (CNNs), and self-attention mechanisms, do not inherently possess explicit sequence awareness. Therefore, positional encoding, by encoding positional information into the input data in some form, enables the model to perceive and utilize this positional information. There are many methods of positional encoding, such as the positional encoding generated by sine and cosine functions used in the Transformer model.
[0069] In this embodiment, the position encoding layer takes the sequentially accumulated cycle time information corresponding to different cycles in a gene sequencing process as known information and inputs it into the position encoding layer. The position encoding layer uses sine and cosine functions of different frequencies to generate a unique vector for each cycle time information, resulting in position encoding. This allows the base recognition model to distinguish inputs from different cycles and capture the differences between inputs collected in different cycles. The order of fluorescence image acquisition (cycle time information) is similar to the position information. By utilizing the same principle of position encoding technology to encode cycle information into vector form and input it into the base recognition model, the discrete cycle index is converted into a continuous vector representation, thereby enhancing the base recognition model's understanding and use of cycle information, and improving the performance and generalization ability of the base recognition model.
[0070] In some embodiments, the loss function of the base recognition model is the cross-entropy loss function, which is expressed as follows:
[0071]
[0072] Where C is the number of categories, y is the one-hot encoding of the true label, and p is the probability distribution predicted by the model.
[0073] The cross-entropy loss function (CE Loss) is used to measure the difference between the base class probability distribution predicted by the base recognition model and the actual base class probability distribution. The goal of the cross-entropy loss function is to minimize the cross-entropy between the predicted probability distribution and the actual probability distribution, so that the base class predicted by the base recognition model is as close as possible to the actual base class.
[0074] In some embodiments, the base recognition method includes, before the base recognition model is applied to the fluorescence images to be tested collected during the gene sequencing process, a step of constructing an initial deep learning model and training it to obtain the base recognition model. The training of the base recognition model includes:
[0075] Obtain the training dataset; wherein each training sample includes a multi-channel sample image formed by multiple original fluorescence images corresponding to sequencing signal responses of different base types for the sequencing chip in each cycle, the cycle time information corresponding to the cycle, and the base type label corresponding to the multi-channel sample image;
[0076] An initial deep learning model is constructed and trained using the training dataset. The deep learning model undergoes supervised learning with the base category label as the training objective until the loss function converges to obtain the trained base recognition model. The deep learning model includes a position encoding layer that takes the cyclic temporal information as input and a convolutional network layer that takes multi-channel input image data formed by superimposing multiple original fluorescence images as input. The convolutional network layer includes a symmetrical encoder and decoder. The encoder includes multiple residual blocks and pooling layers, and the decoder includes multiple residual blocks and deconvolutional layers. Each residual block includes a first convolutional layer and a second convolutional layer that feature the multi-channel input image, and a linear layer located between the first and second convolutional layers and superimposed on the feature channels between the first and second convolutional layers. The linear layer in each residual block is connected to the position encoding layer.
[0077] Please see Figure 9 This is a logical diagram illustrating the training of a base recognition model. The process involves acquiring a training dataset, including obtaining training samples through data annotation. An initial deep learning model is then constructed and iteratively trained using the training dataset. The model architecture of the deep learning model can be combined with… Figure 7 and Figure 8 As shown. For each training sample of the base recognition model, a group is formed by combining multiple raw fluorescence images corresponding to the sequencing signal responses of different base types acquired within a loop. This results in a multi-channel sample image composed of multiple fluorescence images acquired within the corresponding loop, the loop time sequence information corresponding to the loop, and the base type label corresponding to the multi-channel sample image. During the training phase, the base recognition model extracts training samples from the training dataset for iterative training. Please refer to [link to relevant documentation]. Figure 10This diagram illustrates the working principle of the base recognition model. In each iteration of training, the multi-channel sample images corresponding to each cycle in the training samples and their cycle time information are used as input. The base recognition model calculates and predicts the base recognition result of the input sample based on the current weight parameters, determines the recognition error based on the corresponding base type label, and judges whether the error is less than or equal to a set value. If the error is greater than the set value, backpropagation is performed based on the error to optimize the weight parameters of the base recognition model. Then, training samples are repeatedly extracted from the training dataset as input for the next iteration of training. This process is repeated until the recognition error of the base recognition result calculated by the base recognition model based on the current weight parameters and determined based on the corresponding base type label is less than the set value. That is, the base recognition model performs supervised learning with base type label as the training target until the loss function converges to obtain the trained base recognition model.
[0078] In the above embodiments, in each training sample of the base recognition model, the sample image is a multi-channel input formed by multiple fluorescence images corresponding to different base categories collected in each cycle. The cycle time information of each cycle is associated with the sample image. The prediction of the base recognition result can capture the brightness interference of the base signal acquisition unit caused by different cycles, and maintain the relative magnitude relationship of the brightness values of the base signal acquisition unit in multiple channels. That is, it maintains the relative magnitude relationship of the brightness values of multiple fluorescence images corresponding to the sequencing signal response of different base types in the same cycle, so as to obtain more accurate recognition results. It is more targeted at overcoming the spatial crosstalk problem between base signal acquisition units caused by various uncertain factors, and has stronger adaptability to different base signal acquisition unit densities. It can learn richer feature representations, thereby effectively improving the accuracy of base recognition results.
[0079] Optionally, each training sample further includes a mask label image corresponding to the multi-channel sample image; the creation of the training sample includes:
[0080] Multiple raw fluorescence images corresponding to sequencing signal responses of different base types for the sequencing chip are obtained in each cycle. The multiple raw fluorescence images corresponding to sequencing signal responses of different base types are used as multi-channel sample images of training samples for one cycle.
[0081] The original fluorescence image is subjected to primary base identification to obtain base identification results, and a mask image is formed based on the position of the base signal acquisition unit;
[0082] Based on the base identification results of the original fluorescence images continuously acquired by the sequencing chip during gene sequencing, a base sequence is obtained. The base sequence is compared with a standard base sequence in a known gene library. Successfully matched base sequences are selected and corrected according to their respective matching standard base sequences. The corresponding base identification results of the original fluorescence image determined by base identification are corrected according to the corrected base sequence. After correction, the base type label of the multi-channel sample image used as the training sample is obtained.
[0083] The mask image is corrected based on the unmatched base sequence, and the corrected image is the mask label image.
[0084] During the training of a base recognition model, supervision is required using the actual base categories of the input data. The creation of training samples for the base recognition model and base type labels for supervising the training data includes:
[0085] 1. Take the original fluorescence images corresponding to the sequencing signal responses of different base types for each cycle in the gene sequencing process and form multi-channel sample images of the training samples for each cycle.
[0086] 2. Primary base identification is performed on the original fluorescence image used as a sample image to obtain the base identification result. This mainly refers to the identification result of the location information and base type of the base signal acquisition unit in the original fluorescence image obtained by various known algorithms with relatively low accuracy requirements. In an optional example, primary base identification refers to processing the original fluorescence image using any known traditional base signal acquisition unit detection and localization algorithm to obtain the location of the base signal acquisition unit, and then using a traditional base identification algorithm based on the location of the base signal acquisition unit to determine the base type of the base signal acquisition unit in the original fluorescence image acquired in each cycle. The mask image refers to the selected area or template used to occlude the image being processed to control the image processing or the processing procedure. In a single gene sequencing run using the same sequencing chip, the positions of the base signal acquisition units within the chip are identical. This means that the positions of the base signal acquisition units in the fluorescence images acquired in different cycles should be the same. Therefore, in a single gene sequencing run, forming a mask based on the positions of the base signal acquisition units can refer to processing a set of raw fluorescence images corresponding to sequencing signal responses of different base types using a traditional base signal acquisition unit detection and localization algorithm, and then forming a position data matrix or image based on the union of the base signal acquisition unit positions in this set of raw fluorescence images.
[0087] 3. Obtaining a base sequence based on the base identification results of the raw fluorescence images continuously acquired from the sequencing chip during gene sequencing refers to identifying the base type corresponding to the base signal acquisition unit position in the fluorescence images acquired in different cycles during a single gene sequencing process. The base sequence formed based on the base type of the base signal acquisition unit in each cycle corresponds to the position of each base signal acquisition unit. In gene sequencing, the accuracy of detecting and locating base signal acquisition units and identifying the corresponding base types based on the fluorescence intensity at the base signal acquisition unit location in fluorescence images acquired in different cycles is inevitably affected by various factors. After determining the base type through primary base identification, the obtained base sequence is compared with the standard base sequence in the known gene library. In a base sequence, only when more than a proportion of bases are correctly identified compared with the standard base sequence can the comparison be successful. This allows all matching strands in the sample to be found. For the matching strands, the misidentified bases (less than a proportion of mismatched bases) in these matching strands are corrected according to the standard base sequence in the gene library. The corrected base sequence is then used to correct the base category results obtained from primary base identification, thereby correcting and improving the quality of the base type labels of the multi-channel sample images used as training samples.
[0088] In one optional example, after primary base identification, the base signal acquisition unit location points A(2,2) and B(3,3) are obtained. At this time, the mask image is a mask image where location points A(2,2) and B(3,3) are 1, and all other locations are 0. Based on the base identification results of the raw fluorescence images collected in 10 consecutive cycles of gene sequencing, the base sequence at location A is ACGTGTCAGT and the base sequence at location B is ACAGTTCAGT. After comparison with the standard base sequences in the known gene library, the standard base sequence that successfully matched the base sequence at location A is selected as ACCTGTCAGT. The base sequence at location A is corrected to ACCTGTCAGT based on the standard base sequence. Thus, the base identification results of the raw fluorescence images collected in 10 consecutive cycles of gene sequencing are corrected based on the corrected base sequence. In the base identification results of the raw fluorescence image collected in the 3rd cycle, the base type of location A is corrected from the originally identified base type G to base type C. That is, the base type label of the training sample formed by the raw fluorescence image collected in the 3rd cycle is corrected accordingly.
[0089] 4. Based on the determination of base types through primary base identification, the obtained base sequence is compared with standard base sequences in a known gene library. A successful alignment occurs only when more than a proportion of bases in a given base sequence are correctly identified compared to the standard sequence. If no matching standard base sequence is found during the alignment process, it is considered that the detection and localization of the base signal acquisition unit position was incorrect during primary base identification. Accordingly, the base signal acquisition unit positions of the unaligned base sequences are deleted from the mask image. For unaligned base sequences, the information of that chain is removed from the mask image of the training samples. For example, in the mask image formed by the base signal acquisition unit positions obtained through primary base identification, the positions of unaligned chains are replaced with 0 to avoid erroneous data contaminating the training data and improve the quality of the training samples. As in the example above, after comparing with the standard base sequence in the known gene library, no standard base sequence that successfully matches the base sequence at position B is found. Therefore, position B is deleted in the mask image for correction, and the corrected mask label image is a mask image with position A(2,2) and all other positions are 0.
[0090] In the above embodiments, in the training set for base recognition, each training sample includes a multi-channel sample image formed by multiple original fluorescence images corresponding to sequencing signal responses of different base types in each cycle, the cycle time information corresponding to the cycle, base type labels and mask label images of the multi-channel sample images. The introduction of mask label images allows the output of the base recognition model to retain only the prediction results of the center position of the base signal acquisition unit, while the predictions of other positions are set to 0. The mask strategy effectively focuses the attention of the base recognition model on important regions, eliminates possible background noise and interference, and improves the accuracy of prediction. Furthermore, the base type labels and mask label images of the multi-channel sample images are obtained by correcting them using standard base sequences in a known gene library. This not only effectively reduces the labeling difficulty of the training samples but also improves the labeling accuracy of the training samples. A higher-precision training set is beneficial to improving the recognition accuracy of the trained base recognition model.
[0091] In some embodiments, obtaining the base sequence based on the base identification results of the raw fluorescence images continuously acquired for the sequencing chip during gene sequencing includes:
[0092] For the raw fluorescence images continuously acquired from the sequencing chip during gene sequencing, based on the corresponding base signal acquisition unit positions in the mask image, the base signal acquisition units in the raw fluorescence images are identified using a base recognition algorithm to obtain base recognition results. The base sequence is then obtained based on the base recognition results of the continuously acquired raw fluorescence images; or,
[0093] For the raw fluorescence images continuously acquired from the sequencing chip during gene sequencing, a pre-trained base recognition model is used to identify the bases to obtain the base recognition results, and the base sequence is obtained based on the base recognition results of the continuously acquired raw fluorescence images.
[0094] In this embodiment, the training sample production process first uses the base identification results obtained from primary base identification to form a base sequence. Then, it compares the base type labels and correction mask label images of the multi-channel sample images in the training samples with standard base sequences in a known gene library for calibration. Primary base identification can include using a traditional base identification algorithm to identify and determine the base type of each base signal acquisition unit in the original fluorescence image acquired in each cycle, or it can be obtained by using a pre-trained base identification model. By designing the calibration scheme, the accuracy requirements for the position and base type identification of the base signal acquisition unit template points in the original fluorescence image used as the training sample can be reduced. Thus, the training samples obtained from the traditional base identification algorithm can be used to train the initially constructed base identification model. The base identification results can be obtained using the pre-trained base identification model before the training completion conditions are met. Compared to the method where the base identification results of the original fluorescence images used as all training samples are obtained using the traditional base identification algorithm, the production efficiency of training samples can be greatly improved.
[0095] The process involves using a pre-trained base recognition model to identify bases in sample images, obtaining preliminary base recognition results, updating base type labels through self-training, and then correcting the preliminary base recognition results using standard base sequences from a known gene library and updating the base type labels. This process improves the base recognition model's ability to learn from previously unsuccessfully matched image data, thereby enhancing its ability to recognize all base signal acquisition units.
[0096] To gain a more comprehensive understanding of the base recognition method based on fluorescently labeled dNTP gene sequencing provided in the embodiments of this application, please refer to [link to relevant documentation]. Figure 11 The following is a specific example illustrating a base identification method based on fluorescently labeled dNTP gene sequencing. This method includes:
[0097] S11, Creating training samples. This includes creating the label data for the training samples:
[0098] 1. First, a traditional base recognition algorithm is used to perform base call on the training sample images to obtain the base category of each base signal acquisition unit (cluster) in the sample images. The category labels for A, C, G, and T can be 1, 2, 3, and 4, respectively. At the same time, the position of each base signal acquisition unit is also determined, and a mask image with the same size as the original image is obtained, i.e., the center position of a cluster is 1, and the background area without a cluster is 0.
[0099] 2. The base category results obtained by the traditional algorithm are compared with the standard sequences of the known gene library. In a base sequence, the comparison can only be successful when most of the bases are correctly identified. By using this method, all matching chains in the sample image can be found.
[0100] 3. Even in the matched strands, there may be a small number of incorrectly identified bases. These incorrectly identified bases in the strands are corrected according to the standard sequences in the gene library to obtain the corrected strands. In the corrected strands, all base categories are correct and can be used as label data for training.
[0101] 4. After obtaining the category and location information of the base signal acquisition units, a labeled dataset can be created. First, generate a matrix of the same size as the original image. Based on the location and category information of the base signal acquisition units obtained in step one, fill the specified positions in the matrix with the category labels of the base types (A is 1, C is 2, G is 3, T is 4), and fill the remaining positions with 0.
[0102] 5. Labels can be updated using a self-training approach. This involves using a pre-trained base recognition network model to identify bases in the samples and obtain the identification results. Then, steps 2 through 4 are repeated to update the labels. This method can improve the model's ability to learn from previously unsuccessfully aligned data, thereby improving the model's ability to identify all base signal acquisition units.
[0103] S12, construct the initial Unet-like deep learning model, and iteratively train it using training samples to obtain the trained base recognition model. The architecture of the initial deep learning model is as follows: Figure 7 and Figure 8 As shown. The principle of iterative training of deep learning models is as follows. Figure 9 As shown, the training process mainly includes the following parts:
[0104] 1. Input
[0105] The four fluorescence images corresponding to the four base types collected in each loop are stacked in the channel dimension to form a 4-channel input data with dimensions (4, H, W), where H and W are the height and width of the training image.
[0106] The cycle number corresponding to the loop is input into an embedding layer to obtain the encoded vector. This encoded vector is then fed into each residual block for further processing along with image features.
[0107] 2. Network Model
[0108] A Unet-like deep learning model architecture is used to handle the base classification problem in images. This structure consists of two main components: an encoder and a decoder.
[0109] The encoder consists of three residual blocks and a max-pooling layer, which extracts rich and multi-scale feature information from the input image. Each residual block provides an efficient way to learn complex representations of the input data, while the max-pooling layer can reduce the data dimensionality while preserving the most salient features.
[0110] The decoder consists of three residual blocks and a deconvolutional layer, designed to generate images corresponding to the target base classification. Furthermore, to preserve detail in the original image and reduce information loss due to convolution operations, skip connections are used between the encoder and decoder. This method allows intermediate feature maps from the encoding process to be directly concatenated with feature maps of the corresponding scale from the decoding process, based on the channel dimension.
[0111] For the decoder output, a masking strategy is introduced to further optimize base prediction. Specifically, the mask formed by the positions of the template points of the base signal acquisition unit retains only the prediction results at the center position of the base signal acquisition unit, while setting all other positions to 0. This strategy effectively focuses the focus of the prediction process and eliminates unnecessary background interference.
[0112] Finally, the model's predictive performance is quantified by calculating the cross-entropy loss between the decoder output (i.e., the predicted result) and the true label. This loss function is very effective for training the model and can directly measure the consistency between the predicted probability distribution and the actual label.
[0113] 3. ResBlock
[0114] Residual blocks (ResBlocks) are crucial structural units in network models, primarily tasked with mitigating the vanishing and exploding gradient problems during model training. Each ResBlock mainly consists of two convolutional layers, with an innovative linear layer added between them, specifically responsible for encoding the Cycle vector output from the embedding layer. The two convolutional layers extract spatial features from the input multi-channel image data to reveal local structural features within the image.
[0115] The input image data originates from images generated at different reaction cycles, and the quality of these images is significantly correlated with the number of cycles. For example, as the reaction progresses into later stages (i.e., the number of cycles increases), the influence of biochemical interferences such as phasing gradually intensifies. However, conventional convolutional operations, due to their locality and translation invariance, cannot directly capture image features related to the number of cycles. A cycle encoding layer is inserted between the two convolutional layers in each residual block. The main task of the cycle encoding layer is to assign a unique code to each input image data, enabling the model to more accurately capture the correspondence between the input image and the number of cycles. This design ensures computational efficiency while making the model more sensitive to cycle information.
[0116] The inclusion of linear layers in the residual blocks creates a specialized design for time-series problems. This allows the base recognition model to not only effectively learn the complex feature representations of input image data but also accurately capture the temporal relationships between features from different reaction cycles. Many key pieces of information in different images with associated temporal information are contained within the relative temporal relationships of features. This design approach enables the model to better understand and process temporal data, thereby enhancing its ability to handle time-series problems.
[0117] 4. Cycle-Embedding layer:
[0118] The principle of the Cylinder-Embedding layer is essentially the same as positional encoding in deep learning. Positional encoding is an important technique for processing sequential data in natural language processing, especially in Transformer models. In deep learning models, the main goal of positional encoding is to encode sequential information into the model. For example, when processing sequential data (such as text or time series data), the position or order of input elements often has a significant impact on the result. However, many deep learning models, such as fully connected networks, convolutional neural networks (CNNs), and self-attention mechanisms, do not inherently possess explicit sequence awareness. Therefore, positional encoding enables the model to perceive and utilize this positional information by encoding it into the input data in some form.
[0119] There are many ways to perform positional encoding, such as the positional encoding generated by sine and cosine functions used in the Transformer model. This encoding method uses sine and cosine functions of different frequencies to generate a unique vector for each position, enabling the model to distinguish inputs at different positions and capture the relative positional relationships between input elements.
[0120] The order and location information of fluorescence image acquisition in different cycles have similar characteristics. In this embodiment, this principle is used to encode cycle information into vector form and input it into the model. The discrete cycle index corresponding to the cycle is converted into a continuous vector representation, thereby enhancing the model's understanding and use of cycle information, and improving the model's performance and generalization ability.
[0121] 5. Loss Function
[0122] The CE Loss is used as the loss function, as shown in Formula 1 above.
[0123] S13, In the gene sequencing process, multiple fluorescence images corresponding to the sequencing signal responses of different base types are collected for the sequencing chip in each cycle. The multiple fluorescence images in each cycle are formed into a multi-channel input image data. The multi-channel input image data and the cycle time information of the corresponding cycle are input into the base recognition model, and the base recognition model outputs the corresponding base recognition result.
[0124] The base recognition method provided in the above embodiments has at least the following characteristics:
[0125] 1. Through the design of the location encoding layer, the Cycle number is used as a key prior information and encoded into the model. This design makes the model more sensitive to the Cycle information of the image and enables it to effectively learn and understand the changes in the sample image as the Cycle number increases.
[0126] 2. In the preparation of training samples, a self-training approach is introduced to update labels. Initially, traditional base identification methods are used to obtain basic base categories as labels. In subsequent stages, the trained model can be used to obtain basic base categories for label updates. This strategy allows us to continuously expand the number of training samples, enhancing the model's ability to learn from data that initially failed to match, thereby improving the model's classification ability across all base signal acquisition units.
[0127] 3. A masking strategy was introduced to optimize the base prediction results. In this strategy, a mask map formed by the position of the template point of the base signal acquisition unit retains only the prediction result of the center position of the base signal acquisition unit, while the prediction of other positions is set to 0. The masking strategy can effectively focus the model's attention on important regions, eliminate possible background noise and interference, and improve the accuracy of prediction.
[0128] This application also provides a gene sequencer. Please refer to [link / reference]. Figure 12 This is a schematic diagram of an optional hardware structure of a gene sequencer provided in an embodiment of this application, including a processor 111 and a memory 112 connected to the processor 111. The memory 112 stores a computer program for implementing the base recognition method based on fluorescently labeled dNTP gene sequencing provided in any embodiment of this application. When the computer program is executed by the processor, it implements the steps of the base recognition method based on fluorescently labeled dNTP gene sequencing provided in any embodiment of this application and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0129] In another aspect, this application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described embodiments of the base recognition method based on fluorescently labeled dNTP gene sequencing, and achieves the same technical effect. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0130] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0131] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0132] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A base recognition method based on fluorescently labeled dNTP gene sequencing, characterized in that, include: Multiple fluorescence images corresponding to sequencing signal responses of different base types for the sequencing chip are acquired in each cycle of gene sequencing to form multi-channel input image data for each cycle. For each cycle, the input to the base recognition model is formed based on the cycle timing information of the corresponding cycle and the multi-channel input image data. The position encoding layer of the base recognition model generates a position code based on the cycle timing information. The convolutional network layer of the base recognition model performs feature extraction based on the multi-channel input image data. The classification prediction is performed by combining the position code and the feature extraction results, and the base recognition result corresponding to each channel input image data is output. The cycle timing information includes the time point information of the current cycle within the total time period of a base sequencing process, or the cycle number information of the current cycle in the total number of cycles of a base sequencing process.
2. The base recognition method as described in claim 1, characterized in that, The convolutional network layer is a fully convolutional network layer including a symmetrical encoder and decoder; the convolutional network layer of the base recognition model performs feature extraction based on the multi-channel input image data, combines the position encoding and feature extraction results to perform classification prediction, and outputs the base recognition result corresponding to each channel of input image data, including: The encoders of the fully convolutional network layer of the base recognition model sequentially perform feature extraction on the multi-channel input image and linearly superimpose the positional encoding to obtain the corresponding feature map. The decoder and the encoder are connected by a skip connection. Each decoder receives a first feature map output by the encoder that is symmetrical to it, and receives a second feature map of the corresponding scale that is sequentially decoded and passed on during the decoding process. The corresponding base classification prediction information is obtained by splicing the first feature map and the second feature map based on the channel dimension and the position encoding. Based on the base classification prediction information, the base recognition results corresponding to the input image data of each channel are output.
3. The base recognition method as described in claim 2, characterized in that, The encoder includes multiple residual blocks and pooling layers, and the decoder includes multiple residual blocks and deconvolution layers; Each residual block includes a first convolutional layer and a second convolutional layer for feature extraction of the multi-channel input image, and a linear layer located between the first convolutional layer and the second convolutional layer and superimposed on the feature channels between the first convolutional layer and the second convolutional layer; wherein, the linear layer is connected to the position encoding layer and is used to superimpose the position encoding generated by the position encoding layer based on the cyclic temporal information onto the image features extracted by the convolutional layer; In each residual block, the linear layer is connected to the position encoding layer; the second convolutional layer of each residual block takes the feature map output by the first convolutional layer superimposed with the position encoding of the linear layer as input, and takes the union of its own extracted feature map and the feature map input by the first convolutional layer as output.
4. The base recognition method as described in claim 1, characterized in that, The position encoding layer generates position codes using sine and cosine functions; wherein, the position encoding layer generated based on the cyclic time sequence information by the base recognition model includes: During a gene sequencing process, the cyclic timing information corresponding to the fluorescence images to be tested within each cycle is determined by accumulating the acquisition sequence of the fluorescence images to be tested. The cyclic timing information is used as the input to the position coding layer of the base recognition model, and the corresponding position code is generated by using sine and cosine functions.
5. The base recognition method as described in claim 1, characterized in that, The loss function of the base recognition network is the cross-entropy loss function, which is expressed as follows: ; Where C is the number of categories, y is the one-hot encoding of the true label, and p is the probability distribution predicted by the model.
6. The base recognition method as described in claim 1, characterized in that, Also includes: Obtain the training dataset; wherein each training sample includes a multi-channel sample image formed by multiple original fluorescence images corresponding to sequencing signal responses of different base types for the sequencing chip in each cycle, the cycle time information corresponding to the cycle, and the base type label corresponding to the multi-channel sample image; An initial deep learning model is constructed, and the deep learning model is trained using the training dataset. The deep learning model is supervised learning with the base category label as the training target until the loss function converges to obtain the trained base recognition model. The deep learning model includes a position encoding layer that takes the cyclic temporal information as input and a convolutional network layer that takes multi-channel input image data formed by superimposing multiple original fluorescence images as input. The convolutional network layer includes a symmetrical encoder and decoder. The encoder includes multiple residual blocks and pooling layers. The decoder includes multiple residual blocks and deconvolutional layers. Each residual block includes a first convolutional layer and a second convolutional layer that perform feature extraction on the multi-channel input image, and a linear layer located between the first convolutional layer and the second convolutional layer and superimposed on the feature channels between the first convolutional layer and the second convolutional layer. The linear layer in each residual block is connected to the position encoding layer.
7. The base recognition method as described in claim 6, characterized in that, Each training sample further includes a mask label image corresponding to the multi-channel sample image; the creation of the training sample includes: Multiple raw fluorescence images corresponding to sequencing signal responses of different base types for the sequencing chip are obtained in each cycle. The multiple raw fluorescence images corresponding to sequencing signal responses of different base types are used as multi-channel sample images of training samples for one cycle. The original fluorescence image is subjected to primary base identification to obtain base identification results, and a mask image is formed based on the position of the base signal acquisition unit; Based on the base identification results of the original fluorescence images continuously acquired by the sequencing chip during gene sequencing, a base sequence is obtained. The base sequence is compared with a standard base sequence in a known gene library. Successfully matched base sequences are selected and corrected according to their respective matching standard base sequences. The corresponding base identification results of the original fluorescence image determined by base identification are corrected according to the corrected base sequence. After correction, the base type label of the multi-channel sample image used as the training sample is obtained. The mask image is corrected based on the unmatched base sequence, and the corrected image is the mask label image.
8. The base recognition method as described in claim 7, characterized in that, The step of obtaining the base sequence based on the base identification results of the raw fluorescence images continuously acquired for the sequencing chip during gene sequencing includes: For the raw fluorescence images continuously acquired from the sequencing chip during gene sequencing, based on the corresponding base signal acquisition unit positions in the mask image, the base signal acquisition units in the raw fluorescence images are identified using a base recognition algorithm to obtain base recognition results. The base sequence is then obtained based on the base recognition results of the continuously acquired raw fluorescence images; or, For the raw fluorescence images continuously acquired from the sequencing chip during gene sequencing, a pre-trained base recognition model is used to identify the bases to obtain the base recognition results, and the base sequence is obtained based on the base recognition results of the continuously acquired raw fluorescence images.
9. A gene sequencer, characterized in that, The device includes a processor and a memory connected to the processor, wherein the memory stores a computer program that can be executed by the processor, and the computer program, when executed by the processor, implements the base recognition method based on fluorescently labeled dNTP gene sequencing as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the base identification method based on fluorescently labeled dNTP gene sequencing as described in any one of claims 1 to 8.