Base recognition method and training set construction method, gene sequencer and medium

By constructing a training set of multi-channel sample images and mask label images, the base recognition results are corrected using neural network models, and the problem of brightness interference in base recognition is solved and the recognition accuracy of the sequencer is improved.

CN117274739BActive Publication Date: 2025-08-12SHENZHEN SALUS BIOMED CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311222846.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-20
Publication Date
2025-08-12
Estimated Expiration
2043-09-20

AI Technical Summary

Technical Problem

Existing base recognition technology cannot effectively correct unknown brightness interference factors, resulting in serious brightness crosstalk between base clusters in high-density samples, affecting sequencing accuracy.

Method used

By constructing a training set for base recognition, multiple original fluorescence images are used to form multi-channel sample images, combined with mask label images, and training is performed using neural network models to correct base recognition results and eliminate background noise and interference.

Benefits of technology

It improves the accuracy of base recognition, adapts to the density of different base signal acquisition units, overcomes spatial crosstalk, and improves the recognition accuracy of the sequencer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117274739B_ABST
    Figure CN117274739B_ABST
Patent Text Reader

Abstract

The present application provides a base recognition method and a training set construction method thereof, a gene sequencer, and a medium. The base recognition training set construction method includes: using multiple original fluorescence images corresponding to sequencing signal responses of different base types as a multi-channel sample image of a training sample; performing primary base recognition on the original fluorescence images to obtain base recognition results, and forming a mask image according to the position of the base signal acquisition unit; obtaining a base sequence based on the base recognition results of the original fluorescence images continuously collected from the sequencing chip during gene sequencing, comparing the base sequence with a standard base sequence in a known gene library, correcting the successfully matched base sequences according to their respective matched standard base sequences, and obtaining a base type label of the multi-channel sample image as a training sample after correction; and correcting the mask image based on the base sequences that did not successfully match, and obtaining a mask label image after correction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of gene sequencing technology, and in particular to a method for constructing a base recognition training set, a base recognition method, a gene sequencer, and a computer-readable storage medium. Background Art

[0002] Currently, gene sequencing technology can be divided into four generations. The first-generation sequencing technology, the Sanger method, is based on DNA synthesis reactions. It is also known as the SBS method or the terminal termination method. It was proposed by Sanger in 1975, and the first complete genome sequence of an organism was published in 1977. The second-generation sequencing technology, represented by the Illumina platform, has achieved high-throughput sequencing and made revolutionary progress, making large-scale parallel sequencing a reality and greatly promoting the development of genomics in the life sciences. The third-generation sequencing technology is Nanopore sequencing, a new generation of single-molecule real-time sequencing technology. It mainly uses the changes in the electrical signal caused by the passage of ssDNA or RNA template molecules through the nanopore to infer the base composition for real-time sequencing.

[0003] Second-generation gene sequencing technology utilizes fluorescence microscopy to capture the signals of fluorescent molecules into images, which are then decoded to obtain the base sequence. To distinguish between different bases, filters are used to capture images of the fluorescence intensity of the sequencing chip at different frequencies to obtain the spectral characteristics of the fluorescent molecules' emission. Multiple images are captured for the same scene. These images are then positioned and aligned, and the point signals are extracted and brightness information is analyzed and processed to obtain the base sequence. With the advancement of second-generation sequencing technology, sequencers are now equipped with software for real-time processing of sequencing data. Different sequencing platforms utilize different optical systems and fluorescent dyes, resulting in differences in the spectral characteristics of fluorescent molecules' emission. If the algorithm fails to detect the appropriate characteristics or find the right parameters to handle these different characteristics, large errors in base classification can occur, affecting sequencing quality.

[0004] In addition, the second generation sequencing technology uses different fluorescent molecules with different fluorescence emission wavelengths. When these fluorescent molecules are irradiated by laser, they will emit fluorescent signals of corresponding wavelengths, such as Figure 1 By using a filter after laser irradiation to selectively filter out non-specific wavelengths of light, the fluorescence signal of a specific wavelength can be obtained, such as Figure 2 In DNA sequencing, there are four commonly used fluorescent markers. These four fluorescent markers are added to a cycle at the same time, and the image of the fluorescent signal is captured by a camera. Since each fluorescent marker corresponds to a specific wavelength, we can separate the fluorescent signals corresponding to different fluorescent markers in the image and obtain the corresponding fluorescent images, such as Figure 3. During this process, the camera focus can be adjusted and the sampling parameters can be set to ensure that the quality of the obtained TIF grayscale image is optimal. However, in actual applications, the brightness of the base clusters in the fluorescence image is always affected by multiple factors, mainly including the crosstalk between base clusters in the image (Spatial Crosstalk), crosstalk between channels (Crosstalk) and crosstalk between cycles (Phasing, Prephasing). The known base recognition technology mainly normalizes the crosstalk and intensity, but the correction methods are different. The fluorescence intensity value is corrected by the crosstalk matrix and the phasing and prephasing ratio in each cycle to remove the crosstalk noise, and then the base is identified by the light intensity values of the four channels, such as Figure 4 However, existing base recognition technology can only correct known brightness interference factors, such as brightness crosstalk between channels, phasing and prephasing caused by early or delayed reactions between cycles, but cannot correct brightness interference caused by other unknown biochemical or environmental influences, resulting in low recognition accuracy. When the sample density is higher, the base clusters are denser and the brightness crosstalk between base clusters is more severe, resulting in a significant reduction in sequencing accuracy. Summary of the Invention

[0005] In order to solve the existing technical problems, the embodiments of the present application provide a base recognition method, a base recognition training set construction method, a gene sequencer and a computer-readable storage medium that can overcome the spatial crosstalk between base signal acquisition units, adapt to different base signal acquisition unit densities, and thus effectively improve the base recognition accuracy.

[0006] To achieve the above objectives, the technical solution of the embodiment of the present application is implemented as follows:

[0007] In a first aspect, the present invention provides a method for constructing a base recognition training set, comprising:

[0008] Acquire multiple original fluorescence images corresponding to sequencing signal responses of different base types for the sequencing chip, and use the multiple original fluorescence images corresponding to sequencing signal responses of different base types as a multi-channel sample image of a training sample;

[0009] Performing primary base recognition on the original fluorescent image to obtain a base recognition result, and forming a mask image according to the position of the base signal acquisition unit;

[0010] Obtaining a base sequence according to the base recognition results of the original fluorescence images continuously collected by the sequencing chip during gene sequencing, aligning the base sequence with a standard base sequence in a known gene library, screening successfully aligned base sequences, and correcting the successfully aligned base sequences according to their respective matched standard base sequences, correcting the corresponding base recognition results determined by base recognition of the original fluorescence image according to the corrected base sequence, and obtaining a base type label of the multi-channel sample image as the training sample after correction;

[0011] The mask image is corrected according to the base sequences that have not been successfully aligned, and a mask label image is obtained after the correction.

[0012] In a second aspect, an embodiment of the present application provides a base recognition method, comprising:

[0013] Acquiring multi-channel input image data formed by multiple fluorescence images to be tested corresponding to sequencing signal responses of different base types for the sequencing chip;

[0014] The base recognition model uses the multi-channel input image data as input to recognize the original fluorescence image and outputs the base recognition results corresponding to the input image data of each channel; wherein the base recognition model is obtained by training the initial neural network model using the training samples obtained by the base recognition training set construction method described in any embodiment of the present application.

[0015] In a third aspect, an embodiment of the present application provides a gene sequencer, comprising a processor and a memory connected to the processor, wherein the memory stores a computer program executable by the processor, and when the computer program is executed by the processor, the method for constructing a training set for base recognition as described in any embodiment of the present application or the method for base recognition as described in any embodiment of the present application is implemented.

[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for constructing a training set for base recognition as described in any embodiment of the present application, or implements the base recognition method as described in any embodiment of the present application.

[0017] In the above embodiment, in the training set of base recognition, the training samples include a multi-channel sample image formed by multiple original fluorescence images corresponding to sequencing signal responses of different base types, base type labels of the multi-channel sample images, and mask label images. The training samples are a multi-channel input formed by multiple fluorescence images corresponding to different base categories, so that the prediction of the base recognition result can maintain the relative size relationship of the brightness values of the base signal acquisition unit on multiple channels, and has strong adaptability to overcoming the spatial crosstalk between the base signal acquisition units caused by various uncertain factors and adapting to the conditions of different base signal acquisition unit densities. In order to learn richer feature representations, the accuracy of base recognition results can be effectively improved; the base type labels and mask label images of multi-channel sample images are obtained after correction using standard base sequences in the known gene library, which not only effectively reduces the difficulty of labeling training samples, but also improves the labeling accuracy of training samples. The higher-precision training set is conducive to improving the recognition accuracy of the trained base recognition model; the introduction of mask label images and the use of mask strategies can make the output of the base recognition model only retain the predicted results of the base signal acquisition unit position, effectively eliminating background noise and interference, and further helping to improve the accuracy of base recognition.

[0018] In the above embodiments, the base recognition method, gene sequencer and computer-readable storage medium belong to the same concept as the corresponding base recognition training set construction method embodiments, and thus have the same technical effects as the corresponding base recognition training set construction method embodiments, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Schematic diagram of the distribution of fluorescence signal wavelengths of different fluorescent molecules in one embodiment;

[0020] Figure 2 FIG2 is a schematic diagram showing the principle of a camera collecting a fluorescence image in one embodiment, wherein the camera selectively filters out light of non-specific wavelengths using a filter to obtain an image of a fluorescence signal of a specific wavelength;

[0021] Figure 3 Schematic diagram of four fluorescence images corresponding to sequencing signal responses of four base types, A, C, G, and T, and a partially enlarged schematic diagram of one of the fluorescence images in one embodiment;

[0022] Figure 4 is a flowchart of a known base recognition in one embodiment;

[0023] Figure 5 Schematic diagram of a chip and a base signal acquisition unit on the chip in one embodiment;

[0024] Figure 6Flowchart of a method for constructing a training set for base calling in one embodiment;

[0025] Figure 7 is a flow chart of a base recognition method in one embodiment;

[0026] Figure 8 1. A model architecture diagram of a base recognition model in one embodiment;

[0027] Figure 9 Schematic diagram of the working principle of the base recognition model in one embodiment;

[0028] Figure 10 for Figure 9 Schematic diagram of the working principle of the classification prediction network to obtain base recognition results;

[0029] Figure 11 This is a schematic diagram of the structure of a Dense Block in one embodiment;

[0030] Figure 12 This is an overall flow chart of a base calling method and a training set construction method thereof in an optional specific example;

[0031] Figure 13 Schematic diagram of the structure of a gene sequencer in one embodiment. DETAILED DESCRIPTION

[0032] The technical solution of this application is further elaborated in detail below with reference to the accompanying drawings and specific embodiments.

[0033] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0034] In the following description, the expression "some embodiments" is involved, which describes a subset of all possible embodiments. It should be noted that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict.

[0035] In the following description, the terms "first, second, and third" are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first, second, and third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0036] The second generation of gene sequencing technology, also known as next generation sequencing technology (NGS), can sequence hundreds of thousands to millions of DNA molecules at a time. Known second generation sequencers generally record base information with optical signals, which are converted into base sequences through optical signals. The base cluster positions generated by image processing and fluorescence positioning technology are references for the subsequent chip template point positions. Therefore, image processing and fluorescence positioning technology are directly related to the accuracy of base sequence data. The base recognition method provided in the embodiment of the present application is based on the fluorescent image collected for the sequencing chip in the fluorescent labeling dNTP gene sequencing as input data, and is mainly used in the second generation of gene sequencing technology. Among them, fluorescent labeling is a measurement technology that uses optical signals and is commonly used in the industry in the fields of DNA sequencing, cell labeling, and drug research. The gene sequencing optical signal method used by the second generation sequencer is to use different bases with different wavelengths of fluorescence. After filtering through a filter, the successful connection of a specific base will excite light of a specific wavelength, which is finally identified as the DNA base sequence to be measured. This technology of generating images by collecting optical signals and then converting them into base sequences is the main principle of the second generation of gene sequencing technology.

[0037] The second-generation sequencer, taking the Illumina sequencer as an example, its sequencing process mainly includes four stages: sample preparation, cluster generation, sequencing and data analysis.

[0038] Sample preparation, also known as library construction, refers to breaking the basic group DNA to be tested into a large number of DNA fragments, and adding adapters to both ends of each DNA fragment. The adapters contain sequencing binding sites, indices (information identifying the source of the DNA segment), and specific sequences complementary to the oligonucleotides on the sequencing chip (Flowcell).

[0039] Cluster generation is achieved by seeding the library onto a flow cell and utilizing bridge DNA amplification to form a base cluster from a DNA fragment.

[0040] Sequencing refers to the sequencing read for each base cluster on the flow cell. The sequencing is performed by adding a fluorescently labeled dNTP sequencing primer. One end of the dNTP chemical formula is connected to an azide group, which can prevent polymerization during the sequencing chain extension, ensuring that only one base can be extended in a cycle, and a corresponding sequencing read is generated, that is, sequencing by synthesis. In one cycle, a base is identified for each base cluster by a fluorescently labeled dNTP, and the specific color of the fluorescent signal corresponds to the sequencing signal response of different base types. Laser scanning can be used to determine which base corresponds to each base cluster in the current cycle based on the emitted fluorescence color. In one cycle, tens of millions of base clusters are sequenced simultaneously on the flow cell. A fluorescent spot represents the fluorescence emitted by a base cluster, and a base cluster corresponds to a read in fastq. During the sequencing phase, an infrared camera captures a fluorescent image of the flowcell surface. The fluorescent image is processed and the fluorescent spot positions are located for base cluster detection (a traditional base cluster detection and positioning algorithm). Templates are constructed based on the base cluster detection results of multiple fluorescent images corresponding to the sequencing signal responses of different base types, and the positions of all base cluster template points (cluster) on the flowcell are constructed. Based on the base cluster positions in the template, the fluorescence intensity of the filtered image is extracted (a traditional base recognition algorithm), and then the fluorescence intensity is corrected. Finally, the base is identified based on the maximum intensity of each base cluster template point position, and the score is calculated, and the fastq base sequence file is output. Please refer to Figure 5 , respectively, are flowcell schematics ( Figure 5 (a) in the figure), and the fluorescence images taken of the corresponding parts on the flow cell in one cycle ( Figure 5 (b) in the figure), and the schematic diagram of the sequencing results in the fastq file ( Figure 5 (c) in the figure.

[0041] Data analysis, by analyzing millions of reads representing all DNA fragments, for each sample, the base sequences from the same library can be clustered by the unique index in the adapter introduced during the library construction process. The reads are paired to generate continuous sequences, which are compared with the reference genome for mutation identification.

[0042] It should be noted that the above is an illustration of the sequencing process using Illumina sequencing technology as an example of massively parallel sequencing technology (MPS). The DNA molecules to be tested are amplified through a specific amplification technology, and base clusters are amplified for each DNA fragment (single-stranded library molecule). The base cluster detection results are used to construct template points of the base clusters on the sequencing chip, so that subsequent operations such as base recognition can be performed based on the template points of the base clusters, thereby improving the efficiency and accuracy of base recognition. It can be understood that the base recognition method provided in the embodiment of the present application is a strategy of using machine learning to train a neural network model to improve the accuracy of base recognition. The training sample is based on the fluorescence image obtained from the base clusters after amplification of the single-stranded library molecules on the sequencing chip to perform base cluster positioning detection and base type identification. Here, each base cluster refers to a base signal acquisition unit, so it is not limited to which amplification technology is used for the single-stranded library molecules. That is, the base recognition method provided in the embodiment of the present application can also be applied to the base type identification of the base signal acquisition unit of the sequencing chip in other massively parallel sequencing technologies. For example, the base signal acquisition unit can refer to the base cluster obtained by the bridge amplification technology in the Illumina sequencing technology, and also includes the nanospheres obtained by the rolling circle amplification technology (RCA), etc., and the present application does not impose any restrictions on this.

[0043] See also Figure 6 , a method for constructing a base recognition training set provided in an embodiment of the present application, comprising the following steps:

[0044] S201 , obtaining multiple original fluorescence images corresponding to sequencing signal responses of different base types for a sequencing chip, and using the multiple original fluorescence images corresponding to sequencing signal responses of different base types as a multi-channel sample image of a training sample.

[0045] Each fluorescent spot in each raw fluorescence image corresponds one-to-one to each base signal acquisition unit of the corresponding base type. Base types typically refer to the four base types A, C, G, and T. Since different base types correspond to fluorescent signals from different fluorescently labeled dNTPs, there is no overlap between base signal acquisition units for different fluorescently labeled dNTPs. The raw fluorescence image corresponding to the sequencing signal response for each base type refers to the image of the base signal acquisition unit of the same base type contained in the corresponding portion of the sequencing chip after being excited and illuminated by the corresponding fluorescent label. Multiple raw fluorescence images corresponding to the sequencing signal responses of different base types are acquired for the sequencing chip, each raw fluorescence image including positional information for a base signal acquisition unit of a base type. Based on the positional information of the base signal acquisition units contained in each of the multiple raw fluorescence images, the positional information of the complete set of base signal acquisition units of multiple types contained in the target portion of the sequencing chip can be obtained. The target portion can be a localized location on the surface of the sequencing chip or the entire surface of the sequencing chip, and is typically related to the imaging area that can be included in a single fluorescence image.

[0046] The original fluorescence image refers to the original fluorescence image captured on the surface of the sequencing chip during the sequencing phase of the gene sequencing process. In this embodiment, the bases A, C, G, and T correspond to the fluorescence signals of four different fluorescently labeled dNTPs, respectively, and there is theoretically no intersection between the base signal acquisition units of the four different fluorescently labeled dNTPs. Acquiring multiple original images corresponding to the sequencing signal responses of different base types for the sequencing chip refers to capturing the fluorescence images corresponding to the fluorescence signals of four different fluorescently labeled dNTPs for the target site of the same sequencing chip, taking advantage of the different brightness of the four bases A, C, G, and T under light illumination of different wavelengths, and correspondingly capturing the fluorescence images (four original fluorescence images) corresponding to the four bases A, C, G, and T being excited and illuminated by the fluorescence signals (four environments) of the four different fluorescently labeled dNTPs for the same field of view (the same target site of the sequencing chip), as multiple original fluorescence images corresponding to the sequencing signal responses of different base types.

[0047] Multiple raw fluorescence images corresponding to sequencing signal responses of different base types collected within the same cycle of the gene sequencing process are grouped together and stacked along the channel dimension to form a multi-channel sample image for a training sample. For example, four fluorescence images to be tested corresponding to sequencing signal responses of the four base types A, C, G, and T are stacked along the channel dimension to form a 4-channel sample image, whose dimensions can be expressed as (4, H, W), where H and W are the height and width of the fluorescence images to be tested. The training set consists of a large number of training samples.

[0048] S203 , performing primary base recognition on the original fluorescent image to obtain a base recognition result, and forming a mask image according to the position of the base signal acquisition unit.

[0049] The original fluorescence image is subjected to primary base recognition to obtain a base recognition result, which mainly refers to the recognition result of the position information and base type of the base signal acquisition unit in the original fluorescence image obtained by various known algorithms, the accuracy of which may not reach the target. The known algorithm can be a traditional algorithm, or a currently known image recognition neural network model, such as a support vector machine (SVM), a convolutional neural network (CNN), a recurrent neural network (RNN), etc., which detects fluorescence images. In an optional example, initial base recognition refers to processing the original fluorescence image using any known traditional base signal acquisition unit detection and positioning algorithm to obtain the position of the base signal acquisition unit, and determining the base type of the base signal acquisition unit in the original fluorescence image collected in each cycle using a traditional base recognition algorithm based on the position of the base signal acquisition unit. Wherein, a mask image refers to a template selected for shielding the processed image for controlling the area or processing process of image processing. In a gene sequencing run on the same sequencing chip, the positions of the base signal acquisition units in the sequencing chip are the same, that is, the positions of the base signal acquisition units of all base types in the fluorescence images collected in different cycles should be the same. Therefore, in a gene sequencing run, forming a mask image based on the positions of the base signal acquisition units can refer to processing a group of original fluorescence images corresponding to sequencing signal responses of different base types through a traditional base signal acquisition unit detection and positioning algorithm, and forming a position data matrix or image based on the union of the positions of the base signal acquisition units in this group of original fluorescence images.

[0050] S205, obtaining a base sequence based on the base recognition results of the original fluorescence images continuously collected by the sequencing chip during gene sequencing, comparing the base sequence with the standard base sequences in a known gene library, screening the successfully aligned base sequences, and correcting the successfully aligned base sequences according to their respective matching standard base sequences, correcting the corresponding base recognition results determined by base recognition of the original fluorescence image according to the corrected base sequence, and obtaining the base type label of the multi-channel sample image as the training sample after correction.

[0051] Obtaining a base sequence based on the base recognition results of the original fluorescence images continuously collected from the sequencing chip during gene sequencing means that in one gene sequencing, the corresponding base type is identified according to the fluorescence intensity at the position of the corresponding base signal acquisition unit in the fluorescence images collected in different cycles, and the base sequence corresponding to the position of each base signal acquisition unit is formed according to the base type of the base signal acquisition unit in each cycle, that is, the base sequence. In gene sequencing, the detection and positioning of the base signal acquisition unit position and the accuracy of identifying the corresponding base type based on the fluorescence intensity at the base signal acquisition unit position in the fluorescence images collected in different cycles are inevitably affected by various factors. Based on the determination of the base type by primary base recognition, the obtained base sequence is compared with the standard base sequence in the known gene library. In a base sequence, the comparison is successful only when the base recognition is correct in excess of the proportion compared with the standard base sequence. In this way, all matching chains in the sample can be found. For the matching chains, the misidentified bases (mismatched bases below the proportion) in these matching chains are corrected according to the standard base sequence in the gene library. The base category results obtained by primary base recognition are then reversely corrected based on the corrected base sequence to correct and improve the quality of the base type labels of the multi-channel sample images used as training samples.

[0052] In an optional example, after preliminary base recognition, the base signal acquisition unit position points A (2, 2) and B (3, 3) are obtained. At this time, the mask image is a mask image in which position points A (2, 2) and B (3, 3) are 1, and the remaining positions are 0. Based on the base recognition results of the original fluorescence images collected for 10 consecutive cycles in gene sequencing, the base sequence of position A is obtained to be ACGTGTCAGT, and the base sequence of position B is obtained to be ACAGTTCAGT. After comparison with the standard base sequence in the known gene library, the standard base sequence that successfully aligns with the base sequence of position A is screened out as ACCTGTCAGT. Based on the standard base sequence, the base sequence of position A is corrected to ACCTGTCAGT. In this way, the base recognition results of the original fluorescence images collected for 10 consecutive cycles in gene sequencing are corrected according to the corrected base sequence. In the base recognition results of the original fluorescence image collected in the third cycle, the base type of position A is corrected from the originally recognized base type G to the base type C. That is, the base type label of the training sample formed by the original fluorescence image collected in the third cycle is correspondingly corrected.

[0053] S207 , correcting the mask image according to the base sequences that have not been successfully aligned, and obtaining a mask label image after correction.

[0054] On the basis of determining the base type through primary base recognition, the obtained base sequence is compared with the standard base sequence in the known gene library. In a base sequence, the comparison is successful only when more than a proportion of the base recognitions are correct compared to the standard base sequence. If no matching standard base sequence is screened after the comparison, it is considered that the overall recognition rate of the base type of the base signal acquisition unit position corresponding to the base sequence in the primary base recognition does not meet the requirements, and the base signal acquisition unit position of the base sequence that has not been successfully matched is deleted from the mask map. Correcting the mask map means removing the information of this chain from the mask map of the training sample based on the base sequence that has not been successfully matched. For example, the position of the chain that has not been successfully matched is replaced by 0 in the mask map formed by the base signal acquisition unit position obtained by primary base recognition, so as to avoid the contamination of the training data by erroneous data and improve the quality of the training samples. As in the above example, after comparison with the standard base sequences in the known gene library, no standard base sequence that is successfully aligned with the base sequence of position point B is screened out. Therefore, position point B is deleted from the mask image (position point B is changed from 1 to 0) for correction. The corrected mask label image is a mask image in which the value of position point A (2, 2) is 1 and the value of the other positions are 0.

[0055] In the above embodiment, in the training set of base recognition, the training samples include a multi-channel sample image formed by multiple original fluorescence images corresponding to sequencing signal responses of different base types, base type labels of the multi-channel sample images, and mask label images. The training samples are a multi-channel input formed by multiple fluorescence images corresponding to different base categories, so that the prediction of the base recognition result can maintain the relative size relationship of the brightness values of the base signal acquisition unit on multiple channels, and has strong adaptability to overcoming the spatial crosstalk between the base signal acquisition units caused by various uncertain factors and adapting to the conditions of different base signal acquisition unit densities. In order to learn richer feature representations, the accuracy of base recognition results can be effectively improved; the base type labels and mask label images of multi-channel sample images are obtained after correction using standard base sequences in the known gene library, which not only effectively reduces the difficulty of labeling training samples, but also improves the labeling accuracy of training samples. The higher-precision training set is conducive to improving the recognition accuracy of the trained base recognition model; the introduction of mask label images and the use of mask strategies can make the output of the base recognition model only retain the predicted results of the base signal acquisition unit position, effectively eliminating background noise and interference, and further helping to improve the accuracy of base recognition.

[0056] In some embodiments, in step S203, the original fluorescent image is subjected to primary base recognition to obtain a base recognition result, and a mask image is formed according to the position of the base signal acquisition unit, including:

[0057] For at least one training sample, the original fluorescence image is processed by a base signal acquisition unit detection and positioning algorithm to determine the position of the base signal acquisition unit, and a mask image is formed according to the position of the base signal acquisition unit;

[0058] According to the position of the base signal acquisition unit, the base signal acquisition unit in the original fluorescent image is identified by a base recognition algorithm to obtain a base recognition result.

[0059] In an embodiment of the present application, the sample image contained in a training sample is a multi-channel sample image formed by superimposing multiple original fluorescence images corresponding to sequencing signal responses of different base types. The base signal acquisition unit position is obtained according to the union of the base signal acquisition unit positions in a group of original fluorescence images corresponding to different base types. In this way, the mask image can be obtained by using the union of the base signal acquisition unit positions of multiple original fluorescence images in any training sample. In a gene sequencing process, all training samples formed by fluorescence images collected in different cycles can share the same mask image. For each training sample, the base recognition result in multiple original fluorescence images of the same group can be to identify the corresponding base type according to the fluorescence intensity at the base signal acquisition unit position that has been detected and located in the original fluorescence image, and the base type label of each training sample is obtained according to the recognition result of the base type of the original fluorescence image contained in the corresponding training sample. Thus, primary base recognition includes detecting and locating the position information of the base signal acquisition unit in the original fluorescence image collected in one or several cycles, and identifying and determining the base type of each base signal acquisition unit in the original fluorescence image collected in each cycle.

[0060] In some embodiments, in step S205, obtaining a base sequence according to the base recognition results of the original fluorescence images continuously collected from the sequencing chip during gene sequencing includes:

[0061] For the original fluorescence images continuously collected by the sequencing chip during gene sequencing, according to the positions of the base signal collection units in the corresponding mask image, the base signal collection units in the original fluorescence images are respectively identified by a base recognition algorithm to obtain base recognition results, and the base sequence is obtained according to the base recognition results of the continuously collected original fluorescence images; or

[0062] The original fluorescence images continuously collected by the sequencing chip during gene sequencing are identified by a preliminarily trained base recognition model to obtain base recognition results, and the base sequence is obtained according to the base recognition results of the continuously collected original fluorescence images.

[0063] In an embodiment of the present application, in the production process of training samples, the base recognition results obtained by primary base recognition are first used to form a base sequence, and then the base type label and the correction mask label image of the multi-channel sample image in the training sample are corrected by comparing with the standard base sequence in the known gene library. In the calibration scheme, the primary base recognition can include using a traditional base recognition algorithm to identify and determine the base type of each base signal acquisition unit in the original fluorescence image collected in each cycle, or it can be obtained by using a base recognition model that has been preliminarily trained. Among them, through the design of the calibration scheme, the accuracy requirement for the recognition of the base type in the original fluorescence image as the training sample can be reduced. In this way, the training samples obtained by the traditional base recognition algorithm can be used to train the initially constructed base recognition model, and the base recognition results can be obtained by using the preliminarily trained base recognition model before the training completion condition is met. Compared with the method in which the base recognition results of the original fluorescence images of all training samples are obtained by the traditional base recognition algorithm, the production efficiency of the training samples can be greatly improved.

[0064] In some embodiments, in step S201, a plurality of original fluorescence images corresponding to sequencing signal responses of different base types for a sequencing chip are obtained, and the plurality of original fluorescence images corresponding to sequencing signal responses of different base types are used as a multi-channel sample image of a training sample, including:

[0065] In gene sequencing, during multiple cycles corresponding to multiple base identifications, multiple fluorescence images corresponding to sequencing signal responses of different base types are collected from the target sites of the sequencing chip;

[0066] In each cycle, four original fluorescence images corresponding to the sequencing signal responses of the four types of bases A, C, G, and T are taken as a group, and each training sample includes a multi-channel sample image formed by a group of the original fluorescence images.

[0067] In the sequencing reads of the base signal acquisition unit, one cycle corresponds to one base recognition of each base signal acquisition unit. Since different base types correspond to the fluorescence signals of different fluorescent-labeled dNTPs, the four fluorescence images to be tested corresponding to the sequencing signal responses of the four types of bases A, C, G, and T can be the fluorescence signals of four different fluorescent-labeled dNTPs (4 environments) collected separately within one base recognition cycle to excite and illuminate the corresponding fluorescence images. In a base recognition cycle, each four original fluorescence images corresponding to the sequencing signal responses of the four types of bases A, C, G, and T are taken as a group. For each cycle, the brightness of the four base types A, C, G, and T is different under light illumination of different bands. Corresponding fluorescence images (4 grayscale images) of the four bases A, C, G, and T excited by the fluorescence signals of four different fluorescent-labeled dNTPs (4 environments) are collected for the same field of view (same chip target site). Each four fluorescence images corresponding to the four base types A, C, G, and T are taken as a group and serve as a training sample corresponding to one cycle. Each training sample includes a multi-channel sample image formed by stacking the four fluorescence images of the corresponding group.

[0068] See also Figure 7 In another aspect, the present invention provides a base recognition method, comprising:

[0069] S301, acquiring multi-channel input image data formed by multiple fluorescence images to be tested corresponding to sequencing signal responses of different base types for a sequencing chip;

[0070] S303, using the multi-channel input image data as input through a base recognition model, recognizes the original fluorescence image, and outputs a base recognition result corresponding to the input image data of each channel; wherein the base recognition model is obtained by training an initial neural network model using training samples obtained by the base recognition training set construction method described in an embodiment of the present application.

[0071] Among them, the fluorescent image to be tested refers to the original fluorescent image taken of the surface of the sequencing chip during the sequencing stage in the sequencing process. The multi-channel input image data corresponds to the same form as the multi-channel sample image in the training sample for training the base recognition model. The initial neural network model is trained using the training samples obtained by the base recognition training set construction method described in the embodiment of the present application, and the base recognition model is obtained after the training is completed. Each training sample includes a multi-channel sample image formed by multiple fluorescence images corresponding to different base types, a base type label corresponding to the multi-channel sample image, and a mask label image. The base recognition model performs supervised learning with the base type label as the training target. During the training phase, the base recognition model extracts training samples from the training set for iterative training. In each iterative training, multiple original fluorescence images corresponding to the sequencing signal responses of different base types in the training samples are used as a multi-channel input. The classification prediction network calculates and predicts the base recognition results of the input samples based on the current weight parameters and determines the recognition error based on the corresponding base type label. The base type prediction result at the center of the corresponding base signal acquisition unit is quickly extracted from the base recognition result through the mask label image to determine whether it is consistent with the base type label of the corresponding sample and whether the error is less than or equal to the set value. If the error is greater than the set value, , then back propagation is performed based on the error to optimize the weight parameters of the feature extraction network and the classification prediction network; and training samples are repeatedly extracted from the training data set as the input of the model for the next iterative training, and the iterative cycle is repeated to continuously optimize the weight parameters of the base recognition model until the classification prediction network calculates the base recognition result detected based on the current weight parameters. The recognition error of the base type at the center of each base signal acquisition unit is quickly extracted based on the corresponding mask label image and is less than the set value, that is, the classification prediction network performs supervised learning with the base type label and mask label image as training targets until the loss function converges to obtain the trained base recognition model.

[0072] The base recognition model obtained after training is then used to recognize the multi-channel input image data formed by the fluorescent images to be tested collected in each cycle in gene sequencing. This can fully utilize the automatic learning advantages of the neural network model and fully tap into more image information that traditional algorithms cannot extract to improve recognition accuracy. In particular, multiple fluorescent images to be tested corresponding to sequencing signal responses of different base types in the same cycle are formed into a multi-channel input of the base recognition model. This can fully retain the weak brightness difference information between the multiple fluorescent images to be tested. It has strong adaptability in correcting spatial crosstalk between base signal acquisition units caused by various unknown biochemical or environmental influences and adapting to different base signal acquisition unit densities, effectively improving base recognition accuracy.

[0073] In some embodiments, see Figure 8The base recognition model includes a feature extraction network and a classification prediction network; the base recognition model uses the multi-channel input image data as input to recognize the original fluorescence image and outputs a base recognition result corresponding to the input image data of each channel, including:

[0074] Performing feature extraction on the multi-channel input image data respectively through the feature extraction network of the base recognition model to obtain corresponding feature maps;

[0075] The classification prediction network uses the feature map output by the feature extraction network as input, and performs classification prediction on whether each pixel point in the input image data of each channel is the center of the base signal acquisition unit of the corresponding base type based on the feature map. According to the classification prediction result, the base recognition results corresponding to the input image data of each channel are output respectively through the output channel; wherein, the base recognition results include the recognition results of the base types respectively belonging to the positions of the centers of each base signal acquisition unit.

[0076] Please refer to Figure 9 and Figure 10 The feature extraction network uses multi-channel input image data as a multi-channel input and extracts features from the image to produce a feature map. The classification prediction network uses the feature map as input and, based on the feature map, classifies and predicts whether each pixel in each channel's input image data is a base signal acquisition unit center. Based on the classification prediction results, the network outputs the base call results corresponding to each channel's input image data through the output channels.

[0077] The base recognition results output by the base recognition model corresponding to the input image data of each channel can be presented in different forms, such as multi-channel or single-channel output. Multi-channel output refers to outputting multi-channel recognition results that correspond one-to-one to the image data of multiple input channels. For example, the recognition result of channel 1 is the recognition result of the base signal acquisition unit for the A base type in the current cycle, the recognition result of channel 2 is the recognition result of the base signal acquisition unit for the C base type in the current cycle, the recognition result of channel 3 is the recognition result of the base signal acquisition unit for the G base type in the current cycle, and the recognition result of channel 4 is the recognition result of the base signal acquisition unit for the T base type in the current cycle. Single-channel output refers to the output of a single-channel recognition result of the base signal acquisition unit position and its base type including all base types, which is formed based on the multi-channel recognition results corresponding to the image data of multiple input channels. For example, based on the union of the corresponding A base type recognition results, C base type recognition results, G base type recognition results, and T base type recognition results obtained by processing the input image data of each channel, an recognition result that simultaneously includes A, C, G, and T base signal acquisition units is formed in the current cycle.

[0078] Furthermore, the base recognition results can be presented in different forms, which is also reflected in the form of representing the recognition results of the base signal acquisition unit, which can be a data matrix that identifies the base type of each base signal acquisition unit in the current cycle, or an image that identifies the base type of each base signal acquisition unit. Taking multi-channel output as an example, each channel output corresponds to the recognition result of a base signal acquisition unit of a base type. Channel 1 can be a coordinate data matrix of the position information of the center of the base signal acquisition unit of base type A, so that the coordinate data matrix output by channel 1 represents the recognition result of the base signal acquisition unit of base type A in the current cycle; similarly, the coordinate data matrix of channel 2 corresponds to the recognition result of the base signal acquisition unit of base type C, the coordinate data matrix of channel 3 corresponds to the recognition result of the base signal acquisition unit of base type G, and the coordinate data matrix of channel 4 corresponds to the recognition result of the base signal acquisition unit of base type T. Taking single-channel output as an example, based on the recognition results of A, C, G, and T obtained by channels 1, 2, 3, and 4, a coordinate data matrix with base type labels marked at the corresponding positions of the base signal acquisition unit center in the current cycle is formed. It should be noted that although the output of the base recognition model includes the coordinate data matrix of the base signal acquisition unit center, it expresses the recognition of the base type at the center of different base signal acquisition units in the current cycle, and what is achieved is base type recognition.

[0079] The above-mentioned coordinate data matrix can adopt other forms that can characterize the base type at the center of each base signal acquisition unit, such as a probability data matrix indicating whether the pixel point is of a certain base type at the position of the center of the base signal acquisition unit. The probability value at the position of the center of the base signal acquisition unit represents the probability that the base signal acquisition unit belongs to the A, C, G or T base type.

[0080] Other forms of characterizing the base type at the center of each base signal acquisition unit can also be image forms, such as the position of the center of the base signal acquisition unit of the A, C, G, and T base types obtained according to the coordinate data matrix and the probability data matrix, and directly outputting the fluorescent image marked with the base type label at the position of the center of each base signal acquisition unit in the current cycle.

[0081] According to the various possible presentation forms of the base recognition results provided above, it can be seen that the base recognition results output by the base recognition model corresponding to the input image data of each channel are obtained after the base recognition model processes the multiple fluorescent images to be tested collected in the current cycle. The base recognition results can know the base types at the center of each base signal acquisition unit in the current cycle. It may not be limited to a specific form and is not limited here.

[0082] Take the probability data matrix representing the base type at the center of each base signal acquisition unit as an example, Figure 10 As shown, each channel represents a base type. The classification prediction network outputs the base recognition results corresponding to each channel through multiple output channels, including the probability that each pixel in the input image data of the corresponding channel belongs to the center of the base signal acquisition unit of the corresponding type of base. Here, the probability matrix output by each output channel can be directly used as the base recognition result. The probability matrix output by each channel is used to characterize the position of the center of the base signal acquisition unit of the corresponding type of base, as shown in FIG. Figure 10 Only the probability data matrix of channel A is shown. In this example, output channels 1, 2, 3, and 4 correspond to the four base types A, C, G, and T, respectively. The base recognition result of output channel 1 is a probability parameter matrix with a probability value of 0.95 at the position parameter (2, 2) and a probability value of 0.9 at (4, 4) for base type A. In other embodiments, the base types at the center positions of all base signal acquisition units can be further determined based on the probability matrices obtained from multiple channels, and base recognition results in which the positions of the centers of each base signal acquisition unit in the input image data of the corresponding channel are marked with other forms of base type labels are obtained. For example, output channels 1, 2, 3, and 4 correspond to the four base types of A, C, G, and T, respectively, and the base recognition results are coordinate data matrices in which the base type label at the center position point (2, 2) of the base signal acquisition unit is 1, the base type label at the center position point (4, 4) of the base signal acquisition unit is 1, the base type label at the center position point (3, 2) of the base signal acquisition unit is 3, and the base type label at the center position point (1, 4) of the base signal acquisition unit is 4.

[0083] In some embodiments, the classification prediction result includes the probability that each pixel point in the input image data of each channel is the center of the base signal acquisition unit of the corresponding type of base, and the sum of the probabilities of the pixel points at the same position in the multi-channel input image data is 1; and outputting the base recognition results corresponding to the input image data of each channel through the output channel according to the classification prediction result includes:

[0084] Determine the base type of the pixel at the center of the base signal acquisition unit corresponding to the base type of the input image data of each channel according to the classification prediction result;

[0085] The output channels respectively output the coordinate data matrix, probability data matrix or fluorescence image of the base signal acquisition unit center of the base type corresponding to the input image data of each channel; or, the output channels output the coordinate data matrix, probability data matrix or fluorescence image of the base type labels of the base types respectively located at the positions of the centers of the base signal acquisition units according to the input image data of each channel.

[0086] Multiple output channels correspond to multiple base types one by one. For the same set of multi-channel input image data, the pixel points at the same position in multiple images respectively represent the probability of being the center of the base signal acquisition unit of the corresponding base type, and the sum of the probabilities of the pixel points at the position of the center of the base signal acquisition unit in multiple images of the same group is 1. Figure 10 As shown, the classification prediction network outputs the base recognition results corresponding to each channel through the output channels, including the probability that each pixel point in the input image data of the corresponding channel belongs to the center of the base signal acquisition unit of the corresponding type of base. According to the probability, it can be determined whether the position of the center of the base signal acquisition unit is the corresponding type of base, and the coordinate data matrix of the center of the base signal acquisition unit of the corresponding type of base contained in the input image data of the corresponding channel is obtained. At the same time, its base type can be determined according to the maximum value of the probability corresponding to the center of the base signal acquisition unit. In an example, output channels 1, 2, 3, and 4 correspond to the four base types A, C, G, and T, respectively. The corresponding probability of the pixel at the coordinate parameter (2, 2) in the base recognition result of output channel 1 is 0.95, the corresponding probability of the pixel at the coordinate parameter (2, 2) in the base recognition result of output channel 2 is 0, the corresponding probability of the pixel at the coordinate parameter (2, 2) in the base recognition result of output channel 3 is 0.25, and the corresponding probability of the pixel at the coordinate parameter (2, 2) in the base recognition result of output channel 4 is 0.25. Therefore, the base type at the coordinate parameter (2, 2) in the base recognition result of output channel 1 is category label 1, representing the coordinate data matrix of base type A.

[0087] In some embodiments, the feature extraction network includes a primary convolutional layer and a dense block network layer; the feature extraction network of the base recognition model performs feature extraction on the multi-channel input image data respectively to obtain corresponding feature maps, including:

[0088] The multi-channel input image data is subjected to feature extraction through the primary convolutional layer of the feature extraction network; the primary features extracted by the primary convolutional layer are processed through the Dense block network layer, wherein the Dense block network layer includes a plurality of Dense blocks connected in sequence, and each convolutional layer in the Dense block takes the union of the outputs of the previous convolutional layer as input, and outputs a feature map corresponding to the multi-channel input image data through the last Dense block.

[0089] The primary convolution layer uses the number of convolution kernels, stride, and padding to maintain the spatial size of the feature map and extract primary features. In one example, in the primary convolution layer, each convolution kernel is 3x3, the stride is 1, the padding is 1, and the number of output channels is set to 64. The dense block network layer processes the primary feature map extracted by the primary convolution layer and forms a feature extraction network with the primary convolution layer to extract two-level features from multi-channel input images. The dense block network layer consists of six dense blocks connected in sequence. Figure 11 In one example, each Denseblock can contain 6 convolutional layers. Each convolutional layer uses a 3x3 convolution kernel with a stride of 1 and a padding of 1 to maintain the spatial size of the feature map. The number of output channels of each convolutional layer is set to 16, so the number of output channels of each Denseblock is 96. In the Denseblock network layer, each convolutional layer in the Denseblock takes the output union of the previous convolutional layer as input, fully mining the brightness contrast information between the same group of fluorescence images corresponding to different base types to extract image features, and then delivering the extracted high-level features to the classification prediction network. In this way, the image information of all channels can be fully considered, and the original brightness ratio between multiple fluorescence images in the same group can be maintained to obtain more accurate base recognition results.

[0090] In an embodiment of the present application, a multi-channel input is formed by using multiple fluorescent images corresponding to different types of bases. The brightness values of different channels represent different biological information, and the feature extraction of the feature extraction network retains the relative size relationship between the brightness values of different channels, thereby maintaining the original biological information and obtaining more accurate results.

[0091] In some embodiments, the loss function of the base recognition model is a cross-entropy loss function, which is expressed as follows:

[0092]

[0093] Where C is the number of categories, y is the one-hot encoding of the true label, and p is the probability distribution predicted by the model.

[0094] The cross entropy loss function (CE Loss) is used to measure the difference between the base category probability distribution predicted by the base recognition model and the true base category probability distribution. The goal of the cross entropy loss function is to minimize the cross entropy between the predicted probability distribution and the true probability distribution, so that the base category predicted by the classification prediction network is as close to the true base category as possible.

[0095] In order to have a more comprehensive understanding of the base recognition method and training set construction method provided in the examples of this application, please refer to Figure 12 The following is an explanation using a specific example, wherein the method for constructing a training set for base recognition includes:

[0096] S11, generate training samples. The generation of label data in the training samples includes:

[0097] 1. First, a traditional base call algorithm is used to perform base call on the sample images used for training, and the base category of each cluster in the sample image is obtained. The base categories of A, C, G, and T are 1, 2, 3, and 4 respectively.

[0098] 2. Using the preliminarily trained base recognition model, perform base recognition on the sample images used for training to obtain the base category of each cluster in the sample images; the training samples of the preliminarily trained base recognition model can be obtained through method 1, but the training has not yet been completed.

[0099] 3. Compare the base category results obtained by the traditional algorithm in method 1 or the model identified in method 2 with the standard sequence of the known gene library. In a chain, the comparison is successful only when most of the bases are correctly identified. In this way, all matching chains in the sample image can be found.

[0100] 4. Even in the matched chains, there are still a few base recognition errors. The incorrectly recognized bases in these chains are corrected according to the standard sequence in the gene library to obtain the corrected chain. In the corrected chain, all base categories are correct and can be used as base type label data during training.

[0101] 5. Use a mature cluster detection algorithm to detect and locate the positions of clusters in the image and obtain a mask image with the same size as the original image. That is, the center position of the cluster is 1, and the background area without clusters is 0. In the mask image, fill in the base category obtained in the previous step to the corresponding position (the position with 1 in the mask) to obtain the mask label image for training.

[0102] 6. For the chain that is not successfully matched, the information of this chain is removed from the training data set and the label data, that is, its position is replaced by 0 in the mask to avoid the erroneous data from contaminating the overall data.

[0103] 7. The sample image in method 1 or method 2 may be a multi-channel input formed by taking multiple fluorescence images corresponding to four base types for each cycle.

[0104] S12, forming a training set through the training samples. A training set may include training samples generated according to S11 from original fluorescence images collected in multiple consecutive cycles in a gene sequencing process.

[0105] The base recognition method comprises:

[0106] S13, constructing an initial neural network model, and training the neural network model with the training set to obtain a base recognition model. The architecture of the initial neural network model is as follows Figure 7 As shown in Figure 2. The training process of the initial neural network model mainly includes the following parts:

[0107] 1. Input

[0108] The four fluorescence images corresponding to the four base types are stacked in the channel dimension to form a four-channel input data with the dimension of (4, H, W) where H and W are the height and width of the training image.

[0109] 2. Feature extraction

[0110] Primary convolution layer: The input fluorescence image first passes through a convolution layer. A 3x3 convolution kernel with a stride of 1 and a padding of 1 can be used here to maintain the spatial size of the feature map. The output channels of this layer are set to 64 for extracting primary features. Dense blocks: The primary feature map is processed by 6 Dense blocks. In each Denseblock, 6 convolution layers can be included. Each convolution layer uses a 3x3 convolution kernel with a stride of 1 and a padding of 1 to maintain the spatial size of the feature map. The output channels of each convolution layer are set to 16, and the output channels of each Dense block are 96. The high-level features extracted by the Dense blocks are fed into the Basecall network.

[0111] 4. Basecall Network

[0112] Basecall convolutional layer: Contains two convolutional layers to convert high-level feature maps into Basecall results. The convolution kernel size is 3x3, the stride is 1, and the padding is 1. The number of output channels of the last convolutional layer is equal to the number of predicted categories.

[0113] All the above convolution and deconvolution layers use ReLU as the activation function.

[0114] 5. Loss Function

[0115] Use CE Loss as the loss function.

[0116] S14, collecting multiple fluorescence images to be tested corresponding to the sequencing signal responses of different base types for the sequencing chip, forming a multi-channel input image data input into the base recognition model, performing feature extraction through the feature extraction network, and outputting the corresponding base recognition results by the classification prediction network.

[0117] In the above-mentioned embodiment, the construction of the training set introduces the technical idea of first performing primary base recognition and then correcting the base type labels of the training samples using standard base sequences. This can reduce the difficulty of data annotation used for model training, improve the accuracy of training samples, and enhance the model's ability to learn data that has not been successfully aligned, thereby improving the model's ability to recognize all base signal acquisition units. The base recognition model uses a technical idea of stacking multiple fluorescence images corresponding to the sequencing signal responses of different base types to form a multi-channel input for image feature extraction. This can effectively retain the original brightness contrast information between the same group of fluorescence images to improve the accuracy of base type recognition and improve the ability to correct brightness interference between base signal acquisition units caused by unknown biochemical or environmental influences. The strategy of introducing a mask label image into the output of the base recognition results can only retain the prediction results of the center position of the base signal acquisition unit, effectively focusing the model's attention on important areas and eliminating possible background noise and interference. This has a positive effect on improving the accuracy of base results, faster convergence of the loss function in the model training phase, and improving the efficiency of base recognition in the model recognition application phase.

[0118] Another aspect of the present application is to provide a gene sequencer. Figure 13, is an optional hardware structure diagram of a gene sequencer provided in an embodiment of the present application, wherein the gene sequencer includes a processor 111 and a memory 112 connected to the processor 111, and the memory 112 stores a computer program for implementing the method for constructing a training set for base recognition provided in any embodiment of the present application, so that when the corresponding computer program is executed by the processor, the steps of the method for constructing a training set for base recognition provided in any embodiment of the present application are implemented, or the memory 112 stores a computer program for implementing the base recognition method provided in any embodiment of the present application, so that when the corresponding computer program is executed by the processor, the steps of the base recognition method provided in any embodiment of the present application are implemented. The gene sequencer loaded with the corresponding computer program has the same technical effect as the corresponding method embodiment, and to avoid repetition, it is not repeated here.

[0119] On the other hand, the embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-mentioned base recognition training set construction method embodiment or each process of the base recognition method embodiment, and can achieve the same technical effect. To avoid repetition, it is not repeated here. Wherein, the computer-readable storage medium is such as a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.

[0120] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0121] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0122] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for constructing a training set for base recognition, characterized in that: include: Acquire multiple original fluorescence images corresponding to sequencing signal responses of different base types for the sequencing chip, and use the multiple original fluorescence images corresponding to sequencing signal responses of different base types as a multi-channel sample image of a training sample; Performing primary base recognition on the original fluorescent image to obtain a base recognition result, and forming a mask image according to the position of the base signal acquisition unit; Obtaining a base sequence according to the base recognition results of the original fluorescence images continuously collected by the sequencing chip during gene sequencing, aligning the base sequence with a standard base sequence in a known gene library, screening successfully aligned base sequences, and correcting the successfully aligned base sequences according to their respective matched standard base sequences, correcting the corresponding base recognition results determined by base recognition of the original fluorescence image according to the corrected base sequence, and obtaining a base type label of the multi-channel sample image as the training sample after correction; The mask image is corrected according to the base sequences that have not been successfully aligned, and a mask label image is obtained after the correction.

2. The method for constructing a base recognition training set according to claim 1, wherein: The original fluorescent image is subjected to primary base recognition to obtain a base recognition result, and a mask image is formed according to the position of the base signal acquisition unit, including: For at least one training sample, the original fluorescence image is processed by a base signal acquisition unit detection and positioning algorithm to determine the position of the base signal acquisition unit, and a mask image is formed according to the position of the base signal acquisition unit; According to the position of the base signal acquisition unit, the base signal acquisition unit in the original fluorescent image is identified by a base recognition algorithm to obtain a base recognition result.

3. The method for constructing a base recognition training set according to claim 1, wherein: The obtaining of a base sequence according to the base recognition results of the original fluorescent images continuously collected from the sequencing chip during gene sequencing includes: For the original fluorescence images continuously collected by the sequencing chip during gene sequencing, according to the positions of the base signal collection units in the corresponding mask image, the base signal collection units in the original fluorescence images are respectively identified by a base recognition algorithm to obtain base recognition results, and the base sequence is obtained according to the base recognition results of the continuously collected original fluorescence images; or The original fluorescence images continuously collected by the sequencing chip during gene sequencing are identified by a preliminarily trained base recognition model to obtain base recognition results, and the base sequence is obtained according to the base recognition results of the continuously collected original fluorescence images.

4. The method for constructing a base recognition training set according to claim 1, wherein: The step of acquiring multiple original fluorescence images corresponding to sequencing signal responses of different base types from a sequencing chip and using the multiple original fluorescence images corresponding to sequencing signal responses of different base types as a multi-channel sample image of a training sample includes: In gene sequencing, during multiple cycles corresponding to multiple base identifications, multiple fluorescence images corresponding to sequencing signal responses of different base types are collected from the target sites of the sequencing chip; In each cycle, four original fluorescence images corresponding to the sequencing signal responses of the four types of bases A, C, G, and T are taken as a group, and each training sample includes a multi-channel sample image formed by a group of the original fluorescence images.

5. A base recognition method, characterized in that: include: Acquiring multi-channel input image data formed by multiple fluorescence images to be tested corresponding to sequencing signal responses of different base types for the sequencing chip; The base recognition model uses the multi-channel input image data as input to recognize the original fluorescence image and outputs a base recognition result corresponding to the input image data of each channel; wherein the base recognition model is obtained by training an initial neural network model using training samples obtained using the base recognition training set construction method according to any one of claims 1 to 4.

6. The base recognition method according to claim 5, wherein The base recognition model includes a feature extraction network and a classification prediction network; the base recognition model uses the multi-channel input image data as input to recognize the original fluorescence image and outputs a base recognition result corresponding to the input image data of each channel, including: Performing feature extraction on the multi-channel input image data respectively using the base recognition model to obtain corresponding feature maps; The classification prediction network uses the feature map output by the feature extraction network as input, and performs classification prediction on whether each pixel point in the input image data of each channel is the center of the base signal acquisition unit of the corresponding base type based on the feature map. According to the classification prediction result, the base recognition results corresponding to the input image data of each channel are output respectively through the output channel; wherein, the base recognition results include the recognition results of the base types respectively belonging to the positions of the centers of each base signal acquisition unit.

7. The base recognition method according to claim 6, wherein The classification prediction result includes the probability that each pixel point in the input image data of each channel is the center of the base signal acquisition unit of the corresponding type of base, and the sum of the probabilities of the pixel points at the same position in the multi-channel input image data is 1; Outputting the base recognition results corresponding to the input image data of each channel through the output channels according to the classification prediction results includes: Determine the base type of the pixel at the center of the base signal acquisition unit corresponding to the base type of the input image data of each channel according to the classification prediction result; The output channel outputs the coordinate data matrix, probability data matrix or fluorescence image of the base signal acquisition unit center of the base type corresponding to the input image data of each channel; or, the output channel outputs the coordinate data matrix, probability data matrix or fluorescence image of the base type label of the base type at the position of the center of each base signal acquisition unit according to the input image data of each channel.

8. The base recognition method according to claim 6, wherein The feature extraction network includes a primary convolutional layer and a dense block network layer; the feature extraction network of the base recognition model performs feature extraction on the multi-channel input image data respectively to obtain corresponding feature maps, including: The multi-channel input image data is subjected to feature extraction through the primary convolutional layer of the feature extraction network; the primary features extracted by the primary convolutional layer are processed through the Dense block network layer, wherein the Dense block network layer includes a plurality of Dense blocks connected in sequence, and each convolutional layer in the Dense block takes the union of the outputs of the previous convolutional layer as input, and outputs a feature map corresponding to the multi-channel input image data through the last Dense block.

9. The base recognition method according to claim 5, wherein The loss function of the base recognition model is the cross entropy loss function, which is expressed as follows: Where C is the number of categories, y is the one-hot encoding of the true label, and p is the probability distribution predicted by the model.

10. A gene sequencer, characterized in that: It comprises a processor and a memory connected to the processor, wherein the memory stores a computer program executable by the processor, and when the computer program is executed by the processor, it implements the method for constructing a training set for base recognition as described in any one of claims 1 to 4, or implements the base recognition method as described in any one of claims 5 to 9.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the method for constructing a training set for base recognition according to any one of claims 1 to 4, or implements the base recognition method according to any one of claims 5 to 9.

Citation Information

Patent Citations

  • Base category detection method and device, electronic equipment and storage medium

    CN115376613A

  • Base cluster detection method and device, gene sequencer and storage medium

    CN116596933A