Base calling method and apparatus, gene sequencer, and storage medium

Through semi-supervised learning and data augmentation technology, multiple cycle fluorescence images and mask marking are used to solve the problem of insufficient base recognition accuracy in gene sequencing, achieving higher recognition accuracy and model adaptability.

WO2025148599A1PCT designated stage expired Publication Date: 2025-07-17SHENZHEN SALUS BIOMED CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/138250
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-08
Filing Date
2024-12-10
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

In the existing gene sequencing technology, the base recognition accuracy is affected by the background difference of fluorescence image and the brightness crosstalk, which increases the difficulty of model training. The traditional methods cannot effectively correct unknown interference factors, resulting in low recognition accuracy, especially in high-density samples.

Method used

Using semi-supervised learning method, by obtaining fluorescent image data under multiple cycles and using mask diagrams to mark signal acquisition units with base type tags, the training model is constructed, and combined with data enhancement technology, the model adaptability and generalization ability of the model to the brightness relationship between different cycles and channels is improved.

Benefits of technology

It improves the accuracy of base recognition, reduces the risk of overfitting, enhances the adaptability and generalization ability of the model in different situations, and improves the accuracy of gene sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024138250_17072025_PF_FP_ABST
    Figure CN2024138250_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are a base calling method and apparatus based on semi-supervised learning, a gene sequencer, and a storage medium. The method comprises: obtaining fluorescence images to be analyzed corresponding to base signal acquisition units of multiple base types, and on the basis of said fluorescence images, forming input image data to be analyzed; and using said input image data as the input of a trained base calling model, and outputting a base calling result of said input image data by means of the trained base calling model, the trained base calling model being a model obtained by performing semi-supervised learning training on the basis of a training data set, wherein the training data set comprises training samples collected over multiple cycles, each training sample comprises sample fluorescence image corresponding to multiple base types, base type label maps corresponding to the sample fluorescence images, and corresponding first mask maps and corresponding second mask maps.
Need to check novelty before this filing date? Find Prior Art

Description

Base recognition method and device, gene sequencer and storage medium Technical Field

[0001] The present invention relates to the field of gene technology, and in particular to a base recognition method and device based on semi-supervised learning, a gene sequencer, and a computer-readable storage medium. Background Art

[0002] Sequencers are widely used for genome sequencing, enabling rapid and accurate DNA sequencing. Sequencer algorithms have evolved from traditional dataset-independent algorithms to deep learning-based algorithms trained on datasets. Deep learning involves training a deep learning network using a dataset that includes training samples and labels. During the training process, the deep learning network is trained using labels as training targets, and similar labels corresponding to the training samples are obtained through fitting of the deep learning network. Therefore, the effectiveness of deep learning depends on both the dataset and the network model. The dataset is fundamental, and obtaining a complete and representative dataset is more conducive to improving the base calling accuracy of deep learning-based sequencing algorithms.

[0003] Gene sequencing involves analyzing the base sequence of a DNA fragment—specifically, the arrangement of adenine (A), thymine (T), cytosine (C), and guanine (G)—in the data to be tested. The input image for gene sequencing consists of clusters containing multiple base types. After staining the sample, fluorescence is excited by a specific laser and captured by a camera. By using different laser powers to excite the sample, fluorescence of varying brightness is generated, resulting in four fluorescence images captured at different laser powers: an image of A bases, an image of T bases, an image of C bases, and an image of G bases. The brightness of these captured fluorescence images is analyzed to identify the base type of each base cluster in the data to be tested. However, since each of the four images captured at different laser powers contains information about only one base type, the amount of information is limited. Furthermore, due to the varying laser powers, the background brightness of the four images also varies. For example, the image captured at a higher power may appear brighter overall than the image captured at a lower power. This results in significant background differences between the fluorescence images of different base types. When training a deep learning network model, due to the large background differences between training samples, the deep learning network model will pay more attention to the classification results brought by the background differences rather than the classification results brought by the brightness differences of the gene clusters themselves, making it difficult for the deep learning network model to converge, thereby increasing the difficulty of training.

[0004] Currently, gene sequencing technology can be divided into three generations. The first-generation sequencing technology, the Sanger method, is based on DNA synthesis reactions. It is also known as the SBS method or the terminal termination method. It was proposed by Sanger in 1975, and the first complete genome sequence of an organism was published in 1977. The second-generation sequencing technology, represented by the Illumina platform, has achieved high-throughput sequencing and made revolutionary progress, making large-scale parallel sequencing a reality and greatly promoting the development of genomics in the life sciences. The third-generation sequencing technology is Nanopore sequencing, a new generation of single-molecule real-time sequencing technology. It mainly uses the changes in the electrical signal caused by the passage of ssDNA or RNA template molecules through the nanopore to infer the base composition for real-time sequencing.

[0005] Second-generation gene sequencing utilizes fluorescence microscopy to capture the signals of fluorescent molecules in an image. The base sequence is then decoded from the fluorescence signals in the image. To distinguish between bases, filters are used to capture images of the fluorescence intensity of the sequencing chip at different frequencies to obtain the spectral characteristics of the fluorescent molecules' emission. Multiple images are captured for the same scene. These images are then positioned and aligned, and the point signals are extracted and brightness information analyzed to obtain the base sequence. With the advancement of second-generation sequencing technology, sequencers now come equipped with software for real-time processing of sequencing data. Different sequencing platforms utilize different optical systems and fluorescent dyes, resulting in variations in the spectral characteristics of fluorescent molecules' emission. If the algorithm fails to detect the appropriate characteristics or find the right parameters to handle these varying characteristics, significant errors in base classification can occur, impacting sequencing quality.

[0006] Furthermore, next-generation sequencing utilizes different fluorescent molecules with different fluorescence emission wavelengths. When these fluorescent molecules are irradiated by laser light, they emit fluorescence signals of corresponding wavelengths, as shown in Figure 1. After laser irradiation, filters are used to selectively filter out non-specific wavelengths of light, allowing fluorescence signals of specific wavelengths to be captured, as shown in Figure 2. Four fluorescent markers are commonly used in DNA sequencing. These four fluorescent markers are simultaneously added to a cycle, and a camera is used to capture images of the fluorescence signals. Because each fluorescent marker corresponds to a specific wavelength, the fluorescence signals corresponding to the different fluorescent markers can be separated in the image, resulting in the corresponding fluorescence images, as shown in Figure 3. During this process, the camera focus can be adjusted and sampling parameters can be set to ensure optimal TIF grayscale image quality. However, in practical applications, the brightness of base clusters in fluorescence images is often affected by various factors, including spatial crosstalk between base clusters within the image, crosstalk within channels, and crosstalk between cycles (phasing and prephasing). Existing base calling technologies primarily normalize crosstalk and intensity, but correction methods vary. Fluorescence intensity values ​​are corrected using the crosstalk matrix and phasing / prephasing ratio within each cycle to remove crosstalk noise. Bases are then identified based on the intensity values ​​of the four channels, as shown in Figure 4. However, existing base calling technologies can only correct for known brightness interference factors, such as brightness crosstalk between channels and phasing and prephasing caused by premature or delayed reactions between cycles. They are unable to correct for brightness interference caused by other unknown biochemical or environmental factors, resulting in low recognition accuracy. As sample density increases, the denser the base clusters, the more severe the brightness crosstalk between base clusters, significantly reducing sequencing accuracy. Existing machine learning methods typically operate only after brightness extraction, using only the extracted center brightness information of the base clusters as input. These methods fail to fully exploit the image data to leverage the advantages of machine learning and fail to fully utilize information across multiple cycles. Therefore, recognition accuracy needs to be improved. Moreover, the training of existing machine learning models requires precise labels. Due to the limitations of traditional sequencing algorithms, there are always about 10% of base chains that cannot be accurately labeled, which affects the accuracy of model training. Summary of the Invention

[0007] In order to solve the existing technical problems, the embodiments of the present invention provide a base recognition method, device, equipment and computer-readable storage medium based on semi-supervised learning, which can implement a training method based on semi-supervised learning to obtain a base recognition model, so that the model can better understand and generalize to different situations, thereby improving the accuracy of base recognition.

[0008] In a first aspect, a base recognition method based on semi-supervised learning is provided, comprising:

[0009] Acquire fluorescence images to be measured corresponding to base signal acquisition units of multiple base types on a sequencing chip, and form input image data to be measured based on the fluorescence images to be measured; wherein the fluorescence images to be measured include fluorescence images to be measured corresponding to multiple base types;

[0010] Using the input image data to be tested as input to a trained base recognition model, and outputting a base recognition result of the input image data to be tested through the trained base recognition model, wherein the trained base recognition model is a model obtained by semi-supervised learning based on a training data set;

[0011] The training data set includes training samples collected in multiple cycles, each training sample includes sample fluorescence images corresponding to multiple base types, and base type label images corresponding to the sample fluorescence images, and the training labels corresponding to each training sample also include a first mask image corresponding to the sample fluorescence image and a second mask image corresponding to the sample fluorescence image, wherein the first mask image is used to mark the position of the base signal acquisition unit with a base type label in the sample fluorescence image; the second mask image is used to mark the position of the base signal acquisition unit without a base type label in the sample fluorescence image.

[0012] In a second aspect, a base recognition device based on semi-supervised learning is provided, comprising:

[0013] an acquisition module, configured to acquire fluorescence images to be measured corresponding to base signal acquisition units of multiple base types on a sequencing chip, and to form input image data to be measured based on the fluorescence images to be measured; wherein the fluorescence images to be measured include fluorescence images to be measured corresponding to multiple base types;

[0014] a recognition module, configured to use the input image data to be tested as input to a trained base recognition model, and output a base recognition result of the input image data to be tested through the trained base recognition model, wherein the trained base recognition model is a model obtained by semi-supervised learning based on a training data set;

[0015] The training data set includes training samples collected in multiple cycles, each training sample includes sample fluorescence images corresponding to multiple base types, and base type label images corresponding to the sample fluorescence images, and the training labels corresponding to each training sample also include a first mask image corresponding to the sample fluorescence image and a second mask image corresponding to the sample fluorescence image, wherein the first mask image is used to mark the position of the base signal acquisition unit with a base type label in the sample fluorescence image; the second mask image is used to mark the position of the base signal acquisition unit without a base type label in the sample fluorescence image.

[0016] In a third aspect, a gene sequencer is provided, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the base recognition method based on semi-supervised learning provided in the first aspect of the present application.

[0017] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the processor executes the steps of the base recognition method based on semi-supervised learning provided in the first aspect of the present application.

[0018] In a fifth aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the base recognition method based on semi-supervised learning provided in the first aspect of the present application.

[0019] In a sixth aspect, a gene sequencer is provided, comprising: an acquisition module and an identification module;

[0020] The acquisition module is configured to acquire fluorescence images to be tested corresponding to base signal acquisition units of multiple base types on the sequencing chip, and form input image data to be tested based on the fluorescence images to be tested; wherein the fluorescence images to be tested include fluorescence images to be tested corresponding to multiple base types;

[0021] The recognition module is configured to use the input image data to be tested as input to a trained base recognition model, and output a base recognition result of the input image data to be tested through the trained base recognition model, wherein the trained base recognition model is a model obtained by semi-supervised learning based on a training data set;

[0022] The training data set includes training samples collected in multiple cycles, each training sample includes sample fluorescence images corresponding to multiple base types, and base type label images corresponding to the sample fluorescence images, and the training labels corresponding to each training sample also include a first mask image corresponding to the sample fluorescence image and a second mask image corresponding to the sample fluorescence image, wherein the first mask image is used to mark the position of the base signal acquisition unit with a base type label in the sample fluorescence image; the second mask image is used to mark the position of the base signal acquisition unit without a base type label in the sample fluorescence image.

[0023] The base recognition method and device based on semi-supervised learning, genetic tester, and computer-readable storage medium provided in the present application use training samples collected under multiple cycles for training and learning when training the base recognition model, so that the base recognition model learns the brightness relationship characteristics of sample images under different cycles when predicting the base recognition results, which can improve the model's adaptability to early or delayed reactions under different cycles. In addition, each training sample includes fluorescence images corresponding to multiple base channels, so that the base recognition model learns the brightness relationship characteristics between different base channels, which can improve the model's adaptability to brightness crosstalk between different base channels. By marking base clusters that already have real base type labels in the sample fluorescence image through a first mask image, and marking base clusters that do not have real base type labels in the sample fluorescence image through a second mask image, a semi-supervised learning method can be implemented when training the base recognition model. The data with real base type labels in the training sample are supervised learning with the real base type labels as the training target, which can make the model focus on the characteristics of the data with real base type labels, thereby accelerating model convergence. At the same time, base clusters without true base type labels allow the model to focus on the characteristics of data without true base type labels during training, learning more diverse characteristics of data without base type labels, helping the model better understand and generalize to different situations. This allows the model to better balance training data and generalization requirements, reducing the risk of overfitting. Furthermore, the second mask image can be used to integrate base clusters without true base type labels into the training samples, thereby increasing the size of the training sample. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] FIG1 is a schematic diagram showing the distribution of fluorescence signal wavelengths of different fluorescent molecules in one embodiment;

[0025] FIG2 is a schematic diagram showing the principle of a camera capturing a fluorescence image in one embodiment, wherein the camera utilizes a filter to selectively filter out light of non-specific wavelengths to obtain an image of a fluorescence signal of a specific wavelength;

[0026] FIG3 is a schematic diagram of four fluorescence images corresponding to sequencing signal responses of four ATCG base types in one embodiment, and a partially enlarged schematic diagram of one of the fluorescence images;

[0027] FIG4 is a flowchart of a known base recognition method according to an embodiment;

[0028] FIG5 is a schematic diagram of a chip and a base signal acquisition unit on the chip in one embodiment;

[0029] FIG6 is a flow chart of a base recognition method based on semi-supervised learning in one embodiment;

[0030] FIG7 is a schematic diagram of a first mask image and a second mask image in one embodiment;

[0031] FIG8 is a flowchart of training a base recognition model in a base recognition method based on semi-supervised learning in one embodiment;

[0032] FIG9 is a schematic diagram of generating a mask image based on base cluster positions in one embodiment;

[0033] FIG10 is a schematic diagram of the structure of a base recognition model in one embodiment;

[0034] FIG11 is an architecture diagram of a base recognition method based on semi-supervised learning in one embodiment;

[0035] FIG12 is a schematic diagram of calculating a first loss value when training a base recognition model in a base recognition method based on semi-supervised learning in one embodiment;

[0036] FIG13 is a schematic diagram of calculating a second loss value when training a base recognition model in a base recognition method based on semi-supervised learning in one embodiment;

[0037] FIG14 is a schematic diagram of a base recognition device based on semi-supervised learning in one embodiment;

[0038] FIG15 is a schematic diagram of the structure of a gene sequencer in one embodiment. DETAILED DESCRIPTION

[0039] The technical solution of the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the scope of protection of the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0041] In the following description, reference is made to “some embodiments” which describe a subset of all possible embodiments, but it should be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0042] Gene sequencing involves analyzing the base sequence of a DNA fragment—specifically, the arrangement of adenine (A), thymine (T), cytosine (C), and guanine (G)—in order to identify the bases. Fluorescent labeling is currently widely used for gene sequencing. The sequencing optical system uses lasers to excite fluorescent markers on a sequencing chip, which then emit fluorescence. The fluorescence signal is then collected. The four bases bind to different fluorescent markers, producing four distinct fluorescence bands, which are then used to identify the bases.

[0043] Next-generation sequencing technology, such as the Illumina sequencer, utilizes different fluorescent molecules with different fluorescence emission wavelengths. When these fluorescent molecules are irradiated by laser light, they emit fluorescence signals of corresponding wavelengths. After laser irradiation, filters are used to selectively filter out non-specific wavelengths of light to obtain fluorescence signals of specific wavelengths. By acquiring and analyzing the fluorescence signals, base types can be identified. This process mainly includes sample preparation, cluster generation, sequencing, and data analysis.

[0044] Sample preparation: The DNA sample to be sequenced is extracted and purified, followed by DNA fragmentation and adapter ligation. In some embodiments, the DNA sample is cut into a large number of smaller DNA fragments, typically using ultrasound or restriction endonucleases. Adapters containing specific sequences are then ligated to both ends of the DNA fragments for subsequent ligation and sequencing reactions.

[0045] Cluster generation: This process involves amplifying DNA fragments to form fixed DNA fragments, allowing for subsequent clustering of bases. In an alternative example, DNA fragments are amplified using methods such as polymerase chain reaction (PCR) or bridge amplification, resulting in millions of copies of each DNA fragment. The amplified DNA fragments are then fixed to a stationary plate, where each DNA fragment forms a separate cluster.

[0046] Sequencing refers to the sequencing read for each base cluster on the flow cell. Sequencing is performed by adding a fluorescently labeled dNTP sequencing primer. One end of the dNTP chemical formula is connected to an azide group, which can prevent polymerization during the sequencing chain extension, ensuring that only one base can be extended in a cycle, and a corresponding sequencing read is generated, that is, sequencing while synthesis. In one cycle, a fluorescently labeled dNTP is used to identify a base for each base cluster, and the specific color of the fluorescent signal corresponds to the sequencing signal response of different base types. Laser scanning can be used to determine which base corresponds to each base cluster in the current cycle based on the emitted fluorescence color. In one cycle, tens of millions of base clusters are sequenced simultaneously on the flow cell. A fluorescent spot represents the fluorescence emitted by a base cluster, and a base cluster corresponds to a read in fastq. During the sequencing phase, an infrared camera captures a fluorescent image of the flowcell surface, processes the fluorescent image, and locates the fluorescent spot position for base cluster detection. A template is constructed based on the base cluster detection results of multiple fluorescent images corresponding to the sequencing signal responses of different base types, and the positions of all base cluster template points (cluster) on the flowcell are constructed. Based on the template, the fluorescence intensity of the filtered image is extracted, and then the fluorescence intensity is corrected. Finally, the base is identified and the score is calculated based on the maximum intensity of each base cluster template point position, and the fastq base sequence file is output. Please refer to Figure 5, which shows a flowcell schematic (Figure 5 (a)), a fluorescent image taken of the corresponding part of the flowcell during a cycle (Figure 5 (b)), and a schematic diagram of the sequencing results displayed in the fastq file (Figure 5 (c)).

[0047] The gene sequencer may also include an optical platform, which may include an operating table and a camera. The sequencing chip may be placed on the operating table. The gene sequencer uses a laser to excite the fluorescent markers on the sequencing chip to produce fluorescence and collects the fluorescence signal. The four bases bind to different fluorescent markers to produce four different fluorescence bands, which are also called fluorescence images of four base types. The camera takes a picture of the sequencing chip and captures a fluorescence image of the fluorescence signal generated by the charge-coupled device (CCD) on the test chip. In a fluorescence image, there are many fluorescent spots, and each fluorescent spot in the fluorescence image represents the fluorescence emitted by a base cluster.

[0048] The imaging mode of the gene sequencer can be a four-channel imaging system or a two-channel imaging system. For the two-channel imaging system, each camera needs to expose twice at the same position of the test chip. For the four-channel imaging system, the camera of each channel shoots once at the same position of the sample, and obtains fluorescent images of four base types respectively. For example, a fluorescent image of the A base type, a fluorescent image representing the A base type, a fluorescent image of the C base type, a fluorescent image of the G base type, and a fluorescent image of the T base type are obtained respectively. Since a filter is used after laser irradiation to selectively filter out non-specific wavelengths of light to obtain a fluorescent signal of a specific wavelength, each base type corresponds to a different fluorescent signal. In the same cycle reaction, the brightness of the same type of base cluster in its corresponding category of base type is much greater than that of other categories of bases. In theory, there will be no duplication of the base clusters that emit light in each channel.

[0049] After the gene sequencer obtains the fluorescent image, it will reconstruct the collected image, perform gene image registration, and gene base identification (gene basecall) to obtain the gene sequence.

[0050] Gene image reconstruction is used to improve the resolution of fluorescence images, thereby enhancing image clarity and reducing crosstalk between samples. Gene image reconstruction includes, but is not limited to, conventional operations such as deconvolution.

[0051] Gene image registration involves correcting the fluorescence images of four base types so that they overlap. This allows the fluorescence brightness of the four channels at the same location to be extracted, facilitating subsequent base identification. Gene image registration includes, but is not limited to, image registration within the same channel and global or local affine registration.

[0052] The gene recognition process uses the registered image to determine whether the base clusters in the image belong to one of the four bases: A, C, G, or T. After gene recognition, the test data is converted from a digital image into sequence information of the four bases A, C, G, and T, which is the DNA sequence result of the sample for subsequent analysis and evaluation.

[0053] Data Analysis: Analyze and interpret sequencing data based on image data and sequence information. Compare sequence information with the reference genome for mutation identification.

[0054] The process of sequencing a single sample is called a run. The sequencing process consists of multiple cycles, each corresponding to a reaction cycle and, in other words, a single base type identification on the sequencing chip. Sequencing occurs simultaneously through synthesis. In a single cycle, tens of millions of base clusters are sequenced simultaneously.

[0055] A test data set consists of many DNA fragments. During the sequencing process, a base is added to each DNA fragment. Therefore, the length of the base sequence of the test data determines the number of cycles. In each cycle, the gene sequencer generates a fluorescence image for each of the four base types (ACGT). When sequencing the test data, the gene sequencer can obtain fluorescence images of the ACGT channel for multiple cycles.

[0056] It should be noted that the above is an illustration of the sequencing process using Illumina sequencing technology as an example of massively parallel sequencing technology (MPS). The DNA molecules to be tested are amplified through a specific amplification technology, and base clusters are amplified for each DNA fragment (single-stranded library molecule). The base cluster detection results are used to construct template points of the base clusters on the sequencing chip, so that subsequent operations such as base recognition can be performed based on the template points of the base clusters, thereby improving the efficiency and accuracy of base recognition. It is understandable that the base recognition method based on fluorescently labeled dNTP gene sequencing provided in the embodiment of the present application is based on the positioning detection and base type identification of the base clusters after amplification of the single-stranded library molecules on the sequencing chip. Here, each base cluster refers to a base signal acquisition unit, so that it is not limited to which amplification technology is adopted for the single-stranded library molecules. That is, the base recognition method based on semi-supervised learning provided in the embodiment of the present application can also be applied to the positioning detection and base type identification of the base signal acquisition unit of the sequencing chip in other massively parallel sequencing technologies. For example, the base signal acquisition unit can refer to the base cluster obtained by the bridge amplification technology in the Illumina sequencing technology, and also includes the nanospheres obtained by the rolling circle amplification technology (RCA, Rolling Circle Amplification), etc., and the present application does not limit this. In the following embodiments, for ease of understanding, the base signal acquisition unit is described as an example of a base cluster.

[0057] Please refer to Figure 6, which is a flowchart of a base recognition method based on semi-supervised learning provided in one embodiment of the present application. The base recognition method based on semi-supervised learning is applied to a gene sequencer, and the base recognition method based on semi-supervised learning includes the following steps:

[0058] S11. Acquire fluorescence images to be measured corresponding to base signal acquisition units of multiple base types on the sequencing chip.

[0059] In this embodiment, the fluorescence image to be measured includes fluorescence images corresponding to multiple base types. The fluorescence image to be measured can be a fluorescence image collected in one cycle or in multiple cycles.

[0060] The sequencing process for a single genetic sample is called a run. A genetic sample is fragmented into M test base sequences, also called short chains. Each test base sequence consists of N base clusters. In a single cycle, the top base clusters of these M short chains undergo sequencing reactions on the sequencing chip. Each base cluster being sequenced corresponds to a position on the sequencing chip, and tens of millions of base clusters are sequenced simultaneously in a single cycle. N determines the number of cycles used in the test; a larger N indicates a higher number of cycles. Sequencing is performed on each of the M test base clusters in each cycle. For example, if a genetic sample is fragmented into 10,000 short chains, each 100 bases long, 100 cycles of sequencing reactions are required to identify the base types. In each cycle, the top base clusters of these 10,000 short chains undergo sequencing reactions on the sequencing chip.

[0061] During the sequencing reaction, different base clusters on the sequencing chip are each labeled with a different fluorescent marker. During a single cycle, the gene sequencer uses a laser to excite the fluorescent light on the sequencing chip, emitting a fluorescent signal. The gene sequencer's camera then captures a fluorescent image of the target location on the sequencing chip within the field of view corresponding to that cycle. During each cycle, the gene sequencer's camera takes a single image, producing fluorescent images corresponding to various base types, such as an image of an A base type, a C base type, a G base type, and a T base type. For example, if the gene sequencer's imaging system uses a four-channel imaging mode, then during each cycle, a single image is taken within the field of view of that cycle, producing fluorescent images of all four base types. For example, if a gene sample to be tested is fragmented into 10,000 short chains, during each cycle, the gene sequencer's camera adjusts its field of view to capture fluorescent images of the top base clusters on the sequencing chip corresponding to those 10,000 short chains within the field of view of that cycle. If one base cluster corresponds to a read, then there are 10,000 reads.

[0062] S12. Using the input image data to be tested as the input of the trained base recognition model, and outputting the base recognition result of the input image data to be tested through the trained base recognition model, wherein the trained base recognition model is a model obtained by semi-supervised learning based on the training data set.

[0063] In this embodiment, the trained base recognition model is a model obtained by semi-supervised learning based on a training dataset. The training dataset includes training samples collected in multiple cycles, each training sample includes sample fluorescence images corresponding to multiple base types and base type label images corresponding to the sample fluorescence images, and the training labels corresponding to each training sample also include a first mask image corresponding to the sample fluorescence image and a second mask image corresponding to the sample fluorescence image, wherein the first mask image is used to mark the positions of base signal acquisition units with base type labels in the sample fluorescence image; and the second mask image is used to mark the positions of base signal acquisition units without base type labels in the sample fluorescence image.

[0064] The sample fluorescence images corresponding to the multiple base types in one cycle include a sample fluorescence image of the A base type, a sample fluorescence image of the C base type, a sample fluorescence image of the G base type, and a sample fluorescence image of the T base type. The base type label image is used to identify the base type label of the base cluster at the position corresponding to the first mask image in the sample fluorescence image. The image size of the first mask image and the second mask image is the same as the size of the sample fluorescence image. For a gene sequencing on the same sequencing chip, the first mask image corresponding to the sample fluorescence images in multiple cycles is the same, and the second mask image corresponding to the sample fluorescence images in multiple cycles is the same. For example, if 300,000 short chains are sequenced simultaneously in a gene sequencing process, where the base length of each short chain is 100 bases, then the first mask image corresponding to the sample fluorescence images collected in these 100 cycles is the same, and the second mask image corresponding to the sample fluorescence images collected in these 100 cycles is the same.

[0065] The first mask image is used to mark the base cluster positions in the sample fluorescence image that already have true base type labels. For example, in the first mask image, the base cluster positions with true base type labels are marked as "1", and other positions are marked as 0. The second mask image is used to mark the base cluster positions in the sample fluorescence image that do not have true base type labels. For example, in the second mask image, the base cluster positions without true base type labels are marked as "1", and other positions are marked as 0. When training the base recognition model, the first mask image and the second mask image can be used to mark the data with true base type labels and the data without true base type labels in the training samples. For the data with true base type labels, during training, training and learning can be carried out with the true base type labels as the training target. For the data without true base type labels, unlabeled training and learning can be carried out. In this way, the data with true base type labels and the data without true base type labels are combined to train the base recognition model, thereby realizing a semi-supervised learning. Data with true base type labels usually contain features from various situations. However, some collected data may contain missing information, and base clusters at some positions do not have true base type labels. By introducing data without true base type labels into the training set, the model can learn more information, thereby better adapting to various situations and improving generalization performance.

[0066] For example, as shown in Figure 7, in the 4×4 sample fluorescence image, there are base clusters at (1,3), (2,1), (2,3), (3,2), (3,4), and (4,2) in the base cluster position diagram, and other positions are background images. The first mask image, the sample fluorescence image, and the second mask image are all 4×4 matrix images, where the first mask image is marked as "1" at positions (1,3), (2,1), (2,3), and (4,2), indicating that the input sample fluorescence image has true base type labels at positions (1,3), (2,1), (2,3), and (4,2). Base calling is performed on the input sample fluorescence image, generating an output image representing the base types. After processing the output image with the first mask, base call results are found at positions (1,3), (2,1), (2,3), and (4,2). In the output image processed with the first mask, 1 indicates base type A, 2 indicates base type C, 3 indicates base type G, and 4 indicates base type T. In the second mask, positions (3,2) and (3,4) are marked as "1," while all other positions are marked as "0." Base clusters at positions marked as "1" do not have true base type labels. Non-base cluster positions are marked as "0" in both the first and second masks; background values ​​at non-base cluster positions are not included in the calculation.

[0067] When training a base recognition model, training samples collected under multiple cycles are used for training and learning, so that the base recognition model learns the brightness relationship characteristics of sample images under different cycles when predicting base recognition results, which can improve the model's adaptability to early or delayed reactions under different cycles. In addition, each training sample includes fluorescence images corresponding to multiple base channels, so that the base recognition model learns the brightness relationship characteristics between different base channels, which can improve the model's adaptability to brightness crosstalk between different base channels. The base clusters that already have real base type labels in the sample fluorescence image are marked by a first mask image, and the base clusters that do not have real base type labels in the sample fluorescence image are marked by a second mask image. This allows a semi-supervised learning method to be implemented when training the base recognition model. The data with real base type labels in the training samples are trained and learned with the real base type labels as the training targets, which allows the model to focus on the characteristics of the data with real base type labels, thereby accelerating model convergence. At the same time, base clusters without true base type labels allow the model to focus on the characteristics of data without true base type labels during training, learning more diverse characteristics of data without base type labels, helping the model better understand and generalize to different situations. This allows the model to better balance training data and generalization requirements, reducing the risk of overfitting. Furthermore, the second mask image can be used to integrate base clusters without true base type labels into the training samples, thereby increasing the size of the training sample.

[0068] In some embodiments, the method further comprises:

[0069] Get the training dataset;

[0070] Acquire training samples from the training data set as input training samples, process the input training samples based on different data augmentation methods to obtain multiple groups of processed training samples corresponding to the input training samples, and form multiple groups of input data corresponding to the input training samples based on the multiple groups of processed training samples corresponding to the input training samples;

[0071] Constructing an initial base recognition model, using multiple groups of input data corresponding to the input training samples as inputs to the base recognition model, obtaining base recognition data corresponding to each group of input data, iteratively training the initial base recognition model using the training data set until the loss function converges, and obtaining the trained base recognition model;

[0072] The loss functions include:

[0073] calculating a first loss function of a first loss value between adjusted base identification data corresponding to each set of input data and the base type label map corresponding to the input training sample, wherein the adjusted base identification data corresponding to each set of input data is obtained by adjusting the base identification data corresponding to each set of input data based on the first mask map corresponding to the input training sample;

[0074] and a second loss function for calculating a second loss value between base recognition data corresponding to each two sets of processed input data, wherein the base recognition data corresponding to each set of processed input data is obtained by processing the base recognition data corresponding to each set of input data based on the second mask image corresponding to the input training sample.

[0075] In the above embodiment, the input training samples are processed based on different data enhancement methods to obtain multiple groups of processed training samples, thereby increasing the training samples. The multiple groups of input data corresponding to the input training samples formed based on the multiple groups of processed training samples are used as input data of the base model under training. During the iterative training process, the first mask image is used to block the base clusters without real base type labels in the base recognition data corresponding to each group of input data. When calculating the loss using the first loss function, more attention is paid to the base recognition results of the base clusters with real base type labels, thereby reducing the impact of the recognition results of the base clusters without real base type labels, thereby speeding up the training speed of the model. At the same time, a second mask image is introduced to block the base clusters with real base type labels in the base recognition data corresponding to each set of input data. When using the second loss function to calculate the loss, more attention is paid to the consistency loss of the recognition results of the base clusters without real base type labels in each two groups. This allows the model to learn more features of base clusters without real base type labels during training, and learn more diverse features, helping the model to better understand and generalize to different situations. In this way, the model can better balance training data and generalization requirements, and reduce the risk of overfitting.

[0076] As shown in FIG8 , FIG8 is a flowchart of training a base recognition model in a base recognition method based on semi-supervised learning in one embodiment; the flowchart includes:

[0077] S81. Obtain a training data set.

[0078] The training dataset includes training samples collected over multiple cycles. Each training sample includes sample fluorescence images corresponding to multiple base types, a base type label corresponding to the sample fluorescence image, a first mask corresponding to the sample fluorescence image, and a second mask corresponding to the sample fluorescence image. For example, sample fluorescence images of four base types, A, C, G, and T, are collected in one cycle.

[0079] In some embodiments, a conventional base cluster position location algorithm is used to locate the base cluster position representing the center of the base cluster in the sample fluorescence images of multiple base types collected at each cycle. A conventional base recognition algorithm is used to perform base recognition on the base types at the base cluster positions in the sample fluorescence images of multiple base types collected at each cycle, obtaining base recognition results corresponding to the sample fluorescence images at each cycle. A base sequence is obtained based on the base recognition results of the sample fluorescence images collected continuously at multiple cycles on the sequencing chip. The base sequence is compared with a standard base sequence in a known gene library to determine the base sequences that successfully compare with the standard base sequence and the base sequences that fail to compare with the standard base sequence. A first mask image and a second mask image are generated based on the base cluster positions, the base sequences that successfully compare with the standard base sequence, and the base sequences that fail to compare with the standard base sequence. A mask image refers to a template selected to block the area or processing process of the processed image for controlling the image processing. The first mask image is used to mark the base cluster positions with true base type labels and to block the base cluster positions without true base type labels in the processed image. The second mask image is used to mark the base cluster positions without real base type labels and to mask the base cluster positions with real base type labels in the processed image.

[0080] Optionally, a base sequence in which the proportion of correctly identified bases is greater than or equal to a preset ratio is determined as a base sequence that has successfully been compared with the standard base sequence, and a base sequence in which the proportion of correctly identified bases is less than the preset ratio is determined as a base sequence that has failed to be compared with the standard base sequence. The proportion of correctly identified bases in a base sequence is equal to the number of correctly identified bases in the base sequence divided by the total number of bases in the base sequence.

[0081] During a gene sequencing run on a sequencing chip, multiple sample gene sequences, i.e., multiple sample short chains, are input at once. In each cycle, the base clusters at the top of the multiple sample short chains are sequenced, and photos are taken to obtain sample fluorescence images of multiple base types in each cycle. In each cycle, the position of the base cluster corresponding to each sample short chain is fixed on the sequencing chip. Therefore, during a gene sequencing run, the first and second mask images for the gene sequencing run can be obtained based on the base cluster positions on the sequencing chip and the alignment results of the base sequences corresponding to the multiple sample short chains, i.e., the first and second mask images corresponding to the multiple sample short chains. The generated first mask image serves as the first mask image corresponding to the sample fluorescence images collected during multiple cycles of the gene sequencing run, and the generated second mask image serves as the second mask image corresponding to the sample fluorescence images collected during multiple cycles of the gene sequencing run.

[0082] During a gene sequencing process on the same sequencing chip, multiple sample short chains with the same input are sequenced. Therefore, the first mask images corresponding to the sample fluorescence images of multiple base types collected in each cycle are the same, and the corresponding second mask images are also the same.

[0083] As shown in Figure 9, Figure 9 is a schematic diagram of generating a mask map based on base cluster positions in one embodiment; the base cluster position distribution map is a base cluster position distribution map when the base clusters at the top of the four short chains are sequenced in one cycle, and the background position in the base cluster position distribution map is "0". Among them, base cluster A1 represents the base cluster of sample short chain A1, base cluster A2 represents the base cluster of sample short chain A2, and base cluster A3 represents the base cluster of sample short chain A3. The base length of these three sample short chains is 10. The traditional base recognition algorithm is used to perform base recognition on the sample fluorescence images of various base types collected under 10 cycles of sample short chains A1, A2 and A3, and the base sequences of sample short chains A1, A2 and A3 are obtained respectively. After comparison with the standard base sequence, the base sequence corresponding to sample short chain A1 is the base sequence with successful comparison, and the base sequences corresponding to sample short chains A2 and A3 are the base sequences with failed comparison. Therefore, the position of the base cluster of sample short chain A1 in the first mask image generated according to the base cluster position is marked as "1", and the other positions are 0, indicating that the base cluster at the position marked as "1" has a real base type label. The position of the base cluster of the sample short chain A1 in the first mask image generated according to the base cluster position is marked as "0", and the remaining positions are 1, indicating that the base cluster at the position marked as "1" has no real base type label.

[0084] Optionally, in the successfully compared base sequence, some bases may be incorrectly identified based on the traditional base recognition algorithm. The incorrectly identified bases in the successfully compared base sequence are corrected according to the standard base sequence to obtain a corrected base sequence. Based on the corrected base sequence and the located base cluster position on the sequencing chip, the base type label map corresponding to the sample fluorescence image corresponding to the multiple base types in each cycle is determined.

[0085] S82. Obtain training samples from the training data set as input training samples, process the input training samples based on different data enhancement methods to obtain multiple groups of processed training samples corresponding to the input training samples, and form multiple groups of input data corresponding to the input training samples based on the multiple groups of processed training samples corresponding to the input training samples.

[0086] In some embodiments, different data enhancement methods include at least one combination of the following: adding different noises to the input training samples, and performing different brightness processing on the input training samples.

[0087] The image size of each set of processed training samples is the same as the image size of the input training samples. For example, two different random Gaussian noises are added to the input training sample A to obtain two processed training samples A1 and A2.

[0088] By processing the training samples in different data enhancement methods, multiple groups of processed training samples can be obtained, thereby expanding the scale of the training samples. In addition, data enhancement methods such as adding noise to the training samples can improve the diversity and robustness of the training sample data. This allows the model to learn more features in the training samples during training, making the trained base recognition model more adaptable to data of different data types, thereby improving the accuracy of base recognition.

[0089] S83. Construct an initial base recognition model, use multiple groups of input data corresponding to the input training samples as inputs of the base recognition model, obtain base recognition data corresponding to each group of input data, iteratively train the initial base recognition model using the training data set, and calculate the loss value in each iteration process based on the loss function.

[0090] In some embodiments, the base recognition model is a deep learning model based on the Unet (U-shaped Convolutional Neural Network) network, which mainly includes an encoder (Encoder), intermediate connections (Skip Connections), and a decoder (Decoder).

[0091] The encoder consists of four convolutional layers and a pooling layer (MaxPooling). It extracts features from the input image, gradually reducing the resolution of the input image and capturing features at different scales. Skip connections connect the encoder's feature maps to the corresponding decoder's feature maps. These skip connections allow information to flow freely between the encoder and decoder, helping the network better recover detailed information. The decoder converts the features extracted by the encoder into predictions at the same resolution as the input image. The decoder typically consists of deconvolutional layers and upsampling layers. The convolutional layer performs convolution operations on the input image data to extract features. The pooling layer downsamples the output of the convolutional layer, reducing the data dimensionality, model complexity, and computational complexity. The deconvolutional layer uses deconvolution to upsample the encoder input image to produce the decoded image. Furthermore, to preserve the details of the original image and minimize information loss caused by the convolution operation, skip connections are used between the encoder and decoder. This approach allows intermediate feature maps from the encoding process to be directly concatenated and fused with feature maps of the corresponding scale from the decoding process based on the channel dimension. Unet (U-shaped Convolutional Neural Network) is a common deep learning neural network architecture, primarily consisting of an encoder, intermediate connections, and a decoder.

[0092] Figure 10 is a schematic diagram of the composition of the base recognition model in one embodiment; the input image is (12, H, W), where H and W represent the length and width of the training image, respectively. First, the four fluorescence images under each cycle are stacked according to the channel dimension to create a four-channel input data as the data of a cycle. The dimension of this input data is (4, H, W), where H and W represent the height and width of the training image, respectively. The sample fluorescence images under multiple cycles are input at a time. Taking 3 cycles as an example, the input data is (12, H, W), where H is 2160 and W is 4096. For the input image (12, H, W), the encoder performs four convolutions and four downsamplings in succession in the encoding stage, doubling the number of channels each time and halving the length and width. Then the decoder uses an upsampling operation in the decoding stage, and the encoder and decoder are connected using a jump connection.

[0093] In some embodiments, during each iteration, multiple sets of input data corresponding to the input training samples are respectively used as inputs to the base recognition model to obtain base recognition data corresponding to each set of input data. For each set of input data, the base recognition data corresponding to each set of input data are masked using a first mask image corresponding to the input training sample, retaining the base recognition data at base cluster positions with true base type labels. This can reduce the impact of incorrect base recognition results at base cluster positions with true base type labels on the first loss value of the training model. The first loss value for that iteration is then calculated based on the first loss function, combining the base recognition data at base cluster positions with true base type labels during that iteration with the base type label image corresponding to the input training sample of that iteration. In this way, the first loss value corresponding to each set of input data during that iteration can be calculated.

[0094] For each set of input data, the base call data corresponding to each set of input data is masked using the second mask corresponding to the input training sample, retaining the base call data for base cluster positions without true base type labels. This can reduce the impact of incorrect base calls at base cluster positions with true base type labels on the second loss value of the training model. The second loss value between the base call data corresponding to each two sets of input data processed during this iteration is then calculated based on the second loss function, thereby obtaining multiple second loss values ​​during this iteration. For example, if there are two sets of input data, the second loss value between the base call data corresponding to the two processed sets of input data is directly calculated. If there are three sets of input data, the second loss value is calculated separately between the base call data corresponding to each two processed sets of input data.

[0095] The loss value in the iteration process is calculated based on the first loss value and multiple second loss values ​​corresponding to each set of input data in the iteration process.

[0096] Optionally, the loss value FinalLoss in iterative training is calculated based on the loss function as follows:

[0097] Wherein CE_j represents the first loss value corresponding to the j-th group of input data calculated based on the first loss function, n represents n groups of input data, w(t) represents the weight corresponding to the t-th iteration round, MSE_j represents the j-th second loss value calculated based on the second loss function, and k represents the total number of calculated second loss values.

[0098] Optionally, the first loss function is:

[0099] Where CE(p,y) is the cross entropy loss function, C is the number of categories, and y i is the one-hot encoding of the true label of the i-th category, p i is the probability distribution value of the base recognition model predicting that the base cluster type is the i-th type.

[0100] Optionally, the second loss function is:

[0101] Where N represents the number of pixels in the base recognition data corresponding to each set of input data, p i and q i are the base recognition data distributions corresponding to each two sets of input data after processing, where p i represents the probability distribution of the base category of the i-th pixel in the base recognition data corresponding to one set of input data in each two groups, q i It represents the probability distribution of the base category of the i-th pixel in the base recognition data corresponding to the other set of input data in every two groups.

[0102] Optionally, the process of obtaining the trained base recognition model based on the training data set includes multiple rounds of iterations, wherein in one round of iteration, the base recognition model is trained using the training data set based on multiple iteration numbers, and w(t) increases with the increase in the number of iteration rounds.

[0103] Generally, during the training process, multiple rounds of iteration are required to complete model training. In one round of iteration, there are m iterations. The value of m depends on the size of the training dataset. Generally, the larger the training dataset, the larger the value of m. In one round of iteration, there are multiple iterations. Each iteration extracts a portion of training samples from the training dataset as input training samples. One round of iteration is completed until all training samples in the training dataset have been extracted. For example, if m is 4000 and three rounds of iteration are required to complete training, then the first, second, and third rounds will each have 4000 iterations. In the second round of iteration, w(t) is greater than w(t) in the first round of iteration, and w(t) in the third round of iteration is greater than w(t) in the second round of iteration.

[0104] In this embodiment, in the early stage of semi-supervised learning training, the base recognition model needs to learn more from the base type label map as the training target to accelerate the learning ability of the model. Therefore, the value of w(t) in the early stage of iteration is very small. In the later stage of training, the base recognition model tends to be stable and can learn more useful information from the consistency regularization of data without real base type labels. Therefore, it is necessary to increase the value of w(t), that is, to increase the MSE weight, so that the model can learn more features of base clusters without real base type labels during training, and learn more diverse features, which helps the model better understand and generalize to different situations. In this way, the model can better balance the training data and generalization requirements and reduce the risk of overfitting.

[0105] As shown in Figure 11, Figure 11 is an architectural diagram of a base recognition method based on semi-supervised learning in one embodiment. In each iteration, the input training sample is processed using a first data augmentation method to obtain image data X1, and then processed using a second data augmentation method to obtain image data X2. Based on the X1 and X2 image data, the input data for the corresponding base recognition models are obtained. After recognition by the base recognition models, the base recognition results Y1 corresponding to the X1 image data and Y2 corresponding to the X2 image data are obtained. The base recognition results Y1 and Y2 are processed using a first mask image to obtain image data Z1 corresponding to Y1 and image data Z2 corresponding to Y2. Based on a first loss function, the loss value CE_1 is calculated between image data Z1 and the base type label image. Based on the first loss function, the loss value CE_2 is calculated between image data Z2 and the base type label image. Based on a second loss function, the loss value U1 corresponding to Y1 and image data U2 corresponding to Y2 are obtained between U1 and U2. The loss value MSE is calculated between U1 and U2 using a second loss function. Then the total loss value corresponding to the input training sample of this iteration is:

[0106] FIG12 is a schematic diagram illustrating the calculation of a first loss value when training a base recognition model in a base recognition method based on semi-supervised learning in one embodiment. FIG12 takes an input training sample comprising sample fluorescence images of four base types as an example. The center position of the base cluster of the sample fluorescence image of base type A is located at (1, 3). The input training sample is processed by a data augmentation method to form a set of input data for the base recognition model. After recognition by the base recognition model, a base recognition result Y1 of A is obtained as shown in FIG12. After base recognition result Y1 is processed by a first mask image B1, image data Z1 is obtained. The loss is calculated for image data Z1 and base type label image D.

[0107] As shown in Figure 13, Figure 13 is a schematic diagram of calculating the second loss value when training a base recognition model in a base recognition method based on semi-supervised learning in one embodiment; taking the input training sample including sample fluorescence images of four base types as an example, the black dot positions in the base cluster position distribution map in the sample fluorescence image are base cluster positions, and other positions are background positions. The first mask image and the second mask image are obtained based on the base cluster position distribution map. The center position of the base cluster of the sample fluorescence image of base type A is located at (1,3). The input training sample is processed by a data enhancement method to form a set of input data of the base recognition model as an example. After recognition by the base recognition model, a base recognition result Y1 of base A as shown in Figure 13 is obtained. The input training sample is processed by another data enhancement method to form a set of input data of the base recognition model as an example. After recognition by the base recognition model, a base recognition result Y2 of base A as shown in Figure 13 is obtained. After base recognition result Y1 and base recognition result Y2 are processed by the second mask image B2, U1 image data corresponding to base recognition result Y1 and U2 image data corresponding to base recognition result Y2 are obtained. Based on the second loss function, the loss between U1 image data and U2 image data is calculated.

[0108] S84. Determine whether the iteration termination condition is met.

[0109] In some embodiments, the iteration termination condition includes, but is not limited to, the number of iterations and whether the loss value during the iteration is less than a preset loss value. If the iteration termination condition is not met during the iteration process, the process returns to S82 and continues to obtain training samples from the training dataset and train the base recognition model until the iteration termination condition is met. If the iteration condition is met, the process proceeds to S85.

[0110] S85. The base recognition model after the iteration is terminated is used as the trained base recognition model.

[0111] On the other hand, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the base recognition method based on semi-supervised learning described in any embodiment of the present application.

[0112] Among them, in the computer program product, an optional implementation form of the program module architecture of the computer program that implements each step of the base recognition method based on semi-supervised learning can be a base recognition device based on semi-supervised learning.

[0113] Referring to FIG. 14 , an embodiment of the present application provides a base recognition device based on semi-supervised learning, comprising: an acquisition module 21 for acquiring fluorescence images to be tested corresponding to base signal acquisition units of multiple base types on a sequencing chip, and forming input image data to be tested based on the fluorescence images to be tested; wherein the fluorescence images to be tested include fluorescence images to be tested corresponding to multiple base types; a recognition module 22 for using the input image data to be tested as input to a trained base recognition model, and outputting a base recognition result of the input image data to be tested through the trained base recognition model, wherein the trained base recognition model is a model obtained by semi-supervised learning based on a training data set;

[0114] The training data set includes training samples collected in multiple cycles, each training sample includes sample fluorescence images corresponding to multiple base types, and base type label images corresponding to the sample fluorescence images, and the training label corresponding to each training sample also includes a first mask image corresponding to the sample fluorescence image and a second mask image corresponding to the sample fluorescence image, wherein the first mask image is used to mark the position of the base signal acquisition unit with a base type label in the sample fluorescence image; the second mask image is used to mark the position of the base signal acquisition unit without a base type label in the sample fluorescence image.

[0115] Optionally, the identification module 22 is used to:

[0116] Get the training dataset;

[0117] Acquire training samples from the training data set as input training samples, process the input training samples based on different data augmentation methods to obtain multiple groups of processed training samples corresponding to the input training samples, and form multiple groups of input data corresponding to the input training samples based on the multiple groups of processed training samples corresponding to the input training samples;

[0118] Constructing an initial base recognition model, using multiple groups of input data corresponding to the input training samples as inputs to the base recognition model, obtaining base recognition data corresponding to each group of input data, iteratively training the initial base recognition model using the training data set until the loss function converges, and obtaining the trained base recognition model;

[0119] The loss functions include:

[0120] calculating a first loss function of a first loss value between adjusted base identification data corresponding to each set of input data and the base type label map corresponding to the input training sample, wherein the adjusted base identification data corresponding to each set of input data is obtained by adjusting the base identification data corresponding to each set of input data based on the first mask map corresponding to the input training sample;

[0121] and a second loss function for calculating a second loss value between base recognition data corresponding to each two sets of processed input data, wherein the base recognition data corresponding to each set of processed input data is obtained by processing the base recognition data corresponding to each set of input data based on the second mask image corresponding to the input training sample.

[0122] Optionally, the different data enhancement methods include at least one combination of the following: adding different noises to the input training samples, and performing different brightness processing on the input training samples.

[0123] Optionally, the loss value FinalLoss in iterative training is calculated based on the loss function as follows:

[0124] Wherein CE_j represents the first loss value corresponding to the j-th group of input data calculated based on the first loss function, n represents n groups of input data, w(t) represents the weight corresponding to the t-th iteration round, MSE_j represents the j-th second loss value calculated based on the second loss function, and k represents the total number of calculated second loss values.

[0125] Optionally, the process of obtaining the trained base recognition model based on the training data set includes multiple rounds of iterations, wherein in one round of iteration, the base recognition model is trained using the training data set based on multiple iteration numbers, and w(t) increases with the increase in the number of iteration rounds.

[0126] Optionally, the first loss function is:

[0127] Where CE(p,y) is the cross entropy loss function, C is the number of categories, and y i is the one-hot encoding of the true label of the i-th category, p i It is the probability value that the base recognition model predicts that the base signal acquisition unit type is the i-th type.

[0128] Optionally, the second loss function is:

[0129] Where N represents the number of pixels in the base recognition data corresponding to each set of input data, p i and q i are the base recognition data distributions corresponding to each two sets of input data after processing, where p i represents the probability distribution of the base category of the i-th pixel in the base recognition data corresponding to one set of input data in each two groups, q i It represents the probability distribution of the base category of the i-th pixel in the base recognition data corresponding to the other set of input data in every two groups.

[0130] It will be understood by those skilled in the art that the structure of the base recognition device based on semi-supervised learning in Figure 14 does not constitute a limitation on the base recognition device based on semi-supervised learning, and the various modules can be implemented in whole or in part by software, hardware, and a combination thereof. The above modules can be embedded in or independent of the controller in the gene sequencer in the form of hardware, or can be stored in the memory in the gene sequencer in the form of software, so that the controller can call and execute the operations corresponding to the above modules. In other embodiments, the base recognition device based on semi-supervised learning may include more or fewer modules than shown in the figure.

[0131] Referring to FIG. 15 , another embodiment of the present application further provides a gene sequencer 200, comprising a memory 3011 and a processor 3012. The memory 3011 stores a computer program. When the computer program is executed by the processor, the processor 3012 performs the steps of the semi-supervised learning-based base recognition method provided in any of the above embodiments of the present application. The gene sequencer 200 may include a gene sequencer (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a wireless phone, etc.), a wearable device (e.g., a pair of smart glasses or a smart watch), or the like.

[0132] The processor 3012 is the control center, connecting the various components of the entire gene sequencer using various interfaces and circuits. It executes the various functions of the gene sequencer and processes data by running or executing software programs and / or modules stored in the memory 3011 and accessing data stored in the memory 3011. Optionally, the processor 3012 may include one or more processing cores. Preferably, the processor 3012 may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, user interfaces, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor 3012.

[0133] The memory 3011 can be used to store software programs and modules. The processor 3012 executes various functional applications and data processing by running the software programs and modules stored in the memory 3011. The memory 3011 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created based on the use of the gene sequencer, etc. In addition, the memory 3011 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 3011 may also include a memory controller to provide the processor 3012 with access to the memory 3011.

[0134] On the other hand, an embodiment of the present application further provides a storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the base recognition method based on semi-supervised learning provided in any of the above embodiments of the present application.

[0135] Those skilled in the art will appreciate that all or part of the processes in the methods provided in the above embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0136] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. The scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A base recognition method based on semi-supervised learning, characterized in that, Comprising: Obtaining a to-be-tested fluorescence image corresponding to a base signal acquisition unit of multiple base types on a sequencing chip, and forming to-be-tested input image data based on the to-be-tested fluorescence image; Wherein the to-be-tested fluorescence image includes fluorescence images corresponding to multiple base types; Taking the to-be-tested input image data as the input of a trained base recognition model, and outputting a base recognition result of the to-be-tested input image data through the trained base recognition model, where the trained base recognition model is a model obtained by performing semi-supervised learning training based on a training data set; Wherein the training data set includes training samples collected under multiple cycles, each training sample includes a sample fluorescence image corresponding to multiple base types, and a base type label map corresponding to the sample fluorescence image, and the training label corresponding to each training sample further includes a first mask map corresponding to the sample fluorescence image and a second mask map corresponding to the sample fluorescence image, wherein the first mask map is used to mark the positions of the base signal acquisition units with base type labels in the sample fluorescence image; the second mask map is used to mark the positions of the base signal acquisition units without base type labels in the sample fluorescence image.

2. The base recognition method based on semi-supervised learning according to claim 1, wherein, The method further includes: Obtaining a training data set; Obtaining a training sample from the training data set as an input training sample, processing the input training sample based on different data augmentation methods to obtain multiple groups of processed training samples corresponding to the input training sample, and forming multiple groups of input data corresponding to the input training sample based on the multiple groups of processed training samples corresponding to the input training sample; Constructing an initial base recognition model, taking the multiple groups of input data corresponding to the input training sample as the input of the base recognition model respectively to obtain base recognition data corresponding to each group of input data, and performing iterative training on the initial base recognition model through the training data set until the loss function converges to obtain the trained base recognition model; Wherein the loss function includes: A first loss function for calculating a first loss value between the base recognition data corresponding to each group of adjusted input data and the base type label map corresponding to the input training sample, wherein the base recognition data corresponding to each group of adjusted input data is obtained by adjusting the base recognition data corresponding to each group of input data based on the first mask map corresponding to the input training sample; And a second loss function for calculating a second loss value between the base recognition data corresponding to every two groups of processed input data, wherein the base recognition data corresponding to each group of processed input data is obtained by processing the base recognition data corresponding to each group of input data based on the second mask map corresponding to the input training sample.

3. The base recognition method based on semi-supervised learning according to claim 2, characterized in that, The different data augmentation methods include at least one of the following combinations: adding different noises to the input training sample, performing different brightness processing on the input training sample.

4. The base recognition method based on semi-supervised learning according to claim 2, characterized in that , the loss value FinalLoss in iterative training is calculated based on the loss function as follows: Where CE_j represents the first loss value corresponding to the j-th group of input data calculated based on the first loss function, n represents that there are n groups of input data, w(t) represents the weight corresponding to the t-th iteration round, MSE_j represents the j-th second loss value calculated based on the second loss function, and k represents the total number of calculated second loss values.

5. The base recognition method based on semi-supervised learning according to claim 4, wherein , The process of training the trained base recognition model based on the training data set includes multiple rounds of iteration. In one round of iteration, the base recognition model is trained using the training data set based on multiple iteration times, and w(t) increases as the number of iteration rounds increases.

6. The base recognition method based on semi-supervised learning according to claim 2, wherein The first loss function is as follows: where CE(p, y) is the cross-entropy loss function, C is the number of categories, and y i is the one-hot encoding of the true label of the i-th category, and p i is the probability distribution value predicted by the base recognition model that the base signal acquisition unit type is the i-th category.

7. The base recognition method based on semi-supervised learning according to claim 2 or 4, characterized in that, The second loss function is as follows: Where N represents the number of pixels in the base recognition data corresponding to each set of input data, p i and q i are respectively the base recognition data distributions corresponding to every two sets of processed input data, where p i represents the probability distribution of the base category of the i-th pixel in the base recognition data corresponding to one set of input data in every two sets, and q i represents the probability distribution of the base category of the i-th pixel in the base recognition data corresponding to the other set of input data in every two sets.

8. A base recognition device based on semi-supervised learning, characterized in that, Including: An acquisition module, configured to acquire a to-be-tested fluorescence image corresponding to a base signal acquisition unit of multiple base types on a sequencing chip, and form to-be-tested input image data based on the to-be-tested fluorescence image; Where the to-be-tested fluorescence image includes to-be-tested fluorescence images corresponding to multiple base types; An identification module, configured to use the to-be-tested input image data as the input of the trained base recognition model, and output the base recognition result of the to-be-tested input image data through the trained base recognition model. The trained base recognition model is a model obtained by performing semi-supervised learning training based on a training data set; Where the training data set includes training samples collected under multiple cycles. Each training sample includes a sample fluorescence image corresponding to multiple base types and a base type label map corresponding to the sample fluorescence image. The training label corresponding to each training sample further includes a first mask map corresponding to the sample fluorescence image and a second mask map corresponding to the sample fluorescence image. The first mask map is used to mark the positions of the base signal acquisition units with base type labels in the sample fluorescence image; the second mask map is used to mark the positions of the base signal acquisition units without base type labels in the sample fluorescence image.

9. A gene sequencer, characterized in that, Including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 7.

11. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

12. A gene sequencer, characterized in that, Including: An acquisition module and an identification module; The acquisition module is configured to acquire a to-be-tested fluorescence image corresponding to a base signal acquisition unit of multiple base types on a sequencing chip, and form to-be-tested input image data based on the to-be-tested fluorescence image; where the to-be-tested fluorescence image includes to-be-tested fluorescence images corresponding to multiple base types; The identification module is configured to use the to-be-tested input image data as the input of the trained base recognition model, and output the base recognition result of the to-be-tested input image data through the trained base recognition model. The trained base recognition model is a model obtained by performing semi-supervised learning training based on a training data set; Wherein the training data set includes training samples collected under multiple cycles. Each training sample includes a sample fluorescence image corresponding to multiple base types and a base type label map corresponding to the sample fluorescence image. The training label corresponding to each training sample further includes a first mask map corresponding to the sample fluorescence image and a second mask map corresponding to the sample fluorescence image. The first mask map is used to mark the positions of the base signal acquisition units with base type labels in the sample fluorescence image; the second mask map is used to mark the positions of the base signal acquisition units without base type labels in the sample fluorescence image.

Citation Information

Patent Citations

  • Improved variant call procedure using single cell analysis

    CN114766056A

  • Image classification method and system based on meta-learning embedded semi-supervised learning

    CN114821204A

  • Base classification method, gene sequencer and computer readable storage medium

    CN115240189A

  • Multi-enhancement semi-supervised image classification method based on contrast loss

    CN117333709A

  • Base recognition method and device, gene sequencer and storage medium

    CN117523559A