End-to-end gene sequencing method, and gene sequencer and storage medium

Through the end-to-end deep learning super-resolution gene sequencing method, super-resolution image processing is directly performed, solving the problems of image reconstruction time and information redundancy in the prior art, and improving the accuracy and efficiency of sequencing.

WO2025113110A1PCT designated stage expired Publication Date: 2025-06-05SHENZHEN SALUS BIOMED CO LTD

Patent Information

Application Number
PCT/CN2024/129876
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-11-05
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

In the existing gene sequencing technology, the image reconstruction process is time-consuming and not real-time, and depends on the stability of the optical system, and there is information redundancy that affects the sequencing accuracy.

Method used

The end-to-end deep learning super-resolution gene sequencing method is used to directly perform super-resolution image processing from the input end to the output end, and base recognition is achieved through feature extraction networks and base type prediction networks.

Benefits of technology

It improves the accuracy of base recognition and sequencing time, reduces the dependence on the stability of the optical system, and reduces the impact of information redundancy on sequencing accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024129876_05062025_PF_FP_ABST
    Figure CN2024129876_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention are an end-to-end gene sequencing method, and a gene sequencer and a storage medium. The method comprises: acquiring fluorescence images to be tested that correspond to base signal collection units of a plurality of base types on a sequencing chip, wherein the fluorescence images to be tested comprise fluorescence images corresponding to the plurality of base types; on the basis of the fluorescence images to be tested, determining input image data to be tested; and using said input image data as an input for a trained deep learning gene prediction model, and the deep learning gene prediction model outputting a super-resolution feature map by means of a feature extraction network, performing base type recognition by means of a base type prediction network and on the basis of the super-resolution feature map, and outputting a basecall result, wherein the feature extraction network is obtained by means of performing training by using as labels super-resolution feature maps obtained by a trained super-resolution image model.
Need to check novelty before this filing date? Find Prior Art

Description

End-to-end gene sequencing method, gene sequencer, and storage medium

[0001] The present invention claims priority to the Chinese patent application filed with the Patent Office of China on November 28, 2023, with application number 202311597518.3, entitled “End-to-end gene sequencing method and device, gene sequencer and storage medium”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present invention relates to the field of gene technology, and in particular to an end-to-end deep learning-based super-resolution gene sequencing method, a gene sequencer, and a computer-readable storage medium. Background Art

[0003] A sequencer is a widely used sequencing instrument for genome sequencing that can quickly and accurately determine DNA sequences. The entire sequencing process includes acquiring sample images through an optical system, reconstructing the sample images, aligning the gene images, identifying the gene bases (gene basecalling), and obtaining and evaluating the sequencing results. Reconstructing the sample images refers to processing and manipulating the sample images through a reconstruction algorithm after the sequencer acquires them through the optical system to restore some of the information lost in the optical system to ensure the accuracy of the sequencing results. The reconstructed image often includes a super-resolution image output to improve image clarity. The basecalling process based on the super-resolution image can reduce crosstalk between different base types and improve sequencing accuracy. However, the process of reconstructing high-quality images is time-consuming.

[0004] While the reconstructed images obtained during the gene basecalling process have improved image quality, the real-time performance of the image reconstruction cannot meet requirements, and the quality of the image reconstruction is highly dependent on the stability of the optical system. Furthermore, the existing basecalling process based on super-resolution reconstructed images requires image reconstruction based on the image features of the original image, and then base type identification based on the image features of the reconstructed image. This process introduces information redundancy, which can affect the accuracy of gene sequencing.

[0005] Summary of the Invention

[0006] In order to solve the existing technical problems, the embodiments of the present invention provide an end-to-end deep learning-based super-resolution gene sequencing method, device and computer-readable storage medium, which can avoid image reconstruction and realize the super-resolution image-based sequencing process directly from the input end to the output end, thereby improving the accuracy of base recognition and sequencing time.

[0007] In a first aspect, a method for super-resolution gene sequencing based on end-to-end deep learning is provided, comprising:

[0008] Acquire fluorescence images to be measured corresponding to base signal acquisition units of multiple base types on a sequencing chip, wherein the fluorescence images to be measured include fluorescence images corresponding to multiple base types;

[0009] Determining input image data to be measured based on the fluorescent image to be measured;

[0010] The input image data to be tested is used as the input of a trained deep learning gene prediction model. The deep learning gene prediction model outputs a super-resolution feature map through a feature extraction network, and performs base type recognition based on the super-resolution feature map through a base type prediction network, and outputs a base recognition result; wherein the feature extraction network is obtained by training with the super-resolution feature map obtained by the trained super-resolution image model as a label.

[0011] In some embodiments, determining the input image data to be tested based on the fluorescent image to be tested includes:

[0012] Calculating, based on the fluorescence images corresponding to the multiple base types, an average brightness value of the fluorescence image to be measured and a brightness variance value of the fluorescence image to be measured;

[0013] The brightness of each pixel in the fluorescent image to be measured is preprocessed according to the average brightness value of the fluorescent image to be measured and the brightness variance value of the fluorescent image to be measured to obtain input image data to be measured.

[0014] In some embodiments, the calculating, based on the fluorescence images corresponding to the multiple base types, the average brightness value of the fluorescence image to be measured and the brightness variance value of the fluorescence image to be measured comprises:

[0015] Among them, P j i represents the pixel value of the jth pixel in the fluorescence image corresponding to the i-th base type, N represents the total number of pixels, μ represents the average brightness value of the fluorescence image to be measured, and σ represents the brightness variance value of the fluorescence image to be measured.

[0016] In some embodiments, preprocessing the brightness of each pixel in the fluorescent image to be measured according to the average brightness value of the fluorescent image to be measured and the brightness variance value of the fluorescent image to be measured to obtain the preprocessed fluorescent image to be measured includes:

[0017] where x k represents the pixel value of a pixel in the fluorescence image corresponding to the k-th base type, x′ k Represents the pixel value of a pixel in the fluorescence image corresponding to the k-th base type after preprocessing.

[0018] In some embodiments, the input image data to be tested includes multiple fluorescence images to be tested corresponding to multiple base types collected in the same cycle; before performing base type recognition based on the super-resolution feature map by the base type prediction network, the method further includes:

[0019] Registering the super-resolution feature maps corresponding to the multiple fluorescence images to be tested in the same group using a trained registration model;

[0020] Each training sample of the registration model includes feature maps of sample fluorescence images corresponding to multiple base types and registered feature map labels corresponding to the sample fluorescence images.

[0021] In some embodiments, the method further comprises:

[0022] The super-resolution feature maps of the registered images of the fluorescent images to be tested corresponding to the multiple base types collected in the same cycle are used as labels to train the deep learning model to obtain a trained registration model.

[0023] In some embodiments, before inputting the input image data to be tested into the trained deep learning gene prediction model, the process includes:

[0024] An image registration algorithm is used to register the fluorescence images to be tested corresponding to the multiple base types collected in the same cycle.

[0025] In some embodiments, in some embodiments, the method further comprises:

[0026] Obtaining a training sample set; wherein each training sample includes sample fluorescence images corresponding to a plurality of base types, base type labels corresponding to the sample fluorescence images, and super-resolution feature map labels obtained by a trained super-resolution image model for the sample fluorescence images;

[0027] Constructing an initial deep learning gene prediction model, wherein the initial deep learning gene prediction model includes a feature extraction network and a base type prediction network;

[0028] The initial deep learning gene prediction model is iteratively trained using the training sample set until the loss function converges to obtain the trained deep learning gene prediction model; wherein the loss function includes a first loss function that calculates the loss value between the super-resolution feature map output by the feature extraction network and the super-resolution feature map label, and a second loss function that calculates the loss value between the base recognition result output by the base type prediction network and the base type label.

[0029] In some embodiments, before iteratively training the initial deep learning gene prediction model using the training sample set, the method further includes:

[0030] Obtaining a first training sample set; wherein each first training sample includes sample fluorescence images corresponding to a plurality of base types and super-resolution image labels corresponding to each of the sample fluorescence images;

[0031] Constructing an initial neural network model, and training the initial neural network model based on the first training sample set to obtain a trained super-resolution image model;

[0032] Obtain a second training sample set; wherein each second training sample includes sample fluorescence images corresponding to a plurality of base types, and a super-resolution feature map label obtained by a trained super-resolution image model using the sample fluorescence images in the second training sample as input;

[0033] The feature extraction network of the initial deep learning gene prediction model is iteratively trained with the second training sample set as input until the first loss function converges, thereby obtaining a pre-trained feature extraction network.

[0034] In some embodiments, the loss function is:

[0035] S-Loss=λMCE+ηMSE,

[0036] Among them, S-Loss is the loss function, MCE is the second loss function, MSE is the first loss function, λ and η are empirical values, y k is the kth base type, P(y k ) is the probability of being judged as the kth base type, and the value is between 0 and 1.

[0037] Where M is the number of feature maps of the training samples output by the feature extraction network during training, N is the number of pixels in a feature map, and y ij The pixel value of the pixel in the i-th row and j-th column in the feature map of the output training sample of the feature extraction network in training, is the pixel value of the pixel in the i-th row and j-th column in the super-resolution feature map label corresponding to the training sample.

[0038] In some embodiments, determining the input image data to be tested based on the fluorescent image to be tested includes:

[0039] In each cycle, the fluorescent images to be tested corresponding to the four base types of ACGT are grouped together to determine the input image data to be tested.

[0040] In some embodiments, outputting the base call result comprises:

[0041] Output the probability of the base category corresponding to each base signal acquisition unit in the fluorescence image corresponding to each base type.

[0042] In some embodiments, outputting the base call result comprises:

[0043] Output the brightness value at the center position corresponding to each base signal acquisition unit in the fluorescence image corresponding to each base type.

[0044] In a second aspect, a gene sequencer is provided, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the end-to-end deep learning-based super-resolution gene sequencing method provided in an embodiment of the present application.

[0045] In a third aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the processor executes the steps of the end-to-end deep learning-based super-resolution gene sequencing method provided in an embodiment of the present application.

[0046] The above embodiments provide an end-to-end deep learning super-resolution gene sequencing method, a gene tester, and a computer-readable storage medium. In the trained deep learning gene prediction model, the feature extraction network is obtained by training based on the super-resolution feature map obtained by the trained super-resolution image model as a label. The feature extraction network can directly obtain the super-resolution feature map corresponding to the input image data to be tested by taking the input image data to be tested as input. The base type prediction network then performs base recognition based on the super-resolution feature map corresponding to the input image data to be tested obtained by the feature extraction network. In this way, during the gene sequencing process, the input image data to be tested can be directly input into the deep learning gene prediction model. The deep learning gene prediction model does not need to reconstruct the super-resolution image of the input image data to be tested, and obtains a base recognition result equivalent to base type prediction based on the feature extraction result of the super-resolution image, thereby realizing the sequencing process from the input end to the output end, reducing the gene sequencing operation time, and improving the accuracy of base recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] FIG1 is a diagram showing an application environment of an end-to-end deep learning-based super-resolution gene sequencing method according to an embodiment.

[0048] FIG2 is a flowchart of an end-to-end deep learning super-resolution gene sequencing method in one embodiment.

[0049] FIG3 is a schematic diagram of an end-to-end deep learning super-resolution gene sequencing system in one embodiment.

[0050] FIG4 is a schematic diagram of an end-to-end deep learning super-resolution gene sequencing system in another embodiment.

[0051] Figure 5 is a flowchart of training a deep learning gene prediction model in an end-to-end deep learning super-resolution gene sequencing method in one embodiment.

[0052] FIG6 is a schematic diagram of the structure of a gene sequencer in one embodiment. DETAILED DESCRIPTION

[0053] The technical solution of the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the scope of protection of the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0055] In the following description, reference is made to “some embodiments” which describe a subset of all possible embodiments, but it should be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0056] Gene sequencing involves analyzing the base sequence of a DNA fragment—specifically, the arrangement of adenine (A), thymine (T), cytosine (C), and guanine (G)—in order to identify the bases. Fluorescent labeling is currently widely used for gene sequencing. The sequencing optical system uses lasers to excite fluorescent markers on a sequencing chip, which then emit fluorescence. The fluorescence signal is then collected. The four bases bind to different fluorescent markers, producing four distinct fluorescence bands, which are then used to identify the bases.

[0057] Next-generation sequencing technology utilizes different fluorescent molecules with different fluorescence emission wavelengths. When these fluorescent molecules are irradiated by laser light, they emit fluorescent signals of corresponding wavelengths. After laser irradiation, filters are used to selectively filter out non-specific wavelengths of light to obtain fluorescent signals of specific wavelengths. By obtaining and analyzing the fluorescent signals, base types can be identified. The process mainly includes sample preparation, cluster generation, sequencing, and data analysis.

[0058] Sample preparation: The DNA sample to be sequenced is extracted and purified, followed by DNA fragmentation and adapter ligation. In some embodiments, the DNA sample is cut into a large number of smaller DNA fragments, typically using ultrasound or restriction endonucleases. Adapters containing specific sequences are then ligated to both ends of the DNA fragments for subsequent ligation and sequencing reactions.

[0059] Cluster generation: This process involves amplifying DNA fragments to form fixed DNA fragments, allowing for subsequent clustering of bases. In an alternative example, DNA fragments are amplified using methods such as polymerase chain reaction (PCR) or bridge amplification, resulting in millions of copies of each DNA fragment. The amplified DNA fragments are then fixed to a stationary plate, where each DNA fragment forms a separate cluster.

[0060] Sequencing refers to the sequencing read of each base cluster on the sequencing chip. The sequencing is performed by adding a fluorescently labeled dNTP sequencing primer. One end of the dNTP chemical formula is connected to an azide group, which can prevent polymerization during the sequencing chain extension, ensuring that only one base can be extended in one cycle, corresponding to the generation of a sequencing read, that is, sequencing by synthesis.

[0061] When sequencing each base cluster on the sequencing chip flow cell, fluorescently labeled sequencing primers are added during sequencing. Through techniques such as primer recognition and chain extension, the fixed DNA fragments undergo a sequencing reaction. Each DNA fragment will have bases added one by one during the sequencing reaction, and the order of each base is recorded using a fluorescent signal. A gene molecule contains multiple bases. During sequencing, one of these bases is attached to a fluorescent marker, which emits fluorescence when excited by laser light. Different bases produce different light-sensitive signals (e.g., fluorescence signals).

[0062] A camera then captures a fluorescence image of the fluorescence signal generated by the charge-coupled device (CCD) on the test chip. The gene sequencer uses laser light to excite the fluorescent markers on the gene sequencing chip, generating fluorescence and collecting the fluorescence signal. The four bases bind to different fluorescent markers, producing four distinct fluorescence bands. These are also fluorescence images of the four base types.

[0063] The gene sequencer may also include a photographic platform, which may include an operating table and a camera. The sequencing chip can be placed on the operating table, and a fluorescent image can be obtained by photographing the sequencing chip with the camera. There are many fluorescent spots in a fluorescent image, and a fluorescent spot in the fluorescent image represents the fluorescence emitted by a base cluster.

[0064] The imaging mode of the gene sequencer can be a four-channel imaging system or a two-channel imaging system. For the two-channel imaging system, each camera needs to expose twice at the same position of the test chip. For the four-channel imaging system, the camera of each channel shoots once at the same position of the sample, and obtains fluorescent images of four base types respectively. For example, a fluorescent image of the A base type, a fluorescent image representing the A base type, a fluorescent image of the C base type, a fluorescent image of the G base type, and a fluorescent image of the T base type are obtained respectively. Since a filter is used after laser irradiation to selectively filter out non-specific wavelengths of light to obtain a fluorescent signal of a specific wavelength, each base type corresponds to a different fluorescent signal. In the same cycle reaction, the brightness of the same type of base cluster in its corresponding category of base type is much greater than that of other categories of bases. In theory, there will be no duplication of the base clusters that emit light in each channel.

[0065] After the gene sequencer obtains the fluorescent image, it will reconstruct the collected image, perform gene image registration, and gene base identification (gene basecal) to obtain the gene sequence.

[0066] Gene image reconstruction is used to improve the resolution of fluorescence images, thereby enhancing image clarity and reducing crosstalk between samples. Gene image reconstruction includes, but is not limited to, conventional operations such as deconvolution.

[0067] Gene image registration involves correcting the fluorescence images of four base types so that they overlap. This allows the fluorescence brightness of the four channels at the same location to be extracted, facilitating subsequent base identification. Gene image registration includes, but is not limited to, image registration within the same channel and global or local affine registration.

[0068] The gene recognition process uses the registered image to determine whether the base clusters in the image belong to one of the four bases: A, C, G, or T. After gene recognition, the test data is converted from a digital image into sequence information of the four bases A, C, G, and T, which is the DNA sequence result of the sample for subsequent analysis and evaluation.

[0069] Data Analysis: Analyze and interpret sequencing data based on image data and sequence information. Compare sequence information with the reference genome for mutation identification.

[0070] The process of sequencing a single sample is called a run. The sequencing process consists of multiple cycles, each corresponding to a reaction cycle and, in other words, a single base type identification on the sequencing chip. Sequencing occurs simultaneously through synthesis. In a single cycle, tens of millions of base clusters are sequenced simultaneously.

[0071] A test data set consists of many DNA fragments. During the sequencing process, a base is added to each DNA fragment. Therefore, the length of the base sequence of the test data determines the number of cycles. In each cycle, the gene sequencer generates a fluorescence image for each of the four base types (ACGT). When sequencing the test data, the gene sequencer can obtain fluorescence images of the ACGT channel for multiple cycles.

[0072] Refer to Figure 1, which is an application environment diagram of an end-to-end deep learning super-resolution gene sequencing method in one embodiment. The end-to-end deep learning super-resolution gene sequencing method is applied to a gene sequencer, which may also include an operating table and a camera, wherein a sequencing chip can be placed on the operating table, and the gene sequencing chip has a number of base clusters arranged in an array or randomly distributed. Through dyeing reagents, different types of base clusters will be connected to one of the different fluorescent markers during the sequencing reaction. These fluorescent markers will emit fluorescent signals after being irradiated by laser. The fluorescent signals of non-specific wavelengths are selectively filtered out by filters to obtain fluorescent signals of specific wavelengths. The fluorescent molecules in different fluorescent markers have different fluorescence emission wavelengths, so that different base clusters correspond to different fluorescent signals. The fluorescent image is obtained by the camera, and the fluorescent image is analyzed to identify the base category of each base cluster. The camera can be an optical microscope.

[0073] Please refer to Figure 2, which is a flow chart of an end-to-end deep learning super-resolution gene sequencing method provided in one embodiment of the present application. The end-to-end deep learning super-resolution gene sequencing method is applied to a gene sequencer and includes the following steps:

[0074] S11. Acquire fluorescence images to be measured corresponding to base signal acquisition units of multiple base types on a sequencing chip.

[0075] In this embodiment, the base signal acquisition units on the sequencing chip are presented in the form of base clusters. In some other gene sequencing methods, due to different amplification methods, the base signal acquisition units can be presented in the form of nanospheres.

[0076] As shown in Figure 3, the end-to-end deep learning-based super-resolution gene sequencing system includes an image acquisition unit, which includes a staining reagent, a laser, and a microscope. The microscope can be an optical microscope. The staining reagent is used to add a fluorescent label to each base cluster in the sequencing chip. The microscope's objective lens and laser power parameters are initialized, and the laser power is adjusted to ensure that base clusters of various base types emit light uniformly without overexposure. The sequencing chip is then illuminated with a laser, and the fluorescent labels on the base clusters in the sequencing chip generate a fluorescent signal. The sequencing chip can then be photographed using a microscope to obtain a fluorescent image to be measured.

[0077] In this embodiment, the fluorescence image to be measured includes fluorescence images corresponding to multiple base types, for example, a fluorescence image of base type A, a fluorescence image of base type C, a fluorescence image of base type G, and a fluorescence image of base type T. A fluorescent dot in the fluorescence image represents a base cluster. For example, a fluorescence image of base type A contains multiple fluorescent dots representing the A base cluster. A fluorescent dot can be composed of multiple pixels.

[0078] During each cycle, the gene sequencer's camera captures a single image, producing fluorescence images corresponding to multiple base types. For example, if the gene sequencer's imaging system uses a four-channel imaging mode, then a single image capture in one cycle can produce fluorescence images of four base types. During sequencing, one cycle corresponds to one reaction cycle, i.e., one base type identification on the sequencing chip. Fluorescence images captured during one cycle, or multiple cycles, can be used as the fluorescence image to be tested.

[0079] When taking fluorescence images, initialize the objective lens and laser power parameters of the gene sequencer, adjust the laser power so that base clusters of various base types can emit light evenly but not overexposed, and adjust the position of the objective lens until the base clusters in the field of view are clearly visible without blurred edges, so as to obtain a high-quality fluorescence image to be tested.

[0080] S12. Determine input image data to be tested based on the fluorescent image to be tested.

[0081] During gene sequencing, channel crosstalk, fluorophore phasing, fluorophore prephasing, camera accuracy errors when acquiring fluorescence images, and stage movement accuracy can all affect the brightness of fluorescent spots in the fluorescence image. Uneven brightness or high noise levels in the fluorescence image can negatively impact base call accuracy.

[0082] Channel crosstalk is the brightness interference between fluorescence images of different base types. Because the wavelength distributions of fluorescent molecules from different fluorescent markers overlap, light intensity interference can occur between fluorescence images corresponding to different base types. For example, if position B on a sequencing chip contains base cluster A, and position D is adjacent to position B, which contains base cluster C, when the sequencing chip is illuminated by laser, the fluorescence intensity generated by base cluster A at position B may interfere with the fluorescence intensity generated by base cluster C at position D. Consequently, in one cycle, the camera generates four fluorescence images: one corresponding to base type A, one corresponding to base type C, one corresponding to base type G, and one corresponding to base type T. Consequently, the shadow of the fluorescent spot at position B in the fluorescence image corresponding to base type A may appear at the fluorescent spot at position D in the fluorescence image corresponding to base type C.

[0083] Regarding the reaction hysteresis effect of the fluorophore, there is some fluorescence that has not been completely removed in the current cycle due to incomplete fluorescence removal or unclean elution. The fluorescence that has not been completely removed will react in the sequencing reaction of the next cycle, thus interfering with the fluorescence intensity of the fluorescence image collected in the next cycle.

[0084] The fluorophore reaction prephasing effect occurs when a fluorophore reacts in the current cycle, rather than the next. This delay or prephasing reflects the asynchrony and inconsistency of the fluorophore copy reactions, which is the main cause of the base call error rate.

[0085] Therefore, preprocessing the collected fluorescence images can reduce noise interference in the fluorescence images. Preprocessing includes but is not limited to: denoising, brightness adjustment, image background processing, and base channel normalization.

[0086] In some embodiments, determining the input image data to be tested based on the fluorescent image to be tested includes:

[0087] Calculating, based on the fluorescence images corresponding to the multiple base types, an average brightness value of the fluorescence image to be measured and a brightness variance value of the fluorescence image to be measured;

[0088] The brightness of each pixel in the fluorescent image to be measured is preprocessed according to the average brightness value of the fluorescent image to be measured and the brightness variance value of the fluorescent image to be measured to obtain input image data to be measured.

[0089] The step of calculating the average brightness value of the fluorescent image to be measured and the brightness variance value of the fluorescent image to be measured based on the fluorescent images corresponding to the multiple base types includes:

[0090] Among them, P j i represents the pixel value of the jth pixel in the fluorescence image corresponding to the i-th base type, N represents the total number of pixels, μ represents the average brightness value of the fluorescence image to be measured, and σ represents the brightness variance value of the fluorescence image to be measured;

[0091] Preprocessing the brightness of each pixel in the fluorescent image to be measured according to the average brightness value of the fluorescent image to be measured and the brightness variance value of the fluorescent image to be measured to obtain the preprocessed fluorescent image to be measured includes:

[0092] where x k represents the pixel value of a pixel in the fluorescence image corresponding to the k-th base type, x′ k Represents the pixel value of a pixel in the fluorescence image corresponding to the k-th base type after preprocessing.

[0093] In the above embodiment, the fluorescence image to be tested includes fluorescence images of multiple base types. The light intensities of the pixels corresponding to these multiple base types are accumulated to obtain the average brightness of the fluorescence image to be tested. The brightness variance of the fluorescence image to be tested can be calculated based on the light intensities of the pixels in the fluorescence images corresponding to the multiple base types. Each pixel in each fluorescence image corresponding to the multiple base types is processed based on the average brightness of the fluorescence image to be tested and the brightness variance of the fluorescence image to be tested. In the base channel normalization process, when sequencing base clusters of multiple base types on a sequencing chip, the light intensity of each pixel is pre-processed based on the fluorescence intensity in the fluorescence images of the multiple base types. This can reduce brightness interference between base channels and prevent the brightness values ​​corresponding to base clusters at certain locations from being too bright or from decaying too quickly during subsequent processing. When the fluorescence images corresponding to the multiple base types are fluorescence images collected over multiple cycles, the brightness effects caused by the early reaction effect and the delayed reaction of the fluorophore can also be reduced, thereby improving the accuracy of subsequent base recognition.

[0094] For example, during gene sequencing, a fluorescence image of each of the four base types (AC, G, and G) is acquired in a single cycle. The image size is 4 x 4. The average brightness is calculated by summing the light intensities of the pixels in these four fluorescence images. The brightness variance is then calculated based on the light intensities of the pixels in these four fluorescence images. The average brightness and brightness variance are then used to preprocess the 16 pixels in these four fluorescence images.

[0095] S13. Using the input image data to be tested as the input of the trained deep learning gene prediction model, the deep learning gene prediction model outputs a super-resolution feature map through the feature extraction network, and performs base type recognition based on the super-resolution feature map through the base type prediction network, and outputs the base recognition result.

[0096] In this embodiment, the feature extraction network is obtained by training with the super-resolution feature map obtained by the trained super-resolution image model as a label. Therefore, when the input image data to be tested is input into the feature extraction network, the feature extraction network can output the super-resolution feature map of the input image data to be tested.

[0097] Optionally, the base type prediction network performs base type recognition based on the super-resolution feature map, and can perform image registration on the super-resolution feature map of the input image data to be tested by adding a registration model to obtain the registered feature map corresponding to the input image data to be tested, and use the registered feature map corresponding to the input image data to be tested as the input of the base type prediction network, and perform base type recognition on the base clusters in the input image data to be tested through the base type prediction network. In other optional embodiments, before the input image data to be tested is input into the deep learning gene prediction model, a traditional image registration algorithm can be used to register the fluorescence images corresponding to the multiple base types collected in the same cycle, or a registration model obtained after training the deep learning model with the registered images as labels can be used to register the fluorescence images corresponding to the multiple base types collected in the same cycle.

[0098] In this embodiment, for each cycle, the fluorescence image to be tested is the fluorescence image of multiple base types collected in each cycle, and the fluorescence images of multiple base types are input into the deep learning gene prediction model for base recognition, thereby obtaining base recognition results corresponding to the fluorescence images of multiple base types. The deep learning gene prediction model can output the base recognition results in various ways, such as using multi-channel or single-channel output. Multi-channel output means that one channel outputs the recognition result of one base type. For example, the recognition result of channel 1 is the recognition result of the base cluster of the A base type in the current cycle, and the recognition result of channel 2 is the recognition result of the base cluster of the C base type in the current cycle. The single-channel output is the union of the corresponding A base type recognition result, C base type recognition result, G base type recognition result, and T base type recognition result obtained by processing the fluorescence images of multiple base types, forming a recognition result of the current cycle that simultaneously contains the A, C, G, and T base signal acquisition units.

[0099] In one cycle, after the fluorescence images of multiple base types are processed by the deep learning gene prediction model, the deep learning gene prediction model can also output data in multiple ways, including outputting the probability of the base category corresponding to the fluorescence image of each base type, outputting the brightness value at the center position of each base cluster in the fluorescence image of each base type, and so on.

[0100] In the above embodiment, the base recognition result can be obtained by directly inputting the input image data to be tested into the trained deep learning gene prediction model, thereby realizing the sequencing process from the input end to the output end and reducing the gene sequencing running time; and the feature extraction network in the trained deep learning gene prediction model is obtained after training based on the super-resolution feature map obtained by the trained super-resolution image model as a label. The feature extraction network can output the super-resolution feature map corresponding to the input image data to be tested, and the base type prediction network performs base recognition based on the super-resolution feature map corresponding to the input image data to be tested, thereby improving the accuracy of base recognition.

[0101] In some embodiments, the input image data to be tested includes multiple fluorescence images to be tested corresponding to multiple base types collected in the same cycle. Before performing base type recognition based on the super-resolution feature map by the base type prediction network, the method further includes:

[0102] The super-resolution feature maps corresponding to the multiple fluorescence images to be tested in the same group are registered using a trained registration model; wherein each training sample of the registration model includes feature maps of sample fluorescence images corresponding to multiple base types and registered feature map labels corresponding to the sample fluorescence images.

[0103] The super-resolution feature map corresponding to the input image data to be tested is used as the input of the trained registration model. The trained registration model then outputs a registered feature map corresponding to the input image data to be tested, so that the super-resolution feature maps corresponding to the fluorescence images of multiple base types in the same cycle can overlap. This results in a registered feature map corresponding to the fluorescence image to be tested, ensuring that subsequent base calls can be performed.

[0104] Specifically, when training the registration model, a sample fluorescence image can be obtained first, and then the feature map of the sample fluorescence image can be calculated using a traditional image registration algorithm, and the feature map of the sample fluorescence image can be added to the training data set used for the registration model. A training sample is obtained from the training data set, and the training sample is input into the registration model. The supervised learning is performed with the registered feature map label corresponding to the training sample as the training target until the loss function for the registration model converges, thereby obtaining a trained registration model. In each iterative calculation, the registration model under the current iteration is based on the feature map of the training sample under the current iteration, and outputs the registration feature map corresponding to the training sample under the current iteration. Based on the loss function, the feature map loss value between the registration feature map corresponding to the training sample under the current iteration and the registered feature map label corresponding to the training sample is calculated. When the feature map loss value under the current iteration meets the iteration termination condition, the iteration is stopped, and the registration model under the current iteration is used as the trained registration model.

[0105] In the above embodiment, the super-resolution feature map corresponding to the input image data to be tested is used as the input of the trained registration model into the trained registration model, and the registered feature map corresponding to the input image data to be tested can be directly output. There is no need to calculate the registered feature map based on the super-resolution feature map corresponding to the input image data to be tested using the traditional image registration method, thereby reducing the sequencing time.

[0106] As shown in Figure 3, the end-to-end deep learning-based super-resolution gene sequencing system includes an image acquisition unit, which includes a staining reagent, a laser, and a microscope. The microscope can be an optical microscope. The staining reagent is used to add a fluorescent label to each base cluster in the sequencing chip. The microscope objective lens and laser power parameters are initialized, and the laser power is adjusted to ensure that base clusters of various base types can emit light evenly but not overexposed. The sequencing chip is then illuminated with a laser, and the fluorescent labels on the base clusters in the sequencing chip generate a fluorescent signal. The sequencing chip can be photographed using a microscope to obtain a fluorescent image to be tested. Based on the fluorescent image to be tested, the input image data to be tested is obtained.

[0107] The trained deep learning gene prediction model includes a feature extraction network, a trained registration model, and a base type prediction network. The input image data to be tested is input into the feature extraction network, which outputs a super-resolution feature map corresponding to the input image data to be tested. The super-resolution feature map corresponding to the input image data to be tested serves as the input to the trained registration model. The trained registration model outputs a registered feature map corresponding to the input image data to be tested. The registered feature map corresponding to the input image data to be tested serves as the input to the base type prediction network, which outputs the base types of the base clusters in the input image data to be tested.

[0108] The feature extraction network includes an upsampling layer, a first convolutional layer, and an activation layer. Through the upsampling layer, an interpolation method is used to increase the length of the input image data to be measured to a first preset multiple and the width to a second preset multiple, thereby obtaining the interpolated input image data to be measured. The interpolated input image data to be measured is input into the first convolutional layer to perform a convolution operation and extract a feature map of the input image data to be measured. The feature map of the fluorescence image to be measured is input into the activation layer, and a nonlinear mapping is performed using an activation function. Based on the processing of the activation layer, a super-resolution feature map corresponding to the input image data to be measured is obtained. In other embodiments, the feature extraction network may also include normalization, pooling, and fully connected layers. The normalization operation prevents sudden increases or decreases in values ​​during processing. The pooling layer is used to downsample the output of the first convolutional layer, reduce data dimensionality, and reduce model complexity and computational complexity. The fully connected layer can expand the output obtained after processing by the first convolutional layer, activation layer, and pooling layer, and connect the expanded content to the output layer through a fully connected method to obtain a super-resolution feature map corresponding to the input image data to be measured.

[0109] Specifically, when processing a fluorescence image, the upsampling layer first doubles the image's length and width through interpolation. The interpolated image is then fed into the trained first convolutional layer for fine-tuning by the deep learning network. Interpolation methods include, but are not limited to, interpolation and bi-triple interpolation. Increasing the number of pixels and file size through interpolation can eliminate aliasing caused by image magnification, thereby improving the resolution of the fluorescence image and facilitating the subsequent generation of super-resolution feature maps, thereby enhancing base call accuracy.

[0110] The primary function of the first convolutional layer is to extract features from the interpolated input image data. When the convolutional layer processes the interpolated input image data, it slides a convolution kernel through the interpolated input image data to obtain the portion of the interpolated input image data that lies within the convolution kernel. This portion of data is then convolved with the kernel to produce the output of the convolutional layer. In other words, the convolution kernel acts as a feature detector, filtering the interpolated input image data to extract base features from the interpolated input image data.

[0111] The activation layer uses a nonlinear activation function to introduce nonlinear factors for nonlinear mapping, thereby improving the feature expression capability of the super-resolution model. The nonlinear activation function may include the ReLU function.

[0112] The base type prediction network includes a second convolutional layer and a classification network layer. The registered feature map serves as the input to the second convolutional layer, and the output of the second convolutional layer serves as the input to the classification network layer, which then outputs the base call results. The classification network layer can be a Unet-like classification network.

[0113] In some embodiments, it is necessary to obtain samples and train a deep learning gene prediction model based on the samples to obtain a trained deep learning gene prediction model. The trained deep learning gene prediction model includes:

[0114] Obtaining a training sample set; wherein each training sample includes sample fluorescence images corresponding to a plurality of base types, base type labels corresponding to the sample fluorescence images, and super-resolution feature map labels obtained by a trained super-resolution image model for the sample fluorescence images;

[0115] Constructing an initial deep learning gene prediction model, wherein the initial deep learning gene prediction model includes a feature extraction network and a base type prediction network;

[0116] The initial deep learning gene prediction model is iteratively trained using the training sample set until the loss function converges to obtain the trained deep learning gene prediction model; wherein the loss function includes a first loss function that calculates the loss value between the super-resolution feature map output by the feature extraction network and the super-resolution feature map label, and a second loss function that calculates the loss value between the base recognition result output by the base type prediction network and the base type label.

[0117] As shown in FIG5 , FIG5 is a flowchart of training a deep learning gene prediction model in an end-to-end deep learning super-resolution gene sequencing method in one embodiment, specifically including:

[0118] S21. Obtain a training sample set.

[0119] In this embodiment, each training sample includes sample fluorescence images corresponding to multiple base types, base type label images corresponding to the sample fluorescence images, and super-resolution feature map labels obtained by the trained super-resolution image model for the sample fluorescence images. The base type label images are images of the actual base categories corresponding to the sample fluorescence images. The base type label images and super-resolution feature map labels are used to guide the deep learning gene prediction model to continuously learn based on the training samples until the training termination condition is met. The acquisition of the corresponding sample fluorescence images can be similar to step S11. A large number of fluorescence images are acquired using an image acquisition unit as shown in FIG4 . The image acquisition unit includes a staining reagent, a laser, and a microscope. The microscope can be an optical microscope. The staining reagent is used to add a fluorescent label to each base cluster on the sequencing chip. The microscope's objective lens and laser power parameters are initialized, and the laser power is adjusted to ensure uniform illumination of base clusters of multiple base types without overexposure. The sequencing chip is then illuminated with a laser, and the fluorescent labels on the base clusters on the sequencing chip generate a fluorescent signal. The sequencing chip can then be photographed using the microscope to obtain a sample fluorescence image.

[0120] For example, this could be fluorescence images of multiple base types collected over multiple cycles. Alternatively, by diversifying the genetic sample data input to the sequencing chip, fluorescence images of multiple base types can be collected over multiple cycles for different genetic samples. A richer training sample set improves the learning ability of subsequent models.

[0121] For each sample fluorescence image in the training sample, input data of the super-resolution image model is obtained based on each sample fluorescence image, and the input data of each sample fluorescence image is input into the trained super-resolution image model to obtain a super-resolution feature map label corresponding to each sample fluorescence image.

[0122] In some embodiments, before iteratively training the initial deep learning gene prediction model using the training sample set, the method further comprises: obtaining a first training sample set; wherein each first training sample in the first training sample set includes sample fluorescence images corresponding to a plurality of base types and super-resolution image labels corresponding to each of the sample fluorescence images;

[0123] An initial neural network model is constructed, and the initial neural network model is trained based on the first training sample set to obtain a trained super-resolution image model.

[0124] S22. Build an initial deep learning gene prediction model.

[0125] In this embodiment, the initial deep learning gene prediction model includes a feature extraction network and a base type prediction network.

[0126] The feature extraction network includes an upsampling layer, a first convolutional layer, and an activation layer. Through the upsampling layer. In other embodiments, the feature extraction network may also include normalization processing, a pooling layer, a fully connected layer, and the like. The normalization operation avoids a sudden increase or decrease in the value during the processing process, and the pooling layer is used to downsample the output of the first convolutional layer, reduce the data dimension, and reduce the model complexity and computational complexity. The fully connected layer can expand the output obtained after processing the first convolutional layer, the activation layer, and the pooling layer, and connect the expanded content to the output layer in a fully connected manner to obtain the output data of the feature network.

[0127] The base type prediction network includes a second convolutional layer and a classification network layer.

[0128] In some embodiments, the initial deep learning gene prediction model includes a feature extraction network, a trained alignment model, and a base type prediction network. Sample input data is obtained based on the sample fluorescence image, and the sample input data is input into the feature extraction network to obtain a super-resolution feature map corresponding to the sample input data. The super-resolution feature map corresponding to the sample input data serves as the input of the trained alignment model. The trained alignment model outputs a registered feature map corresponding to the sample input data. The registered feature map corresponding to the sample input data serves as the input of the base type prediction network, and the base type prediction network outputs the base types of the base clusters in the sample input data.

[0129] S23. Train the deep learning gene prediction model using the training sample set, and calculate the loss value of the loss function during the training process.

[0130] In some embodiments, channel crosstalk, fluorophore reaction phasing, fluorophore reaction prephasing, camera accuracy errors when acquiring fluorescence images, and the movement accuracy of the operating table may affect the brightness of the fluorescent spots in the fluorescence image. When the brightness of the fluorescence image is uneven or there is a lot of noise, it will affect the deep learning gene prediction model. Therefore, it is necessary to process the sample fluorescence images in the training samples to obtain sample input data, and then train the deep learning gene prediction model based on the sample input data.

[0131] The sample fluorescence images in the training samples are processed to obtain sample input data including:

[0132] Based on the sample fluorescence images corresponding to the multiple base types corresponding to each training sample, the average brightness value of each training sample and the brightness variance value of each training sample are calculated;

[0133] According to the average brightness value and the brightness variance value of each training sample, the brightness of each pixel in the sample fluorescence image corresponding to the multiple base types in each training sample is calculated to obtain the sample input data of each training sample.

[0134] In the above embodiment, each training sample includes fluorescence images of multiple base types. The light intensities of the pixels corresponding to the multiple base types in each training sample are accumulated to obtain the average brightness of each training sample. The brightness variance of each training sample can be calculated based on the light intensities of the pixels in the fluorescence images corresponding to the multiple base types. Each pixel in each fluorescence image corresponding to the multiple base types in each training sample is processed based on the average brightness of each training sample and the brightness variance of each training sample, thereby reducing brightness interference between base channels and preventing the brightness values ​​corresponding to base clusters at certain positions from being too bright or decaying too quickly during subsequent processing.

[0135] In some embodiments, the loss function includes a first loss function that calculates the loss value between the super-resolution feature map output by the feature extraction network and the super-resolution feature map label, and a second loss function that calculates the loss value between the base recognition result output by the base type prediction network and the base type label.

[0136] The loss function is:

[0137] S-Loss=λMCE+ηMSE,

[0138] Among them, S-Loss is the loss function, cross entropy loss (Minimum Classification Error Loss, MCE) is the second loss function, Mean Square Error (MSE) MSE is the first loss function, λ and η are empirical values, yk is the kth base type, P(y k ) is the probability of being judged as the kth base type, and the value is between 0 and 1.

[0139] Where M is the number of feature maps of the training samples output by the feature extraction network in training, N is the number of pixels in a feature map, and yij is the pixel value of the pixel in the i-th row and j-th column of the feature map of the training samples output by the feature extraction network in training. is the pixel value of the pixel in the i-th row and j-th column in the super-resolution feature map label corresponding to the training sample.

[0140] When inputting sample input data into a deep learning gene prediction model for training, the deep learning gene prediction model is iteratively trained based on the training sample set. In each iteration, the sample input data for the current iteration is input into a feature extraction network within the deep learning gene prediction model for the current iteration to obtain a super-resolution feature map of the sample input data for the current iteration. A first loss value is calculated between the super-resolution feature map of the sample input data for the current iteration and the super-resolution feature map label corresponding to the sample input data based on a first loss function. The base type prediction network within the deep learning gene prediction model for the current iteration obtains a base classification result corresponding to the sample input data for the current iteration based on the super-resolution feature map of the sample input data for the current iteration. A second loss value is calculated between the base classification result corresponding to the sample input data for the current iteration and the base type label corresponding to the sample input data for the current iteration based on a second loss function. A loss value corresponding to the sample input data for the current iteration is calculated based on the loss function, the first loss value for the current iteration, and the second loss value for the current iteration. Based on the loss value corresponding to the sample input data for the current iteration, it is determined whether the current iteration meets an iteration termination condition. If the current iteration meets the iteration termination condition, the deep learning gene prediction model for the current iteration is deemed the trained deep learning gene prediction model. If the current iteration does not meet the iteration termination condition, back propagation is performed according to the loss value corresponding to the sample input data under the current iteration to optimize the parameters in the deep learning gene prediction model; and training samples are repeatedly extracted from the training sample set to obtain the input of the deep learning gene prediction model for the next iterative training. The iteration cycle is repeated to continuously optimize the parameters of the deep learning gene prediction model until the iteration termination condition is met and the iterative training is stopped.

[0141] S24: Determine whether the training reaches the training termination condition.

[0142] In this embodiment, the training termination conditions include, but are not limited to, when the loss value is less than a preset error or when the number of iterations exceeds a preset number. When the training termination conditions are met, the training is terminated. If the training termination conditions are not met, the training continues by returning to obtain training samples and continuing to train the deep learning gene prediction model.

[0143] In the above embodiment, the deep learning gene prediction model is trained and learned based on each training sample in the training sample set, and the base type label corresponding to each training sample and the super-resolution feature map label corresponding to each training sample. The deep learning gene prediction model can use the super-resolution feature map label corresponding to each training sample as the training target, thereby outputting the super-resolution feature map, which facilitates the deep learning gene prediction model to accurately perform base recognition in the subsequent training process. During the training process, the deep learning gene prediction model can use the base type label corresponding to each training sample as the training target, thereby outputting accurate base recognition results.

[0144] In some embodiments, when training a deep learning gene prediction model, the feature extraction network in the deep learning gene prediction model may be a pre-trained feature extraction network. Before iteratively training the initial deep learning gene prediction model using the training sample set, the method further includes:

[0145] Obtain a second training sample set; wherein each second training sample includes sample fluorescence images corresponding to a plurality of base types, and a super-resolution feature map label obtained by the trained super-resolution image model using the sample fluorescence images in the second training sample as input;

[0146] The feature extraction network of the initial deep learning gene prediction model is iteratively trained with the second training sample set as input until the first loss function converges, thereby obtaining a pre-trained feature extraction network.

[0147] In the above embodiment, by pre-training the feature extraction network to obtain a pre-trained feature extraction network, when training the deep learning gene prediction model, the pre-trained feature extraction network is used to accelerate the convergence of the deep learning gene prediction model, thereby improving the training speed.

[0148] In some embodiments, as shown in FIG4 , a schematic diagram of training a deep learning gene prediction model is provided. The schematic diagram includes an image acquisition unit and an initial deep learning gene prediction model for training. The image acquisition unit is used to acquire sample fluorescence images of multiple base types in each training sample in the training sample set. In each iteration, sample input data is obtained based on the sample fluorescence image, the sample input data is input into a feature extraction network to obtain a super-resolution feature map corresponding to the sample input data, and the sample input data is input into a trained super-resolution image model to obtain a super-resolution feature map label corresponding to the sample input data. Based on a first loss function, a first loss value is calculated between the super-resolution feature map corresponding to the sample input data and the super-resolution feature map label corresponding to the sample input data. Inputting the super-resolution feature map corresponding to the sample input data into the trained registration model to obtain a registered feature map corresponding to the sample input data, inputting the registered feature map corresponding to the sample input data into the base type prediction network to obtain a base call result corresponding to the sample input data, calculating a second loss value between the base call result corresponding to the sample input data and the base type label corresponding to the sample input data based on a second loss function, calculating a total loss value based on the loss function, the first loss value, and the second loss value, and determining whether the current iteration meets the iteration termination condition based on the total loss value. If the current iteration meets the iteration termination condition, the deep learning gene prediction model in the current iteration is used as the trained deep learning gene prediction model. If the current iteration does not meet the iteration termination condition, backpropagation is performed based on the total loss value corresponding to the sample input data in the current iteration to optimize the parameters of the deep learning gene prediction model; and repeatedly extracting training samples from the training sample set to obtain inputs to the deep learning gene prediction model, and performing the next iterative training. The iterative cycle is repeated to continuously optimize the parameters of the deep learning gene prediction model until the iteration termination condition is met and the iterative training is terminated.

[0149] Referring to FIG6 , another aspect of an embodiment of the present application further provides a gene sequencer, comprising a memory 3011 and a processor 3012. The memory 3011 stores a computer program, and when the computer program is executed by the processor, the processor 3012 executes the steps of the end-to-end deep learning super-resolution gene sequencing method provided in any of the above embodiments of the present application. The gene sequencer may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a wireless phone, etc.), a wearable device (e.g., a pair of smart glasses or a smart watch), or a similar device.

[0150] The processor 3012 is the control center, connecting the various components of the entire computer device using various interfaces and lines. It executes the various functions of the computer device and processes data by running or executing software programs and / or modules stored in the memory 3011 and accessing data stored in the memory 3011. Optionally, the processor 3012 may include one or more processing cores. Preferably, the processor 3012 may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, user interfaces, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor 3012.

[0151] The memory 3011 can be used to store software programs and modules. The processor 3012 executes various functional applications and data processing by running the software programs and modules stored in the memory 3011. The memory 3011 may mainly include a program storage area and a data storage area. The program storage area may store an operating system, at least one application required for a function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 3011 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 3011 may also include a memory controller to provide the processor 3012 with access to the memory 3011.

[0152] On the other hand, an embodiment of the present application further provides a storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the end-to-end deep learning-based super-resolution gene sequencing method provided in any of the above embodiments of the present application.

[0153] Those skilled in the art will appreciate that all or part of the processes in the methods provided in the above embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0154] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. The scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A super-resolution gene sequencing method based on end-to-end deep learning, characterized in that: include: Acquire fluorescent images to be tested corresponding to base signal acquisition units of multiple base types on a sequencing chip, wherein the fluorescent images to be tested include fluorescent images corresponding to multiple base types; Determining input image data to be tested based on the fluorescent image to be tested; The input image data to be tested is used as the input of a trained deep learning gene prediction model, and the deep learning gene prediction model outputs a super-resolution feature map through a feature extraction network, and performs base type recognition based on the super-resolution feature map through a base type prediction network, and outputs a base recognition result; wherein the feature extraction network is obtained after training using the super-resolution feature map obtained by the trained super-resolution image model as a label.

2. The end-to-end deep learning super-resolution gene sequencing method according to claim 1, characterized in that: The determining of the input image data to be tested based on the fluorescent image to be tested comprises: Based on the fluorescence images corresponding to the multiple base types, calculating the average brightness value of the fluorescence image to be tested and the brightness variance value of the fluorescence image to be tested; According to the average brightness value of the fluorescent image to be tested and the brightness variance value of the fluorescent image to be tested, the brightness of each pixel in the fluorescent image to be tested is preprocessed to obtain the input image data to be tested.

3. The end-to-end deep learning super-resolution gene sequencing method according to claim 2, characterized in that: The step of calculating the average brightness value of the fluorescent image to be tested and the brightness variance value of the fluorescent image to be tested based on the fluorescent images corresponding to the multiple base types includes: Where P j i represents the pixel value of the jth pixel in the fluorescence image corresponding to the ith base type, N represents the total number of pixels, μ represents the average brightness value of the fluorescence image to be tested, and σ represents the brightness variance value of the fluorescence image to be tested.

4. The end-to-end deep learning super-resolution gene sequencing method according to claim 3, characterized in that: Preprocessing the brightness of each pixel in the fluorescent image to be tested according to the average brightness value of the fluorescent image to be tested and the brightness variance value of the fluorescent image to be tested to obtain the preprocessed fluorescent image to be tested includes: where x k represents the pixel value of a pixel in the fluorescence image corresponding to the kth base type, x′ k Represents the pixel value of a pixel in the fluorescence image corresponding to the k-th base type after preprocessing.

5. The end-to-end deep learning super-resolution gene sequencing method according to claim 1, characterized in that: The input image data to be tested includes a plurality of fluorescence images to be tested corresponding to a plurality of base types collected in the same cycle; Before performing base type recognition based on the super-resolution feature map by the base type prediction network, the method further includes: Registering the super-resolution feature maps corresponding to the plurality of fluorescence images to be tested in the same group through the trained registration model; Each training sample of the registration model includes feature maps of sample fluorescence images corresponding to multiple base types and registered feature map labels corresponding to the sample fluorescence images.

6. The end-to-end deep learning super-resolution gene sequencing method according to claim 5, characterized in that: The method further comprises: The super-resolution feature map of the registration image of the fluorescence image to be tested corresponding to multiple base types collected in the same cycle is used as a label to train the deep learning model to obtain a trained registration model.

7. The end-to-end deep learning super-resolution gene sequencing method according to claim 1, characterized in that: Before inputting the input image data to be tested into the trained deep learning gene prediction model, the method includes: An image registration algorithm is used to register the fluorescent images to be tested corresponding to the multiple base types collected in the same cycle.

8. The end-to-end deep learning super-resolution gene sequencing method according to claim 1, characterized in that: The method further comprises: Acquire a training sample set; wherein each training sample includes a plurality of sample fluorescence images corresponding to base types, a base type label corresponding to the sample fluorescence image, and a base type label corresponding to the sample fluorescence image. The super-resolution feature map label obtained by the trained super-resolution image model of the optical image; Constructing an initial deep learning gene prediction model, wherein the initial deep learning gene prediction model includes a feature extraction network and a base type prediction network; The initial deep learning gene prediction model is iteratively trained using the training sample set until the loss function converges to obtain the trained deep learning gene prediction model; wherein the loss function includes a first loss function that calculates the loss value between the super-resolution feature map output by the feature extraction network and the super-resolution feature map label, and a second loss function that calculates the loss value between the base recognition result output by the base type prediction network and the base type label.

9. The end-to-end deep learning super-resolution gene sequencing method according to claim 8, characterized in that: Before iteratively training the initial deep learning gene prediction model through the training sample set, the method further includes: Acquire a first training sample set; wherein each first training sample includes sample fluorescence images corresponding to a plurality of base types and super-resolution image labels corresponding to each of the sample fluorescence images; Constructing an initial neural network model, and training the initial neural network model based on the first training sample set to obtain a trained super-resolution image model; Acquire a second training sample set; wherein each second training sample includes sample fluorescence images corresponding to a plurality of base types, and a super-resolution feature map label obtained by a trained super-resolution image model using the sample fluorescence images in the second training sample as input; The feature extraction network of the initial deep learning gene prediction model is iteratively trained with the second training sample set as input until the first loss function converges to obtain a pre-trained feature extraction network.

10. The end-to-end deep learning super-resolution gene sequencing method according to claim 9, characterized in that: The loss function is: S-Loss = λMCE + ηMSE, Among them, S-Loss is the loss function, MCE is the second loss function, MSE is the first loss function, λ and η are empirical values, y k is the kth base type, P(y k ) is the kth The probability of a base type, the value is between 0-1, Where M is the number of feature maps of the training samples output by the feature extraction network during training, N is the number of pixels in a feature map, and y ij is the pixel value of the pixel in the i-th row and j-th column in the feature map of the output training sample of the feature extraction network in training, is the pixel value of the pixel in the i-th row and j-th column in the super-resolution feature map label corresponding to the training sample.

11. The end-to-end deep learning super-resolution gene sequencing method according to claim 1, characterized in that: The step of determining input image data to be tested based on the fluorescent image to be tested comprises: In each cycle, the fluorescent images to be tested corresponding to the four base types of ACGT are grouped as a group to determine the input image data to be tested.

12. The end-to-end deep learning super-resolution gene sequencing method according to claim 1, characterized in that: The output of the base recognition result includes: Output the probability of the base category corresponding to each base signal acquisition unit in the fluorescence image corresponding to each base type.

13. The end-to-end deep learning super-resolution gene sequencing method according to claim 1, characterized in that: The output of the base recognition result includes: Output the brightness value at the center position corresponding to each base signal acquisition unit in the fluorescent image corresponding to each base type.

14. A gene sequencer, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 13.

15. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Super-resolution-based vehicle detection method, device and equipment, and storage medium

    CN112016507A

  • Base classification method, gene sequencer and computer readable storage medium

    CN115240189A

  • Base identification method and device based on multi-task combination, gene sequencer and medium

    CN116994246A

  • End-to-end gene sequencing method and device, gene sequencer and storage medium

    CN117315654A

  • Deep learning based methods and systems for nucleic acid sequencing

    US11580641B1

Cited By

  • Gene chip fluorescence signal analysis method and system based on deep learning

    CN121767987A