Nucleic acid molecule sequencing method and related device

Super-resolution processing of nucleic acid molecular sequencing images using an image processing model solves the problem of fluorescence signal crosstalk when the distance between nucleic acid molecules is smaller than the resolution of the imaging system, thus improving the accuracy of base interpretation.

WO2025137825A9PCT designated stage Publication Date: 2026-05-15SHENZHEN HUADA GENE INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SHENZHEN HUADA GENE INST
Filing Date
2023-12-25
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies, when the distance between nucleic acid molecules is smaller than the resolution of the imaging system, crosstalk occurs in the fluorescence signals of adjacent nucleic acid molecules, leading to a significant decrease in the accuracy of base interpretation.

Method used

Image processing models are used to perform super-resolution processing on sequencing images. By constructing models trained from real images of nucleic acid molecules and simulated nucleic acid images, the optical conditions of a preset optical sequencing system are simulated to reduce the influence of crosstalk on fluorescence signals.

Benefits of technology

It improves the accuracy of base interpretation, especially when the distance between nucleic acid molecules is smaller than the resolution of the imaging system, effectively reduces fluorescence signal crosstalk, and improves the accuracy of sequencing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023141583_15052026_PF_FP_ABST
    Figure CN2023141583_15052026_PF_FP_ABST
Patent Text Reader

Abstract

A nucleic acid molecule sequencing method of the present disclosure, comprising: first acquiring a sequencing image of a target nucleic acid sample; performing image processing on the sequencing image by means of an image processing model, so as to obtain a target image, the image processing model being a model obtained by training constructed nucleic acid molecule true value images and corresponding simulation nucleic acid images; and performing nucleic acid molecule sequencing on the basis of the target image, so as to obtain a sequencing result corresponding to the target nucleic acid sample.
Need to check novelty before this filing date? Find Prior Art

Description

Nucleic acid molecular sequencing methods and related devices Technical Field

[0001] This disclosure relates to the fields of image processing and biomolecular sequencing, and in particular to a nucleic acid molecular sequencing method and related apparatus. Background Technology

[0002] High-throughput sequencing is a technology for sequencing nucleic acid molecules, enabling the parallel determination of a large number of nucleic acid molecules in a single operation. Specifically, high-throughput sequencing immobilizes nucleic acid molecules on an arrayed sequencing chip. Through each round of reactions between the nucleic acid molecules, specific enzymes, and fluorescent probes, different bases emit fluorescent signals of different wavelengths. This process is captured by an imaging system, and by reconstructing and identifying the acquired images, the base sequence can be determined.

[0003] According to relevant technologies, in the process of determining base sequences, reducing the spacing between adjacent nucleic acid molecules and increasing the density of nucleic acid molecules on the chip can effectively increase the number of bases per unit field of view, thereby increasing sequencing throughput and reducing the cost per unit throughput sequencing. However, limited by the optical diffraction limit, when the spacing between nucleic acid molecules is smaller than the resolution of the imaging system, crosstalk will occur between the fluorescence signals of adjacent nucleic acid molecules, which will significantly affect the accuracy of base interpretation. Therefore, how to effectively improve the accuracy of base interpretation during nucleic acid sequencing has become a major problem that urgently needs to be solved in the industry.

[0004] Summary of the Invention

[0005] This disclosure aims to address at least one of the technical problems existing in the prior art. To this end, this disclosure proposes a nucleic acid molecular sequencing method and related apparatus, which can effectively improve the accuracy of base interpretation during the sequencing of nucleic acid molecules.

[0006] A nucleic acid molecular sequencing method according to a first aspect of this disclosure includes:

[0007] Acquire sequencing images of the target nucleic acid sample, wherein the sequencing images are obtained by image acquisition of the target nucleic acid sample using a preset optical sequencing system;

[0008] The sequencing image is processed using an image processing model to obtain a target image. The image processing model is a model trained using a constructed true image of a nucleic acid molecule and a corresponding simulated nucleic acid image. The simulated nucleic acid image is an image obtained by simulating sequencing of a simulated nucleic acid sample corresponding to the true image of the nucleic acid molecule using a simulated optical system. The simulated optical system is an optical system obtained by simulating the preset optical sequencing system.

[0009] Nucleic acid molecular sequencing is performed based on the target image to obtain the sequencing results corresponding to the target nucleic acid sample.

[0010] According to some embodiments of this disclosure, before performing image processing on the sequencing image based on the image processing model to obtain the target image, the method further includes training the image processing model, specifically including:

[0011] Construct the true image of the nucleic acid molecule based on the simulated nucleic acid sample;

[0012] The simulated nucleic acid image is obtained by performing simulated sequencing based on the simulated nucleic acid sample using the simulated optical system.

[0013] The true image of the nucleic acid molecule and the simulated nucleic acid image corresponding to the true image of the nucleic acid molecule are input into the original image processing model, and the image processing model is iteratively trained.

[0014] When the image processing model meets the first predetermined condition during iterative training, the trained image processing model is obtained.

[0015] According to some embodiments of this disclosure, the step of performing simulated sequencing based on the simulated nucleic acid sample using the simulated optical system to obtain the simulated nucleic acid image includes:

[0016] The optical calibration information of the preset optical sequencing system is simulated;

[0017] The nucleic acid distribution information is simulated based on the simulated optical calibration information to obtain the simulated nucleic acid image.

[0018] According to some embodiments of this disclosure, the nucleic acid molecule in the simulated nucleic acid sample includes multiple bases, and each base is labeled with a fluorescent signal;

[0019] The simulation of the optical calibration information of the preset optical sequencing system includes:

[0020] Determine the pixel size of the preset optical sequencing system; wherein, the pixel size is the size of a single pixel in the preset optical sequencing system on the imaging plane;

[0021] Optical imaging analysis is performed based on the fluorescence signal corresponding to the base in the simulated nucleic acid sample to obtain the optical transfer function.

[0022] By integrating the pixel size with the optical transfer function, the simulated optical calibration information is obtained.

[0023] According to some embodiments of this disclosure, the step of performing optical imaging analysis processing based on the fluorescence signal corresponding to the bases in the simulated nucleic acid sample to obtain the optical transfer function includes:

[0024] The simulated nucleic acid sample was scanned to obtain a scanned nucleic acid image;

[0025] Based on the scanned nucleic acid image, a local maximum search is performed to obtain a fluorescence image reflecting each of the fluorescence signals;

[0026] The fluorescence images corresponding to multiple fluorescence signals are subjected to Gaussian fitting and averaging to obtain the point spread function;

[0027] The optical transfer function is obtained by performing a Fourier transform on the point spread function.

[0028] According to some embodiments of this disclosure, integrating the pixel size with the optical transfer function to obtain the simulated optical calibration information includes:

[0029] Noise is extracted from the blank regions between nucleic acid molecules in the simulated nucleic acid sample to obtain simulated noise information;

[0030] The pixel size, the optical transfer function, and the simulated noise information are integrated to obtain the simulated optical calibration information.

[0031] According to some embodiments of this disclosure, the step of simulating the nucleic acid distribution information based on the simulated optical calibration information to obtain the simulated nucleic acid image includes:

[0032] Based on the pixel size, the nucleic acid distribution information is mapped onto the imaging space to obtain a first simulated image of the simulated nucleic acid sample in the imaging space;

[0033] Based on the optical transfer function, the imaging performance of the first simulated image is simulated to obtain the second simulated image;

[0034] Based on the simulated noise information, environmental noise simulation is performed on the second simulated image to obtain the simulated nucleic acid image.

[0035] According to some embodiments of this disclosure, the step of inputting the ground truth image of the nucleic acid molecule and the simulated nucleic acid image corresponding to the ground truth image of the nucleic acid molecule into the original image processing model, and iteratively training the image processing model, includes:

[0036] In each round of iterative training, the simulated nucleic acid image is input into the image processing model for image processing training to obtain the processing result of this round;

[0037] The results of this round of processing are compared with the true image of the nucleic acid molecule to obtain training bias data;

[0038] The weight parameters of the image processing model are updated based on the training bias data.

[0039] According to some embodiments of this disclosure, obtaining the trained image processing model when the image processing model meets a first predetermined condition during iterative training includes:

[0040] When the training bias data reflects that the image processing model converges during iterative training, it is determined that the image processing model meets the first predetermined condition during iterative training, and the trained image processing model is obtained.

[0041] A nucleic acid molecular sequencing apparatus according to a second aspect embodiment of the present disclosure includes:

[0042] The image acquisition module is used to acquire sequencing images of the target nucleic acid sample, wherein the sequencing images are obtained by acquiring images of the target nucleic acid sample using a preset optical sequencing system;

[0043] The image processing module is used to process the sequencing image based on an image processing model to obtain a target image. The image processing model is a model trained using a constructed true image of a nucleic acid molecule and a corresponding simulated nucleic acid image. The simulated nucleic acid image is an image obtained by simulating sequencing of a simulated nucleic acid sample corresponding to the true image of the nucleic acid molecule based on a simulated optical system. The simulated optical system is an optical system obtained by simulating the preset optical sequencing system.

[0044] The sequencing module is used to perform nucleic acid molecular sequencing based on the target image to obtain the sequencing results corresponding to the target nucleic acid sample.

[0045] Thirdly, embodiments of this disclosure provide an electronic device, including: a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the nucleic acid molecular sequencing method as described in any one of the embodiments of the first aspect of this disclosure.

[0046] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing a program that is executed by a processor to implement the nucleic acid molecular sequencing method as described in any one of the embodiments of the first aspect of this disclosure.

[0047] The nucleic acid molecular sequencing method and related apparatus according to the embodiments of this disclosure have at least the following beneficial effects:

[0048] This disclosed nucleic acid molecule sequencing method first acquires a sequencing image of the target nucleic acid sample. The sequencing image is obtained by image acquisition of the target nucleic acid sample using a pre-defined optical sequencing system. Then, the sequencing image is processed using an image processing model to obtain the target image. This image processing model is trained using a constructed ground truth image of the nucleic acid molecule and its corresponding simulated nucleic acid image. The simulated nucleic acid image is obtained by simulating sequencing of a simulated nucleic acid sample corresponding to the ground truth image using a simulated optical system, which is an optical system simulated from the pre-defined optical sequencing system. Finally, nucleic acid molecule sequencing is performed based on the target image to obtain the sequencing result corresponding to the target nucleic acid sample. Because the image processing model is trained using the ground truth image of the nucleic acid molecule and its corresponding simulated nucleic acid image, and the simulated nucleic acid image is obtained by simulating sequencing of a simulated nucleic acid sample corresponding to the ground truth image using a simulated optical system, the image processing model can be used to process the sequencing image, thereby improving the resolution and effectively increasing the accuracy of base identification during nucleic acid molecule sequencing.

[0049] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description

[0050] The above and / or additional aspects and advantages of this disclosure will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0051] Figure 1 is a flowchart of the nucleic acid molecular sequencing method provided in an embodiment of this disclosure;

[0052] Figure 2 is a schematic diagram of the image processing model of the deep residual channel attention neural network structure provided in the embodiments of this disclosure;

[0053] Figure 3 is a schematic diagram of the principle of the nucleic acid molecular sequencing method provided in the embodiments of this disclosure;

[0054] The flowchart in Figure 4 shows the training process of the image processing model before step S102.

[0055] Figure 5 is a flowchart of step S402 in Figure 4;

[0056] Figure 6 is a flowchart of step S501 in Figure 5;

[0057] Figure 7 is a flowchart of step S602 in Figure 6;

[0058] Figures 8(a) to 8(g) are schematic diagrams of obtaining the optical transfer function according to embodiments of the present disclosure;

[0059] Figure 9 is a flowchart of step S603 in Figure 6;

[0060] Figure 10 is a schematic diagram of determining the average value, standard deviation, and maximum average value of the blank region signal provided in the embodiments of this disclosure;

[0061] Figure 11 is a flowchart of step S502 in Figure 5;

[0062] Figures 12(a) to 12(c) are schematic diagrams of obtaining simulated nucleic acid images provided in the embodiments of this disclosure;

[0063] Figure 13 is a flowchart of step S403 in Figure 4;

[0064] Figure 14 shows the sequencing results of a multi-spacing nucleic acid molecular chip using the sequencer provided in Specific Embodiment 1 of this disclosure;

[0065] Figure 15 shows the sequencing results of a multi-spacing nucleic acid molecular chip using a high-sampling optical engine provided in Specific Embodiment 2 of this disclosure;

[0066] Figure 16 is a schematic diagram of the structure of the nucleic acid molecular sequencing device provided in the embodiments of this disclosure;

[0067] Figure 17 is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0068] Embodiments of this disclosure are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this disclosure, and should not be construed as limiting this disclosure.

[0069] In the description of this disclosure, "multiple" means two or more; "greater than," "less than," and "exceeding" are understood to exclude the number itself; "above," "below," and "within" are understood to include the number itself. The use of "first" and "second" is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of technical features indicated, or implicitly indicating the order of the indicated technical features.

[0070] In the description of this disclosure, it is understood that the orientation descriptions, such as up, down, left, right, front, back, etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the accompanying drawings and are only for the convenience of describing this disclosure and simplifying the description, and are not intended to indicate or imply that the device or element referred to has a specific orientation, or is constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this disclosure.

[0071] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0072] In the description of this disclosure, it should be noted that, unless otherwise expressly defined, terms such as "setup," "installation," and "connection" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this disclosure in conjunction with the specific content of the technical solution. Furthermore, the identification of specific steps in the following text does not imply a limitation on the order of steps or execution logic; the execution order and logic between each step should be understood and inferred from the content described in the embodiments.

[0073] High-throughput sequencing is a technology for sequencing nucleic acid molecules, enabling the parallel determination of a large number of nucleic acid molecules in a single operation. Specifically, in high-throughput sequencing, nucleic acid molecules can be linked to a solid surface, with complementary probes carrying fluorescent groups attached to the molecules, allowing for sequential confirmation of the base sequence via fluorescence imaging.

[0074] According to relevant technologies, nucleic acid molecules are immobilized on an arrayed sequencing chip. Through each round of reactions between nucleic acid molecules, specific enzymes, and fluorescent probes, different bases emit fluorescence signals of different wavelengths. This process is captured by an imaging system. Based on this, the acquired images are reconstructed and identified to determine the base sequence. Reducing the spacing between adjacent nucleic acid molecules and increasing the density of nucleic acid molecules on the chip can effectively increase the number of bases per unit field of view, thereby increasing sequencing throughput and reducing the cost per unit throughput. However, limited by the optical diffraction limit, when the spacing between nucleic acid molecules is smaller than the resolution of the imaging system, crosstalk occurs between the fluorescence signals of adjacent nucleic acid molecules, significantly affecting the accuracy of base interpretation. Therefore, how to effectively improve the accuracy of base interpretation during nucleic acid sequencing has become a major challenge that urgently needs to be addressed in the industry.

[0075] This disclosure aims to address at least one of the technical problems existing in the prior art. To this end, this disclosure proposes a nucleic acid molecular sequencing method and related apparatus, which can effectively improve the accuracy of base interpretation during the sequencing of nucleic acid molecules.

[0076] The following explanation is based on the accompanying drawings.

[0077] Referring to Figure 1, the nucleic acid molecular sequencing method provided according to the embodiments of this disclosure may include, but is not limited to, the following steps S101 to S103.

[0078] Step S101: Obtain the sequencing image of the target nucleic acid sample. The sequencing image is an image obtained by acquiring the target nucleic acid sample using a preset optical sequencing system.

[0079] Step S102: The sequencing image is processed based on the image processing model to obtain the target image. The image processing model is a model trained using the constructed true image of nucleic acid molecules and the corresponding simulated nucleic acid image. The simulated nucleic acid image is an image obtained by simulating sequencing of the simulated nucleic acid sample corresponding to the true image of nucleic acid molecules based on the simulated optical system. The simulated optical system is an optical system obtained by simulating a preset optical sequencing system.

[0080] Step S103: Perform nucleic acid molecular sequencing based on the target image to obtain the sequencing results corresponding to the target nucleic acid sample.

[0081] The nucleic acid sequencing method described in steps S101 to S103 of this disclosure first acquires a sequencing image of the target nucleic acid sample. This sequencing image is obtained by acquiring the target nucleic acid sample using a preset optical sequencing system. The sequencing image is then processed using an image processing model to obtain the target image. This image processing model is trained using a constructed true nucleic acid molecule image and its corresponding simulated nucleic acid image. The simulated nucleic acid image is obtained by simulating sequencing of a simulated nucleic acid sample corresponding to the true nucleic acid molecule image using a simulated optical system. The simulated optical system is an optical system obtained by simulating the preset optical sequencing system. Finally, nucleic acid sequencing is performed based on the target image to obtain the sequencing result corresponding to the target nucleic acid sample. Since the image processing model is trained using the true nucleic acid molecule image and its corresponding simulated nucleic acid image, and the simulated nucleic acid image is obtained by simulating sequencing of a simulated nucleic acid sample corresponding to the true nucleic acid molecule image using a simulated optical system, and the simulated optical system is an optical system obtained by simulating the preset optical sequencing system, using the image processing model to process the sequencing image can improve resolution. This effectively improves the accuracy of base identification during nucleic acid sequencing.

[0082] In some embodiments, step S101 involves acquiring a sequencing image of the target nucleic acid sample. The sequencing image is obtained by using a preset optical sequencing system to acquire images of the target nucleic acid sample. It can be explained that, in order to sequence nucleic acid molecules, a sequencing image of the target nucleic acid sample can be acquired first. Here, the target nucleic acid sample refers to the nucleic acid sample that serves as the sequencing target. Acquiring an image of the target nucleic acid sample allows for the acquisition of its sequencing image. It should be understood that image acquisition of the target nucleic acid sample can rely on a preset optical sequencing system, which refers to a nucleic acid image acquisition and sequencing system pre-set with certain optical conditions. It can be pointed out that acquiring images of the target nucleic acid sample using a preset optical sequencing system under different optical conditions will produce different results.

[0083] In step S102 of some embodiments, the sequencing image is processed based on an image processing model to obtain a target image. The image processing model is a model trained using a constructed ground truth image of nucleic acid molecules and the corresponding simulated nucleic acid image. The simulated nucleic acid image is an image obtained by simulating sequencing of a simulated nucleic acid sample corresponding to the ground truth image of nucleic acid molecules using a simulated optical system. The simulated optical system is an optical system obtained by simulating a preset optical sequencing system. It can be emphasized that by fixing nucleic acid molecules on an arrayed sequencing chip, different bases will emit fluorescence signals of different wavelengths through each round of reactions between nucleic acid molecules, specific enzymes, and fluorescent probes. This process is acquired by an imaging system, and based on this, the acquired image is reconstructed and identified to determine the base sequence. It can be noted that both the target nucleic acid sample and the simulated nucleic acid sample can be prepared using the above method. The difference is that the target nucleic acid sample is the nucleic acid sample used as the sequencing target in actual applications, while the simulated nucleic acid sample is used as training data for constructing the image processing model.

[0084] Some implementations can effectively increase the number of bases per unit field of view by reducing the spacing between adjacent nucleic acid molecules and increasing the arrangement density of nucleic acid molecules on the chip, thereby increasing sequencing throughput and reducing the cost per unit throughput sequencing. However, due to the optical diffraction limit, when the spacing between nucleic acid molecules is smaller than the resolution of the imaging system, crosstalk will occur between the fluorescence signals of adjacent nucleic acid molecules, which will significantly affect the accuracy of base interpretation.

[0085] To address this issue, step S102 of this embodiment can perform image processing on the sequencing image based on an image processing model to obtain the target image. It should be noted that image processing of the sequencing image based on the image processing model aims to improve the resolution of the sequencing image through the image processing model, i.e., super-resolution processing. Super-resolution processing improves the resolution of the original image through hardware or software methods; the process of obtaining a high-resolution image from a series of low-resolution images is called super-resolution reconstruction.

[0086] It is clear that the image processing model is trained using constructed ground truth images of nucleic acid molecules and corresponding simulated nucleic acid images. The simulated nucleic acid images are obtained by simulating sequencing of simulated nucleic acid samples corresponding to the ground truth images using a simulated optical system. The simulated optical system is an optical system simulated from a pre-defined optical sequencing system. It should be understood that images obtained by simulating sequencing of simulated nucleic acid samples corresponding to the ground truth images using a simulated optical system can realistically simulate the actual acquisition of sequencing images. Therefore, using ground truth images of nucleic acid molecules and corresponding simulated nucleic acid images to train the image processing model can improve the model's super-resolution processing capabilities. After image processing of the sequencing images based on the image processing model, a higher resolution target image can be obtained, facilitating subsequent nucleic acid molecule sequencing based on the target image. In this way, even when the spacing between nucleic acid molecules is smaller than the resolution of the imaging system, the crosstalk caused by the fluorescence signals of adjacent nucleic acid molecules can be reduced, effectively improving the accuracy of base identification during nucleic acid molecule sequencing.

[0087] Referring to Figure 2, in some specific embodiments, the image processing model can be a deep residual channel attention neural network structure. Based on this model, the sequencing image is processed to obtain the target image. Specifically, shallow feature extraction is first performed on the sequencing image to obtain sequencing image features. These features are then sequentially processed through multiple RG layers and Conv layers before being output. This output, along with the initial sequencing image features, is input into the reconstruction module to generate the target image. Here, RG represents the residual shallow feature extraction module, FCAB represents the channel attention module, Conv represents the convolutional layer, ReLU represents the linear rectified activation layer, FFT represents the Fourier transform layer, and Sigmoid represents the sigmoid activation layer.

[0088] In step S103 of some embodiments, nucleic acid molecule sequencing is performed based on the target image to obtain the sequencing results corresponding to the target nucleic acid sample. It can be noted that, due to image processing of the sequencing image using an image processing model, the resolution of nucleic acid molecules in the target image is improved compared to the resolution of nucleic acid molecules in the sequencing image. Based on this, performing nucleic acid molecule sequencing based on the target image helps reduce crosstalk caused by fluorescence signals from adjacent nucleic acid molecules when the spacing between nucleic acid molecules is smaller than the resolution of the imaging system. This effectively improves the accuracy of base identification during nucleic acid molecule sequencing.

[0089] In some embodiments, gene sequencers can be used for nucleic acid molecular sequencing. It should be noted that a gene sequencer, also known as a DNA sequencer, is an instrument used to determine the base sequence, type, and quantification of DNA fragments. Its main applications include human genome sequencing, gene diagnosis of human genetic diseases, infectious diseases, and cancer, forensic paternity testing and individual identification, screening of bioengineering drugs, and animal and plant hybridization breeding. Currently, the working principle of DNA sequencers is mainly based on the dideoxy chain termination method or chemical degradation method. Although these two methods differ in principle, both rely on the extension of nucleotide chains starting at fixed sites and randomly terminating at a specific base, producing a series of nucleotide chains of four different lengths ending in A, T, C, and G. These fragments are then separated and detected by electrophoresis on a denaturing polyacrylamide gel to obtain the DNA sequence. Because the dideoxy chain termination method is simpler and more suitable for optical automated detection, it is widely used in fully automated DNA sequencers solely for determining DNA sequences. The chemical degradation method, on the other hand, has important applications in studying the secondary structure of DNA and protein-DNA interactions. It should be understood that there are many ways to perform nucleic acid molecular sequencing based on target images, including, but not limited to, the specific embodiments mentioned above.

[0090] Refer to Figure 3 for a schematic diagram of the nucleic acid molecular sequencing method. A simulated nucleic acid sample corresponding to the ground truth image of the nucleic acid molecule is used to obtain a simulated nucleic acid image after simulated sequencing. Then, an image processing model is trained based on the ground truth image and the corresponding simulated nucleic acid image. During nucleic acid molecular sequencing, the sequencing image of the target nucleic acid sample is first acquired. Then, the image processing model is used to process the sequencing image to obtain the target image. Nucleic acid molecular sequencing is then performed based on the target image to obtain the sequencing results corresponding to the target nucleic acid sample. Using the image processing model to process the sequencing image can improve resolution, thus effectively improving the accuracy of base identification during nucleic acid molecular sequencing.

[0091] Since steps S101 to S103 of this embodiment have been described in detail above, the steps that may precede step S102 will be described in detail below.

[0092] The following describes an embodiment of the training of the image processing model of this disclosure before step S102.

[0093] Referring to FIG4, according to some embodiments of the present disclosure, before performing image processing on the sequencing image based on the image processing model to obtain the target image in step S102, the image processing model may be trained, specifically including, but not limited to, the following steps S401 to S404.

[0094] Step S401: Construct a true image of nucleic acid molecules based on simulated nucleic acid samples;

[0095] Step S402: Simulated sequencing is performed on simulated nucleic acid samples using a simulated optical system to obtain simulated nucleic acid images;

[0096] Step S403: Input the true image of the nucleic acid molecule and the simulated nucleic acid image corresponding to the true image of the nucleic acid molecule into the original image processing model, and perform iterative training on the image processing model;

[0097] Step S404: When the image processing model meets the first predetermined condition during iterative training, the trained image processing model is obtained.

[0098] In some embodiments, step S401 involves constructing a true image of the nucleic acid molecule based on a simulated nucleic acid sample. It is important to emphasize that by immobilizing nucleic acid molecules on an arrayed sequencing chip, different bases will emit fluorescence signals of different wavelengths through each round of reactions between the nucleic acid molecules, specific enzymes, and fluorescent probes. This process is captured by an imaging system, and the acquired image is reconstructed and identified to determine the base sequence. It can be noted that the simulated nucleic acid sample can be prepared in the above manner and used as training data for constructing an image processing model. It can also be noted that the base sequence of the simulated nucleic acid sample is known; therefore, a true image of the nucleic acid molecule can be constructed based on the known base sequence of the simulated nucleic acid sample.

[0099] According to some specific embodiments, each nucleic acid molecule in the simulated nucleic acid sample is regularly loaded into a sequencing chip for arrangement. The arrangement unit of nucleic acid molecules is called a block. Each imaging field of view contains multiple blocks, and the spacing between adjacent blocks is called a tracing line. It can be noted that information such as the number of blocks, their arrangement order, and the spacing between nucleic acid molecules in each imaging field of view corresponding to the simulated nucleic acid sample is recorded in a mask file. Since the mask file contains the known nucleic acid arrangement in the simulated nucleic acid sample, a true image of the nucleic acid molecules can be constructed based on the mask file. A set of simulations can contain multiple mask files describing different imaging fields.

[0100] According to some more specific embodiments, each nucleic acid molecule consists of a single base sequence, which can be randomly generated or derived from standard genomic libraries such as E. coli or human genomes. In constructing a true-value image of the nucleic acid molecule based on a mask file, multiple imaging rounds can be performed, sequentially imaging one base at a time. After several rounds of imaging in each imaging field, the bases of the nucleic acid molecule in each imaging field are clearly presented in the true-value image of the nucleic acid molecule obtained from simulated sequencing. Since different bases emit fluorescence signals of different wavelengths, the brightness of the fluorescence signal corresponding to each base can be represented by a set of normalized four-dimensional vectors [i_a, i_c, i_g, i_t]. When the base type is one of A, C, G, or T, the brightness at the corresponding position is set to a certain range of values; if there is no base at a position, all four elements of the vector are 0. In this way, the true value of the simulated nucleic acid sample can be clearly presented in the true-value image of the nucleic acid molecule.

[0101] In step S402 of some embodiments, simulated sequencing is performed based on simulated nucleic acid samples using a simulated optical system to obtain simulated nucleic acid images. It should be noted that the simulated optical system is an optical system obtained by simulating a preset optical sequencing system. Therefore, by simulating the optical conditions of the preset optical sequencing system, simulated sequencing can be performed based on simulated nucleic acid samples to obtain simulated nucleic acid images.

[0102] In the execution of the nucleic acid molecular sequencing method disclosed herein, the sequencing image is obtained by acquiring images of the target nucleic acid sample using a pre-set optical sequencing system. Therefore, in order to enable the trained image processing model to adapt to super-resolution processing of sequencing images, during the training process of the image processing model, simulated sequencing can be performed based on simulated nucleic acid samples using a simulated optical system to obtain simulated nucleic acid images.

[0103] Referring to FIG5, according to some embodiments of the present disclosure, step S402 performs simulated sequencing based on a simulated nucleic acid sample using a simulated optical system to obtain a simulated nucleic acid image, which may include, but is not limited to, steps S501 to S502 below.

[0104] Step S501: Simulate the optical calibration information of the preset optical sequencing system;

[0105] Step S502: Simulate the nucleic acid distribution information based on the simulated optical calibration information to obtain a simulated nucleic acid image.

[0106] In some embodiments, step S501 involves simulating the optical calibration information of a preset optical sequencing system. It can be explained that since the simulated optical system is an optical system obtained by simulating the preset optical sequencing system, by simulating the optical conditions of the preset optical sequencing system, simulated sequencing can be performed based on simulated nucleic acid samples to obtain simulated nucleic acid images. Specifically, simulating the optical conditions of the preset optical sequencing system can be achieved by simulating the optical calibration information of the preset optical sequencing system. The so-called optical calibration information refers to the optical physical quantity information that the simulated optical system can obtain to simulate the image acquisition process of the preset optical sequencing system.

[0107] In some embodiments, since the preset optical sequencing system is a nucleic acid image acquisition and sequencing system with certain optical conditions set in advance, the optical calibration information of the preset optical sequencing system can be simulated by querying its preset parameters.

[0108] Referring to Figure 6, according to some specific embodiments of this disclosure, the nucleic acid molecule in the simulated nucleic acid sample includes multiple bases, each base being labeled with a fluorescent signal. Step S501, which simulates the optical calibration information of a preset optical sequencing system, may include, but is not limited to, steps S601 to S603 described below.

[0109] Step S601: Determine the pixel size of the preset optical sequencing system; wherein, the pixel size is the size of a single pixel in the preset optical sequencing system on the imaging plane;

[0110] Step S602: Optical imaging analysis is performed based on the fluorescence signals corresponding to the bases in the simulated nucleic acid sample to obtain the optical transfer function;

[0111] Step S603: Integrate the pixel size with the optical transfer function to obtain simulated optical calibration information.

[0112] In step S601 of some embodiments, the pixel size of a preset optical sequencing system is determined; wherein, the pixel size is the size of a single pixel in the preset optical sequencing system corresponding to the imaging plane. It can be explained that the preset optical sequencing system acquires images through a camera, and the pixel size represents the size of a single pixel in the camera corresponding to the imaging plane. For example, if the physical size of a camera pixel is 2 μm and the magnification of the preset optical sequencing system is 20x, then the pixel size corresponding to the preset optical sequencing system is 0.1 μm. It can be noted that determining the pixel size of the preset optical sequencing system aims to establish a mapping from the simulated nucleic acid sample to the camera imaging space.

[0113] In some more specific embodiments, determining the pixel size of a preset optical sequencing system can be achieved by using a high-precision stage to move a sequencing chip loaded with a simulated nucleic acid sample a considerable distance (e.g., 100 μm). During this process, the number of displacement pixels at corresponding positions is measured from the image. The system pixel size is then calculated by dividing the distance moved by the number of displacement pixels. It should be understood that there are various ways to determine the pixel size of a preset optical sequencing system, including, but not limited to, the specific embodiments described above.

[0114] In step S602 of some embodiments, optical imaging analysis is performed based on the fluorescence signals corresponding to the bases in the simulated nucleic acid sample to obtain the optical transfer function. It can be explained that the optical transfer function is used to model the imaging performance of the optical system, and the modeled imaging performance can be used to generate simulated nucleic acid images. By performing optical imaging analysis based on the fluorescence signals corresponding to the bases in the simulated nucleic acid sample, that is, modeling the imaging performance of the optical system, the corresponding optical transfer function can be obtained.

[0115] Referring to FIG7, according to some more specific embodiments of the present disclosure, step S602 performs optical imaging analysis processing based on the fluorescence signal corresponding to the bases in the simulated nucleic acid sample to obtain the optical transfer function, which may include, but is not limited to, steps S701 to S704 below.

[0116] Step S701: Scan the simulated nucleic acid sample to obtain a scanned nucleic acid image;

[0117] Step S702: Perform local maximum search based on the scanned nucleic acid image to obtain a fluorescence image reflecting each fluorescence signal;

[0118] Step S703: Gaussian fitting and averaging are performed on the fluorescence images corresponding to multiple fluorescence signals to obtain the point spread function;

[0119] Step S704: Perform a Fourier transform on the point spread function to obtain the optical transfer function.

[0120] In step S701 of some embodiments, the simulated nucleic acid sample is scanned to obtain a scanned nucleic acid image. It should be noted that scanning imaging refers to the process of capturing a physical object to accurately represent its geometry in a digital environment. Since the optical transfer function is used to model the imaging performance of an optical system, some embodiments can scan and image regions with sparse dots in the simulated nucleic acid sample. This allows for more efficient acquisition of scanned nucleic acid images from sparsely dotted simulated nucleic acid samples. It should be understood that sparsely dotted regions in a simulated nucleic acid sample refer to regions where the spacing between nucleic acid molecules is much greater than the optical resolution.

[0121] In steps S702 to S704 of some embodiments, a local maximum search is first performed based on the scanned nucleic acid image to obtain a fluorescence image reflecting each fluorescence signal; then, the fluorescence images corresponding to multiple fluorescence signals are subjected to Gaussian fitting and averaging to obtain a point spread function; and a Fourier transform is performed on the point spread function to obtain the optical transfer function. It can be noted that the scanned nucleic acid image can also be filtered using a Gaussian difference function, and then, through local maximum search, the individual fluorescent microspheres corresponding to each base can be extracted to form a fluorescence image containing multiple fluorescent microspheres. After Gaussian fitting and averaging, the point spread function is obtained, and then a Fourier transform is performed to obtain the system's optical transfer function.

[0122] According to the embodiments of this disclosure shown in steps S701 to S704, a simulated nucleic acid sample is first scanned to obtain a scanned nucleic acid image; then, a local maximum search is performed based on the scanned nucleic acid image to obtain a fluorescence image reflecting each fluorescence signal; the fluorescence images corresponding to multiple fluorescence signals are Gaussian fitted and averaged to obtain a point spread function; and then a Fourier transform is performed on the point spread function to obtain an optical transfer function. The optical transfer function obtained in this way can more accurately model the imaging performance of the optical system.

[0123] Some embodiments of obtaining optical transfer functions according to this disclosure are shown with reference to Figures 8(a) through 8(g). It can be emphasized that the nucleic acid molecules in the simulated nucleic acid samples comprise multiple bases, each of which is labeled with a fluorescent signal.

[0124] In Figure 8(a), a scanned nucleic acid image is obtained by scanning the simulated nucleic acid sample. The scanned nucleic acid image shown in Figure 8(a) contains multiple fluorescent microspheres discretely distributed on the sequencing chip loaded with the simulated nucleic acid sample.

[0125] In Figure 8(b), the region where the fluorescent microspheres are located is extracted based on the scanned nucleic acid image;

[0126] In Figure 8(c), a filtering algorithm (e.g., local maximum search) is used to extract sparse, discrete individual fluorescent microsphere regions that meet the requirements, and a fluorescence image reflecting each fluorescence signal is obtained.

[0127] Figure 8(d) shows one of the fluorescence images of a single fluorescent microsphere, and Figure 8(e) shows another fluorescence image of a single fluorescent microsphere;

[0128] The image of a single fluorescent sphere is fitted with a Gaussian function to obtain the image center. Multiple fluorescent sphere images are then aligned and averaged, which means that the fluorescent images corresponding to multiple fluorescent signals are averaged by Gaussian fitting to obtain the point spread function.

[0129] Figures 8(f) and 8(g) are schematic diagrams of obtaining the optical transfer function by performing a Fourier transform on the point spread function.

[0130] In step S603 of some embodiments, the pixel size and optical transfer function are integrated to obtain simulated optical calibration information. It can be explained that the pixel size is used to establish the mapping from the simulated nucleic acid sample to the camera imaging space, and the optical transfer function is used to model the imaging performance of the optical system. By integrating the pixel size and the optical transfer function, the simulated optical system can be used to simulate the image acquisition process of a preset optical sequencing system, thereby determining the possible simulated optical calibration information.

[0131] Referring to FIG9, according to some embodiments of the present disclosure, step S603 integrates the pixel size with the optical transfer function to obtain simulated optical calibration information, which may include, but is not limited to, steps S901 to S902.

[0132] Step S901: Extract noise from the blank areas between nucleic acid molecules in the simulated nucleic acid sample to obtain simulated noise information;

[0133] Step S902: Integrate pixel size, optical transfer function and simulation noise information to obtain simulated optical calibration information.

[0134] Through the embodiments shown in steps S901 to S902, noise is extracted from the blank areas between nucleic acid molecules in the simulated nucleic acid sample to obtain simulated noise information. Then, the pixel size, optical transfer function, and simulated noise information are integrated to obtain simulated optical calibration information. It can be noted that the blank areas between nucleic acid molecules in the simulated nucleic acid sample do not contain sequenceable nucleic acids; therefore, these blank areas are suitable for noise extraction to obtain simulated noise information. This disclosure aims to incorporate optical noise during image acquisition by a pre-defined optical sequencing system into the simulation modeling considerations, improving the accuracy of simulating nucleic acid distribution information using optical calibration information, and obtaining higher-quality simulated nucleic acid images.

[0135] In some specific embodiments, optical noise can be extracted manually or algorithmically. Manual extraction uses image processing software to read the average signal value (bg_ave), standard deviation (bg_std), and average maximum value (sig) of the nucleic acid molecule region in the blank area of ​​the tracking line. Algorithmic extraction automatically locates the tracking line and blank area and extracts the corresponding signals. During noise simulation, normalized Gaussian noise, independent for each pixel, can be generated first. The average value of the Gaussian noise is bg_ave / sig, and the standard deviation is bg_std / sig. Then, the Gaussian noise is added to the simulated signal, and the pixel value of the superimposed shot noise is calculated using the Poisson distribution function.

[0136] Referring to the embodiment shown in Figure 10, a more specific embodiment is provided for determining the average signal value (bg_ave), standard deviation (bg_std), and average maximum value (sig) of the nucleic acid molecule region signal in this disclosure. In Figure 10, the white box represents the background area of ​​the tracing line, and the average signal value and standard deviation of this area are measured as bg_ave and bg_std. The white cross indicates the bright spots as nucleic acid molecule signals, and the brightness of the center positions of 10 bright spots is measured and the average value is taken as sig.

[0137] In the embodiments of this disclosure shown in steps S601 to S603, the pixel size of a preset optical sequencing system is first determined; wherein, the pixel size is the size of a single pixel in the preset optical sequencing system corresponding to the imaging plane; optical imaging analysis processing is performed based on the fluorescence signal corresponding to the bases in the simulated nucleic acid sample to obtain the optical transfer function; then, the pixel size and the optical transfer function are integrated to obtain optical calibration information. Furthermore, the pixel size in the optical calibration information can be used to establish a mapping from the simulated nucleic acid sample to the camera imaging space, and the optical transfer function in the optical calibration information can be used to model the imaging performance of the optical system. It should be understood that the optical calibration information obtained in this way helps to more accurately simulate the nucleic acid distribution information and obtain high-quality simulated nucleic acid images.

[0138] In some embodiments, step S502 involves simulating nucleic acid distribution information based on simulated optical calibration information to obtain a simulated nucleic acid image. It can be explained that after simulating the optical calibration information of a preset optical sequencing system, the nucleic acid distribution information can be simulated based on the simulated optical calibration information to obtain a simulated nucleic acid image. Since the simulated optical calibration information is specifically obtained by simulating the preset optical sequencing system, the simulated nucleic acid image obtained can more realistically simulate the actual acquisition of sequencing images.

[0139] Referring to FIG11, according to some embodiments of the present disclosure, step S502 simulates the nucleic acid distribution information based on simulated optical calibration information to obtain a simulated nucleic acid image, which may include, but is not limited to, steps S1101 to S1103 below.

[0140] Step S1101: Based on the pixel size, the nucleic acid distribution information is mapped to the imaging space to obtain the first simulated image of the simulated nucleic acid sample in the imaging space;

[0141] Step S1102: Simulate the imaging performance of the first simulated image based on the optical transfer function to obtain the second simulated image;

[0142] Step S1103: Based on the simulation noise information, perform environmental noise simulation on the second simulation image to obtain a simulated nucleic acid image.

[0143] In some embodiments, steps S1101 to S1103 involve first mapping nucleic acid distribution information to the imaging space based on pixel size to obtain a first simulated image of the simulated nucleic acid sample in the imaging space; then, performing imaging performance simulation on the first simulated image based on the optical transfer function to obtain a second simulated image; and finally, performing environmental noise simulation on the second simulated image based on simulated noise information to obtain a simulated nucleic acid image. It can be explained that pixel size is used to establish the mapping from the simulated nucleic acid sample to the camera imaging space. Therefore, based on pixel size, nucleic acid distribution information can be mapped to the imaging space to obtain a first simulated image of the simulated nucleic acid sample in the imaging space. Since the optical transfer function is used to model the imaging performance of the optical system, imaging performance simulation can be performed on the first simulated image based on the optical transfer function to obtain a second simulated image. Furthermore, since the simulated noise information is the optical noise during image acquisition by the preset optical sequencing system, environmental noise simulation can be performed on the second simulated image based on the simulated noise information, ultimately obtaining a simulated nucleic acid image. It should be understood that, based on the fact that optical calibration information includes pixel size, optical transfer function, and simulation noise information, optical calibration information helps to more accurately simulate the preset optical sequencing system. In this way, the simulated nucleic acid image obtained can more realistically simulate the actual acquisition of sequencing images.

[0144] Referring to some specific embodiments shown in Figures 12(a) to 12(c), during simulation, the distribution of nucleic acid molecules in each imaging field of view can be determined first, and a simulated list of nucleic acid molecules can be generated, as shown in Figure 12(a). The list of nucleic acid molecules is then mapped to a two-color or four-color imaging space, as shown in Figure 12(b), which shows the mapping of the list of nucleic acid molecules to a four-color imaging space. The mapped image is then low-pass filtered using the measured optical transfer function, and finally, simulated noise is added to obtain a low-resolution sequencing image that is close to that during nucleic acid molecule sequencing, i.e., a simulated nucleic acid image, as shown in Figure 12(c).

[0145] In some more specific embodiments, during multiple rounds of simulation, the actual signal and noise levels are extracted based on the actual images captured by the system, which can also simulate the signal-to-noise ratio decline in long-term sequencing.

[0146] In the embodiments of this disclosure shown in steps S501 to S502, the optical calibration information of a preset optical sequencing system is first simulated, and then the nucleic acid distribution information is simulated based on the simulated optical calibration information to obtain a simulated nucleic acid image. Using the optical calibration information obtained by simulating the preset optical sequencing system for the simulation of the nucleic acid image helps to improve the simulation effect, enabling the simulated nucleic acid image to more realistically simulate the actual acquisition of sequencing images.

[0147] In step S403 of some embodiments, the ground truth image of the nucleic acid molecule and the corresponding simulated nucleic acid image are input into the original image processing model, and the image processing model is iteratively trained. It can be explained that the simulated nucleic acid image is an image obtained based on a simulated nucleic acid sample corresponding to the ground truth image of the nucleic acid molecule. It can realistically simulate the acquisition of sequencing images. Inputting the simulated nucleic acid image into the original image processing model for training helps the trained image processing model adapt to super-resolution processing of sequencing images. The ground truth image of the nucleic acid molecule can reflect the true and accurate arrangement of the base sequence in the simulated nucleic acid sample. Inputting the simulated nucleic acid image into the original image processing model for training can improve the super-resolution processing capability of the image processing model. Therefore, inputting the ground truth image of the nucleic acid molecule and the corresponding simulated nucleic acid image into the original image processing model for iterative training can continuously improve the super-resolution processing capability of the image processing model during iterative training, and enable the image processing model to adapt to super-resolution processing of sequencing images.

[0148] Referring to Figure 13, according to some embodiments of this disclosure, step S403, which involves inputting the true image of the nucleic acid molecule and the simulated nucleic acid image corresponding to the true image of the nucleic acid molecule into the original image processing model and iteratively training the image processing model, may include, but is not limited to, the following steps S1301 to S1303:

[0149] Step S1301: In each round of iterative training, the simulated nucleic acid image is input into the image processing model for image processing training to obtain the processing result of this round;

[0150] Step S1302: Compare the results of this round of processing with the true image of nucleic acid molecules to obtain training bias data;

[0151] Step S1303: Update the weight parameters of the image processing model based on the training bias data.

[0152] In step S1301 of some embodiments, in each round of iterative training, a simulated nucleic acid image is input into the image processing model for image processing training to obtain the processing result for that round. It can be explained that since the simulated nucleic acid image is an image obtained by simulating a simulated nucleic acid sample corresponding to the true image of the nucleic acid molecule, it can realistically simulate the acquisition of sequencing images. Furthermore, since the multiple rounds of iterative training aim to optimize the super-resolution processing capability of the image processing model, the simulated nucleic acid image can be input into the image processing model for image processing training in each round of iterative training to obtain the processing result for that round.

[0153] In step S1302 of some embodiments, the processing result of the current round is compared with the ground truth image of the nucleic acid molecule to obtain training bias data. It can be explained that the ground truth image of the nucleic acid molecule can reflect the true and accurate arrangement of the base sequence in the simulated nucleic acid sample. Therefore, the processing result obtained in each round of iterative training can be compared with the ground truth image of the nucleic acid molecule. This yields the training bias data between the processing result of the current round and the ground truth image of the nucleic acid molecule. It can be noted that the closer the training bias data between the processing result of the current round and the ground truth image of the nucleic acid molecule reflects, the better the resolution improvement effect of the image processing model, and the stronger the super-resolution processing capability of the image processing model.

[0154] In some embodiments, step S1303 updates the weight parameters of the image processing model based on the training bias data. It can be explained that the weight parameters of the image processing model can be updated after obtaining the training bias data. It can also be explained that neural network models generally have two types of parameters: one type is the tuning parameters in machine learning algorithms, which can be flexibly set based on existing or current experience; these are also called hyperparameters. Examples include the regularization coefficient λ and the depth of the tree in a decision tree model. Hyperparameters are also a type of parameter, possessing the characteristics of parameters, such as uncertainty—that is, they are not known constants but configurable values. A "correct" value can be assigned to them based on existing or current experience; they are flexibly set values ​​and are not learned through system learning. The other type of parameter can be learned and estimated from data; these are called model parameters, which are the learnable parameters of the model itself. For example, the weighting coefficients (slope) and the bias term (intercept) of a linear regression line are model parameters. Learnable parameters specifically refer to the parameter values ​​learned during the training process of a neural network model. For learnable parameters, they typically start with a set of random values ​​and are then updated iteratively as the neural network model learns. In fact, "when the neural network model learns" more accurately means that the parameters of the neural network model are in the process of iterative updating, gradually determining appropriate values ​​for these parameters. It can be noted that appropriate values ​​can be those that minimize or converge the loss function. Therefore, in some embodiments of this disclosure, in order to gradually determine appropriate values ​​for these model parameters during the iterative updating process of the data classification model, the weight parameters of the image processing model can be updated based on training bias data.

[0155] In the embodiments of this disclosure shown in steps S1301 to S1303, in each round of iterative training, a simulated nucleic acid image is input into the image processing model for image processing training to obtain the processing result of this round. Then, the processing result of this round is compared with the ground truth image of the nucleic acid molecule to obtain training bias data. Based on the training bias data, the weight parameters of the image processing model are updated. In this way, the super-resolution processing capability of the image processing model can be continuously optimized over several rounds of iterative training.

[0156] In step S404 of some embodiments, when the image processing model meets a first predetermined condition during iterative training, a trained image processing model is obtained. It can be explained that when the image processing model meets the first predetermined condition during iterative training, it means that the image processing model's super-resolution processing capability for images has reached the expected level for practical application, and the image processing model is adaptable to super-resolution processing of sequencing images. In this case, the iterative training of the image processing model can be terminated, and the trained image processing model is obtained.

[0157] In some embodiments, the image processing model meets a first predetermined condition during iterative training. This can be that after the image processing model performs image processing training on simulated nucleic acid images during iterative training, the processing result reaches the resolution level of the true image of nucleic acid molecules. Alternatively, the image processing model meets the first predetermined condition during iterative training, and its super-resolution processing capability for images converges to a certain level.

[0158] According to some specific embodiments of this disclosure, step S404, where the image processing model meets the first predetermined condition during iterative training, to obtain the trained image processing model, may include, but is not limited to: when the training bias data reflects that the image processing model has converged during iterative training, determining that the image processing model meets the first predetermined condition during iterative training, and obtaining the trained image processing model. It can be explained that when the training bias data reflects that the image processing model has converged during iterative training, it means that the image processing model's super-resolution processing capability for images has been optimized to a current limit and is difficult to further improve. At this point, it can be determined that the image processing model meets the first predetermined condition during iterative training, and the trained image processing model is obtained.

[0159] It should be understood that there are many ways to determine whether an image processing model meets the first predetermined condition during iterative training, including, but not limited to, the specific embodiments described above.

[0160] According to the embodiments of this disclosure shown in steps S401 to S404, a true image of nucleic acid molecules is first constructed based on simulated nucleic acid samples; then, simulated sequencing is performed based on the simulated nucleic acid samples using a simulated optical system to obtain a simulated nucleic acid image; the true image of nucleic acid molecules and the simulated nucleic acid image corresponding to the true image of nucleic acid molecules are input into the original image processing model, and the image processing model is iteratively trained; when the image processing model meets the first predetermined condition during iterative training, the trained image processing model is obtained. In this way, the super-resolution processing capability of the image processing model can be improved during the training process until the super-resolution processing capability of the image processing model reaches the expected level for practical application, and the image processing model can adapt to super-resolution processing of sequencing images. On this basis, using the trained image processing model to process sequencing images helps to improve image resolution and helps to reduce the crosstalk effect caused by the fluorescence signals of adjacent nucleic acid molecules when the spacing between nucleic acid molecules is smaller than the resolution of the imaging system. Thus, the accuracy of base interpretation can be effectively improved during the sequencing of nucleic acid molecules.

[0161] The following two specific embodiments are provided to illustrate the technical effects of the nucleic acid molecular sequencing method disclosed herein.

[0162] Specific Implementation Example 1: Sequencing of multi-spacing nucleic acid molecular chips on a sequencer optical engine.

[0163] The sequencer uses a 0.8NA air microscope with an optical resolution of approximately 500-600nm and a sampling resolution of 566nm. Under normal use, the inter-molecule spacing of nucleic acids is 715nm. When sequencing an *E. coli* genome library, the alignment rate for 100 rounds (SE100) of a single strand can reach over 90%. Based on this, this specific embodiment will perform sequencing on a multi-interval sequencing chip (576nm, 480nm, 450nm).

[0164] First, nucleic acid molecule patterns with spacings of 576nm, 480nm, and 450nm are generated based on the chip template. This generates ground-value images of nucleic acid molecules from multiple rounds of sequencing and corresponding simulated nucleic acid images, training the image processing model. Taking 480nm as an example, each imaging field of view (FoV) of the chip template contains 10*10 blocks, totaling 100 blocks. The number of rows / columns of nucleic acid molecules in each block is 145, 120, 180, 210, 240, 240, 210, 180, 120, and 145, with a tracing line width of 1440nm between adjacent blocks. First, a list of nucleic acid molecules required for simulation is generated based on the template. Then, based on the measured pixel size of 566nm, the nucleic acid molecules are mapped to the corresponding four-color imaging space.

[0165] Nucleic acid molecules on a sequencing chip are sequenced using a sequencer, and multiple rounds of sequencing images are captured. The signal-to-noise ratio (SNR) of each image is extracted, and simulation data pairs are generated based on the system's optical transfer function and SNR to train the image processing model.

[0166] After the resequencing images are processed by the image processing model, the target images are subjected to base identification and comparison.

[0167] Referring to Figure 14, the sequencing results show that after sequencing nucleic acid molecules on multi-spacing chips using the nucleic acid molecule sequencing method disclosed in this paper, the alignment rate of nucleic acid molecules with a spacing of 576 nm can reach 90%, and the alignment rate of nucleic acid molecules with a spacing of 480 nm can reach 75%, which is an improvement of 18% compared with traditional algorithms.

[0168] Specific Implementation Example 2: Sequencing of multi-spacing nucleic acid molecular chips on a self-built high-sampling optical engine.

[0169] The self-built high-sampling optical engine uses a 0.8NA air mirror with an optical resolution of approximately 500-600nm and a sampling resolution of 260nm. The intermolecular distance of nucleic acids used in normal use is 715nm.

[0170] When sequencing an E. coli genome library using the nucleic acid molecular sequencing method disclosed herein, the single-stranded 100-round (SE100) alignment rate can reach over 90%. The same nucleic acid molecular sequencing method can be used to sequence multi-interval sequencing chips (576nm, 480nm, 450nm, 400nm, 360nm). The implementation steps are as described in the aforementioned specific embodiment one.

[0171] Referring to Figure 15, the final sequencing results show that, compared with traditional algorithms, the alignment rates of nucleic acid molecules with spacings of 576nm, 480nm, 450nm, 400nm, and 360nm are improved by 10%, 14%, 18%, 17%, and 24%, respectively, using the algorithm described in this technical solution.

[0172] Referring to Figure 16, according to some embodiments, this disclosure provides a nucleic acid molecular sequencing device 1600, comprising:

[0173] The image acquisition module 1601 is used to acquire the sequencing image of the target nucleic acid sample. The sequencing image is an image obtained by acquiring the target nucleic acid sample using a preset optical sequencing system.

[0174] The image processing module 1602 is used to process the sequencing image based on the image processing model to obtain the target image. The image processing model is a model trained using the constructed true image of nucleic acid molecules and the corresponding simulated nucleic acid image. The simulated nucleic acid image is an image obtained by simulating sequencing of the simulated nucleic acid sample corresponding to the true image of nucleic acid molecules based on the simulated optical system. The simulated optical system is an optical system obtained by simulating a preset optical sequencing system.

[0175] The sequencing module 1603 is used to perform nucleic acid molecular sequencing based on the target image to obtain the sequencing results corresponding to the target nucleic acid sample.

[0176] It is evident that the content of the above-described nucleic acid molecular sequencing method embodiments is applicable to the embodiments of this nucleic acid molecular sequencing device. The specific functions implemented by this nucleic acid molecular sequencing device embodiment are the same as those of the above-described nucleic acid molecular sequencing method embodiments, and the beneficial effects achieved are also the same as those achieved by the above-described nucleic acid molecular sequencing method embodiments.

[0177] Referring to Figure 17, which illustrates the hardware structure of an electronic device according to another embodiment, the electronic device includes:

[0178] The processor 1701 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits to execute related programs in order to achieve the technical solutions provided in the embodiments of this disclosure.

[0179] The memory 1702 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1702 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1702 and is called and executed by the processor 1701 to execute the nucleic acid molecular sequencing method of the embodiments of this disclosure.

[0180] The input / output interface 1703 is used to implement information input and output;

[0181] The communication interface 1704 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0182] Bus 1705 transmits information between each component of the device (e.g., processor 1701, memory 1702, input / output interface 1703, and communication interface 1704);

[0183] The processor 1701, memory 1702, input / output interface 1703 and communication interface 1704 are connected to each other within the device via bus 1705.

[0184] This disclosure also provides a computer program product comprising a computer program. A processor of a computer device reads and executes the computer program, causing the computer device to perform the nucleic acid molecular sequencing method described above.

[0185] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It is understood that such data can be interchanged where appropriate so that embodiments of this disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “including,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.

[0186] It should be understood that in this disclosure, "at least one item" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0187] It should be understood that in the description of the embodiments disclosed herein, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.

[0188] In the several embodiments provided in this disclosure, it can be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0189] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Depending on the actual situation, some or all of the units can be selected to achieve the purpose of this embodiment.

[0190] Furthermore, the functional units in each embodiment of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units.

[0191] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of each embodiment of the method of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0192] It should also be understood that the various implementation methods provided in this disclosure can be combined arbitrarily to achieve different technical effects.

[0193] The above is a detailed description of the embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.

Claims

1. [Corrected according to Rule 91, July 15, 2024] A nucleic acid molecular sequencing method, characterized in that, include: Acquire sequencing images of the target nucleic acid sample, wherein the sequencing images are obtained by image acquisition of the target nucleic acid sample using a preset optical sequencing system; The sequencing image is processed using an image processing model to obtain a target image. The image processing model is a model trained using a constructed true image of a nucleic acid molecule and a corresponding simulated nucleic acid image. The simulated nucleic acid image is an image obtained by simulating sequencing of a simulated nucleic acid sample corresponding to the true image of the nucleic acid molecule using a simulated optical system. The simulated optical system is an optical system obtained by simulating the preset optical sequencing system. Nucleic acid molecular sequencing is performed based on the target image to obtain the sequencing results corresponding to the target nucleic acid sample.

2. The method according to claim 1, characterized in that, Before processing the sequencing image using an image processing model to obtain the target image, the process further includes training the image processing model, including: Construct the true image of the nucleic acid molecule based on the simulated nucleic acid sample; The simulated nucleic acid image is obtained by performing simulated sequencing based on the simulated nucleic acid sample using the simulated optical system. The true image of the nucleic acid molecule and the simulated nucleic acid image corresponding to the true image of the nucleic acid molecule are input into the original image processing model, and the image processing model is iteratively trained. When the image processing model meets the first predetermined condition during iterative training, the trained image processing model is obtained.

3. The method according to claim 2, characterized in that, The step of performing simulated sequencing based on the simulated nucleic acid sample using the simulated optical system to obtain the simulated nucleic acid image includes: The optical calibration information of the preset optical sequencing system is simulated; The nucleic acid distribution information is simulated based on the simulated optical calibration information to obtain the simulated nucleic acid image.

4. The method according to claim 3, characterized in that, The nucleic acid molecules in the simulated nucleic acid sample include multiple bases, and each base is labeled with a fluorescent signal. The simulation of the optical calibration information of the preset optical sequencing system includes: Determine the pixel size of the preset optical sequencing system; wherein, the pixel size is the size of a single pixel in the preset optical sequencing system on the imaging plane; Optical imaging analysis is performed based on the fluorescence signal corresponding to the base in the simulated nucleic acid sample to obtain the optical transfer function. By integrating the pixel size with the optical transfer function, the simulated optical calibration information is obtained.

5. The method according to claim 4, characterized in that, The optical transfer function is obtained by performing optical imaging analysis based on the fluorescence signal corresponding to the bases in the simulated nucleic acid sample, including: The simulated nucleic acid sample was scanned to obtain a scanned nucleic acid image; Based on the scanned nucleic acid image, a local maximum search is performed to obtain a fluorescence image reflecting each of the fluorescence signals; The fluorescence images corresponding to multiple fluorescence signals are subjected to Gaussian fitting and averaging to obtain the point spread function; The optical transfer function is obtained by performing a Fourier transform on the point spread function.

6. The method according to claim 4, characterized in that, The step of integrating the pixel size with the optical transfer function to obtain the simulated optical calibration information includes: Noise is extracted from the blank regions between nucleic acid molecules in the simulated nucleic acid sample to obtain simulated noise information; The pixel size, the optical transfer function, and the simulated noise information are integrated to obtain the simulated optical calibration information.

7. The method according to claim 6, characterized in that, The simulation of the nucleic acid distribution information based on the simulated optical calibration information to obtain the simulated nucleic acid image includes: Based on the pixel size, the nucleic acid distribution information is mapped onto the imaging space to obtain a first simulated image of the simulated nucleic acid sample in the imaging space; Based on the optical transfer function, the imaging performance of the first simulated image is simulated to obtain the second simulated image; Based on the simulated noise information, environmental noise simulation is performed on the second simulated image to obtain the simulated nucleic acid image.

8. The method according to claim 2, characterized in that, The step of inputting the ground truth image of the nucleic acid molecule and the simulated nucleic acid image corresponding to the ground truth image of the nucleic acid molecule into the original image processing model, and iteratively training the image processing model, includes: In each round of iterative training, the simulated nucleic acid image is input into the image processing model for image processing training to obtain the processing result of this round; The results of this round of processing are compared with the true image of the nucleic acid molecule to obtain training bias data; The weight parameters of the image processing model are updated based on the training bias data.

9. The method according to claim 8, characterized in that, The step of obtaining the trained image processing model when the image processing model meets a first predetermined condition during iterative training includes: When the training bias data reflects that the image processing model converges during iterative training, it is determined that the image processing model meets the first predetermined condition during iterative training, and the trained image processing model is obtained.

10. A nucleic acid molecular sequencing device, characterized in that, include: The image acquisition module is used to acquire sequencing images of the target nucleic acid sample, wherein the sequencing images are obtained by acquiring images of the target nucleic acid sample using a preset optical sequencing system; The image processing module is used to process the sequencing image based on an image processing model to obtain a target image. The image processing model is a model trained using a constructed true image of a nucleic acid molecule and a corresponding simulated nucleic acid image. The simulated nucleic acid image is an image obtained by simulating sequencing of a simulated nucleic acid sample corresponding to the true image of the nucleic acid molecule based on a simulated optical system. The simulated optical system is an optical system obtained by simulating the preset optical sequencing system. The sequencing module is used to perform nucleic acid molecular sequencing based on the target image to obtain the sequencing results corresponding to the target nucleic acid sample.

11. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the nucleic acid molecular sequencing method according to any one of claims 1 to 9.

12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the nucleic acid molecular sequencing method according to any one of claims 1 to 9.

13. A computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the nucleic acid molecular sequencing method according to any one of claims 1 to 9.