Sequencing image processing method, computer equipment and storage medium

By performing image reconstruction and spot prediction model processing on gene sequencing images, the error problem caused by noise in sequencing images was solved, achieving higher sequencing accuracy and base recognition stability.

CN120997248APending Publication Date: 2025-11-21SIKUN LIFE SCIENCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511093025.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

During gene sequencing, the presence of fluorescence noise leads to large errors in spot recognition in sequencing images, resulting in low accuracy and poor recognition stability.

Method used

Multi-frame sequencing images are processed using image reconstruction and spot prediction models, including image reconstruction, alignment, spot prediction and reconstruction, to eliminate fluorescence noise and improve the accuracy of spot recognition.

Benefits of technology

By using deep learning models to process sequencing images, noise can be removed to the greatest extent, thereby improving sequencing accuracy and the stability of base identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997248A_ABST
    Figure CN120997248A_ABST
Patent Text Reader

Abstract

The invention provides a sequencing image processing method, computer equipment and a storage medium, and the method comprises the steps: carrying out the image reconstruction of a plurality of frames of original sequencing images from a plurality of sequencing cycles through an image reconstruction model, and obtaining a plurality of frames of reconstructed sequencing images; determining an alignment template for image alignment from the multiple frames of reconstructed sequencing images, and performing image alignment on the multiple frames of original sequencing images by using the alignment template to obtain multiple frames of aligned sequencing images; light spot prediction is carried out on at least part of sequencing images in the multiple frames of aligned sequencing images through a light spot prediction model, a predicted light spot probability graph is obtained, and the light spot probability graph comprises information of light spots in the aligned sequencing images; performing light spot reconstruction processing on the light spot probability graph to obtain a template image; and based on the template image, performing base recognition on the multiple frames of original sequencing images to obtain a base sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of gene sequencing technology, and more specifically, to a sequencing image processing method, computer equipment, and storage medium. Background Technology

[0002] Next-generation sequencing (NGS), also known as high-throughput sequencing, pioneered the introduction of reversible end-termination sequencing schemes, thereby enabling sequencing while synthesizing. It creatively realizes the simultaneous replication, extension, and luminescence of bases on a single DNA strand, using imaging technology to capture base information on the DNA strand in real time and determine the base sequence of the DNA strand.

[0003] In sequencing synthesis technology, single-stranded DNA fragments (often referred to as libraries) enter the flow cell of a sequencing chip. There, they bind to oligonucleotides seeded on the surface of the flow cell and become immobilized. Each DNA fragment is amplified, such as by PCR, resulting in millions of copies of the single-stranded DNA fragment in its vicinity, forming clusters. These clustered DNA fragments emit detectable fluorescent signals during subsequent sequencing. These fluorescent signals are captured by a high-resolution camera (such as an industrial camera), forming a spot in the sequencing image. This spot allows identification of the base type corresponding to the fluorescent signal, thus revealing the base sequence of the DNA fragment.

[0004] During gene sequencing, due to various external or internal factors, there is often a lot of fluorescence noise in the sequencing images. This includes background noise caused by cluster overlap, fluid reagent reactions in the flow cell of the sequencing chip, and mechanical deviation caused by the reciprocating motion of the camera to capture different areas of the sequencing chip. The presence of this noise causes a large error in spot recognition of the sequencing images, resulting in low sequencing accuracy and poor recognition stability. Summary of the Invention

[0005] This disclosure provides at least one method for processing sequencing images, a computer device, and a storage medium.

[0006] In a first aspect, embodiments of this disclosure provide a method for processing sequencing images, including:

[0007] Image reconstruction models were used to reconstruct multiple frames of raw sequencing images from multiple sequencing cycles, resulting in reconstructed sequencing images.

[0008] An alignment template for image alignment is determined from the reconstructed sequencing images of the multiple frames. The alignment template is then used to align the original sequencing images of the multiple frames to obtain the aligned sequencing images of the multiple frames.

[0009] A spot prediction model is used to predict the spot patterns in at least a portion of the sequenced images in the multi-frame aligned sequenced images to obtain a predicted spot probability map, which contains information about the spots in the aligned sequenced images.

[0010] The light spot probability map is subjected to light spot reconstruction processing to obtain a template image;

[0011] Based on the template image, base identification is performed on the multiple frames of original sequencing images to obtain the base sequence.

[0012] Optionally, the light spot probability map is a binary classification light spot probability map, which includes light spot region information and background information;

[0013] The step of performing spot reconstruction processing on the spot probability map to obtain a template image includes:

[0014] The light spot probability map is segmented into foreground and background and then thresholded to obtain a first light spot map; wherein the foreground is the light spot region;

[0015] Based on the intensity value of the spot region in the first spot image, the centroid of the spot is determined from the first spot image;

[0016] Using the centroid of the light spot as the center, the light spot to which the centroid belongs is reconstructed using an objective function to obtain the template image;

[0017] The objective function includes: a point spread function or a two-dimensional Gaussian function.

[0018] Optionally, determining the centroid of the light spot from the first light spot image based on the intensity value of the light spot region in the first light spot image includes:

[0019] The pixels belonging to the foreground in the first light spot image are traversed, and it is determined whether all other pixels in the target neighborhood corresponding to the traversed pixel belong to the foreground and whether the intensity value of the traversed pixel is greater than the gray value of the other pixels in the corresponding target neighborhood. If all other pixels in the target neighborhood corresponding to the traversed pixel belong to the foreground and the gray value of the traversed pixel is greater than the gray value of the other pixels in the corresponding target neighborhood, the traversed pixel is determined as the centroid of the light spot.

[0020] or,

[0021] From the foreground pixels in the first light spot image, determine the pixels to be traversed; wherein, all pixels located within the target neighborhood of the pixel to be traversed are foreground pixels; traverse the pixel to be traversed, and compare the gray value of the traversed pixel with the gray value of other pixels in the corresponding target neighborhood; if the intensity value of the traversed pixel is greater than the gray value of other pixels in the corresponding target neighborhood, determine the traversed pixel as the centroid of the light spot.

[0022] Optionally, the method further includes:

[0023] Using the pixel at the centroid of the light spot as the center, Gaussian fitting is performed in the neighborhood corresponding to the pixel to determine the sub-pixel of the centroid of the light spot.

[0024] The step of reconstructing the light spot to its centroid using an objective function, with the centroid of the light spot as the center, to obtain the template image, includes:

[0025] Using the sub-pixel of the centroid of the light spot as the center, the light spot to which the centroid belongs is reconstructed using the objective function to obtain the template image.

[0026] Optionally, using the centroid of the light spot as the center, a two-dimensional Gaussian function is used to reconstruct the light spot to which the centroid belongs, including:

[0027] Using the centroid of the light spot as the center of a pixel or subpixel, and determining the spot region based on the size of the light spot, the position coordinates of the pixels within the spot region and their approximate standard deviation are substituted into the two-dimensional Gaussian function to obtain the reconstructed grayscale values ​​of the pixels within the spot region; wherein, the standard deviation of the two-dimensional Gaussian function is determined based on the full width and half height of the sequencing image.

[0028] The reconstructed gray values ​​of pixels at the same position within the spot area in different sequencing cycles are summed to obtain the reconstructed spot.

[0029] Optionally, the multi-frame aligned sequencing image includes multiple image patches;

[0030] The process of reconstructing the light spot to which the centroid of the light spot belongs, using a two-dimensional Gaussian function with the centroid of the light spot as the center, includes:

[0031] Using the centroid of the light spot in the plurality of image blocks as the center, a two-dimensional Gaussian function is used to reconstruct the light spot to which the centroid of the light spot in the plurality of image blocks belongs, thereby obtaining the template image corresponding to the plurality of image blocks respectively;

[0032] The approximate standard deviation used when reconstructing the centroid of the light spot in different image blocks using a two-dimensional Gaussian function is determined based on the full width and half height of each image block.

[0033] Optionally, the light spot probability map includes: multiple light spot probability maps of different classifications; any one of the light spot probability maps of classification includes at least one of the following: light spot centroid, light spot boundary, and light spot region, as well as background information;

[0034] The step of performing spot reconstruction processing on the spot probability map to obtain a template image includes:

[0035] By overlaying multiple light spot probability maps of different categories, a multi-category overlaid light spot probability map is obtained.

[0036] The image reconstruction model is used to perform image reconstruction processing on the multiple classification superimposed light spot probability maps to obtain a reconstructed image;

[0037] The reconstructed image is subjected to denoising and threshold segmentation to obtain the template image.

[0038] Optionally, the method further includes: performing upsampling processing on the multi-frame aligned sequencing images to obtain multi-frame upsampled sequencing images;

[0039] The step of using a spot prediction model to predict spot patterns in at least a portion of the sequenced images in the multi-frame aligned sequencing images, to obtain a predicted spot probability map, includes:

[0040] A spot prediction model is used to predict spot patterns in at least a portion of the sequencing images in the multi-frame upsampled sequencing images, resulting in a predicted spot probability map.

[0041] After performing spot reconstruction processing on the spot probability map to obtain the template image, the method further includes:

[0042] The template image is downsampled to restore its resolution to match that of the multi-frame aligned sequencing image.

[0043] Optionally, the step of performing base identification on the multiple frames of original sequencing images based on the template image to obtain the base sequence includes:

[0044] Using the template image as a reference template, the intensity value corresponding to the centroid position of the light spot indicated in the template image is extracted from the original sequencing images corresponding to each sequencing cycle, based on the centroid position of the light spot indicated in the template image.

[0045] Based on the intensity value corresponding to the centroid position of the light spot indicated in the extracted template image, the base type corresponding to the centroid position of the light spot indicated in the template image is identified.

[0046] Optionally, the step of using the template image as a reference template and extracting the intensity value corresponding to the centroid position of the light spot indicated in the template image from the original sequencing images corresponding to each sequencing cycle, based on the centroid position of the light spot indicated in the template image, includes:

[0047] The template image and the original sequencing images corresponding to each sequencing cycle are aligned to obtain the affine transformation matrix between the template image and the original sequencing images.

[0048] According to the affine transformation matrix, the centroid position of the light spot indicated in the template image is projected onto the original sequencing image to obtain the projection point of the centroid position of the light spot in the original sequencing image, and the intensity value of the projection point in the original sequencing image is determined as the intensity value corresponding to the centroid position of the light spot.

[0049] In a second aspect, an optional implementation of this disclosure also provides a computer device, a processor, and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the processor is configured to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the steps of the first aspect above, or any possible implementation of the first aspect, are performed.

[0050] Thirdly, an optional implementation of this disclosure also provides a computer-readable storage medium storing a computer program that, when run, performs the steps of the first aspect or any possible implementation of the first aspect.

[0051] The sequencing image processing method provided in this disclosure, after obtaining multiple frames of raw sequencing images acquired during gene sequencing, first uses an image reconstruction model to reconstruct the multiple frames of raw sequencing images, obtaining reconstructed sequencing images. Then, it determines an alignment template for image alignment from the reconstructed sequencing images and uses the alignment template to align the multiple frames of raw sequencing images, obtaining aligned sequencing images. In this way, the fluorescence noise present in the original sequencing images is eliminated using the image reconstruction model, and the light spots contained in the obtained alignment template are more accurate. Therefore, the alignment template can be used to perform more precise alignment of the multiple frames of raw sequencing images. More precise alignment is achieved to minimize errors between sequencing images from multiple sequencing cycles. Then, a spot prediction model is used to predict spots in at least a portion of the aligned sequencing images, resulting in a predicted spot probability map. This map contains information about the spots in the aligned sequencing images, enabling spot prediction within the aligned images. The spot probability map is then reconstructed to obtain a template image. This template image contains more accurate spot information, allowing for more accurate base identification of multiple raw sequencing images based on the template image, resulting in higher precision of the obtained base sequences. Through this process, a deep learning model is used to process multiple raw sequencing images acquired during gene sequencing, ensuring maximum removal of noise caused by various external or internal factors during image capture, thereby improving sequencing accuracy and ensuring the stability of base identification.

[0052] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0053] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.

[0054] Figure 1 Schematic diagrams of sequencing systems provided in some embodiments of this disclosure are shown;

[0055] Figure 2 A schematic diagram of a sequencing chip provided in some embodiments of this disclosure is shown;

[0056] Figure 3 Schematic diagrams of optical detection systems provided in some embodiments of this disclosure are shown;

[0057] Figure 4 A flowchart illustrating a method for processing sequencing images provided in some embodiments of this disclosure is shown;

[0058] Figure 5a One specific example of the raw sequencing images provided in some embodiments of this disclosure is shown;

[0059] Figure 5b This is a second specific example of a probability map of spot information predicted by a model provided in some embodiments of this disclosure;

[0060] Figure 5c This is a third specific example of a probability diagram of filtered spot information provided in some embodiments of the present disclosure;

[0061] Figure 6a The present disclosure provides specific examples of spot images formed by segmenting the foreground and background in the fusion region of a sharpened sequencing image according to some embodiments of the present disclosure.

[0062] Figure 6b The present disclosure provides some embodiments of the present disclosure, which use candidate spot centroids as seed points for region growing and obtain spot region images according to the region growing algorithm.

[0063] Figure 6c Specific examples are shown of images formed by using boundary finding algorithms to calculate the boundaries of each unfiltered spot, as provided in some embodiments of this disclosure.

[0064] Figure 6d Specific examples of images provided in some embodiments of this disclosure after overlaying an image of the spot region with the spot boundary image are shown;

[0065] Figure 6e This illustration shows a superimposed diagram of the spot background, spot boundary, spot region, and spot centroid provided in some embodiments of this disclosure;

[0066] Figure 7 This is a fourth specific example of multiple registered images provided in some embodiments of this disclosure;

[0067] Figure 8 This is the fifth specific example of a template image provided in some embodiments of this disclosure;

[0068] Figure 9 A schematic diagram of a computer device provided in some embodiments of the present disclosure is shown. Detailed Implementation

[0069] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown herein can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0070] To facilitate understanding of the technical solutions disclosed herein, the technical terms used in the embodiments of this disclosure will first be explained:

[0071] Library Construction:

[0072] The genomic DNA or RNA molecules to be sequenced are broken down using physical or chemical methods, such as by ultrasound. After breaking, DNA or RNA fragments are formed. First, enzymes are used to fill in the ends of the DNA or RNA fragments. Then, specific enzymes are used to ligate specific DNA or RNA sequences (usually, this specific DNA or RNA sequence is also called a linker) to the ends of the fragments, forming a mixture of DNA or RNA. This mixture of DNA or RNA is also known in the industry as a library.

[0073] To save sequencing costs, multiple samples are typically sequenced simultaneously in a sequencer. To differentiate the sequencing results of different samples, when preparing libraries for different samples, a DNA or RNA sequence (usually containing 6-8 bases) is included in the adapters to identify the source of that sample. This DNA or RNA sequence can also be called a sample tag (or index, barcode). Understandably, each sample's library adapter contains its own unique sample tag.

[0074] Typically, library construction is performed outside of the sequencer, for example, by obtaining the library through experimental procedures in the laboratory.

[0075] Amplification reaction:

[0076] Taking the amplification of a DNA library via bridge PCR (polymerase chain reaction) as an example, after constructing the library, it can be seeded onto a sequencing chip for amplification. The adapters at both ends of the library are complementary to the first type of amplification primer on the sequencing chip; therefore, complementary hybridization allows the library to be seeded onto the sequencing chip.

[0077] After the library is seeded onto the sequencing chip, it can be used as a template for amplification. For example, the amplification process can begin by adding dNPs and polymerase to the sequencing chip. The polymerase will synthesize a new DNA strand along the template strand, starting with the first amplification primer. This new DNA strand is completely complementary to the template strand and is therefore called the complementary strand of the template strand. The complementary strand is covalently linked to the sequencing chip. Next, NaOH solution is added to the sequencing chip for rinsing. In the presence of NaOH solution, the template strand and the complementary strand unwind, and the template strand is washed away with the alkali solution, while the complementary strand covalently linked to the sequencing chip is retained. Next, a neutral liquid is added to the sequencing chip to neutralize the NaOH solution. The entire environment within the sequencing chip becomes neutral, allowing the other end of the complementary strand to continue complementary hybridization with the second amplification primer on the sequencing chip. Then, dNP and polymerase are added. The polymerase synthesizes a completely new DNA strand, starting from the second amplification primer and continuing along the complementary strand. At this point, the new DNA strand is completely complementary to the complementary strand and identical to the template strand. NaOH solution is then added again to untie the two strands. This yields two covalently linked and complementary strands for the sequencing chip. Repeating this process will result in an exponential increase in the number of DNA strands.

[0078] After amplification, the sequencing chip retains two identical DNA double strands, one matching the template strand and the other the complementary strand. Specific reaction reagents are then added to the chip to cleave the DNA strand synthesized from one of the amplification primers. For example, the DNA strand matching the complementary strand is cleaved, leaving the strand matching the template strand. NaOH solution is then added to the chip for rinsing. The alkali solution unwinds the DNA double strands, and the cleaved strands are washed away, leaving only single-stranded DNA on the chip. The number of single-stranded DNA strands on the chip at this point is exponentially higher than at the start of amplification, forming DNA clusters. All single-stranded DNA strands within a cluster are identical. Adding a neutral solution allows for sequencing of all single-stranded DNA within the cluster under neutral conditions.

[0079] It should be noted that the amplification reaction described above can be performed outside the sequencer, such as amplifying the library in the laboratory through experimental procedures, or it can be performed inside the sequencer. When the amplification reaction is performed inside the sequencer, the library and various reaction reagents (e.g., dNP, polymerase, NaOH alkaline solution, neutral solution, etc.) participating in the amplification reaction can be added to the sequencing chip through the sequencer's liquid circuit system or a liquid circuit system. In addition, the above amplification reaction is only exemplary, and this application is not limited to the bridge PCR amplification method. Other amplification methods can also be used, such as loop-mediated isothermal amplification (LAMP), nucleic acid sequence-based amplification (NASBA), rolling circle amplification (RCA), multiplex probe amplification (MPA), etc.

[0080] Sequencing reaction:

[0081] Taking the sequencing-by-synthesis principle as an example, during sequencing, four dNTPs with fluorescent groups are added to the sequencing chip through a liquid circuit system. Each dNTP can only be synthesized with one of the four bases ATCG, and the 3' end of each dNTP is blocked by a blocking group (including but not limited to azide). Then, polymerase is added to the sequencing chip through the liquid circuit system. Through the action of polymerase, one of the four dNTPs will be synthesized with a complementary base on the single strand being sequenced. Since the 3' end of the dNTP is blocked by a blocking group, only one dNTP can be extended on the single strand being sequenced at a time. After synthesis, specific chemical reagents are added to the sequencing chip through the liquid circuit system to flush away excess dNTPs and polymerase. Next, the fluorescent groups of the dNTPs synthesized on the single strands can be excited by the optical detection system, causing the fluorescent groups to emit fluorescent signals. Since the fluorescent groups of the dNTPs on each single strand in a cluster will emit the same fluorescent signal, the fluorescent signal is amplified. Therefore, the optical detection system can collect the fluorescent signal and generate a fluorescent image.

[0082] By processing and analyzing fluorescence images using a computer system, it can be determined which type of dNTP was synthesized onto the sequenced single strand. Then, based on the complementarity principle, it can be deduced which base on the sequenced single strand was synthesized with the dNTP. This completes one sequencing cycle.

[0083] Next, specific chemical reagents are added to the sequencing chip through a liquid circuit system to remove the blocking and fluorescent groups, thereby exposing the 3'-terminal hydroxyl groups of the dNTPs.

[0084] Next, proceed to the next sequencing cycle and repeat the above process.

[0085] As we can understand, one sequencing cycle can detect one base, and through multiple sequencing cycles, multiple bases in the sequenced single strand can be detected. Specifically, the number of sequencing cycles can be determined based on the set sequencing read length, for example, 150 or 300 sequencing cycles.

[0086] Of course, it should also be noted that the above sequencing reactions are merely illustrative, and this application is not limited to using the principle of sequencing by synthesis; other sequencing principles may also be used.

[0087] Sequencing system:

[0088] See Figure 1 As shown, the sequencing system includes: a sequencing chip 10, a chip platform 20, a reagent storage container 30, a liquid flow system (or diversion system) 40, an optical detection system 50, a computer system 60, and a waste liquid storage container 70. Among them:

[0089] Sequencing chip 10 is configured to provide reaction regions for amplification and sequencing reactions;

[0090] Chip platform 20, configured to fix and support sequencing chip 10;

[0091] The reagent storage container 30 is configured to store one or more mixed sample libraries and one or more reagents;

[0092] The liquid circuit system 40 is configured to controllably deliver one or more mixed sample libraries and one or more reagents from the reagent storage container 30 to the sequencing chip 10 for amplification and sequencing reactions in the sequencing chip 10, and to controllably deliver the waste liquid after the reaction from the sequencing chip 10 to the waste liquid storage container 70.

[0093] The optical detection system 50 is configured to excite and acquire fluorescence signals during the sequencing reaction, and generate a fluorescence image based on the fluorescence signals;

[0094] Computer system 60 is configured to acquire fluorescence images from optical detection system 50 and identify the base sequences of sample libraries based on the fluorescence images;

[0095] Waste liquid storage container 70 is configured to store waste liquid generated after the reaction.

[0096] Sequencing chip:

[0097] Sequencing chips, serving as carriers for amplification and sequencing reactions, provide reaction regions, known as channels. Typically, a sequencing chip 10 may contain one or more channels 11 (e.g., 2, 4, 6, or 8 channels), which are isolated from each other. See also... Figure 2 As shown, taking four channels as an example, each channel 11 has a small hole 12 at each end for the flow of fluids (e.g., biological samples, reaction reagents) into and out. The upper and lower surfaces of each channel are chemically modified and covalently seeded with two amplification primers. These two amplification primers are complementary to the adapters at both ends of the library to achieve amplification of the library.

[0098] During gene sequencing, due to various external or internal factors, there is often a lot of fluorescence noise in the sequencing images. This includes background noise caused by cluster overlap, fluid reagent reactions in the flow cell of the sequencing chip, and mechanical deviation caused by the reciprocating motion of the camera to capture different areas of the sequencing chip. The presence of this noise causes a large error in spot recognition of the sequencing images, resulting in low sequencing accuracy and poor recognition stability.

[0099] Based on the above research, this disclosure provides a method for processing sequencing images. By using a deep learning model, multiple frames of raw sequencing images acquired during gene sequencing are processed to ensure that noise caused by various external or internal factors during the acquisition of sequencing images is removed to the greatest extent possible, thereby improving the accuracy of sequencing and ensuring the stability of base recognition.

[0100] The shortcomings of the above solutions are the result of the inventor's practical experience and careful research. Therefore, the discovery process of the above problems and the solutions proposed in this disclosure below should be considered as the inventor's contribution to this disclosure.

[0101] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0102] To facilitate understanding of this embodiment, the sequencing image processing method disclosed in this disclosure will first be described in detail. The execution entity of the sequencing image processing method provided in this disclosure is generally a computer device with certain computing power, such as the computer system 60 in the above example. In some possible implementations, the sequencing image processing method can be implemented by a processor calling computer-readable instructions stored in memory.

[0103] The following describes the processing method for sequencing images provided in the embodiments of this disclosure.

[0104] See Figure 4 The diagram shows a flowchart of a sequencing image processing method provided in this embodiment of the present disclosure. The method includes steps S401 to S405, wherein:

[0105] S401: Use an image reconstruction model to reconstruct multiple frames of raw sequencing images from multiple sequencing cycles to obtain reconstructed sequencing images.

[0106] S402: Determine an alignment template for image alignment from the reconstructed sequencing images of the multiple frames, and use the alignment template to align the original sequencing images of the multiple frames to obtain the aligned sequencing images of the multiple frames.

[0107] S403: Use a spot prediction model to predict the spot patterns in at least a portion of the sequencing images in the multi-frame aligned sequencing images to obtain a predicted spot probability map, wherein the spot probability map contains information about the spots in the aligned sequencing images.

[0108] S404: Perform spot reconstruction processing on the spot probability map to obtain a template image;

[0109] S405: Based on the template image, base identification is performed on the multiple frames of original sequencing images to obtain the base sequence.

[0110] The sequencing image processing method provided in this disclosure, after obtaining multiple frames of raw sequencing images acquired during gene sequencing, first uses an image reconstruction model to reconstruct the multiple frames of raw sequencing images, obtaining reconstructed sequencing images. Then, it determines an alignment template for image alignment from the reconstructed sequencing images and uses the alignment template to align the multiple frames of raw sequencing images, obtaining aligned sequencing images. In this way, the fluorescence noise present in the original sequencing images is eliminated using the image reconstruction model, and the light spots contained in the obtained alignment template are more accurate. Therefore, the alignment template can be used to perform more precise alignment of the multiple frames of raw sequencing images. Alignment is performed to minimize errors between sequencing images from multiple sequencing cycles. Then, a spot prediction model is used to predict spots in at least a portion of the aligned sequencing images, resulting in a predicted spot probability map. This probability map contains information about the spots in the aligned sequencing images, enabling spot prediction within the aligned images. Next, spot reconstruction processing is performed on the probability map to obtain a template image. The template image contains more accurate information about the spots, allowing for more accurate base identification of the original sequencing images from each cycle, resulting in higher precision of the obtained base sequences. Through this process, a deep learning model is used to process multiple original sequencing images acquired during gene sequencing, ensuring maximum removal of noise caused by various external or internal factors during image capture, resulting in more accurate template images, thus improving sequencing accuracy and ensuring the stability of base identification.

[0111] The following provides a detailed explanation of S401 to S405 respectively.

[0112] Regarding the above S401:

[0113] In practical implementation, during gene sequencing, mechanical errors caused by the relative movement of the optical detection system and the gene chip in the sequencing system can easily lead to deviations in the light spots located at the same position in multiple sequencing images of the same DNA cluster. Therefore, it is necessary to align sequencing images from different sequencing cycles obtained at the same shooting position. The alignment template and the sequencing image to be aligned typically need to have clear features and a high signal-to-noise ratio. In sequencing images, this corresponds to high grayscale values ​​of the light spots, clear boundaries, and a clear distinction between the foreground and background. Using image reconstruction techniques can reduce image noise and highlight the feature information of the light spots in the sequencing image. Selecting the reconstructed image as the alignment template for image alignment is more conducive to improving alignment accuracy.

[0114] For sequencing images, the clearer the light spot, the higher the resolution of the sequencing image. Since the intensity distribution of an ideal light spot follows a Gaussian distribution, the closer the probability distribution of the reconstructed light spot is to the probability distribution of the ideal light spot, the higher the resolution of the reconstructed sequencing image is considered to be.

[0115] In practical implementation, a convolutional neural network can be used to construct an image reconstruction model, and this model can be trained using a large number of sequencing images and their corresponding spot features as label data. After inputting the sequencing images into the image reconstruction model, a classification probability map of the spot feature information of the sequencing images can be obtained. A segmentation algorithm is used to remove the background. In the classification probability map after background removal, the higher the probability value of pixels belonging to the centroid of the fluorescence spot, the lower the probability of pixels farther from the centroid. The probability distribution of the spots in the classification probability map is similar to the Gaussian distribution of the ideal spot. This classification probability map is used as the reconstructed sequencing image for subsequent image alignment, improving the overall accuracy of sequencing.

[0116] The reconstructed sequencing image contains the probability distribution of each spot in the original sequencing image. To obtain the probability distribution of each spot, an existing segmentation network (such as the classic U-net and its improved derivatives) is used to predict the probability distribution of spots in the sequencing image pixel by pixel. Since the reconstructed sequencing image only cares about the probability distribution of spots in the sequencing image, the centroids of spots and non-spot centroids in the sequencing image can be used as labels to train a convolutional neural network.

[0117] Image reconstruction models can be trained in the following ways:

[0118] 1. Determine the sample image and its corresponding label:

[0119] Specifically, sample images are, for example, sequencing images obtained from multiple sequencing cycles during historical gene sequencing.

[0120] When determining labels for each sample image, at least one of the following two methods (1) and (2) can be used:

[0121] (1): Sharpen the sample image to obtain the sharpened sample image; perform foreground and background segmentation on the sharpened sample image to obtain the segmented image; perform local maximization search based on the pixel value of each pixel in the segmented image to obtain the pixel belonging to the centroid of the spot; generate a label based on the pixel position of the pixel belonging to the centroid of the spot.

[0122] Specifically, the sharpened sample image increases the contrast between the foreground and background. Then, segmentation algorithms such as OTSU, K-means, and support vector machines can be used to segment the foreground and background of the sharpened sample image to remove the background. The intensity value of the fluorescent spot centroid is the highest, and the centroid is found using local maxima. Pixels at the spot centroid are classified into one class, and pixels not at the spot centroid are classified into another class, forming labels.

[0123] When generating labels for sample images using the above method, each sample image is processed separately. Therefore, the image reconstruction model is trained using a single sequencing image and its corresponding label. The model predicts the reconstructed image corresponding to a single sequencing cycle and single channel.

[0124] (2): For the same shooting area, select one sample image from the sample images of multiple sequencing cycles as the alignment image, and the remaining sample images as the images to be aligned. Align the images of multiple sample images of multiple sequencing cycles, and denoise the aligned sample images to obtain the denoised image.

[0125] The denoised image is segmented into foreground and background to obtain the regions of interest (ROIs) corresponding to the denoised images of multiple sequencing cycles. The ROIs of the denoised images of multiple sequencing cycles are then fused together to obtain the fused region. The ROI corresponding to each denoised image is the region where the light spot is located in that denoised image.

[0126] Sharpening is performed on multiple frames of denoised images corresponding to multiple sequencing cycles to obtain the intensity value of each pixel in the fusion region of each sharpened denoised image (for example, the intensity value can be the pixel value of the pixel).

[0127] Based on the intensity values ​​of each pixel in the fusion region of the sharpened and denoised image, foreground and background segmentation is performed on the fusion region in each sharpened and denoised image to obtain the foreground and background of the fusion region in each sharpened and denoised image. The base type is identified by the base detection method in the fusion region of each sharpened and denoised image. By traversing multiple sequencing cycles, the base sequence corresponding to each pixel position in the fusion region can be obtained.

[0128] The base sequences corresponding to each pixel location in the fusion region are compared with the reference genome for consistency, or the base sequences corresponding to each pixel location in the fusion region are compared with the base sequences corresponding to neighboring pixel locations for consistency. Based on the consistency comparison results, the pixel locations corresponding to the base sequences with better consistency are retained from the pixel locations belonging to the fusion region as the pixel locations of candidate spot centroids, and labels are constructed based on the pixel locations corresponding to the spot centroids.

[0129] In the scheme described in (2) above, the sequencing images from multiple sequencing cycles contain more information about the light spots, thus the constructed labels are richer and include the light spot positions from multiple sequencing cycles. When training the convolutional neural network, the intensity values ​​of multiple sequencing cycles corresponding to the same pixel position are weighted and accumulated. A weighted average is used to fuse the sequencing images from multiple sequencing cycles into a single image. The fused single image and the labels constructed based on the sequencing images from multiple sequencing cycles are then fed into the convolutional neural network for training. The model's prediction result is a probability map of the light spot information corresponding to the fused image of the sequencing images from multiple sequencing cycles.

[0130] During model inference, a single sequencing image is used as input to obtain a probability map of light spot information. The pixel value at each position in this probability map represents the probability that the pixel position belongs to a light spot. Denoising and foreground / background segmentation are performed on the probability map of light spot information. Background pixels are filtered out from the probability map of light spot information, and the probability of pixels located in the foreground is multiplied by a fixed value to obtain the reconstructed sequencing image.

[0131] like Figure 5a As shown, a specific example of a raw sequencing image is illustrated; such as Figure 5b As shown, a specific example of a probability map of light spot information predicted by a model is illustrated; for example... Figure 5c As shown, a specific example of the probability map of filtered spot information is illustrated.

[0132] from Figures 5a-5c As can be seen, the filtered probability map shows clear light spots with less noise, obvious distinction between light spots, and high distinction between light spots and background.

[0133] Regarding the above S402:

[0134] After obtaining multiple reconstructed sequencing images, an alignment template for image alignment can be determined from the reconstructed sequencing images. The alignment template is then used to align the multiple original sequencing images to obtain the aligned sequencing images.

[0135] Specifically, during gene sequencing, it is necessary to align the reconstructed sequencing images obtained from different sequencing cycles at the same imaging position to obtain the sub-pixel offset between the reconstructed sequencing images from different sequencing cycles. Then, an affine matrix is ​​constructed based on the sub-pixel offset, and the original sequencing image is affinely transformed using the affine matrix to complete the alignment of the original sequencing image.

[0136] Specifically, the reconstructed sequencing image from the first sequencing cycle can be selected as the alignment template, or the reconstructed sequencing image with the highest signal-to-noise ratio or high resolution from multiple sequencing cycles can be selected as the alignment template. Then, using the alignment template, the reconstructed sequencing image is registered with the alignment template to construct an affine matrix. Using the affine matrix, multiple frames of original sequencing images are aligned to obtain multiple frames of aligned sequencing images.

[0137] When aligning multiple reconstructed sequencing images, for example, the scheme described in the patent application with application number "CN202310031236.0" entitled "A method, apparatus, device and storage medium for registering fluorescence images" can be used to register the multiple reconstructed sequencing images. The specific registration process will not be described in the embodiments of this disclosure.

[0138] Regarding the above S403:

[0139] In the field of gene sequencing, not all clusters produce fluorescence in each image. A single sequencing image may contain clusters that do not produce fluorescence. Therefore, it is necessary to use sequencing images from more sequencing cycles to help determine the location and shape of the clusters in order to avoid cluster loss.

[0140] In this embodiment of the disclosure, a spot prediction model is used to predict the spots in at least a portion of the sequencing images after multi-frame alignment, and a predicted spot probability map is obtained. Then, spot reconstruction processing is performed based on the spot probability map to obtain a template image, so that the obtained template image includes the spot information of the spots in at least the multi-frame aligned sequencing images. Therefore, cluster loss can be avoided.

[0141] Specifically, when training the spot prediction model, multiple sequencing images aligned after sequencing cycles are used as input. The more sequencing images involved in constructing the template image, the richer the clusters and the clearer the edges of the constructed base template, resulting in a higher image signal-to-noise ratio. Furthermore, the construction of the base template aims to identify the location information of all clusters within the photographed area, accurately extract the cluster intensity values, and perform base signal analysis.

[0142] Because gene sequencing requires real-time processing, too many sequencing images cannot be selected for template image construction. Considering that sequencing images need to be temporarily stored in computer memory or hard drive after acquisition for rapid digital image parsing to meet the real-time requirements of sequencing, the template image can be constructed using sequencing images obtained from 3-10 sequencing cycles.

[0143] By constructing labels based on the relevant information of light spots corresponding to multiple frames of sequencing images from multiple sequencing cycles after alignment, a convolutional neural network is trained to obtain a light spot prediction model. Common convolutional neural networks such as FCN, SegNet, U-Net, and DeepLab can be used.

[0144] In particular, the labeling process for constructing the light spot prediction model is similar to that used in image reconstruction models:

[0145] For example, multiple sequencing images from multiple sequencing cycles can be used as samples to train a spot prediction model. One sequencing image can be selected from the multiple sequencing images as an alignment template, and the remaining sequencing images can be used as images to be aligned. The multiple sequencing images from multiple sequencing cycles can then be aligned. Subsequently, the aligned sequencing images are denoised, and then a segmentation algorithm is used to obtain the regions of interest from the multiple sequencing images from multiple sequencing cycles. Pixels in the regions of interest usually have a high probability of belonging to the region where the spot is located.

[0146] Subsequently, all regions of interest (ROIs) in the multi-frame sequencing images from multiple sequencing cycles can be fused together to obtain a fused region, and the positions of each pixel within the fused region can be obtained. Furthermore, the denoised multi-frame sequencing images are sharpened to obtain the grayscale values ​​of each pixel position in the fused region on each sharpened sequencing image frame. These grayscale values ​​are then used for image segmentation of the fused region to determine the foreground and background within the fused region in each sharpened sequencing image frame. Finally, a base detection method is used to identify the base types of the pixels in the fused region. By traversing multiple sequencing cycles, the base sequence at each pixel position in the fused region can be obtained.

[0147] Next, a consistency comparison is performed between the base sequences of each pixel in the foreground and the base sequences of neighboring pixels, or between the base sequences of each pixel and the base sequences of the reference genome. Pixels with good consistency are retained as candidate spot centroids. These candidate spot centroids are used as seed points for region growing. According to the region growing algorithm, pixels that meet the base sequence similarity requirements are identified as being in the same spot as the seed point, thus obtaining the spot region image.

[0148] Then, the light spots in the image are filtered out, removing either too small or too large spots. After removing abnormally sized spots, a boundary search algorithm is used to calculate the boundaries of each unfiltered spot. The pixels within the spot boundaries are then filled using the spot segmentation boundaries to obtain the filled spot range. The spot centroid is calculated using gray-level weighted averaging of the filled spot range, or directly calculated using the spot boundaries. Non-spot areas are marked as background, and the background value is set to a fixed value. Labels are generated based on different categories of pixels, such as spot centroid, background, spot boundary, and spot region.

[0149] like Figure 6a As shown, a specific example of a spot image is formed after segmenting the foreground and background in the fusion region of a sharpened sequencing image.

[0150] like Figure 6b As shown, a specific example of a spot region image obtained by using a candidate spot centroid as a seed point for region growing according to a region growing algorithm is illustrated.

[0151] like Figure 6c As shown, a specific example is illustrated using a boundary finding algorithm to calculate the image formed by the boundaries of each unfiltered spot.

[0152] like Figure 6d As shown, it illustrates Figure 6c The light spot boundaries shown are superimposed. Figure 6b The image following the image of the light spot region shown, where the green boxes indicate the light spots that were filtered out because their size and shape did not meet the threshold.

[0153] like Figure 6e As shown, a schematic diagram of the superposition of the light spot background, light spot boundary, light spot region and light spot centroid is presented.

[0154] For example, the spot prediction model in this embodiment is a binary classification network, a three-class classification network, or a four-class classification network. Depending on the spot prediction model, binary classification labels (such as spot background and spot centroid) can be selected from spot background, spot boundary, spot region, and spot centroid to train a binary classification network; or three-class labels (background, spot region, and spot centroid) can be selected to train a three-class classification network; or four-class labels (spot background, spot boundary, spot region, and spot centroid) can be selected to train a four-class classification network.

[0155] The specific classification result type of the above-mentioned spot prediction model can be determined according to actual needs, and this disclosure does not limit it.

[0156] Based on the above method, a spot prediction model is trained. This model can then be used to predict spots in at least a portion of the aligned sequencing images, resulting in a predicted spot probability map. The spot probability map includes relevant information about the spots in the aligned sequencing images.

[0157] Furthermore, because DNA clusters are very small, the fluorescence emitted by these clusters, when projected onto the image coordinate system, forms a very small spot, which in some cases may only occupy a few pixels. Therefore, to improve the accuracy of spot recognition and to identify spots from sequencing images with higher accuracy, thereby improving the accuracy of the spot probability map, in another embodiment of this disclosure, after obtaining multiple aligned sequencing images, the aligned sequencing images can be upsampled to obtain multiple upsampled sequencing images. By upsampling the aligned sequencing images, the upsampled sequencing images have higher resolution, allowing the spots to occupy more pixels in the upsampled sequencing images. Thus, when using the spot prediction model to predict the spot of at least a portion of the sequencing images in the multi-frame aligned sequencing images to obtain the predicted spot probability map, the spot prediction model can be used to predict the spot of at least a portion of the sequencing images in the multi-frame upsampled sequencing images to obtain the predicted spot probability map. Since the spot occupies more pixel positions in the upsampled sequencing image, it is possible to identify the spot in the sequencing image with higher accuracy and improve the detection rate of subsequent base sequences.

[0158] Regarding the above S404:

[0159] When the sequenced image after multi-frame alignment is put into the spot prediction model for prediction, a predicted spot probability map containing the probability of spot multi-classification (binary, tri-class, and quadri-classification) is obtained.

[0160] Depending on the specific circumstances of the light spot probability map, the light spot reconstruction process will also vary.

[0161] Solution A: For cases where the light spot probability map includes a binary classification light spot probability map, the binary classification light spot probability map includes light spot region information and background information.

[0162] In this case, for example, the spot probability map can be reconstructed using the following method to obtain the template image:

[0163] The light spot probability map is segmented into foreground and background and then thresholded to obtain a first light spot map; wherein the foreground is the light spot region;

[0164] Based on the intensity value of the spot region in the first spot image, the centroid of the spot is determined from the first spot image;

[0165] Using the centroid of the light spot as the center, the light spot to which the centroid belongs is reconstructed using an objective function to obtain the template image; the objective function includes: a point spread function or a two-dimensional Gaussian function.

[0166] Here, when upsampling the aligned sequencing image, it is necessary to first downsample the centroid of the spot, and then use the centroid of the downsampled spot as the center to reconstruct the spot using the objective function, thereby obtaining the template image, so as to facilitate the subsequent registration of the template image with the original sequencing image and base detection.

[0167] In practical implementation, when segmenting the light spot probability map into foreground and background, segmentation algorithms such as OTSU, K-means, and support vector machines can be used to segment the light spot probability map. This process determines the pixels belonging to the light spot region and the pixels belonging to the background. Then, the probability value of the pixels belonging to the foreground is multiplied by a fixed value to obtain the intensity value of each pixel in the first light spot map. Finally, the centroid of the light spot can be determined based on the intensity values ​​of the pixels in the light spot region in the first light spot map.

[0168] Specifically, in this embodiment of the present disclosure, the centroid of the light spot can be determined from the first light spot image in the following manner:

[0169] The pixels belonging to the foreground in the first light spot image are traversed, and it is determined whether all other pixels in the target neighborhood corresponding to the traversed pixel belong to the foreground and whether the intensity value of the traversed pixel is greater than the gray value of the other pixels in the corresponding target neighborhood. If all other pixels in the target neighborhood corresponding to the traversed pixel belong to the foreground and the gray value of the traversed pixel is greater than the gray value of the other pixels in the corresponding target neighborhood, the traversed pixel is determined as the centroid of the light spot.

[0170] Here, we can directly traverse all foreground pixels in the first light spot image and determine other pixels within the target neighborhood corresponding to the traversed pixel based on its position. We then check if all pixel values ​​within the target neighborhood are the first value. If all pixel values ​​within the target neighborhood are the first value, then all other pixels within the target neighborhood belong to the foreground. If there are pixels with second values ​​within the target neighborhood, further checks are unnecessary. Finally, we classify the traversed pixel as a non-centroid pixel.

[0171] After determining that all other pixels in the target neighborhood of a given pixel belong to the foreground, the grayscale value of that pixel can be compared with the grayscale values ​​of other pixels in the target neighborhood. If the grayscale value of that pixel is the largest among all pixels in the target neighborhood, then that pixel can be determined as the centroid of the light spot.

[0172] Alternatively, a specific region, such as the central region, can be defined within the area where the light spot is located, or a region smaller than the light spot region can be defined based on the boundary of the light spot; the pixels in this region can be defined as the pixels to be traversed.

[0173] Other pixels in the target neighborhood of the pixel to be traversed are usually foreground pixels.

[0174] Here, the target neighborhood includes, for example, a four-neighbor neighborhood, an eight-neighbor neighborhood, or a region whose distance to the pixel to be traversed is less than a preset distance threshold. The specific neighborhood can be determined according to actual needs, and this embodiment does not impose any limitations.

[0175] After determining the pixel that belongs to the centroid of the light spot, the light spot to which the centroid belongs can be reconstructed using the pixel that is the centroid of the light spot as the center, and the point spread function or the Gaussian two-dimensional function can be used to obtain the template image.

[0176] In addition, in another embodiment of this disclosure, after obtaining the pixel of the centroid of the light spot, Gaussian fitting can be performed in the neighborhood corresponding to the pixel of the centroid of the light spot to determine the sub-pixel of the centroid of the light spot.

[0177] Using the centroid of the light spot as the center and employing an objective function, the light spot to which the centroid belongs is reconstructed to obtain the template image, including:

[0178] Using the sub-pixel of the centroid of the light spot as the center, the light spot to which the centroid belongs is reconstructed using the objective function to obtain the template image.

[0179] In practical implementation, determining the sub-pixel points of the centroid of the light spot allows for more precise determination of the centroid. When reconstructing the light spot belonging to the centroid using these sub-pixel points as the center, reconstruction can be performed at the sub-pixel level. This results in a template image that more accurately describes the spot's position within the sequencing image. Furthermore, using the template image to perform base identification on multiple frames of the original sequencing image improves the accuracy of base identification, leading to more accurate base sequences.

[0180] When multiple registered images from multiple sequencing cycles are fed into a convolutional neural network for prediction, the result is a template prediction image containing probabilities for various light spot classifications (binary, tri-, and quadri-classification). Further processing of this template prediction image is needed to find the precise centroid of the light spot and construct an accurate template image. For example... Figure 7 The examples shown illustrate specific examples of multiple registered images.

[0181] Since the shape and boundary of the light spot affect its centroid position, multi-class training is more accurate than training using only the light spot centroid when training the template-based model. Because template construction first requires finding the accurate light spot centroid, for ease of calculation, we can simply process the probability map corresponding to the light spot centroid.

[0182] This disclosure also provides a specific method for reconstructing the spot to which the centroid of the spot belongs using an objective function.

[0183] Since the light spot ideally follows a Gaussian distribution, the light spot can be reconstructed using a two-dimensional Gaussian function with the centroid of the filtered light spot as the center.

[0184] Furthermore, when reconstructing the spot to which the centroid belongs using a Gaussian two-dimensional function centered on the spot's centroid, the following method can be used, for example:

[0185] Using the centroid of the light spot as the center of a pixel or subpixel, and determining the light spot region based on the size of the light spot, the position coordinates of the pixels within the light spot region and their approximate standard deviation are substituted into the two-dimensional Gaussian function to obtain the reconstructed grayscale values ​​of the pixels within the light spot region; wherein, the standard deviation of the two-dimensional Gaussian function is determined based on the full width and half height of the sequencing image.

[0186] Curve fitting was performed on the spot area using a Gaussian model. Since the brightness of the area farther away from the center of the spot is smaller in the Gaussian model, and the gray value of the dark area around the center of the spot is not zero due to the presence of background noise, the background noise around the spot is taken into account in the function corresponding to the Gaussian model, as shown in equation (1) below:

[0187]

[0188] Where A represents the amplitude, which can generally be set to a fixed value; x and y represent the pixel coordinates corresponding to the spot area; x0 and y0 represent the coordinates of the center point of the spot area; σ x ,σ yThis represents the standard deviation of the Gaussian function. This standard deviation is the standard deviation of the grayscale values ​​of the pixels in the x and y directions of the spot region; G represents the background, which can generally be set to a fixed value, and g(x,y) represents the grayscale value corresponding to the pixel at coordinates (x,y).

[0189] Assuming the size of the light spot in the sequencing image is Size, choose a Gaussian function of size Size*Size;

[0190] Using the centroid of the light spot as the center of the Gaussian function, the pixel coordinates within the range of (Size)*(Size) are substituted into the two-dimensional Gaussian function to calculate the numerical value, which is then labeled as the reconstructed gray value of each pixel. The gray values ​​at the same pixel position are accumulated. Two-dimensional Gaussian function reconstruction is performed on all light spots, and the reconstructed image is the template image, such as... Figure 8 As shown, this disclosure provides a specific example of a template image.

[0191] The standard deviation of the Gaussian function is determined based on the full width at half maximum (FWHM) of the sequencing image.

[0192] Full width at half maximum (FWHM), also known as half-maximum full width, half-peak full width, or half-height width, refers to the distance between two points whose function values ​​are equal to half the peak value in a Gaussian function obtained by fitting a Gaussian to a light spot. It serves as an indicator of the degree of energy concentration of a light spot; the more concentrated the energy of a single light spot, the smaller its full width at half maximum.

[0193] During the imaging process, the intensity of the light spots in different regions of the sequencing image varies; for example, the intensity of the light spots is higher in the central part of the sequencing image, while the intensity of the light spots in the edge regions of the sequencing image is attenuated to some extent compared to the intensity of the light spots in the central region. Therefore, in order to better reconstruct the light spots in different regions, in another embodiment of this disclosure, the multi-frame aligned sequencing image can be divided into multiple image blocks.

[0194] During sequencer assembly, the method disclosed in patent application CN115937324B can be used to test the camera's functionality. Therefore, before sequencing, the FWHM of different image patches can be directly obtained from the output files of the testing process. The FWHM is then converted into an approximate standard deviation as the empirical standard deviation value for different image patches. Each spot in different image patches is then reconstructed to obtain a template image.

[0195] For any image patch, its corresponding FWHM satisfies, for example, the following formula:

[0196]

[0197] σ x =σ y =σ

[0198] Where σ represents the standard deviation.

[0199] Alternatively, a two-dimensional Gaussian curve can be fitted to the light spots in the multi-cycle aligned image input to the light spot prediction model, and the standard deviation of the light spots can be calculated directly.

[0200] When reconstructing the light spot to which the centroid of the light spot belongs using a two-dimensional Gaussian function with the centroid of the light spot as the center, for example, the light spot to which the centroid of the light spot belongs in the multiple image blocks is reconstructed using a two-dimensional Gaussian function with the centroid of the light spot as the center.

[0201] Specifically, taking the division of a multi-frame aligned sequencing image into multiple image blocks as an example, this disclosure provides a specific method for performing Gaussian fitting on the light spot to obtain the standard deviation of the light spot:

[0202] Gaussian fitting is performed sequentially on the light spots in the image patch. The standard deviation of the corresponding image patch is determined based on the Gaussian fitting results of each light spot. This includes: sequentially performing Gaussian fitting on each light spot in each image patch to obtain the standard deviation of each light spot in the horizontal and vertical directions; and using the median of the standard deviations in the horizontal and vertical directions of all light spots in a single image patch as the standard deviation of that single image patch in the horizontal and vertical directions. For example, assuming that an image patch includes m light spots, curve fitting is performed sequentially on each of the m light spots to obtain m standard deviations in the horizontal and vertical directions of the image patch; the median of the horizontal standard deviation is selected from the m horizontal standard deviations, and the median of the vertical standard deviation is selected from the m vertical standard deviations. The median of these two standard deviations is used as the standard deviation of the image patch in the horizontal and vertical directions.

[0203] After multi-frame alignment, sequencing images exhibit multiple horizontal and vertical standard deviations at the same image patch location. Each horizontal and vertical standard deviation corresponds to an image patch at that location within the aligned sequencing image. For each image patch location, the average of the multiple horizontal standard deviations is used to obtain the horizontal standard deviation for that location. Similarly, the average of the multiple vertical standard deviations for the same image patch location is used to obtain the vertical standard deviation. This process is repeated for all image patch locations to obtain the standard deviation for all image patch locations.

[0204] In addition, there may be cases where the distance between the centroids of the light spots is relatively close. If the distance between the centroids of two light spots is smaller than the light spot size used for two-dimensional Gaussian function light spot reconstruction, an overlapping area will appear between the light spots corresponding to the centroids of the two light spots. This will result in multiple reconstructed gray values ​​at the pixel positions in the overlapping area. For the pixel positions in the overlapping area, the multiple reconstructed gray values ​​corresponding to the pixel positions in the overlapping area can be accumulated to obtain the reconstructed light spot.

[0205] Option B: For cases where the light spot probability map includes multiple different categories of light spot probability maps, any category of light spot probability map includes at least one of the following: light spot centroid, light spot boundary, and light spot region, and includes background information that does not belong to the light spot.

[0206] In this case, for example, the spot probability map can be reconstructed using the following method to obtain the template image:

[0207] By overlaying multiple light spot probability maps of different categories, a multi-category overlaid light spot probability map is obtained.

[0208] The image reconstruction model is used to perform image reconstruction processing on the multiple classification superimposed light spot probability maps to obtain a reconstructed image;

[0209] The reconstructed image is subjected to denoising and threshold segmentation to obtain the template image.

[0210] Then, the probability map of the superimposed light spots (excluding the background) can be multiplied by a fixed value and input into a pre-trained image reconstruction model to obtain the reconstructed image.

[0211] When the aligned sequencing image is upsampled and input into the spot prediction model, in order to use the template image to perform base identification processing on the sequencing image, after the spot probability map is reconstructed to obtain the template image, the template image can also be downsampled to restore the resolution of the template image to be consistent with the multi-frame aligned sequencing image, so as to ensure that the template image and the sequencing image can be aligned in size, which facilitates the base detection process.

[0212] Here, the image reconstruction model refers to the image reconstruction model used in the above embodiments when reconstructing the original sequencing images derived from multiple sequencing images. The training method of this image reconstruction model can be found in the above embodiments, and will not be repeated here.

[0213] like Figure 8 As shown, a specific example of a template image constructed using scheme A is illustrated.

[0214] Regarding the aforementioned S405:

[0215] When performing base identification processing on the multiple frames of original sequencing images based on the template image to obtain the base sequence, the following method can be used, for example:

[0216] Using the template image as a reference template, the intensity value corresponding to the centroid position of the light spot indicated in the template image is extracted from the original sequencing images corresponding to each sequencing cycle, based on the centroid position of the light spot indicated in the template image.

[0217] Based on the intensity value corresponding to the centroid position of the light spot indicated in the extracted template image, the base type corresponding to the centroid position of the light spot indicated in the template image is identified.

[0218] In practice, for example, the template image and the original sequencing images corresponding to each sequencing cycle can be registered first to obtain the affine transformation matrix between the template image and the original sequencing images.

[0219] Here, the method for registering the template image and the original measurement image can be referred to, for example, the scheme described in the patent application with application number "CN202310031236.0" entitled "A method, apparatus, device and storage medium for registering fluorescence images", and will not be repeated here.

[0220] Subsequently, based on the affine transformation matrix, the centroid position of the light spot indicated in the template image is projected onto the original sequencing image to obtain the projection point of the centroid position in the original sequencing image. The intensity value of the projection point in the original sequencing image is then determined as the intensity value corresponding to the centroid position of the light spot. Alternatively, in another embodiment of this disclosure, all pixels in the template image can be projected onto the original sequencing image based on the affine transformation matrix to obtain a projected image. Then, the pixels belonging to the centroid of the light spot are determined from the projected image, and the intensity value corresponding to the centroid position of the light spot is obtained based on the intensity value of the pixels belonging to the centroid of the light spot in the original sequencing image.

[0221] Furthermore, in another embodiment of this disclosure, an affine transformation can be performed on each pixel in the original sequencing image according to the affine transformation matrix to project the original sequencing image onto the template image, thereby obtaining a projected image. Then, the pixels belonging to the centroid of the light spot are determined according to the template image, and the intensity value corresponding to the centroid position is obtained based on the intensity value of the pixels belonging to the centroid of the light spot in the projected image.

[0222] Compared to the latter two methods, the first method saves the computational cost of registering each frame of the original sequencing image, thus saving more time.

[0223] After extracting the intensity values ​​belonging to the centroid of the light spot from each frame of the original sequencing image according to the pixel position of the centroid of the light spot indicated in the template image, the base detection is performed using clustering algorithms such as k-means and Gaussian mixture distribution based on the intensity values ​​belonging to the centroid of the light spot, so as to obtain the base sequence corresponding to each light spot in multiple sequencing cycles.

[0224] This disclosure provides a scheme for constructing template images for sequencing images based on deep learning. The model is trained offline using multiple types of sequencing images, which can adapt to sequencing problems such as fluorescence noise, small fluorescence spots, cluster overlap, background noise caused by liquid reagent reactions, and deviations caused by mechanical reciprocating motion. It can find a more accurate spot centroid, construct a more accurate template image, improve the base detection rate and accuracy, and has higher reliability.

[0225] The following analysis compares existing image analysis methods (neither image reconstruction nor base template construction is based on AI models) and the image analysis method provided in this embodiment (integrating image reconstruction and base template construction models). Existing image analysis methods are labeled SIA, while the image analysis method provided in this embodiment is labeled AI.SIA. The SIA process is replaced with AI.SIA, and tests are conducted using phage libraries and whole-genome sequencing (WGS) libraries to compare changes in data accuracy.

[0226] The accuracy of phage library testing is shown in Table 1 below:

[0227] Table 1

[0228]

[0229]

[0230] The accuracy of the WGS library test is shown in Table 2 below:

[0231] Table 2

[0232]

[0233] Based on phage library and WGS library accuracy tests, the method provided in this disclosure can significantly improve sequencing accuracy.

[0234] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0235] This disclosure also provides a computer device, such as... Figure 9 The diagram shown is a schematic representation of a computer device structure provided in an embodiment of this disclosure, including:

[0236] A processor 91 and a memory 92; the memory 92 stores machine-readable instructions executable by the processor 91, and the processor 91 executes the machine-readable instructions stored in the memory 92. When the machine-readable instructions are executed by the processor 91, the processor 91 performs the following steps:

[0237] Image reconstruction models were used to reconstruct multiple frames of raw sequencing images from multiple sequencing cycles, resulting in reconstructed sequencing images.

[0238] An alignment template for image alignment is determined from the reconstructed sequencing images of the multiple frames. The alignment template is then used to align the original sequencing images of the multiple frames to obtain the aligned sequencing images of the multiple frames.

[0239] A spot prediction model is used to predict the spot patterns in at least a portion of the sequenced images in the multi-frame aligned sequenced images to obtain a predicted spot probability map, which contains information about the spots in the aligned sequenced images.

[0240] The light spot probability map is subjected to light spot reconstruction processing to obtain a template image;

[0241] Based on the template image, base identification is performed on the multiple frames of original sequencing images to obtain the base sequence.

[0242] The aforementioned memory 92 includes a main memory 921 and an external memory 922; the main memory 921, also known as internal memory, is used to temporarily store the computational data in the processor 91, as well as the data exchanged with external memory 922 such as a hard disk. The processor 91 exchanges data with the external memory 922 through the main memory 921.

[0243] The specific execution process of the above instructions can be referred to the steps of the sequencing image processing method described in the embodiments of this disclosure, and will not be repeated here.

[0244] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the sequencing image processing method described in the above-described method embodiments. The storage medium may be a volatile or non-volatile computer-readable storage medium.

[0245] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the sequencing image processing method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.

[0246] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0247] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. A sequencing image processing method, characterized in that, include: Image reconstruction models were used to reconstruct multiple frames of raw sequencing images from multiple sequencing cycles, resulting in reconstructed sequencing images. An alignment template for image alignment is determined from the reconstructed sequencing images of the multiple frames. The alignment template is then used to align the original sequencing images of the multiple frames to obtain the aligned sequencing images of the multiple frames. A spot prediction model is used to predict the spot patterns in at least a portion of the sequenced images in the multi-frame aligned sequenced images to obtain a predicted spot probability map, which contains information about the spots in the aligned sequenced images. The light spot probability map is subjected to light spot reconstruction processing to obtain a template image; Based on the template image, base identification is performed on the multiple frames of original sequencing images to obtain the base sequence.

2. The method according to claim 1, characterized in that, The light spot probability map is a binary classification light spot probability map, which includes light spot region information and background information. The step of performing spot reconstruction processing on the spot probability map to obtain a template image includes: The light spot probability map is segmented into foreground and background and then thresholded to obtain a first light spot map; wherein the foreground is the light spot region; Based on the intensity value of the spot region in the first spot image, the centroid of the spot is determined from the first spot image; Using the centroid of the light spot as the center, the light spot to which the centroid belongs is reconstructed using an objective function to obtain the template image; The objective function includes: a point spread function or a two-dimensional Gaussian function.

3. The method according to claim 2, characterized in that, The step of determining the centroid of the light spot from the first light spot image based on the intensity value of the light spot region in the first light spot image includes: The pixels belonging to the foreground in the first light spot image are traversed, and it is determined whether all other pixels in the target neighborhood corresponding to the traversed pixel belong to the foreground and whether the intensity value of the traversed pixel is greater than the gray value of the other pixels in the corresponding target neighborhood. If all other pixels in the target neighborhood corresponding to the traversed pixel belong to the foreground and the gray value of the traversed pixel is greater than the gray value of the other pixels in the corresponding target neighborhood, the traversed pixel is determined as the centroid of the light spot. or, From the foreground pixels in the first light spot image, determine the pixels to be traversed; wherein, all pixels located within the target neighborhood of the pixel to be traversed are foreground pixels; traverse the pixel to be traversed, and compare the gray value of the traversed pixel with the gray value of other pixels in the corresponding target neighborhood; if the intensity value of the traversed pixel is greater than the gray value of other pixels in the corresponding target neighborhood, determine the traversed pixel as the centroid of the light spot.

4. The method according to claim 3, characterized in that, The method further includes: Using the pixel at the centroid of the light spot as the center, Gaussian fitting is performed in the neighborhood corresponding to the pixel to determine the sub-pixel of the centroid of the light spot. The step of reconstructing the light spot to its centroid using an objective function, with the centroid of the light spot as the center, to obtain the template image, includes: Using the sub-pixel of the centroid of the light spot as the center, the light spot to which the centroid belongs is reconstructed using the objective function to obtain the template image.

5. The method according to any one of claims 2-4, characterized in that, Using the centroid of the light spot as the center, a two-dimensional Gaussian function is used to reconstruct the light spot to which the centroid belongs, including: Using the centroid of the light spot as the center of a pixel or subpixel, and determining the spot region based on the size of the light spot, the position coordinates of the pixels within the spot region and their approximate standard deviation are substituted into the two-dimensional Gaussian function to obtain the reconstructed grayscale values ​​of the pixels within the spot region; wherein, the standard deviation of the two-dimensional Gaussian function is determined based on the full width and half height of the sequencing image. The reconstructed gray values ​​of pixels at the same position within the spot area in different sequencing cycles are summed to obtain the reconstructed spot.

6. The method according to claim 5, characterized in that, The multi-frame aligned sequencing image comprises multiple image blocks; The process of reconstructing the light spot to which the centroid of the light spot belongs, using a two-dimensional Gaussian function with the centroid of the light spot as the center, includes: Using the centroid of the light spot in the plurality of image blocks as the center, a two-dimensional Gaussian function is used to reconstruct the light spot to which the centroid of the light spot in the plurality of image blocks belongs, thereby obtaining the template image corresponding to the plurality of image blocks respectively; The approximate standard deviation used when reconstructing the centroid of the light spot in different image blocks using a two-dimensional Gaussian function is determined based on the full width and half height of each image block.

7. The method according to claim 1, characterized in that, The light spot probability map includes: multiple different categories of light spot probability maps; any category of light spot probability map includes at least one of the following: light spot centroid, light spot boundary, and light spot region, as well as background information; The step of performing spot reconstruction processing on the spot probability map to obtain a template image includes: By overlaying multiple light spot probability maps of different categories, a multi-category overlaid light spot probability map is obtained. The image reconstruction model is used to perform image reconstruction processing on the multiple classification superimposed light spot probability maps to obtain a reconstructed image; The reconstructed image is subjected to denoising and threshold segmentation to obtain the template image.

8. The method according to claim 1, characterized in that, The method further includes: performing upsampling processing on the multi-frame aligned sequencing images to obtain multi-frame upsampled sequencing images; The step of using a spot prediction model to predict spot patterns in at least a portion of the sequenced images in the multi-frame aligned sequencing images, to obtain a predicted spot probability map, includes: A spot prediction model is used to predict spot patterns in at least a portion of the sequencing images in the multi-frame upsampled sequencing images, resulting in a predicted spot probability map. After performing spot reconstruction processing on the spot probability map to obtain the template image, the method further includes: The template image is downsampled to restore its resolution to match that of the multi-frame aligned sequencing image.

9. A computer device, characterized in that, include: The processor and the memory, the memory storing machine-readable instructions executable by the processor, the processor executing the machine-readable instructions stored in the memory, wherein when the machine-readable instructions are executed by the processor, the processor performs the steps of the sequencing image processing method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer device, performs the steps of the sequencing image processing method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • A fluorescence image registration method, device, equipment and storage medium

    CN115937282B

  • Assembly quality evaluation method, device, equipment and storage medium

    CN115937324B