Method for identifying fluorescent light spots in fluorescent image, computer equipment and computer readable storage medium

By performing image alignment and region growth processing on fluorescent images in high-throughput gene sequencing, accurately identifying and labeling fluorescent spots, the problem that traditional methods cannot identify small, dark, irregular fluorescent spots is solved, and the accuracy of sequencing data is improved.

CN119963536APending Publication Date: 2025-05-09SIKUN LIFE SCIENCE CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510121905.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In high-throughput gene sequencing, the fluorescent spots on fluorescent images are small in size, dark in shape, and irregular in shape, and traditional image annotation tools cannot accurately identify and label fluorescent spots.

Method used

By performing image alignment on multiple frames of fluorescence images of multiple sequencing cycles, the base sequence of the spot candidate pixel positions is determined, and the pixel position of the spot center region is determined based on the consistency information of the base sequence, and the spot region is then identified using the region growth algorithm.

Benefits of technology

Accurate identification and labeling of fluorescent spots is achieved, the clarity of the spot edges and image signal-to-noise ratio are improved, and the accuracy of base signal analysis is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963536A_ABST
    Figure CN119963536A_ABST
Patent Text Reader

Abstract

The invention provides a method for recognizing fluorescent light spots in a fluorescent image, computer equipment and a computer readable storage medium, and the method comprises the steps: carrying out the image alignment of a plurality of frames of original fluorescent images from a plurality of sequencing cycles, and obtaining a plurality of frames of aligned fluorescent images; determining a base sequence corresponding to a light spot candidate pixel position according to a gray value of each pixel position in the aligned multi-frame fluorescence image; the light spot candidate pixel positions comprise at least part of pixel positions in the fluorescence image; on the basis of consistency information between the base sequence of the light spot candidate pixel position and the base sequence of the light spot candidate pixel position in a first neighborhood corresponding to the light spot candidate pixel position, determining a first pixel position belonging to a fluorescent light spot central area from the light spot candidate pixel position; and taking at least part of the first pixel positions as seed points, and performing region growth processing on the seed points to obtain a light spot region of a light spot where the seed points are located.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of gene sequencing technology, and in particular to a method for identifying fluorescent spots in fluorescent images, a computer device, and a computer-readable storage medium. Background Art

[0002] During the synthesis sequencing process, since not all clusters after library amplification produce fluorescent spots each time the optical detection system collects fluorescent images, there may be clusters that do not produce fluorescent spots in a single-frame fluorescent image. Therefore, it is necessary to use more frames of fluorescent images from more sequencing cycles to assist in judging the spot information of all clusters (such as spot position, spot shape, and spot centroid) to avoid losing the spot information of the cluster.

[0003] In order to obtain accurate cluster spot information, the model can be trained with multiple frames of fluorescence images of multiple sequencing cycles and labeled fluorescence spots as input. The model will output a probability map, which predicts the spot information of all possible clusters. In principle, the more fluorescence images are involved in model training or model prediction, the more accurate the trained model is, the richer the spots are in the probability map output by the model, the clearer the edges of the spots are, and the higher the image signal-to-noise ratio is. Through the probability map, the spot information of all clusters can be accurately identified. Further, when parsing the base signal, the spot intensity values ​​at all cluster positions can be accurately extracted and the parsing can be completed.

[0004] However, the fluorescent spots on the fluorescence images obtained under high-throughput gene sequencing are small in size (generally ranging from 1 to 10 pixels), dark in shape, irregular in shape, and with unclear boundaries. Traditional image annotation tools cannot meet the requirements of marking fluorescent spots in fluorescence images under gene sequencing. Therefore, a method that can accurately identify and mark fluorescent spots from fluorescent images has become a problem that needs to be solved urgently. Summary of the invention

[0005] The embodiments of the present disclosure at least provide a method for identifying a fluorescent spot in a fluorescent image, a computer device, and a computer-readable storage medium.

[0006] In a first aspect, an embodiment of the present disclosure provides a method for identifying fluorescent spots in a fluorescent image, characterized in that it includes: performing image alignment on multiple frames of original fluorescent images derived from multiple sequencing cycles to obtain an aligned fluorescent image; determining the base sequence of a candidate spot pixel position among the multiple pixel positions according to the grayscale value of each pixel position in the aligned multi-frame fluorescent image; determining a first pixel position belonging to the central area of ​​the fluorescent spot from the candidate spot pixel positions based on consistency information between the base sequence of the candidate spot pixel position and the base sequence of the candidate spot pixel position in its corresponding first neighborhood; using at least part of the first pixel position as a seed point, performing region growing processing on the seed point, and obtaining a spot area of ​​the spot where the seed point is located.

[0007] Optionally, the image alignment of multiple frames of original fluorescence images derived from multiple sequencing cycles includes: performing image reconstruction processing on the multiple frames of original fluorescence images derived from multiple sequencing cycles to obtain multiple frames of reconstructed fluorescence images; determining an alignment template for image alignment from the multiple frames of reconstructed fluorescence images; and performing image alignment processing on the multiple frames of original fluorescence images based on the alignment template to obtain multiple frames of aligned fluorescence images.

[0008] Optionally, the method of determining the base sequence of the candidate spot pixel positions in the multiple pixel positions according to the grayscale value of each pixel position in the aligned multi-frame fluorescence images includes: performing segmentation processing on the aligned multi-frame fluorescence images based on the grayscale value of each pixel point in the aligned fluorescence images to obtain the candidate fluorescence spot area in the aligned multi-frame fluorescence images; and performing fusion processing on the candidate fluorescence spot area in the aligned multi-frame fluorescence images to obtain the fused fluorescence spot area in the aligned multi-frame fluorescence images; and performing image sharpening processing on the aligned multi-frame fluorescence images to obtain the sharpened fluorescence image; taking each pixel position located in the fused fluorescence spot area as the candidate spot pixel position, and determining the base sequence corresponding to the candidate spot pixel position according to the grayscale value of the candidate spot pixel position in the sharpened fluorescence image.

[0009] Optionally, the following method is used to determine the consistency information between the base sequence of the light spot candidate pixel position and the base sequence of the pixel position in its corresponding first neighborhood: determine the first neighborhood of each of the light spot candidate pixel positions; determine the degree of difference between the base sequence of each of the light spot candidate pixel positions and the base sequences of each light spot candidate pixel position in its corresponding first neighborhood; the degree of difference includes: the number of different bases between the base sequence of each of the light spot candidate pixel positions and the base sequence of the light spot candidate pixel position in its corresponding first neighborhood, or the percentage of the number of different bases in the base sequence of the light spot candidate pixel position; determine the consistency information based on the degree of difference.

[0010] Optionally, taking at least part of the first pixel position as a seed point, performing region growing processing on the seed point to obtain a spot region of the spot where the seed point is located, includes: executing multiple iteration cycles, and performing the following region growing process in each iteration cycle: determining a target seed point corresponding to a current iteration cycle from a first pixel position to which the light spot has not been determined; determining a candidate light spot pixel position within a second neighborhood of the target seed point based on the position information of the target seed point; performing region growing processing on the target seed point based on the similarity between the base sequence of the target seed point and the base sequence of the candidate light spot pixel position within its corresponding second neighborhood to obtain a spot region of the light spot where the target seed point is located.

[0011] Optionally, the region growing process is performed on the target seed point based on the similarity between the base sequence of the target seed point and the base sequence of the pixel position in the corresponding second neighborhood to obtain the spot region of the spot where the target seed point is located, including: determining a second pixel position belonging to the same spot region as the target seed point according to the similarity between the base sequence of the target seed point and the base sequence of the candidate pixel position of the spot in the corresponding second neighborhood, and determining the growable pixel position corresponding to the target seed point from the second pixel position; looping the following steps until no new growable pixel position appears: determining a new second pixel position belonging to the same spot region as the target seed point by using the similarity between the base sequence of the growable pixel position and the base sequence of the candidate pixel position of the spot in the corresponding third neighborhood; and determining a new growable pixel position corresponding to the target seed point in the new second pixel position; and constituting the spot region based on the target seed point and the second pixel position.

[0012] Optionally, the method utilizes the similarity between the base sequence of the growable pixel position and the base sequence of the pixel position in the corresponding third neighborhood to determine a new second pixel position belonging to the same light spot area as the target seed point; and determines a new growable pixel position corresponding to the target seed point in the new second pixel position, including: comparing the similarity with a first similarity threshold; if the similarity is greater than the first similarity threshold, determining the light spot candidate pixel position in the second neighborhood as the second pixel position; and in a case where the similarity is greater than the first similarity threshold, comparing the similarity with a second similarity threshold; if the similarity is greater than the second similarity threshold, determining the corresponding second pixel position as the growable pixel position corresponding to the target seed point.

[0013] Optionally, the method further includes: determining whether the second pixel position belongs to a candidate center set consisting of the first pixel positions; if the second pixel position belongs to the candidate center set, deleting the second pixel position from the candidate center set.

[0014] Optionally, the method also includes: determining a spot boundary of the spot area, and / or determining a spot centroid of the spot area based on position information of each pixel position in the spot area; generating a spot label of the original fluorescence image based on the spot boundary and / or the spot centroid, wherein the original fluorescence image with the spot label is used to train a model, and the trained model is used to predict the spot information of all clusters participating in the sequencing reaction.

[0015] Optionally, after determining the spot boundary of the spot area, and / or determining the spot centroid of the spot area according to the position information of each pixel position in the spot area, the method further includes: determining shape characteristic information of the spot area; the shape characteristic information includes at least one of the following: spot size, aspect ratio of the spot, area ratio of the effective area in the spot, whether the spot centroid is within the spot, and the distance between other spot centroids; when the confidence of the corresponding spot represented by the shape characteristic information is less than a preset confidence threshold, the spot area is screened out.

[0016] Optionally, after determining the spot boundary of the spot area and / or determining the spot centroid of the spot area according to the position information of each pixel position in the spot area, the method further includes: matching the base sequence corresponding to the spot centroid with a reference gene sequence to obtain the difference information between the base sequence corresponding to the spot centroid and the reference gene sequence; and screening out the spot area when the difference information is greater than or equal to a preset difference threshold.

[0017] In a second aspect, an optional implementation of the present disclosure further provides a computer device, a processor, and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the processor is used to execute the machine-readable instructions stored in the memory, and when the machine-readable instructions are executed by the processor, the machine-readable instructions perform the steps in the above-mentioned first aspect, or any possible implementation of the first aspect.

[0018] In a third aspect, an optional implementation of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, the steps of the above-mentioned first aspect, or any possible implementation of the first aspect are executed.

[0019] The disclosed embodiment performs base recognition on the fluorescent image to be marked according to the pixel position, selects the pixel position with high consistency with the field pixel position as the candidate spot center, and determines the pixel position of the base sequence that meets the pixel requirements as being located in the same spot as the spot center according to the region growing algorithm, thereby obtaining the spot boundary, completing the segmentation of the spot, and obtaining a spot area with higher accuracy.

[0020] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following is a brief introduction to the drawings required for use in the embodiments. The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can also be obtained based on these drawings without creative work.

[0022] Figure 1 A schematic diagram of a sequencing system provided by some embodiments of the present disclosure is shown;

[0023] Figure 2 A schematic diagram of a sequencing chip provided in some embodiments of the present disclosure is shown;

[0024] Figure 3 A flow chart showing a method for identifying a fluorescent spot in a fluorescent image provided by some embodiments of the present disclosure is shown;

[0025] Figure 4A flowchart showing a specific method for obtaining an original fluorescent image provided by some embodiments of the present disclosure;

[0026] Figure 5 One example of base sequences corresponding to each pixel position provided by some embodiments of the present disclosure is shown;

[0027] Figure 6 The second example of the relative position relationship between the pixel position N and its corresponding first neighborhood pixel position N' provided in some embodiments of the present disclosure is shown;

[0028] Figure 7 A flowchart showing a specific method of determining seed points and performing region generation processing on the seed points provided in some embodiments of the present disclosure;

[0029] Figure 8 The third specific example of four neighborhoods and eight neighborhoods involved in some embodiments of the present disclosure is shown;

[0030] Fig. 9 The fourth example of the spot region obtained by performing region growing processing using the second neighborhood pixel positions of four neighborhoods and eight neighborhoods respectively involved in some embodiments of the present disclosure is shown;

[0031] Fig.10 The fifth example of performing region growing processing using the second neighborhood pixel position of four neighborhoods provided by some embodiments of the present disclosure is shown;

[0032] Fig.11 A flowchart showing a specific example of performing region growing processing on seed points provided by some embodiments of the present disclosure;

[0033] Fig.12 The sixth example of the light spot label provided by some embodiments of the present disclosure is shown;

[0034] Fig.13 A schematic diagram of a computer device provided by some embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0035] In order to make the purpose, technical scheme and advantages of the embodiments of the present disclosure clearer, the technical scheme in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. The components of the embodiments of the present disclosure generally described and shown here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the present disclosure claimed for protection, but merely represents the selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present disclosure.

[0036] To facilitate the understanding of the technical solution of the present disclosure, the technical terms in the embodiments of the present disclosure are first explained:

[0037] Library construction:

[0038] The genomic DNA or RNA molecules to be sequenced are broken by physical or chemical means, for example, by ultrasound, to form DNA or RNA fragments. The two ends of the DNA or RNA fragments are first filled with enzymes, and then the two ends of the fragments are connected to a specific DNA or RNA sequence (usually, this specific DNA sequence or RNA is also called a linker) with a specific enzyme to form a mixture of DNA or RNA. This mixture of DNA or RNA is also called a library in the industry.

[0039] In order to save sequencing costs, generally, multiple samples will be sequenced in the sequencer at the same time. In order to distinguish the sequencing results of different samples, when preparing libraries of different samples, the connectors will contain a DNA or RNA sequence (usually containing 6-8 bases) that can identify the source of the sample. This DNA or RNA sequence that can identify the source of the sample can also be called a sample tag (or Index, Barcode). It can be understood that each sample library connector contains its own sample tag.

[0040] Typically, library construction is done outside of the sequencer, for example, by experimental manipulation in the laboratory to obtain the library.

[0041] Amplification reaction:

[0042] Taking the amplification of a DNA library by bridge PCR (polymerase chain reaction) as an example, after the library is constructed, the library can be inoculated onto a sequencing chip and amplified on the sequencing chip. The adapters at both ends of the library are complementary to the first amplification primer on the sequencing chip, so the library can be inoculated onto the sequencing chip through complementary hybridization.

[0043] After the library is inoculated onto the sequencing chip, an amplification reaction can be performed using the library as a template chain. For example, the amplification reaction process can be to first add dNP and polymerase to the sequencing chip. The polymerase will start from the first amplification primer and synthesize a new DNA chain along the template chain. The new DNA chain is completely complementary to the template chain, so it is also called the complementary chain of the template chain. The complementary chain is covalently linked to the sequencing chip. Next, a NaOH alkaline solution is added to the sequencing chip for washing. The template chain and the complementary chain are mutually untied in the presence of the NaOH alkaline solution, and the template chain is washed away with the alkaline solution, while the complementary chain covalently linked to the sequencing chip is retained. Then add neutral liquid to the sequencing chip to neutralize the NaOH alkaline solution. The entire environment inside the sequencing chip becomes neutral, and the other end of the complementary chain will continue to hybridize with the second amplification primer on the sequencing chip. Add dNP and polymerase, and the polymerase will start from the second amplification primer and synthesize a new DNA chain along the complementary chain. At this time, the new DNA chain is completely complementary to the complementary chain and is exactly the same as the template chain. Then add NaOH alkaline solution to untie the two chains from each other, and then you can get two chains that are covalently connected to the sequencing chip and complementary to each other. Repeat this process, and the number of DNA chains will grow exponentially.

[0044] After amplification, the sequencing chip will retain the DNA double strands identical to the template strand and the complementary strand, respectively. Then, a specific reaction reagent is added to the sequencing chip to cut off the DNA strand synthesized from one of the amplification primers, for example, the DNA strand identical to the complementary strand is cut off, and the DNA strand identical to the template strand is retained. Then, a NaOH alkaline solution is added to the sequencing chip for washing. The alkaline solution disentangles the DNA double strands from each other, and the cut DNA strands are also washed away with the alkaline solution, ultimately leaving only single DNA strands on the sequencing chip. At this time, the number of single DNA strands retained on the sequencing chip is exponentially times the number at the beginning of amplification, thereby forming a DNA cluster, and all single DNA strands in a DNA cluster are identical. Then, a neutral solution is added, and all single DNA strands in the DNA cluster can be sequenced in a neutral solution environment.

[0045] It should be noted that the above-mentioned amplification reaction can be completed outside the sequencer, such as amplifying the library through experimental operations in the laboratory, or it can be completed inside the sequencer. When the amplification reaction is completed inside the sequencer, the library involved in the amplification reaction and various reaction reagents (for example, dNP, polymerase, NaOH alkaline solution, neutral solution, etc.) can be added to the sequencing chip through the liquid circuit system or liquid circuit system of the sequencer. In addition, the above amplification reaction is only exemplary, and the present application is not limited to the use of bridge PCR amplification mode, and other amplification modes can also be used for amplification, for example, loop-mediated isothermal amplification (LAMP), nucleic acid-dependent amplification (NASBA, Nuclear acid sequence-based amplification), rolling circle amplification (RCA, Rolling Circle Amplification), multiplex probe amplification (MPA, Multiplex Probe Amplification), etc.

[0046] Sequencing reaction:

[0047] Taking the principle of sequencing by synthesis as an example, when sequencing is performed, four dNTPs with fluorescent groups are added to the sequencing chip through the liquid system, each dNTP can only be synthesized with one of the four bases of ATCG, and the 3' end of the dNTP has been blocked by a blocking group (the blocking group includes but is not limited to an azide group), and then polymerase is added to the sequencing chip through the liquid system. Through the action of the polymerase, one of the four dNTPs will be synthesized with the complementary base on the single strand being sequenced, and because the 3' end of the dNTP is blocked by a blocking group, only one dNTP can be extended on the single strand being sequenced each time. After synthesis, specific chemical reagents are added to the sequencing chip through the liquid system to flush out excess dNTPs and polymerase. Next, the optical detection system can be used to excite the fluorescent group of the dNTP that has been synthesized on the single chain, causing the fluorescent group to emit a fluorescent signal. Since the fluorescent group of each dNTP on a single chain in a cluster will emit the same fluorescent signal, the fluorescent signal is amplified. Therefore, the optical detection system can collect the fluorescent signal and generate a fluorescent image.

[0048] The computer system processes and analyzes the fluorescence image to determine which dNTP is synthesized on the sequenced single strand, and then based on the principle of complementarity, it can be inferred which base is synthesized with the dNTP on the sequenced single strand. At this point, a sequencing cycle is completed.

[0049] Next, specific chemical reagents are added to the sequencing chip through the liquid system to cut off the blocking group and the fluorescent group, thereby exposing the hydroxyl group at the 3' end of the dNTP.

[0050] Then enter the next sequencing cycle and repeat the above process.

[0051] It is understood that one sequencing cycle can detect one base, and after multiple sequencing cycles, multiple bases in the sequenced single strand can be detected. Specifically, the number of sequencing cycles can be determined according to the set sequencing read length, for example, 150 or 300 sequencing cycles.

[0052] Of course, it should be noted that the above sequencing reactions are only exemplary, and the present application is not limited to the sequencing-by-synthesis principle, and other sequencing principles may also be used.

[0053] Sequencing System:

[0054] See also Figure 1 As shown, the sequencing system includes: a sequencing chip 10, a chip platform 20, a reagent storage container 30, a liquid path system (or a flow guide system) 40, an optical detection system 50, a computer system 60 and a waste liquid storage container 70. Among them:

[0055] The sequencing chip 10 is configured to provide a reaction area for amplification reaction and sequencing reaction;

[0056] A chip platform 20 configured to fix and support the sequencing chip 10;

[0057] A reagent storage container 30, configured to store one or more mixed sample libraries, one or more reagents;

[0058] The liquid circuit system 40 is configured to controllably transport one or more mixed sample libraries and one or more reagents from the reagent storage container 30 to the sequencing chip 10 so as to perform an amplification reaction and a sequencing reaction in the sequencing chip 10, and controllably transport waste liquid after the reaction from the sequencing chip 10 to the waste liquid storage container 70;

[0059] An optical detection system 50 is configured to excite and collect fluorescent signals during a sequencing reaction and generate a fluorescent image based on the fluorescent signals;

[0060] A computer system 60 is configured to obtain a fluorescent image from the optical detection system 50 and identify a base sequence of the sample library based on the fluorescent image;

[0061] The waste liquid storage container 70 is configured to store the waste liquid generated after the reaction.

[0062] Sequencing chip:

[0063] As a carrier of amplification reaction and sequencing reaction, the sequencing chip can provide a reaction area for these reactions, and this area is a channel. Generally, a sequencing chip 10 can include one or more channels 11 (for example, 2, 4, 6, 8), and the channels are isolated from each other. Figure 2 As shown, taking four channels as an example, each channel 11 has a small hole 12 at both ends for fluid (e.g., biological samples, reaction reagents) to flow in and out. The upper and lower surfaces in each channel are chemically modified and inoculated with two amplification primers in a covalent bond manner, and the two amplification primers are complementary to the adapters at both ends of the library to achieve amplification of the library.

[0064] In the synthesis sequencing process, not all clusters after amplification of the sample library generate fluorescent spots each time the optical detection system 50 collects fluorescent images, that is, there may be clusters that do not generate fluorescent spots in a single-frame fluorescent image. When the computer system 60 identifies the base sequence of the sample library based on the fluorescent image, it is necessary to use multiple frames of fluorescent images from more sequencing cycles to assist in determining the spot information (e.g., spot position, spot shape) of all clusters to avoid cluster loss.

[0065] In order to obtain accurate cluster spot information, multiple frames of fluorescence images from multiple sequencing cycles are usually superimposed and fused, and then traditional methods for calculating the centroid coordinates of the fluorescence spots are used, such as grayscale centroid weighting method, Gaussian surface fitting method, parabola fitting method, etc. However, traditional methods require that the grayscale distribution of the fluorescence spot approximates a two-dimensional Gaussian distribution, while the shapes and sizes of the fluorescence spots in the fluorescence images in high-throughput gene sequencing are different, and the grayscale distribution of some spots does not conform to the two-dimensional Gaussian distribution, which leads to errors in the spot information of the cluster finally extracted.

[0066] Based on the model, more accurate cluster spot information can be obtained, but the model-based training process needs to establish the spot label of the fluorescent image used for model training, such as clear spot centroid, spot boundary (constituting the spot shape), etc. With multiple frames of fluorescent images of multiple sequencing cycles and labeled fluorescent spots as input, the model output result includes the spot information of the cluster identified in the fluorescent image of this sequencing. Further, when the base signal is parsed, the spot intensity values ​​at all cluster positions can be accurately extracted and the parsing is completed. Among them, the model can be trained offline or online, and the trained model is integrated in the computer system 60. For all supervised model training, whether it is a deep learning or machine learning model, the label accuracy corresponding to the training data largely determines the model accuracy. In order to obtain more accurate model prediction results, it is necessary to first identify and mark the fluorescent spots of the fluorescent images in the training data to construct accurate spot labels.

[0067] Based on the above research, the present disclosure provides a method for identifying fluorescent spots in fluorescent images for training models, firstly aligning multiple frames of original fluorescent images from multiple sequencing cycles to obtain an aligned fluorescent image, then determining the base sequence of the candidate pixel position of the spot in the fluorescent image based on the grayscale value of each pixel position in the aligned multiple frames of fluorescent images, and based on the consistency information between the base sequence of the candidate pixel position of the spot and the base sequence of the pixel position in its corresponding first neighborhood, determining the first pixel position belonging to the center area of ​​the fluorescent spot from the candidate pixel position of the spot, and then using at least part of the first pixel position as a seed point, performing regional growth processing on the seed point to obtain the spot area of ​​the spot where the seed point is located, and realizing accurate identification of the spot area. Furthermore, filtering is performed based on the obtained spot area to determine the information of the spot label such as the high-quality spot boundary and the spot centroid.

[0068] The defects existing in the above solutions are the results obtained by the inventor after practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the present disclosure for the above problems below should be the contributions made by the inventor to the present disclosure during the disclosure process.

[0069] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0070] To facilitate understanding of this embodiment, a method for identifying fluorescent spots in a sequencing fluorescent image disclosed in an embodiment of the present disclosure is first introduced in detail. The execution subject of the method for identifying fluorescent spots in a fluorescent image provided in the embodiment of the present disclosure is generally a computer device with certain computing capabilities, and the computer device includes, for example, a computer device that can be used for model training and development (e.g., offline training model scenario, the trained model is then integrated into the computer system 60) or a computer system 60 in a gene sequencing system (e.g., online training model scenario), etc. In some possible implementations, the method for identifying and marking fluorescent spots in a fluorescent image can be implemented by a processor calling computer-readable instructions stored in a memory.

[0071] The following is an explanation of the method for identifying and marking fluorescent spots in a fluorescent image provided by an embodiment of the present disclosure.

[0072] See also Figure 3 FIG. 1 is a flow chart of a method for identifying a fluorescent spot in a fluorescent image provided by an embodiment of the present disclosure, wherein the method includes steps S301 to S304, wherein:

[0073] S301: performing image alignment on multiple frames of original fluorescence images from multiple sequencing cycles to obtain aligned fluorescence images.

[0074] In a specific implementation, the original fluorescence image includes fluorescence images corresponding to multiple sequencing cycles; and the fluorescence image corresponding to each sequencing cycle usually has multiple color channels. The mainstream second-generation sequencing uses two channels or four channels for taking pictures (different channels are for different wavelengths of fluorescence emitted by different base types), and takes pictures under different wavelengths of excitation light to form different color channels in the fluorescence image. In the process of high-throughput gene sequencing, for the same cluster, each sequencing cycle needs to be chemically reacted and detected by a high-resolution optical detection system 50, and the fluorescence image is the original fluorescence image.

[0075] During the sequencing process, in the process of collecting fluorescence signals and generating fluorescence images, the optical detection system 50 and the chip platform 20 are constantly moving relative to each other. For example, the chip platform 20 moves relative to the optical detection system 50, and this movement can easily cause mechanical errors, resulting in deviations in the light spot positions of the same cluster in the fluorescence images of different sequencing cycles or even the same sequencing cycle, making it impossible to accurately align the base sequences of the same cluster, thereby affecting the accuracy of identifying the cluster base type. Therefore, in order to reduce the problems caused by position deviation, after obtaining multiple frames of original fluorescence images of multiple sequencing cycles, it is necessary to first perform image alignment on the multiple frames of original fluorescence images of multiple sequencing cycles.

[0076] Specifically, Figure 4 As shown, the embodiment of the present disclosure provides an implementation method for aligning multiple frames of original fluorescence images derived from multiple sequencing cycles, specifically including:

[0077] S401: performing image reconstruction processing on multiple frames of original fluorescence images from multiple sequencing cycles to obtain multiple frames of reconstructed fluorescence images.

[0078] S402: Determine an alignment template for image alignment from the multiple frames of reconstructed fluorescence images.

[0079] In a specific implementation, after each nucleic acid fragment (such as a DNA fragment) is clustered through an amplification reaction, in the subsequent sequencing reaction, each sequencing cycle will collect fluorescent signals for the current state of the cluster, so a series of sequencing images will be generated during the sequencing reaction. Since the sequencing image includes fluorescent spots, the sequencing image is also called a fluorescent image. Due to the mechanical errors of the sequencing equipment itself, the same photographing area has position deviations in the fluorescent images obtained in different sequencing cycles, which makes the fluorescent spots of the same cluster have position deviations in the fluorescent images taken in different sequencing cycles or even the same sequencing cycle, resulting in a decrease in the accuracy of the positions of each cluster located during the sequencing process, which will in turn cause the sequencing quality of the corresponding nucleic acid fragments to decrease. In order to reduce the problems caused by mechanical errors, it is necessary to construct an alignment template of the fluorescent image in the current sequencing process in real time during the sequencing process, and use the alignment template to calibrate and align the positions of the fluorescent spots corresponding to the same cluster in different sequencing cycles.

[0080] Image alignment requires the selection of an alignment template as a reference. Both the alignment template and the image to be aligned need to have clear features and a high signal-to-noise ratio. For fluorescent images, this means high spot grayscale values, clear boundaries, and clear distinction between foreground and background. Image reconstruction technology can reduce noise in the original fluorescent image and highlight the feature information of the fluorescent spot in the image. Selecting the reconstructed image as the alignment template and performing image alignment is more conducive to improving alignment accuracy.

[0081] In S401 of the present embodiment, any fluorescence image reconstruction method in the prior art may be used. For example, the reconstruction method may be: sharpening multiple frames of original fluorescence images respectively to obtain sharpened images; determining a segmentation threshold according to intensity values ​​corresponding to respective pixel positions in the sharpened images, and taking the light spot in the sharpened image as the foreground, segmenting the sharpened image based on the segmentation threshold to obtain a segmented image; performing parabolic interpolation processing on the segmented image to obtain a parabolic interpolation image, and obtaining a reconstructed fluorescence image based on the parabolic interpolation image.

[0082] In a specific implementation, for example, the GL (Grünwald-Letnikov) fractional differential method may be used to sharpen the original fluorescence image, and the specific process may be as follows:

[0083] (1): The original fluorescence image is grayed and normalized in turn, and the gray value of each pixel in the original fluorescence image is normalized to the interval [0,1].

[0084] (2): Fractional differential mask construction.

[0085] (3): The constructed fractional differential mask is used as the convolution kernel to convolve the normalized original fluorescence image, and the convolution result is denormalized to obtain a sharpened image.

[0086] After obtaining the sharpened image, the segmentation threshold can be determined according to the intensity value (i.e., grayscale value) corresponding to each pixel position in the sharpened image, such as taking the mean of the intensity values ​​corresponding to each pixel position as the segmentation threshold, or determining the segmentation threshold based on any one of the Otsu method, K-means, support vector machine algorithm, etc., and performing segmentation processing on the sharpened image block based on the segmentation threshold.

[0087] During the segmentation process, for example, each pixel position in the sharpened image block can be traversed, and the intensity value of the traversed pixel position can be compared with the segmentation threshold; when the intensity value of the traversed pixel position is greater than or equal to the segmentation threshold, the traversed pixel position is determined as a pixel position belonging to the light spot; when the intensity value of the traversed pixel position is less than the segmentation threshold, the traversed pixel position is determined as the background to obtain a segmented image. The pixel value of each pixel position in the segmented image is used to characterize whether the pixel position belongs to the area where the light spot is located. The segmented image is subjected to parabolic interpolation processing to obtain a reconstructed fluorescence image. Afterwards, the first frame image can be selected from the reconstructed fluorescence image as an alignment template for image alignment, or the image with the highest signal-to-noise ratio can be selected from the reconstructed fluorescence image according to the signal-to-noise ratios corresponding to the reconstructed fluorescence images.

[0088] S403: Based on the alignment template, perform image alignment processing on multiple frames of the original fluorescence images to obtain multiple frames of aligned fluorescence images.

[0089] Here, when performing image alignment processing on multiple frames of original fluorescent images based on the alignment template, for example, the normalized mutual power spectrum matrix between each reconstructed fluorescent image and the alignment template can be calculated, and then the normalized power method processing is performed to obtain the offset for alignment, and the original fluorescent images to be aligned are aligned using the offset to obtain the aligned multiple frames of fluorescent images. A series of fluorescent image sets obtained during the cycle sequencing process are sequentially obtained to obtain high-resolution fluorescent images aligned with the alignment template. A series of fluorescent image sets are, for example, fluorescent images of 10 sequencing cycles, 15 sequencing cycles, 20 sequencing cycles, 25 sequencing cycles, and 30 sequencing cycles. The number of sequencing cycles used in image alignment may be consistent with or inconsistent with the number of sequencing cycles used in image reconstruction.

[0090] After aligning the multiple frames of original fluorescence images, the resolution of the aligned multiple frames of fluorescence images can also be adjusted, such as adjusting the low-resolution fluorescence images to high-resolution fluorescence images; in the resolution adjustment process, for example, each pixel in the aligned multiple frames of fluorescence images can be divided into multiple sub-pixels (i.e., sub-pixels), and the pixel values ​​of each sub-pixel are interpolated according to the pixel values ​​corresponding to each pixel and the pixel values ​​corresponding to the adjacent pixels, so as to obtain the pixel values ​​corresponding to each sub-pixel. In the subsequent processing, the above sub-pixel can be regarded as a new pixel and marked.

[0091] Here, when dividing each pixel into multiple sub-pixels, specifically, one-quarter of the sub-pixels can be used. In this case, each pixel in the fluorescent image is divided into sixteen sub-pixels. Alternatively, one-eighth of the sub-pixels can be used. In this case, each pixel in the fluorescent image is divided into sixty-four sub-pixels. The specific division method is not limited in the embodiments of the present disclosure.

[0092] Following the above S301, the method for marking a fluorescent spot in a fluorescent image provided by the embodiment of the present disclosure further includes:

[0093] S302: Determine the base sequence of the candidate pixel position of the light spot among the plurality of pixel positions according to the gray value of each pixel position in the aligned multiple frames of fluorescence images.

[0094] In a specific implementation, the aligned multiple frames of fluorescence images correspond to a unified image coordinate system, and in the image coordinate system, the same pixel position has different pixel points or sub-pixel points in different fluorescence images. For each pixel position, its grayscale value in different original fluorescence images can be obtained by reading the grayscale value of the pixel point or sub-pixel point. The candidate pixel position of the light spot is, for example, a pixel position in the fluorescence image that has a high probability of belonging to the light spot area.

[0095] When determining the base sequence, for example, the following method can be used:

[0096] Based on the grayscale value of each pixel position in the aligned multi-frame fluorescence image, the aligned multi-frame fluorescence image is segmented into foreground and background respectively to obtain the foreground area in the aligned multi-frame fluorescence image as a candidate fluorescence spot area; and the candidate fluorescence spot area in the aligned multi-frame fluorescence image is fused to obtain a fused fluorescence spot area of ​​the aligned multi-frame fluorescence image; and the aligned multi-frame fluorescence image is sharpened respectively to obtain a sharpened multi-frame fluorescence image; each pixel position located in the fused fluorescence spot area is used as the spot candidate pixel position, and the base sequence corresponding to the spot candidate pixel position is determined according to the grayscale value of the spot candidate pixel position in the sharpened multi-frame fluorescence image.

[0097] In a specific implementation, when performing foreground and background segmentation processing on the aligned multiple-frame fluorescence images based on the pixel values ​​of each pixel in the aligned multiple-frame fluorescence images, for example, a grayscale value threshold can be determined, and the grayscale value of each pixel can be compared with the grayscale value threshold; if the grayscale value of a certain pixel is greater than or equal to the grayscale value threshold, the pixel is used as a pixel in the foreground area, that is, a pixel in the candidate spot area; if the grayscale value of a certain pixel is less than the grayscale value threshold, the pixel is used as a pixel in the background area, thereby realizing foreground and background segmentation processing of the fluorescence image, and obtaining the candidate spot areas of each of the multiple-frame fluorescence images. The above-mentioned grayscale value threshold can be pre-set and determined based on experience, or can be determined using the Otsu method, Kmeans classification algorithm, etc., which is not limited in the embodiments of the present disclosure.

[0098] Then, the union of the candidate spot regions of the multiple frames of fluorescence images is determined to obtain a fused fluorescence spot region. After obtaining the fused fluorescence spot region, each pixel position in the fused fluorescence spot region is used as a spot candidate pixel position to determine the base sequence corresponding to each spot candidate pixel position.

[0099] When determining the base sequence corresponding to each candidate light spot pixel position, the base type corresponding to each candidate light spot pixel position in multiple sequencing cycles can be determined based on the grayscale value of each candidate light spot pixel position in the sharpened multi-frame fluorescence image; then the base type corresponding to each candidate light spot pixel position in multiple sequencing cycles is used to form the base sequence of each candidate light spot pixel position.

[0100] like Figure 5In the example shown, assuming that there are N sequencing cycles, there are N corresponding fluorescence images; for pixel 1 located at the same candidate pixel position of the light spot in the N frames of fluorescence images, the base types of the pixel 1 in multiple sequencing cycles can be obtained; assuming that the corresponding base type in the first sequencing cycle is A, the corresponding base type in the second sequencing cycle is T, the corresponding base type in the third sequencing cycle is G, ..., and the corresponding base type in the Nth sequencing cycle is A, then the base sequence corresponding to the pixel 1 is: ATG ... A. The base sequences corresponding to the multiple candidate pixel positions of the light spot constitute a pixel sequence set.

[0101] Here, when determining the base type of each candidate pixel position of each spot in different sequencing cycles, for example, a real-time unsupervised classification model can be used, or a target classification model can be obtained through multiple iterations of training to identify the base type. The classification model can use K-nearest neighbor, decision tree, random forest, SVM, Xgboost and other algorithms, which are not limited in the embodiments of the present disclosure.

[0102] Following the above S302, the method for identifying a fluorescent spot in a fluorescent image provided by the embodiment of the present disclosure further includes:

[0103] S303: Determine a first pixel position belonging to the center area of ​​the fluorescent spot from the candidate light spot pixel positions based on consistency information between the base sequence of the candidate light spot pixel position and the base sequence of the candidate light spot pixel position in the first neighborhood.

[0104] In a specific implementation, after the DNA molecules are bound to the surface of the circulation pool, thousands of DNA copies will be generated around them through PCR amplification. In theory, these copies are considered to be exactly the same, but due to the influence of the reaction environment such as enzymes, errors will gradually accumulate during the amplification clustering process. Therefore, under normal circumstances, the base sequence corresponding to the pixels in the central area of ​​the light spot (the centroid of the light spot is located in the central area of ​​the light spot) has a higher consistency. The closer the pixel is to the centroid of the light spot, the higher the consistency of the base sequence corresponding to the pixel is. Conversely, the closer the pixel is to the edge of the light spot, the lower the consistency of the base sequence corresponding to the pixel is. Therefore, in the embodiment of the present disclosure, from a plurality of candidate light spot pixel positions, according to the consistency of the base sequence of each candidate light spot pixel position in the fluorescent image and the base sequence of its neighboring pixel position, the candidate light spot pixel position with higher consistency is selected as the pixel position belonging to the central area of ​​the fluorescent light spot (i.e., the first pixel position).

[0105] In the disclosed embodiment, the seed point is the starting point for region growing. Ideally, the base sequence corresponding to the pixel point located at the centroid of the light spot in the fluorescent image has the highest consistency. In the same light spot, the farther from the centroid of the light spot, the lower the consistency of the base sequence corresponding to the pixel point. Using a method of light spot shape recognition based on base sequence consistency, the signal recognition accuracy of the central area of ​​the light spot is high. The farther from the central area of ​​the light spot, the lower the recognition accuracy. Therefore, when performing regional growth processing, it is necessary to select high-quality (i.e., more consistent) points as seed points. The disclosed embodiment uses high-resolution fluorescent images of multiple sequencing cycles to perform base recognition for each light spot candidate pixel position in the high-resolution image aligned with the alignment template, and outputs the base sequence of each light spot candidate pixel position to obtain a base sequence set of each light spot candidate pixel position. The base sequence of the candidate pixel position N(i, j) of the light spot is compared with the base sequences of the surrounding candidate pixel positions of the light spot, and the boundary distance is calculated. The base sequence with better consistency, that is, the base sequence with the number of mismatches with the base sequences of the pixel positions in the surrounding neighborhood is lower than the threshold, is selected as the candidate seed point set, that is, the set consisting of the first pixel position.

[0106] Specifically, an embodiment of the present disclosure provides a specific method for determining the consistency information of the base sequence between each of the light spot candidate pixel positions and each of the light spot candidate pixel positions in its corresponding first neighborhood, including: determining the first neighborhood of each of the light spot candidate pixel positions; determining the degree of difference between the base sequence corresponding to each of the light spot candidate pixel positions and the base sequences corresponding to each of the light spot candidate pixel positions in its corresponding first neighborhood; the degree of difference includes: the number of different bases between the base sequence of each of the light spot candidate pixel positions and the base sequence of the light spot candidate pixel positions in its corresponding first neighborhood, or the percentage of the number of different bases in the base sequence of the light spot candidate pixel position; determining the consistency information according to the degree of difference.

[0107] In a specific implementation, the first neighborhood corresponding to each candidate light spot pixel position may be, for example, other candidate light spot pixel positions adjacent to the candidate light spot pixel position. i ,y j ), the other candidate pixel positions of the light spot located in the first neighborhood include the following pixel positions: (x i-1 ,y j )、(x i ,y j-1 )、(x i+1 ,y j )、(x i ,y j+1 ).

[0108] Alternatively, the candidate pixel positions of the light spot may be within an area with the candidate pixel position of the light spot as the center and k pixel positions as the radius. For example, when k is 1, the candidate pixel positions of the light spot located in the first neighborhood include the following pixel positions: i-1 ,y j )、(x i ,y j-1 )、(x i+1 ,y j )、(x i ,y j+1 )、(x i-1 ,y j-1 )、(x i+1 ,y j-1 )、(x i+1 ,y j+1 )、(x i-1 ,y j+1 ).

[0109] like Figure 6 As shown, the embodiment of the present disclosure also provides an example of the relative position relationship between a spot candidate pixel position N and other spot candidate pixel positions N' in its corresponding first neighborhood. In this example, k=1, and all spot candidate pixel positions within a distance of 1 near the spot candidate pixel position can be used as the domain pixel position corresponding to the spot candidate pixel position N. For example, one of the first domain pixel positions is Figure 6 N' in.

[0110] After determining the first neighborhood pixel position corresponding to each pixel position, determine the degree of difference between the base sequence corresponding to the pixel position and the base sequence corresponding to the first neighborhood pixel position. Here, for example, the Euclidean distance, cosine similarity, Hamming distance, Sorensen-Dice index, dynamic time warping (DTW), etc. can be used to determine the degree of difference between each pixel position and the corresponding first neighborhood pixel position.

[0111] The consistency information between each light spot candidate pixel position and other light spot candidate pixel positions in its first neighborhood is determined according to the degree of difference.

[0112] Here, for example, at least one of the sum, average, etc. of the degree of difference between each light spot candidate pixel position and the surrounding light spot candidate pixel positions in the first neighborhood can be calculated, and the sum or average value can be used as a value to measure the consistency information. The larger the sum, the lower the consistency between each light spot candidate pixel position and other light spot candidate pixel positions in the corresponding first neighborhood; the smaller the sum, the higher the consistency between each light spot candidate pixel position and other light spot candidate pixel positions in the corresponding first neighborhood.

[0113] When determining the first pixel position belonging to the center of the fluorescent spot from multiple spot candidate pixel positions based on consistency information, for example, the above sum value can be compared with a preset threshold; if the sum value corresponding to a certain spot candidate pixel position is less than the threshold, the spot candidate pixel position is determined as the first pixel position.

[0114] Through the above process, all the candidate pixel positions of the light spots in the image coordinate system are traversed, and the candidate pixel positions of the light spots that meet the above consistency requirements are all taken as the first pixel positions and put into the candidate light spot center set for subsequent use.

[0115] Following the above S303, the method for identifying a fluorescent spot in a fluorescent image provided by the embodiment of the present disclosure further includes:

[0116] S304: Use at least part of the first pixel positions as seed points, perform region growing processing on the seed points, and obtain a light spot region of the light spot where the seed points are located.

[0117] In a specific implementation, when determining the first pixel position belonging to the center of the fluorescent light spot, multiple adjacent or close candidate pixel positions of the light spot may be determined as the first pixel position; if one of them is used as a seed point for region growing processing, the first pixel position adjacent or close to the seed point may be determined as belonging to the same light spot region as the seed point during the growing process, and thus will be deleted from the candidate center set, and will no longer be used as a seed point in the future. Therefore, when this situation exists, the first pixel position used as the seed point only includes some of the first pixel positions in the candidate center set.

[0118] When performing region growing processing on a seed point based on the similarity between the base sequence of a seed point and the base sequence of the corresponding candidate pixel position of the spot in the second neighborhood, for example, the candidate pixel position of the spot whose similarity satisfies a certain similarity condition can be determined from the candidate pixel positions of the spot in the second neighborhood corresponding to the seed point as the pixel position belonging to the same spot region as the seed point. In addition, the candidate pixel position of the spot that satisfies the similarity condition can be used as the growable pixel position corresponding to the seed point, and the region growing is continued to obtain the secondary spot region corresponding to the seed point.

[0119] See also Figure 7 As shown, the specific method of determining the seed point and performing region generation processing on the seed point provided by the embodiment of the present disclosure is as follows:

[0120] Perform multiple iterations and perform the following region growing process in each iteration:

[0121] S701: Determine a target seed point corresponding to a current iteration cycle from a first pixel position where the light spot has not yet been determined to belong.

[0122] Here, if the current iteration cycle is the first iteration cycle, the first pixel positions whose light spot ownership is not determined are all the first pixel positions in the candidate center set determined in the above step S303. Any of the first pixel positions can be determined as the target seed point corresponding to the current iteration cycle. It is also possible to sort the first pixel positions according to the consistency information corresponding to each first pixel position, from high to low, and determine the first pixel position with the highest corresponding consistency as the target seed point of the first iteration cycle. If the current iteration cycle is not the first iteration cycle, since some of the first pixel positions determined in the above step S303 have been used as seed points for regional growth processing in each iteration cycle before the current iteration cycle, or in the process of regional growth processing of the seed points, the light spot ownership consistent with the seed point will be determined for some of the first pixel positions, so the first pixel positions whose light spot ownership has been determined can be deleted from the candidate center set; in the non-first iteration cycle, the first pixel positions whose light spot ownership is not determined include the remaining first pixel positions in the candidate center set. At this time, any of the first pixel positions can be used as the seed point of the current iteration cycle, or the first pixel position with the highest consistency can be determined from the remaining first pixel positions in the candidate center set in order of consistency from high to low as the target seed point of the current iteration cycle.

[0123] S702: Determine the candidate light spot pixel position within the second neighborhood of the target seed point according to the position information of the target seed point.

[0124] In a specific implementation, the second neighborhood may be in the image coordinate system, and the distance from the target seed point is less than a preset distance range. It may be the same as the first neighborhood, or it may be different from the first neighborhood. When determining the candidate light spot pixel position in the second neighborhood according to the position information corresponding to the target seed point, for example, the candidate light spot pixel position whose light spot has not been determined in the current iteration cycle is determined from the second neighborhood as the second neighborhood pixel position corresponding to the target seed point.

[0125] Specifically, the second neighborhood may be, for example, a four-neighborhood or an eight-neighborhood of the target seed point; Figure 8As shown, a specific example of a four-neighborhood and an eight-neighborhood is shown. Figure 8 In the figure, the yellow position indicates the target seed point; Figure 8 The blue pixel position in a is the second neighborhood pixel position when the second neighborhood is a four-neighborhood neighborhood; Figure 8 The blue pixel position in b is the second neighborhood pixel position when the second neighborhood is an eight-neighborhood. The spot area is obtained by performing region growing processing on the target seed point using the second neighborhood pixel positions of the four-neighborhood and the eight-neighborhood, as shown in Fig. 9 shown.

[0126] S703: Based on the similarity between the base sequence of the target seed point and the base sequence of the candidate light spot pixel position corresponding to the second neighborhood, a region growing process is performed on the target seed point to obtain a light spot region of the light spot where the target seed point is located.

[0127] In a specific implementation, based on whether the similarity between the base sequence of the target seed point and the base sequence between the candidate light spot pixel position in the corresponding second neighborhood meets the preset similarity condition, the second pixel position belonging to the same light spot area as the target seed point is determined from the second neighborhood pixel positions; and the following steps are executed repeatedly until no new second pixel position is generated: the second pixel position is used as a growable pixel position, and based on whether the similarity between the base sequence of the growable pixel position and the base sequence between the candidate light spot pixel position in the corresponding third neighborhood meets the preset similarity condition, the second pixel position belonging to the same light spot area as the target seed point is determined from the third neighborhood pixel positions.

[0128] For example, Fig.10 As shown, assuming that the target seed point is N, the four spot candidate pixel positions N1, N2, N3 and N4 with a distance of 1 from N are used as the spot candidate pixel positions in the second neighborhood; assuming that the similarities between N1 and N2 and the target seed point N meet the preset similarity conditions, while the similarities between N3 and N4 and the target seed point N do not meet the preset similarity conditions, then N1 and N2 are used as the second pixel positions of the target seed point N, and N1 and N2 are deleted from the candidate center set. The first cycle: N1 and N2 are used as growable pixel positions respectively, and the growth process continues.

[0129] For N1, the candidate pixel positions of the light spot within the third neighborhood of N1 include: N5, N6, N7 and N. Since N is the target seed point, N will not be used as the candidate pixel position of the light spot within the third neighborhood of N1, but N5 to N7 will be used as the candidate pixel positions of the light spot within the third neighborhood of N1; assuming that the similarity of the base sequence between N5 and N1 meets the preset similarity condition, and the similarity of the base sequence between N6 and N7 and N1 respectively does not meet the preset similarity condition, then N5 will be used as the new second pixel position, and N5 will be deleted from the candidate center set. Optionally, N5 to N7 can also be judged for similarity with the base sequence of the target seed point N.

[0130] For N2, the candidate pixel positions of the light spot within the third neighborhood of N2 include: N7, N8, N9 and N. Since N is the target seed point, N will not be used as the candidate pixel position of the light spot within the third neighborhood of N2, but N7 to N9 will be used as the candidate pixel positions of the light spot within the third neighborhood of N2. Assuming that the similarity of the base sequence between N8 and N9 and N2 respectively meets the preset similarity condition, and the similarity of the base sequence between N7 and N2 does not meet the preset similarity condition, N8 and N9 will be used as the new second pixel position, and N8 and N9 will be deleted from the candidate center set. Optionally, N7 to N9 can also be judged for similarity with the base sequence of the target seed point N.

[0131] Second cycle: N5, N8 and N9 are respectively used as growable pixel positions and the growth process is continued.

[0132] For N5, the candidate pixel positions of the light spot within the third neighborhood of N5 include: N10, N11, N12 and N1. Since the light spot of N1 has been determined, N1 will not be used as the candidate pixel position of the light spot within the third neighborhood of N5, but N10~N12 will be used as the candidate pixel positions of the light spot within the third neighborhood of N5; assuming that the similarity between the base sequences of N10, N11, N12 and N5 respectively does not meet the preset similarity condition, the growth process based on N5 is terminated. Optionally, N10~N12 can also be judged for similarity with the base sequence of the target seed point N.

[0133] For N8, the candidate pixel positions of the light spot within the third neighborhood of N8 include: N13, N14, N15 and N2. Since the light spot of N2 has been determined, N2 will not be used as the candidate pixel position of the light spot within the third neighborhood of N8, but N13 to N15 will be used as the candidate pixel positions of the light spot within the third neighborhood of N8; assuming that the similarity of the base sequence between N13 and N8 meets the preset similarity condition, and the similarity of the base sequence between N14 and N15 and N8 does not meet the preset similarity condition, then N13 will be used as the new second pixel position, and N13 will be deleted from the candidate center set. Optionally, N13 to N15 can also be judged for similarity with the base sequence of the target seed point N.

[0134] For N9, the candidate pixel positions of the light spot within the third neighborhood of N9 include: N2, N3, N15 and N16. Since the light spot of N2 has been determined, N2 will not be used as the candidate pixel position of the light spot within the third neighborhood of N9, but N3, N15 and N16 will be used as the candidate pixel positions of the light spot within the third neighborhood of N9; assuming that the similarity between the base sequences of N3, N15 and N16 and N9 does not meet the preset similarity condition, the growth process based on N9 is terminated. Optionally, the similarity can also be judged with the base sequence of the target seed point N.

[0135] Third cycle:

[0136] Take N13 as the new growable pixel position and continue the growth process. In this way, after multiple cycles, the region growth process of the target seed point is completed. Assuming that the new second pixel position is not determined during the third cycle, the cycle process ends. Finally, the spot areas to which the target seed point belongs include: N, N1, N2, N5, N8, N9 and N13.

[0137] In addition, in order to obtain better regional growing results, in another embodiment of the present disclosure, not all second pixel positions determined for target seed points may be used as growable pixel positions. Instead, more stringent screening conditions may be set for them, and only the second pixel positions that meet certain screening conditions among the second pixel positions may be used as growable pixel positions for region growing processing. Based on this, in another embodiment of the present disclosure, when performing region growing processing on the target seed point based on the similarity between the base sequence of the target seed point and the base sequence of the corresponding second neighborhood pixel position, for example, the following method can also be adopted: according to the similarity between the base sequence of the target seed point and the base sequence of the corresponding spot candidate pixel position in the second neighborhood, determine the second pixel position belonging to the same spot area as the target seed point, and determine the growable pixel position corresponding to the target seed point from the second pixel position; loop the following steps until no new growable pixel position appears: determine a new second pixel position belonging to the same spot area as the target seed point by using the similarity between the base sequence of the growable pixel position and the base sequence of the corresponding spot candidate pixel position in the third neighborhood; and determine the new growable pixel position corresponding to the target seed point in the new second pixel position; and construct the spot area based on the target seed point and the second pixel position.

[0138] Here, when determining a second pixel position belonging to the same light spot area as the target seed point based on the similarity between the base sequence of the target seed point and the base sequence of the candidate pixel position of the light spot in the corresponding second neighborhood, and determining the growable pixel position corresponding to the target seed point from the second pixel position, for example, the following method can be adopted: comparing the similarity with a first similarity threshold; if the similarity is greater than the first similarity threshold, determining the second neighborhood pixel position as the second pixel position; and when the similarity is greater than the first similarity threshold, comparing the similarity with the second similarity threshold; if the similarity is greater than the second similarity threshold, determining the corresponding second pixel position as the growable pixel position corresponding to the target seed point.

[0139] The second similarity threshold is greater than the first similarity threshold, so that the determination of the position of the growable pixel has stricter conditions than the determination of the second pixel position. Fig.10In the example shown, assuming that the target seed point is N, the four pixel positions N1, N2, N3 and N4 with a distance of 1 from N are taken as the second neighborhood pixel positions. The similarity between N1 to N4 and N is compared with the first similarity threshold L1. Assuming that the similarity between N1 and N2 and the target seed point N is greater than the first similarity threshold L1, and the similarity between N3 and N4 and the target seed point N is less than or equal to the first similarity threshold L1, N1 and N2 are taken as the second pixel position of the target seed point N, and N1 and N2 are deleted from the candidate center set. Afterwards, the similarity between N1 and N2 and the target seed point N is compared with the second similarity threshold L2. Assuming that the similarity between N1 and N2 and the target seed point N is greater than the second similarity threshold, N1 and N2 are both taken as growable pixel positions.

[0140] First cycle: N1 and N2 are respectively used as growable pixel positions and the growth process continues.

[0141] For N1, the candidate pixel positions of the light spot within the third neighborhood of N1 include: N5, N6, N7 and N. Since N is the target seed point, N will not be used as the candidate pixel position of the light spot within the third neighborhood of N1, but N5 to N7 will be used as the candidate pixel positions of the light spot within the third neighborhood of N1; assuming that the similarity of the base sequence between N5 and N1 is greater than the first similarity threshold, and the similarity of the base sequence between N6 and N7 and N1 is less than or equal to the similarity threshold, then N5 will be used as the new second pixel position, and N5 will be deleted from the candidate center set. Compare the similarity of the base sequence between N5 and N1 with the second similarity threshold L2. If the similarity of the base sequence between N5 and N1 is greater than the second similarity threshold L2, then N5 will be used as the new growable pixel position. Optionally, N5 can also be judged for similarity with the base sequence of the target seed point N.

[0142] For N2, the candidate pixel positions of the light spot within the third neighborhood of N2 include: N7, N8, N9 and N. Since N is the target seed point, N will not be used as the candidate pixel position of the light spot within the third neighborhood of N2, but N7 to N9 will be used as the candidate pixel positions of the light spot within the third neighborhood of N2. Assuming that the similarity between the base sequences of N8 and N9 and N2 is greater than the first similarity threshold, and the similarity between the base sequences of N7 and N2 is less than the first similarity threshold, N8 and N9 will be used as the new second pixel positions, and N8 and N9 will be deleted from the candidate center set. The similarity between the base sequences of N8 and N9 and N2 and the second similarity threshold L2 are compared. If the similarity between the base sequences of N8 and N2 is greater than L2, and the similarity between the base sequences of N9 and N2 is less than or equal to L2, N8 will be used as the new growable pixel position.

[0143] Second cycle: N5 and N8 are respectively used as growable pixel positions and the growth process is continued.

[0144] For N5, the candidate pixel positions of the light spot within the third neighborhood of N5 include: N10, N11, N12 and N1. Since the light spot affiliation of N1 has been determined, N1 will not be used as the candidate pixel position of the light spot within the third neighborhood of N5, but N10~N12 will be used as the candidate pixel positions of the light spot within the third neighborhood of N5; assuming that the similarities of the base sequences between N10, N11, N12 and N5 are respectively less than or equal to the first similarity threshold, the growth processing based on N5 is terminated.

[0145] For N8, the candidate pixel positions of the light spot within the third neighborhood of N8 include: N13, N14, N15 and N2. Since the light spot of N2 has been determined, N2 will not be used as the candidate pixel position of the light spot within the third neighborhood of N8, but N13 to N15 will be used as the candidate pixel positions of the light spot within the third neighborhood of N8; assuming that the similarity of the base sequence between N13 and N8 is greater than the first similarity threshold, and the similarity of the base sequence between N14 and N15 and N8 is less than or equal to the first similarity threshold, then N13 will be used as the new second pixel position, and N13 will be deleted from the candidate center set. The similarity of the base sequence between N13 and N8 is compared with the second similarity threshold. If the similarity of the base sequence between N13 and N8 is less than the second similarity threshold, no new growable pixel position will be generated based on N8.

[0146] Since no new pixel position is generated, the loop process ends. In this way, after multiple loops, the target seed point is processed for regional growth. Finally, the spot area to which the target seed point belongs includes: N, N1, N2, N5, N8, N9, and N13.

[0147] The above implementation only uses the similarity between base sequences as an example of a comparison condition. The similarity of base sequences can be characterized, for example, by the distance between base sequences corresponding to different candidate pixel positions of light spots. It can be understood that the larger the distance, the lower the similarity; the smaller the distance, the higher the similarity. Therefore, in addition to the similarity comparison method, the distance between base sequences can also be used as a comparison condition. The distance here can be, for example, the edit distance, Euclidean distance, Hamming distance, Sorensen-Dice index, dynamic time warping (DTW), etc. Among them, the edit distance between base sequence A and base sequence B refers to the minimum number of bases that need to be adjusted to adjust base sequence A to be consistent with base sequence B.

[0148] like Fig.11As shown, the embodiment of the present disclosure provides a specific example of using at least part of the first pixel position as a seed point, and performing a region growing process on the seed point based on the similarity of the base sequence between the seed point and the corresponding second neighborhood pixel position. In this example, a second pixel position N is selected from the candidate center set as a seed point, and the base sequence corresponding to the pixels near it (and the second neighborhood pixel position corresponding to the seed point) is compared with the base sequence corresponding to the pixel N (calculating the distance), and the pixel point x1 whose distance to the base sequence corresponding to the pixel N is less than the threshold value L2 is added to the spot area Z1 where the seed point is located, and the pixel point x1 and its corresponding base sequence are removed from the base sequence set and the candidate center set. If the pixel point x1 is not in the candidate center set, this step does not remove the candidate center set. If the distance between the base sequence corresponding to the pixel point x1 and the base sequence corresponding to the pixel N is less than the threshold value L3, the pixel point x1 is used as a growable pixel position and added to the extendable area Z2 of the seed point. The region Z1 and the region Z2 are in a contained relationship, that is, the region Z2 belongs to the region Z1. Select a single growable pixel N' in the extendable area Z2 of the seed point, add the neighboring pixel x2 that meets the similarity criterion (i.e., the distance is less than the threshold L2) to the spot area Z1 where the seed point pixel N is located, and remove the neighboring pixel x2 that meets the rule from the base sequence set, and remove the pixel N' from the extendable area Z2 of the seed point. If the distance between the base sequence corresponding to the pixel x2 and the base sequence corresponding to the pixel N is less than the threshold L3, then the pixel x2 is added to the extendable area Z2 of the seed point. If the distance between the base sequence corresponding to the pixel x2 and the base sequence corresponding to the pixel N is not less than the threshold L3, then there is no need to add the pixel x2 to the extendable area Z2 of the seed point. Repeat the above process until the extendable area Z2 of the seed point is empty, and then the cluster search process is completed. Traverse all the candidate cluster centers in the candidate center set to complete the search of all clusters on the fluorescence image. In the process of judging the neighborhood pixels, the seed pixel N selected from the candidate center set may be fixedly used for comparison with the subsequent neighborhood pixels, or the pixel N' selected from the extendable area Z2 may be used for comparison with the subsequent neighborhood pixels.

[0149] Through the above process, the light spot area in the original fluorescence image can be determined.

[0150] After the spot area in the original fluorescence image is determined, various applications can be performed on the original fluorescence image. For example, the original fluorescence image with the spot area determined can be directly processed for base sequence recognition to obtain the base sequence of the gene sample to be tested. In addition, the original fluorescence image can be labeled according to the spot area to obtain the label of the original fluorescence image. The original fluorescence image and its label can be used as the training of the deep learning model in the base recognition scenario.

[0151] In another embodiment of the present disclosure, after obtaining the spot area of ​​the spot where the seed point is located, the method may further include: determining the spot boundary of the spot area, and / or, according to the position information of each pixel position in the spot area, determining the spot centroid of the spot area; and generating a spot label of the original fluorescent image based on the spot boundary and / or the spot centroid. The original fluorescent image with the spot label is used to train the model, and the trained model is used to predict the spot information of all clusters participating in the sequencing reaction.

[0152] In a specific implementation, a binary image can be generated based on the original fluorescent image and the cluster's spot information, and the outline of each spot can be found in the binary image to determine the spot boundary. Here, the cluster's spot information includes, for example, information on whether each pixel position belongs to a spot.

[0153] In addition, the centroid of each spot can be calculated based on the spot boundary of each spot, and the position of the pixel where the centroid of the spot is located is recorded, and this position is the position of the centroid of the spot. The method for determining the centroid of the spot can be the arithmetic mean of the positions of all pixels in the spot area, or the weighted average of the similarity between the positions of each pixel in the spot and the sequence.

[0154] After obtaining the spot boundary and / or the spot centroid, it also includes: determining the shape feature information of the spot area; the shape feature information includes at least one of the following: the spot size, the aspect ratio of the spot, the area ratio of the effective area in the spot, whether the spot centroid is in the spot, and the distance between other spot centroids; when the confidence of the corresponding spot represented by the shape feature information is less than a preset confidence threshold, the spot area is subjected to a first screening process.

[0155] Specifically, the shape feature information of the spot area can be determined, for example, based on the spot boundary and / or the spot centroid. After obtaining the shape feature information of the spot area, if the shape feature information of a certain spot area represents that the confidence of the corresponding spot is less than a preset confidence threshold, the spot area with a confidence less than the preset confidence threshold is removed.

[0156] Specifically, the confidence threshold may be represented, for example, according to a threshold corresponding to the shape feature information.

[0157] Exemplarily, assuming that the shape feature information includes the size of the spot area, the confidence threshold includes: a size threshold; wherein the size of the spot area refers to the number of effective pixels in the spot area, and the number of effective pixels of the spot may be required to be greater than 2*2, 3*3, 4*4, etc., and 2*2, 3*3, 4*4 are the confidence thresholds in this case.

[0158] For another example, the effective area ratio in the spot region refers to the ratio of the number of pixels in the spot region to the circumscribed rectangle of the spot. The effective area ratio in the spot region may be required to be greater than 0.5, 0.6, 0.7, etc. The 0.5, 0.6, 0.7 are the confidence thresholds in this case. The larger the threshold, the fuller the spot. The circumscribed rectangle may be the minimum circumscribed rectangle or the maximum circumscribed rectangle. If the distance between two spots is relatively close, the spot with a relatively low score may be removed. The scoring rule may be based on the weighted average of the similarity between pixels in the spot and the distance from the spot mass center, the spot size, the aspect ratio of the spot, the effective area ratio in the spot, etc. One or more indicators may be selected for weighted calculation of the score. Spots with scores below the threshold are low-confidence spots and are removed. Spots with a distance between the spot mass centers less than the threshold are selected as low-confidence spots and are removed. The distance between the spot mass centers may be determined by dividing the image into grids.

[0159] In another embodiment of the present disclosure, the method may further include:

[0160] Matching the base sequence corresponding to the centroid of the light spot with the reference gene sequence to obtain difference information between the base sequence corresponding to the centroid of the light spot and the reference gene sequence;

[0161] When the difference information is greater than or equal to a preset difference threshold, a second screening process is performed on the light spot area.

[0162] Here, the difference information may include, for example, the number of different bases between the base sequence corresponding to the spot mass center and the reference gene sequence; or the percentage of the number of different bases in the base sequence corresponding to the spot mass center. If the difference information indicates that the difference between the base sequence corresponding to the spot mass center and the difficult-to-test gene sequence is too large, then the confidence of the spot area where the spot mass center is located is less than the preset confidence threshold, so the spot area where the spot mass center is located needs to be screened out. Afterwards, the spot label of the original fluorescent image is constructed based on the spot boundary and the spot mass center corresponding to the spot area in the original fluorescent image.

[0163] like Fig.12 In the example shown, the original fluorescence image after spot segmentation is as follows: Fig.12 As shown in a, in the original fluorescence image, three spot areas are identified, and the cluster boundary (spot boundary) and cluster centroid (spot centroid) of each cluster are shown in Fig.12 As shown in a.

[0164] Afterwards, the spot label is constructed according to the needs of deep learning model training. For example, if the trained model is a three-classification model, the background information, spot boundary information and spot centroid information can be predicted, and the constructed spot label includes background, spot boundary and spot centroid; if the trained model is a two-classification model, and the background information and spot centroid information can be predicted, the constructed spot label includes background and spot centroid.

[0165] In the disclosed embodiments, the spot labels constructed for the fluorescence images can be used to train a deep learning model, which can be used to construct templates for the fluorescence images to obtain reference templates for base signal recognition. The reference template predicts the information of all clusters in the fluorescence images of multiple sequencing cycles.

[0166] Specifically, during high-throughput sequencing, DNA sequence information is spread across multiple channels. For example, two-channel sequencing requires fluorescent images of two channels to determine the base type, and four-channel sequencing requires fluorescent images of four channels to determine the base type. Since a single-frame fluorescent image may have clusters that do not generate detectable signals (such as the G signal in the two-color sequencing method, which usually does not emit light), and the cluster also needs to be determined to have a base type, it is necessary to use multiple frames of original fluorescent images corresponding to multiple sequencing cycles to construct a template. In the reference template for base signal recognition obtained by template construction, fluorescent signals that can be detected by multiple sequencing cycles will be included to complete the position of cluster signals that may be missing in certain fluorescent images, to solve the uncertainty of the spot position, and to avoid missing cluster signals. In this way, in the process of base signal recognition of each frame of fluorescent image, the spot information of the cluster in the reference template is used as a reference. If there is a lack of spot information of the cluster at a certain position in the fluorescent image, it can also be detected by the reference template to obtain a higher base signal recognition result.

[0167] When performing the template construction process, the purpose is to enable the obtained reference template to complete the position of the cluster signal that may be missing in some fluorescence images, and the position of the missing cluster signal is usually different in different original fluorescence images. Therefore, the template reconstruction process can be performed based on at least part of the original fluorescence image, that is, the above purpose can be achieved. The original fluorescence image used for template construction can be the original fluorescence image taken by 5, 10, 15, 20, 25, or 30 sequencing cycles, which can be determined according to actual needs, and the embodiment of the present disclosure does not limit it.

[0168] In the field of edge synthesis sequencing, since not all clusters produce fluorescence each time a photo is taken, the location of the cluster cannot be accurately determined by a fluorescence image of only one cycle, and more fluorescence images are needed to assist in determining the location and shape of the cluster. There are clusters in a single-frame fluorescence image that do not produce detectable signals, and the base type of the cluster also needs to be determined, so multiple imaging cycles are needed to resolve the uncertainty of the cluster location to avoid missing cluster signals. Since the DNA consistency at the center of the cluster is higher and the consistency is lower toward the edge, the DNA location information obtained by locating the cluster coordinates as the center of the cluster will be more accurate. Using more cycle images can build a more accurate reference template, and an accurate reference template helps to find a more accurate spot location and spot shape. As shown in Table 1 below, it can be seen that with the increase of superposition cycles, the filtering ratio of the cluster is increasing, and the number of clusters obtained in the end is slightly reduced. However, with the increase of superposition images, the alignment rate is increasing, the strict alignment rate is also increasing, and the error rate is decreasing. Therefore, it is believed that superposition of more cycles of fluorescence images can produce a more accurate reference template, and then find a more accurate spot location, thereby improving the quality of the final data.

[0169] Table 1

[0170]

[0171] In the above Table 1, T10 represents the fluorescence image of the template constructed using 10 sequencing cycles; T15 represents the fluorescence image of the template constructed using 15 sequencing cycles; T20 represents the fluorescence image of the template constructed using 20 sequencing cycles.

[0172] like Fig.12 As shown in b, it is the binary classification spot label constructed by the deep learning model to predict the spot center of mass and background, where the green pixel position is the location of the spot center of mass;

[0173] like Fig.12 As shown in c, it is the multi-classification label constructed by the deep learning model for predicting the background, the center of mass of the light spot, the area of ​​the light spot, and the boundary of the light spot. Fig.12 The middle c is marked with a different green color.

[0174] like Fig.12 As shown in (d), it is the binary classification label constructed by the deep learning model to predict the background and light spot area; the green pixel position is the light spot area.

[0175] The labels constructed according to the spot position and spot shape on the fluorescence image can be used for supervised classification, regression and other neural network or machine learning model training. The input of the neural network or machine learning model is the image of multiple sequencing cycles, and the output can be, for example, a reference template. The images of multiple sequencing cycles can be the original pixel level or the image after amplification or reduction. The images of multiple sequencing cycles are used to solve the uncertainty of the spot position to avoid missing cluster signals. The classification network can choose UNet, encoder-decoder, deep convolution inversion image network, etc.

[0176] Those skilled in the art will appreciate that, in the above method of specific implementation, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of the steps should be determined by their functions and possible internal logic.

[0177] The present disclosure also provides a computer device, such as Fig.13 FIG. 1 is a schematic diagram of a computer device structure provided in an embodiment of the present disclosure, including:

[0178] Processor 131 and memory 132; the memory 132 stores machine-readable instructions executable by the processor 131, and the processor 131 is used to execute the machine-readable instructions stored in the memory 132. When the machine-readable instructions are executed by the processor 131, the processor 131 performs the following steps:

[0179] Performing image alignment on multiple frames of original fluorescence images from multiple sequencing cycles to obtain aligned multiple frames of fluorescence images;

[0180] Determine the base sequence of the candidate pixel position of the light spot according to the gray value of each pixel position in the aligned multiple frames of fluorescent images; the candidate pixel position of the light spot includes at least part of the pixel positions in the fluorescent images;

[0181] Based on consistency information between the base sequence of the candidate light spot pixel position and the base sequence of the candidate light spot pixel position in the first neighborhood corresponding thereto, determining a first pixel position belonging to the center of the fluorescent light spot from the candidate light spot pixel positions;

[0182] At least part of the first pixel positions are used as seed points, and a region growing process is performed on the seed points to obtain a light spot region of the light spot where the seed points are located.

[0183] The above-mentioned memory 132 includes internal memory 1321 and external memory 1322; the memory 1321 here is also called internal memory, which is used to temporarily store the calculation data in the processor 131, as well as the data exchanged with the external memory 1322 such as the hard disk. The processor 131 exchanges data with the external memory 1322 through the internal memory 1321.

[0184] The specific execution process of the above instructions can refer to the steps of the method for identifying fluorescent spots in fluorescent images described in the embodiment of the present disclosure, which will not be repeated here.

[0185] The present disclosure also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps of the method for identifying fluorescent spots in fluorescent images described in the above method embodiment are executed. The storage medium can be a volatile or non-volatile computer-readable storage medium.

[0186] The present disclosure also provides a computer program product that carries a program code. The program code includes instructions that can be used to execute the steps of the method for identifying fluorescent spots in fluorescent images described in the above method embodiment. For details, please refer to the above method embodiment, which will not be repeated here.

[0187] The computer program product may be implemented in hardware, software or a combination thereof. In one optional embodiment, the computer program product is implemented as a computer storage medium. In another optional embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0188] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than to limit them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure is described in detail with reference to the aforementioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the aforementioned embodiments within the technical scope disclosed in the present disclosure, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be based on the protection scope of the claims.

Claims

1. A method for identifying a fluorescent spot in a fluorescent image, characterized in that: include: Performing image alignment on multiple frames of original fluorescence images from multiple sequencing cycles to obtain aligned multiple frames of fluorescence images; Determining the base sequence of the candidate pixel position of the light spot among the plurality of pixel positions according to the gray value of each pixel position in the aligned multiple frames of fluorescent images; Based on the consistency information between the base sequence of the candidate light spot pixel position and the base sequence of the candidate light spot pixel position in the first neighborhood corresponding thereto, determining the first pixel position belonging to the central area of ​​the fluorescent light spot from the candidate light spot pixel positions; At least part of the first pixel positions are used as seed points, and a region growing process is performed on the seed points to obtain a light spot region of the light spot where the seed points are located.

2. The identification method according to claim 1, characterized in that: The image alignment of multiple frames of original fluorescence images derived from multiple sequencing cycles includes: Performing image reconstruction processing on multiple frames of original fluorescence images from multiple sequencing cycles to obtain multiple frames of reconstructed fluorescence images; Determining an alignment template for image alignment from the multiple frames of reconstructed fluorescence images; Based on the alignment template, image alignment processing is performed on multiple frames of original fluorescence images to obtain multiple frames of aligned fluorescence images.

3. The identification method according to claim 1, characterized in that: The step of determining the base sequence of the candidate pixel positions of the light spots in the plurality of pixel positions according to the grayscale value of each pixel position in the aligned multiple frames of fluorescent images comprises: Based on the grayscale value of each pixel point in the aligned multi-frame fluorescence image, the foreground and background of the aligned multi-frame fluorescence image are segmented to obtain candidate fluorescence spot areas in the aligned multi-frame fluorescence image; and the candidate fluorescence spot areas in the aligned multi-frame fluorescence image are fused to obtain fused fluorescence spot areas in the aligned multi-frame fluorescence image; and performing image sharpening processing on the aligned multiple frames of fluorescence images respectively to obtain sharpened multiple frames of fluorescence images; Each pixel position in the fused fluorescent spot area is used as the spot candidate pixel position, and the base sequence corresponding to the spot candidate pixel position is determined according to the gray value of the spot candidate pixel position in the sharpened multi-frame fluorescent image.

4. The identification method according to any one of claims 1 to 3, characterized in that: The consistency information between the base sequence of the candidate pixel position of the light spot and the base sequence of the pixel position in the corresponding first neighborhood is determined in the following manner: Determine a first neighborhood of each candidate pixel position of the light spot; Determining the degree of difference between the base sequence of each of the light spot candidate pixel positions and the base sequences of each of the light spot candidate pixel positions in the corresponding first neighborhood; The degree of difference includes: the number of different bases between the base sequence of each of the candidate light spot pixel positions and the base sequence of the candidate light spot pixel position in the first neighborhood, or the percentage of the number of different bases in the base sequence of the candidate light spot pixel position; The consistency information is determined according to the degree of difference.

5. The identification method according to any one of claims 1 to 4, characterized in that: The step of taking at least part of the first pixel positions as seed points and performing region growing processing on the seed points to obtain a light spot region of the light spot where the seed points are located includes: Perform multiple iterations and perform the following region growing process in each iteration: Determine the target seed point corresponding to the current iteration cycle from the first pixel position where the light spot has not been determined to belong to; Determine, according to the position information of the target seed point, a candidate light spot pixel position within a second neighborhood of the target seed point; Based on the similarity between the base sequence of the target seed point and the base sequence of the candidate light spot pixel position in the corresponding second neighborhood, the target seed point is subjected to region growing processing to obtain the light spot region of the light spot where the target seed point is located.

6. The identification method according to claim 5, characterized in that: The method of performing a region growing process on the target seed point based on the similarity between the base sequence of the target seed point and the base sequence of the pixel position in the corresponding second neighborhood to obtain a spot region of the spot where the target seed point is located includes: Determine, according to the similarity between the base sequence of the target seed point and the base sequence of the candidate pixel position of the light spot in the corresponding second neighborhood, a second pixel position belonging to the same light spot area as the target seed point, and determine, from the second pixel position, a growable pixel position corresponding to the target seed point; The following steps are executed repeatedly until no new growable pixel position appears: using the similarity between the base sequence of the growable pixel position and the base sequence of the spot candidate pixel position in the corresponding third neighborhood, a new second pixel position belonging to the same spot area as the target seed point is determined; and in the new second pixel position, a new growable pixel position corresponding to the target seed point is determined; The light spot area is constructed based on the target seed point and the second pixel position.

7. The identification method according to claim 6, characterized in that: Determining a new second pixel position that belongs to the same light spot area as the target seed point by using the similarity between the base sequence of the growable pixel position and the base sequence of the pixel position in the corresponding third neighborhood; And determining a new growable pixel position corresponding to the target seed point in the new second pixel position, including: Comparing the similarity with a first similarity threshold; If the similarity is greater than the first similarity threshold, determining the candidate pixel position of the light spot in the second neighborhood as the second pixel position; and when the similarity is greater than the first similarity threshold, comparing the similarity with a second similarity threshold; If the similarity is greater than the second similarity threshold, the corresponding second pixel position is determined as the growable pixel position corresponding to the target seed point.

8. The identification method according to claim 6 or 7, characterized in that: The method further comprises: determining whether the second pixel position belongs to a candidate center set consisting of the first pixel positions; If the second pixel position belongs to the candidate center set, the second pixel position is deleted from the candidate center set.

9. A computer device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the processor is used to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the processor executes the steps of the method for identifying fluorescent spots in a fluorescent image as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program. When the computer program is executed by a computer device, the computer device executes the steps of the method for identifying fluorescent spots in a fluorescent image according to any one of claims 1 to 8.

Citation Information

Cited By

  • Single machine offline updating method of deep learning model for gene sequencing, computer equipment and storage medium

    CN120560684A

  • High-voltage power supply DNA sequencing visual detection method and system

    CN121280430A

  • Fluorescence image signal enhancement system and method based on three-dimensional DNA nanostructure

    CN122024232A