Single molecule recognition, counting method and device and processing system

By performing line graph transformation and histogram statistics on the time series of image bright spot intensity, the problems of slow speed and low accuracy of single molecule recognition and counting in existing technologies are solved, and fast and efficient single molecule recognition and counting are achieved.

CN108229097BActive Publication Date: 2026-01-16GENEMIND BIOSCIENCES CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN201710607570.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-12-09
Filing Date
2017-07-24
Publication Date
2026-01-16
Estimated Expiration
2037-07-24

AI Technical Summary

Technical Problem

Existing single-molecule recognition and counting methods rely on human eye recognition, which is slow and labor-intensive. Methods based on HMM and machine learning require a large number of samples for training and are not efficient, and they also suffer from recognition errors and inaccurate counting.

Method used

By converting the time series of image bright spot intensity into a line graph, performing grid division and histogram statistics, and using the threshold condition of the maximum point to identify and count single molecules, including input unit, transformation unit, grid statistics unit, histogram statistics unit and decision unit, fast and high-precision single molecule identification and counting is achieved.

Benefits of technology

It achieves rapid and high-precision single-molecule identification and counting, reduces the need for manual intervention and sample training, and improves the efficiency and accuracy of identification and counting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN108229097B_ABST
    Figure CN108229097B_ABST
Patent Text Reader

Abstract

The application discloses a single molecule recognition and counting method and device. The single molecule recognition method comprises the following steps: inputting a time sequence of image bright point intensity; forming a time and intensity broken line graph of the image bright point according to the time sequence, wherein the broken line graph is composed of multiple line segments; performing grid division on the broken line graph to form multiple grids arranged in an array, and counting the number of line segments and / or endpoints of the line segments falling in each grid; grouping based on the size of the intensity, frequency counting of the number to obtain a histogram; finding a maximum value point of the histogram, and determining that a peak where the maximum value point is located corresponds to a single molecule when the value of the maximum value point is greater than a first set threshold value and the width of the peak where the maximum value point is located is greater than a second set threshold value. The single molecule recognition method can quickly recognize single molecules by converting the broken line graph of the time sequence of the bright point intensity into image processing to obtain a histogram, and the recognition accuracy is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of gene sequencing technology, in particular to a single molecule recognition and counting method, a recognition and counting device and a processing system. BACKGROUND

[0002] In the related art, the third generation sequencing technology is single molecule sequencing, and the single molecule sequencing technology based on imaging optical detection is a base recognition technology relying on optical signals and electrical signals. Among them, the base group relying on fluorescence recognition carries fluorescence which is emitted light intensity from the excited state to the ground state under the irradiation of a specific power laser. However, the different lengths of time of different fluorescent groups, the different light intensities emitted, and the existence of background noise will all cause errors in single molecule recognition. At the same time, uneven distribution of DNA chains and base group clustering will also cause a decrease in effective single molecules.

[0003] The existing method mainly relies on the human eye to recognize and count single molecules on the collected fluorescence image, but such a method consumes manpower and is also slow. Referring to voice recognition, a method based on HMM and machine learning is used, which not only needs a large number of sample training, but also has low running efficiency. SUMMARY

[0004] The embodiments of the present application aim to at least solve one of the technical problems existing in the prior art. To this end, the embodiments of the present application need to provide a single molecule recognition and counting method, a recognition and counting device and a processing system.

[0005] The single molecule recognition method of the embodiments of the present application comprises the steps of: inputting a time sequence of image bright spot intensity; forming a time and intensity broken line graph of the image bright spot according to the time sequence, the broken line graph being composed of a plurality of line segments; performing grid division on the broken line graph to form a plurality of grids arranged in an array, and counting the number of line segments and / or end points of the line segments falling in each grid; grouping based on the size of the intensity, frequency counting the number to obtain a histogram; and finding a maximum value point of the histogram, and determining that a peak where the maximum value point is located corresponds to one single molecule if the value of the maximum value point is greater than a first set threshold and the width of the peak where the maximum value point is located is greater than a second set threshold. The above single molecule recognition method can quickly recognize single molecules by converting the broken line graph of the time sequence of bright spot intensity into image processing to obtain a histogram, and the recognition accuracy is also high.

[0006] The single molecule counting method of the embodiment of the present application comprises the steps of: inputting a time sequence of image bright spot intensity; forming a time-intensity line graph of the image bright spot according to the time sequence, the line graph being composed of a plurality of line segments; performing grid division on the line graph to form a plurality of grids arranged in an array, and counting the number of line segments and / or end points of the line segments falling in each grid; grouping based on the intensity, and frequency counting the number to obtain a histogram; finding a maximum value point of the histogram, and determining that a peak where the maximum value point is located corresponds to a single molecule when the value of the maximum value point is greater than a first set threshold and the width of the peak where the maximum value point is located is greater than a second set threshold; and calculating to obtain a number S1 of single molecules. The single molecule counting method can quickly count single molecules with high accuracy by converting the line graph of the time sequence of bright spot intensity into image processing to obtain a histogram.

[0007] The single molecule counting method of the embodiment of the present application comprises the steps of: inputting a time sequence of image bright spot intensity; forming a time-intensity line graph of the image bright spot according to the time sequence, the line graph being composed of a plurality of line segments; performing grid division on the line graph to form a plurality of grids arranged in an array, and counting the number of line segments and / or end points of the line segments falling in each grid; grouping based on the intensity, and frequency counting the number to obtain a histogram; finding a maximum value point of the histogram, and determining that a peak where the maximum value point is located corresponds to a single molecule when the value of the maximum value point is greater than a first set threshold and the width of the peak where the maximum value point is located is greater than a second set threshold; and calculating to obtain a number S1 of single molecules. The single molecule counting method can quickly count single molecules with high accuracy by converting the line graph of the time sequence of bright spot intensity into image processing to obtain a histogram.

[0008] The single molecule recognition device of the embodiment of the present application is used to implement part or all steps of the single molecule recognition method of the above-mentioned aspect of the present application, and comprises: an input unit configured to input a time sequence of image bright spot intensity; a conversion unit configured to form a time-intensity line graph of the image bright spot according to the time sequence in the input unit, the line graph being composed of a plurality of line segments; a grid statistics unit configured to perform grid division on the line graph from the conversion unit to form a plurality of grids arranged in an array, and count the number of the line segments and / or the end points of the line segments falling in each grid; a histogram statistics unit configured to group based on the size of the intensity, and frequency count the number from the grid statistics unit to obtain a histogram; and a determination unit configured to find a maximum value point of the histogram from the histogram statistics unit, and determine that a peak where one maximum value point is located corresponds to one single molecule if the value of the maximum value point is greater than a first set threshold value and the width of the peak where the maximum value point is located is greater than a second set threshold value. The single molecule recognition device described above can quickly recognize single molecules by converting the line graph of the time sequence of bright spot intensity into image processing to obtain a histogram, and the recognition accuracy is also high.

[0009] The single molecule counting device of the embodiment of the present application is used to implement part or all steps of the single molecule counting method of the above-mentioned aspect of the present application, and comprises: an input unit configured to input a time sequence of image bright spot intensity; a conversion unit configured to form a time-intensity line graph of the image bright spot according to the time sequence in the input unit, the line graph being composed of a plurality of line segments; a grid statistics unit configured to perform grid division on the line graph from the conversion unit to form a plurality of grids arranged in an array, and count the number of the line segments and / or the end points of the line segments falling in each grid; a histogram statistics unit configured to group based on the size of the intensity, and frequency count the number from the grid statistics unit to obtain a histogram; a determination unit configured to find a maximum value point of the histogram from the histogram statistics unit, and determine that a peak where one maximum value point is located corresponds to one single molecule if the value of the maximum value point is greater than a first set threshold value and the width of the peak where the maximum value point is located is greater than a second set threshold value; and a calculation unit configured to calculate the number S1 of single molecules. The single molecule counting device described above can quickly count single molecules by converting the line graph of the time sequence of bright spot intensity into image processing to obtain a histogram, and the counting accuracy is also high.

[0010] The single-molecule counting device of the embodiment of the present application is used to implement part or all of the steps of the single-molecule counting method of the aspect of the present application, and includes: an input unit configured to input a time sequence of image bright spot intensity; a conversion unit configured to form a time-intensity line graph of the image bright spot according to the time sequence in the input unit, the line graph being composed of a plurality of line segments; a grid counting unit configured to perform grid division on the line graph from the conversion unit to form a plurality of grids arranged in an array, and count the number of line segments and / or end points of the line segments falling in each grid; a histogram counting unit configured to group based on the size of the intensity, and frequency count the number from the grid counting unit to obtain a histogram; and a determination unit configured to find a maximum value point of the histogram from the histogram counting unit, and determine to add 1 to the count of single molecules when the value of the maximum value point is greater than a first set threshold value and the width of the peak where the maximum value point is located is greater than a second set threshold value. The single-molecule counting device described above can quickly count single molecules by converting the time sequence of bright spot intensity into image processing to obtain a histogram, and the accuracy of the counting is also high.

[0011] The single-molecule processing system of the embodiment of the present application includes: a data input device configured to input data; a data output device configured to output data; a storage device configured to store data, the data including a computer executable program; and a processor configured to execute the computer executable program, the execution of the computer executable program including the completion of the method of any of the above embodiments. The single-molecule processing system can realize single-molecule recognition and / or single-molecule counting.

[0012] The computer readable storage medium of the embodiment of the present application is used to store a program for execution by a computer, and the execution of the program includes the completion of the method of any of the above embodiments. The computer readable storage medium can include a read-only memory, a random access memory, a magnetic disk or an optical disk, etc.

[0013] Additional aspects and advantages of the embodiments of the present application will be in part apparent and in part pointed out hereinafter in the description of the embodiments of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0014] The above and / or additional aspects and advantages of the embodiments of the present application will become apparent and be readily appreciated from the description of the embodiments of the present application, taken in conjunction with the following drawings:

[0015] Figure 1 is a flowchart of the single-molecule recognition method of the embodiment of the present application.

[0016] Figure 2is another flowchart of the single molecule recognition method of the embodiment of the present application.

[0017] Figure 3 is still another flowchart of the single molecule recognition method of the embodiment of the present application.

[0018] Figure 4 is still another flowchart of the single molecule recognition method of the embodiment of the present application.

[0019] Figure 5 is still another flowchart of the single molecule recognition method of the embodiment of the present application.

[0020] Figure 6 is still another flowchart of the single molecule recognition method of the embodiment of the present application.

[0021] Figure 7 is still another flowchart of the single molecule recognition method of the embodiment of the present application.

[0022] Figure 8 is still another flowchart of the single molecule recognition method of the embodiment of the present application.

[0023] Figure 9 is a curve diagram of the Mexican hat filtering of the single molecule recognition method of the embodiment of the present application.

[0024] Figure 10 is still another flowchart of the single molecule recognition method of the embodiment of the present application.

[0025] Figure 11 is a diagram of 8-connected pixels in the single molecule recognition method of the embodiment of the present application.

[0026] Figure 12 is a diagram of a line graph of the single molecule recognition method of the embodiment of the present application.

[0027] Figure 13 is a diagram of grid division of the line graph in the single molecule recognition method of the embodiment of the present application.

[0028] Figure 14 is a diagram of the line graph before filtering in the single molecule recognition method of the embodiment of the present application.

[0029] Figure 15 is a diagram of the line graph after filtering in the single molecule recognition method of the embodiment of the present application.

[0030] Figure 16 is another diagram of the line graph of the single molecule recognition method of the embodiment of the present application.

[0031] Figure 17is a schematic diagram of a histogram after equalization in a single molecule recognition method of an embodiment of the present application.

[0032] Figure 18 is another schematic diagram of a flowchart of a single molecule recognition method of an embodiment of the present application.

[0033] Figure 19 is a schematic diagram of a line erosion process in a single molecule recognition method of an embodiment of the present application.

[0034] Figure 20 is another schematic diagram of a line erosion process in a single molecule recognition method of an embodiment of the present application.

[0035] Figure 21 is a schematic diagram of 8-connected windows in a single molecule recognition method of an embodiment of the present application.

[0036] Figure 22 is a schematic diagram of identifying connected regions in a single molecule recognition method of an embodiment of the present application.

[0037] Figure 23 is a schematic diagram of a flowchart of a single molecule counting method of an embodiment of the present application.

[0038] Figure 24 is another schematic diagram of a flowchart of a single molecule counting method of an embodiment of the present application.

[0039] Figure 25 is yet another schematic diagram of a flowchart of a single molecule counting method of an embodiment of the present application.

[0040] Figure 26 is still another schematic diagram of a flowchart of a single molecule counting method of an embodiment of the present application.

[0041] Figure 27 is a schematic diagram of a module of a single molecule recognition apparatus of an embodiment of the present application.

[0042] Figure 28 is another schematic diagram of a module of a single molecule recognition apparatus of an embodiment of the present application.

[0043] Figure 29 is yet another schematic diagram of a module of a single molecule recognition apparatus of an embodiment of the present application.

[0044] Figure 30 is another schematic diagram of a module of a single molecule recognition apparatus of an embodiment of the present application.

[0045] Figure 31 is yet another schematic diagram of a module of a single molecule recognition apparatus of an embodiment of the present application.

[0046] Figure 32is a further schematic view of a module of the single molecule recognition device of the present application.

[0047] Figure 33 is a further schematic view of a module of the single molecule recognition device of the present application.

[0048] Figure 34 is a further schematic view of a module of the single molecule recognition device of the present application.

[0049] Figure 35 is a further schematic view of a module of the single molecule recognition device of the present application.

[0050] Figure 36 is a further schematic view of a module of the single molecule recognition device of the present application.

[0051] Figure 37 is a schematic view of a module of the single molecule counting device of the present application.

[0052] Figure 38 is a further schematic view of a module of the single molecule counting device of the present application.

[0053] Figure 39 is a further schematic view of a module of the single molecule counting device of the present application.

[0054] Figure 40 is a further schematic view of a module of the single molecule counting device of the present application.

[0055] Figure 41 is a further schematic view of a module of the single molecule counting device of the present application.

[0056] Figure 42 is a further schematic view of a module of the single molecule counting device of the present application.

[0057] Figure 43 is a schematic view of a module of the single molecule processing system of the present application. DETAILED DESCRIPTION

[0058] Embodiments of the present application are described below in detail, examples of which are shown in the drawings, wherein the same or similar reference numerals used throughout refer to the same or similar elements or elements having the same or similar functions. The embodiments described below by reference to the drawings are exemplary only, for the purpose of explanation, and are not to be understood as limiting the present application.

[0059] In the description of the application, it needs to be understood that the terms "first", "second" are only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the application, the meaning of "a plurality of" is two or more, unless otherwise explicitly specified and limited.

[0060] In the description of the application, it needs to be noted that, unless otherwise explicitly specified and limited, "connection" should be understood broadly, for example, it can be fixed connection, or detachable connection, or integrally connected; it can be mechanical connection, or electrical connection or can communicate with each other; it can be directly connected, or indirectly connected through intermediate medium, it can be the internal communication of two elements or the interaction relationship between two elements. For those skilled in the art, the specific meaning of the above terms in the application can be understood according to the specific circumstances.

[0061] The following disclosure provides many different embodiments or examples for implementing different structures of the application. In order to simplify the disclosure of the application, the components and settings of specific examples are described below. In addition, the present application can repeatedly refer to numbers and / or letters in different examples, and such repetition is for the purpose of simplification and clarity, which itself does not indicate the relationship between the various embodiments and / or settings discussed.

[0062] The single molecule recognition method and counting method of the embodiments of the application can be applied to gene sequencing. The "gene sequencing" referred to in the embodiments of the application is the same as nucleic acid sequencing, including DNA sequencing and / or RNA sequencing, including long fragment sequencing and / or short fragment sequencing.

[0063] Please refer to Figure 1The single-molecule recognition method of the embodiment of the present application comprises the following steps: S01, inputting a time sequence of image bright spot intensity; S02, forming a time-intensity broken line graph of the image bright spot according to the time sequence, the broken line graph being composed of a plurality of line segments; S03, performing grid division on the broken line graph to form a plurality of grids arranged in an array, and counting the number of line segments and / or end points of the line segments falling in each grid; S04, grouping based on the size of the intensity, and frequency counting the number to obtain a histogram; and S05, finding a maximum value point of the histogram, and determining that a peak where the maximum value point is located corresponds to a single molecule if the maximum value point satisfies the following conditions: the value of the maximum value point is greater than a first set threshold value, and the width of the peak where the maximum value point is located is greater than a second set threshold value. The single-molecule recognition method described above can quickly recognize single molecules by converting the broken line graph of the time sequence of bright spot intensity into image processing to obtain a histogram, and the recognition accuracy is also high. The single-molecule recognition method based on histogram statistics can accurately recognize single molecules according to the time sequence data of the intensity of the bright spot, and is particularly suitable for the case where the number of single molecules contained in the bright spot is greater than 3.

[0064] Specifically, in step S01, when the image bright spot is formed, a specific wavelength laser is used to irradiate the test sample to make the test sample emit fluorescence, and then a camera is used to collect the fluorescence to form an image. In the image, there are image bright spots corresponding to the part (nucleic acid molecules) of the test sample emitting fluorescence. The so-called "bright spot" refers to a light-emitting point on the image, and one light-emitting point occupies at least one pixel point. The so-called "pixel point" is the same as "pixel".

[0065] In an embodiment of the present application, the image comes from a single-molecule sequencing platform, such as the sequencing platform of Helicos or Pacific Biosciences (PacBio) Company. The input raw data is the parameters of the pixel points of the image, and the detection of the so-called "bright spot" is the detection of the single-molecule optical signal.

[0066] In some embodiments, please refer to Figure 2 The single-molecule recognition method further comprises: an image preprocessing step S31, the image preprocessing step S31 analyzes the input image to be processed to obtain a first image, the image to be processed contains at least one image bright spot, and the image bright spot has at least one pixel point; a bright spot detection step S32, the bright spot detection step S32 comprises the following steps: S321, analyzing the first image to calculate a bright spot determination threshold value, S322, analyzing the first image to obtain a candidate bright spot, S323, determining whether the candidate bright spot is an image bright spot according to the bright spot determination threshold value, if the determination result is yes, S324, then obtaining a time sequence of image bright spot intensity, and if the determination result is no, S325, discarding the candidate bright spot.

[0067] Therefore, the denoising processing of the to-be-processed image by the image preprocessing step can reduce the calculation amount of the bright spot detection step, and the bright spot judgment threshold can be used to judge whether the candidate bright spot is an image bright spot, thereby improving the accuracy of judging the image bright spot.

[0068] Specifically, in one example, the input to-be-processed image can be a 512*512 or 2048*2048 16-bit tiff format image, and the tiff format image can be a grayscale image. In this way, the processing process of the single molecule recognition method can be simplified.

[0069] In some embodiments, referring to Figure 3 The image preprocessing step S31 includes: performing background subtraction processing on the to-be-processed image to obtain a first image. In this way, the noise of the to-be-processed image can be further reduced, and the accuracy of the single molecule recognition and / or counting method is higher.

[0070] In some embodiments, referring to Figure 4 The image preprocessing step S31 includes: performing simplification processing on the to-be-processed image after the background subtraction processing to obtain a first image. In this way, the calculation amount of the subsequent single molecule recognition and / or counting method can be reduced.

[0071] In some embodiments, referring to Figure 5 The image preprocessing step S31 includes: performing filtering processing on the to-be-processed image to obtain a first image. In this way, the filtering processing on the to-be-processed image can obtain the first image while preserving the image detail features as much as possible, thereby improving the accuracy of the single molecule recognition and / or counting method.

[0072] In some embodiments, referring to Figure 6 The image preprocessing step S31 includes: performing filtering processing on the to-be-processed image after the background subtraction processing to obtain a first image. In this way, the filtering processing on the to-be-processed image after the background subtraction processing can further reduce the noise of the to-be-processed image, and the accuracy of the single molecule recognition and / or counting method is higher.

[0073] In some embodiments, referring to Figure 7 The image preprocessing step S31 includes: performing simplification processing on the to-be-processed image after the background subtraction processing and the filtering processing to obtain a first image. In this way, the calculation amount of the subsequent image processing method can be reduced.

[0074] In some embodiments, referring to Figure 8 The image preprocessing step S31 includes: performing simplification processing on the to-be-processed image to obtain a first image. In this way, the calculation amount of the subsequent single molecule recognition and / or counting method can be reduced.

[0075] In some embodiments, the background subtraction processing on the image to be processed comprises: determining the background of the image to be processed by using the open operation, and performing the background subtraction processing on the image to be processed according to the background. In this way, the open operation is used to eliminate small objects, separate objects at fine points, and smooth the boundaries of large objects without significantly changing the area of the image, so that the image after the background subtraction processing can be more accurately obtained.

[0076] Specifically, in the embodiments of the present application, a window (for example, a 15*15 window) is moved in the image to be processed f(x, y) (for example, a gray image), and the background of the image to be processed is estimated by using the open operation (first erosion and then dilation), as shown in the following formula 1 and formula 2:

[0077] g(x, y) = erode [f(x, y), B] = min{f(x+x', y+y')-B(x', y')|(x', y')∈D b} Formula 1,

[0078] wherein g(x, y) is the eroded gray image, f(x, y) is the original gray image, and B is the structural element.

[0079] g(x, y) = dilate [f(x, y), B] = max{f(x-x', y-y')-B(x', y')|(x', y')∈D b} Formula 2.

[0080] wherein g(x, y) is the dilated gray image, f(x, y) is the original gray image, and B is the structural element.

[0081] Therefore, the background noise g = imopen (f(x, y), B) = dilate [erode (f(x, y), B)] can be obtained. Formula 3.

[0082] The original image is subjected to background subtraction:

[0083] f = f-g = {f(x, y)-g(x, y)|(x, y)∈D} Formula 4.

[0084] It can be understood that the specific method of the present embodiment for performing the background subtraction processing on the image to be processed can be applied to the step of performing the background subtraction processing on the image to be processed mentioned in any of the above embodiments.

[0085] In some embodiments, the filtering processing is a Mexican hat filtering processing. The Mexican hat filtering is easy to implement, which reduces the cost of the single molecule recognition and / or counting method, and at the same time, the Mexican hat filtering can improve the contrast between the foreground and the background, so that the foreground is brighter and the background is darker.

[0086] In performing the Mexican hat filtering, a m*m window is used to perform Gaussian filtering on the to-be-processed image before the filtering, and two-dimensional Laplacian sharpening is performed on the to-be-processed image after the Gaussian filtering, where m is a natural number and an odd number greater than 1. In this way, the Mexican hat filtering is realized through two steps.

[0087] Specifically, please refer to Figure 9 , the Mexican hat kernel can be expressed as:

[0088] where x and y represent the coordinates of the pixel points.

[0089] First, a m*m window is used to perform Gaussian filtering on the to-be-processed image, as shown in the following formula 6:

[0090]

[0091] where t1 and t2 represent the position of the filtering window, and wt1,t2 represents the weight of the Gaussian filtering.

[0092] Then, two-dimensional Laplacian sharpening is performed on the to-be-processed image, as shown in the following formula 7:

[0093]

[0094] where K and k both represent the Laplacian operator, which is related to the sharpening target. If it is necessary to strengthen the sharpening or weaken the sharpening, K and k are modified.

[0095] In one example, m = 3, so m*m = 3*3. When performing Gaussian filtering, formula 6 becomes:

[0096]

[0097] It can be understood that the specific method of the Mexican hat filtering of the present embodiment can be applied to the step of performing filtering processing on the to-be-processed image mentioned in any of the above embodiments.

[0098] In some embodiments, the simplified image is a binary image. In this way, the binary image is easy to process and has a wide range of applications.

[0099] Specifically, in one example, the binary image can contain two numerical values of 0 and 1 representing different attributes of the pixel points, and the binary image can be expressed as:

[0100] In some embodiments, when performing the simplification processing, a signal-to-noise ratio matrix is obtained according to the to-be-processed image before the simplification processing, and the to-be-processed image before the simplification processing is simplified according to the signal-to-noise ratio matrix to obtain a first image.

[0101] In one specific example, the background subtraction processing can be performed on the to-be-processed image first, and then the SNR matrix can be obtained according to the to-be-processed image after the background subtraction processing. In this way, information can be obtained from the image with less noise, and the accuracy of the processing result of the single molecule recognition and / or counting method can be improved.

[0102] Specifically, in one example, the SNR matrix can be represented as: wherein x and y represent the coordinates of the pixel points, h represents the height of the image, w represents the width of the image, i∈w, j∈h.

[0103] In one example, the simplified image is a binary image, and the binary image can be obtained according to the SNR matrix, and the binary image is shown in formula 9:

[0104]

[0105] In the calculation of the SNR matrix, the background subtraction processing and / or the filtering processing can be performed on the to-be-processed image first, such as the background subtraction processing step and the filtering processing step of the above embodiment, and then the ratio matrix of the to-be-processed image after the background subtraction processing and the background can be obtained according to formula 4:

[0106] R = f / g = {f(x, y) / g(x, y) | (x, y) ∈ D} formula 10, wherein D represents the dimension (height * width) of the image f.

[0107] Thus, the SNR matrix can be obtained:

[0108] In some embodiments, the step of analyzing the first image to calculate the bright spot determination threshold value includes: processing the first image by the Otsu method to calculate the bright spot determination threshold value. In this way, the bright spot determination threshold value is found by a more mature and simple method, thereby improving the accuracy of the single molecule recognition and / or counting method and reducing the cost of the single molecule recognition and / or counting method. At the same time, the finding of the bright spot determination threshold value by using the first image can improve the efficiency and accuracy of the single molecule recognition and / or counting method.

[0109] Specifically, the Otsu method (OTSU algorithm) can also be called the maximum inter-class variance method, which uses the maximum inter-class variance to segment the image, which means the minimum misclassification probability and high accuracy. Assuming that the segmentation threshold value of the foreground and the background of the to-be-processed image is T, the proportion of the number of pixel points belonging to the foreground in the entire image is ω0, and the average gray value is μ0; the proportion of the number of pixel points belonging to the background in the entire image is ω1, and the average gray value is μ1. The total average gray value of the to-be-processed image is denoted as μ, and the inter-class variance is denoted as var, and then:

[0110] μ = ω0*μ0 + ω1*μ1 Equation 11

[0111] var = ω0(μ0 - μ) 2 + ω1(μ1 - μ) 2 Equation 12.

[0112] Substitute Equation 11 into Equation 12 to obtain equivalent Equation 13:

[0113] var = ω0ω1(μ1 - μ0) 2 Equation 13.

[0114] The traversal method is used to obtain the segmentation threshold T that maximizes the inter-class variance, i.e., the obtained highlight determination threshold T.

[0115] In some embodiments, referring to Figure 10 , the step of determining whether the candidate highlight is an image highlight according to the highlight determination threshold includes:

[0116] Step S41, find the pixel points greater than (h*h-1) connected in the first image and take the found pixel points as the center of the candidate highlight, h*h is in one-to-one correspondence with the highlight, each value in h*h corresponds to a pixel point, h is a natural number and an odd number greater than 1;

[0117] Step S42, determine whether the center of the candidate highlight satisfies the condition: I max *A BI *ceof guass > T, wherein I max is the strongest intensity of the center of the h*h window, A BI is the ratio of the set value in the h*h window in the first image, ceof guass is the correlation coefficient of the pixels of the h*h window and the two-dimensional Gaussian distribution, and T is the highlight determination threshold.

[0118] If the above condition is satisfied, S43, determine that the highlight corresponding to the center of the candidate highlight is an image highlight contained in the image to be processed.

[0119] If the above condition is not satisfied, S44, discard the highlight corresponding to the center of the candidate highlight. In this way, the detection of the image highlight is realized.

[0120] Specifically, I max can be understood as the strongest intensity of the center of the candidate highlight. In an example, h = 3, find the pixel points greater than 8 connected, as shown in Figure 11 . Take the found pixel points as the pixel points of the candidate highlight. I max is the strongest intensity of the center of the 3*3 window, A BIceof is the ratio of the number of pixels in the 3*3 window in the first image that have the set value. guass ceof is the correlation coefficient of the 3*3 window and the two-dimensional Gaussian distribution.

[0121] The first image is a simplified image. For example, the first image can be a binary image, that is, the set value in the binary image can be a value corresponding to a pixel point satisfying a set condition. In another example, the binary image can contain two values of 0 and 1 representing different attributes of the pixel points, and the set value is 1. BI BI is the ratio of the number of pixels in the h*h window in the binary image that have the value of 1. For example, please refer to formula 9, when SNR<=mean(SNR), BI=1.

[0122] In addition, in some embodiments, the value of h can be equal to the value of m selected when performing the Mexican hat filtering, that is, h=m.

[0123] In some embodiments, when the above-mentioned images are collected, the camera can sequentially collect fluorescence of multiple fields of view (FOV) in time sequence. Therefore, when the image data is obtained, the image data contains image bright spot intensity corresponding to the time sequence of camera collection.

[0124] In step S02, after obtaining the required image bright spot, the image bright spot intensity corresponding to the adjacent collection time is connected by a line to form a time-intensity line graph of the image bright spot, as shown in Figure 12 Figure 12 In the figure, the horizontal axis represents the time of collecting fluorescence, with the unit of millisecond (ms), and the vertical axis represents the image bright spot intensity. In one example, the time interval between two adjacent fluorescence collection is 20 ms.

[0125] The vertical axis is the corresponding bright spot intensity value. In the embodiments of the present application, the bright spot intensity value is the bright spot pixel value. For a 16-bit tiff image, the bright spot pixel value is in the range of 0-65535, and for an 8-bit grayscale image, the bright spot pixel value is in the range of 0-255. The 16-bit tiff image is used in the embodiments of the present application.

[0126] In step S03, the waveform of the line graph is converted into image processing for subsequent histogram statistics. The image processing of the line graph includes grid division of the line graph.

[0127] ​In some embodiments, the grid division of the line graph is performed according to the number of time frames of the collected intensity and the size of the intensity. In this way, the grid division of the line graph can be obtained by simpler processing, thereby reducing the cost of the single molecule recognition method. Specifically, the number of time frames can be divided into M and the size of the intensity can be divided into N, i.e. M*N grids are formed. The number of time frames of the collected intensity is the time interval between two adjacent collected fluorescence. In one embodiment, one grid can be referred to as the length direction along the horizontal axis and the height direction along the vertical axis. The length of one grid can be set as a multiple of the number of time frames, such as 1 times, 2 times, 2.5 times, etc. The height of one grid can be flexibly set, for example, for a 16-bit tiff image, the value of the vertical axis is 0-65535, and when the grid is divided, the value of the vertical axis can be normalized and divided into 50 parts, and then the height of one grid is set to 0.02, i.e. N=50.

[0128] In one example, the time interval between two adjacent collected fluorescence is 20 ms, the length of one grid is equal to one time interval, and the height is 0.02. Please refer to FIG. 6A. Figure 16 In such an example, the number of line segments falling into one grid can be 0, 1 or 2. Figure 16 The black dots in FIG. 6B represent the time sequence of the intensity of the image bright spot.

[0129] In one example, please refer to FIG. 6B. Figure 13 The line graph is divided into an 8*6 grid, and the number of line segments and / or the number of endpoints of the line segments falling into each grid is counted. In Figure 13 In FIG. 6C, the number of line segments falling into each grid (i.e. the number of times each grid is passed by the line segment) is counted, and the number in the grid represents the number of line segments falling into each grid. Figure 13 The black dots in FIG. 6B represent the time sequence of the intensity of the image bright spot.

[0130] In some embodiments, step S04 includes the step of dividing the intensity into N groups according to the size of the intensity, and counting the frequency of the number of times falling into the N groups: wherein n i represents the sum of the frequency of the number of times falling into the i-th row of the grid, j represents the number of time frames, g ij represents the frequency of the number of times falling into the grid (i,j), and M represents the number of time frames. In this way, the number of times can be converted into a histogram of intensity and number of times, making the subsequent single molecule recognition method operation simpler.

[0131] Specifically, in one example, the horizontal axis of the histogram represents the number of groups, and the vertical axis represents the frequency of the number of times falling into the corresponding number of groups. It should be noted that the value of N is equal to the value of N in the formation of M*N grids described above. The value of M is equal to the value of M in the formation of M*N grids described above.

[0132] In some embodiments, the step of grouping based on the size of the intensity and counting the number of times to obtain a histogram comprises the step of performing histogram equalization by L window: wherein n p represents the equalization of n i , n i represents the sum of the equalization results of n i , and p is an integer related to the size of the window L and the i-th row. In this way, the distribution of the histogram can be more uniform and easy to identify. The L window is used for histogram equalization, and the value of L is related to the size of the decay rate of single molecule fluorescence. Generally, if the single molecule fluorescence emission is quenched quickly, the value of L should not be too large. The precision of the histogram is affected by the size of the L window, and the value of L can be flexibly set to select the appropriate precision of the histogram. In one example, the value of L is in the range of [5, 15].

[0133] Please refer to Figure 17 , Figure 17 is the equalized histogram, and the horizontal axis of the histogram represents the number of groups, and the vertical axis represents the frequency of the number of times falling in the corresponding group number.

[0134] In step S05, all the maximum points of the histogram can be obtained by derivation. The first set threshold Q and the second set threshold H are related to the shape of the peak of the line graph. The sharper the peak, the larger the first set threshold Q and the smaller the second set threshold H; the fatter the peak, the smaller the first set threshold Q and the larger the second set threshold H. In one example, the value of the first set threshold Q is in the range of [2, 6], and the value of the second set threshold H is in the range of [4, 10].

[0135] In some embodiments, before the line graph is divided into a grid, the single molecule recognition method further comprises the step of filtering the line graph. In this way, the sudden error caused by light intensity flicker and camera sampling can be eliminated, and the waveform of the line graph is smoother. Specifically, the waveform modification can use median filtering based on an L2 size window: R = medium(Z i ). In one example, please refer to Figure 14 and Figure 15 , Figure 14 is the line graph before filtering, Figure 15 is the line graph after filtering, and as can be seen from the figure, the waveform of the line graph after filtering is smoother, which is conducive to improving the accuracy and efficiency of single molecule recognition.

[0136] In the embodiments of the present application, the maximum point is the peak point, and the maximum point is the vertex (inflection point) of the peak, that is, the peak corresponding to a single molecule is determined by a maximum point satisfying the condition.

[0137] In some embodiments, please refer to Figure 18The single-molecule recognition method further comprises the steps of: S51, performing line erosion on the grid-divided line graph to convert the grid-divided line graph into a simplified graph according to the number of times corresponding to each grid; S52, performing run-length encoding on the simplified graph to identify a connected region; and S53, calculating the area of each connected region to determine that one connected region corresponding to one single molecule satisfies the following condition: the area of the connected region is greater than a third set threshold.

[0138] Thus, the single-molecule recognition and / or counting method can be applied in a wider range. The single-molecule recognition method based on run-length encoding can accurately recognize single molecules according to the time series data of the intensity of the bright spot, and is particularly suitable for the case where the number of single molecules contained in one bright spot is not greater than 3. In the embodiment, the single molecules are recognized by combining the histogram-based method and the run-length encoding-based method, so that single molecules in line graphs (time series of bright spot intensity) of various waveforms can be accurately recognized.

[0139] In the line erosion, the following formula can be used for morphological erosion operation: g(x, y) = erode[f(x, y), B] = min{f(x+x', y+y')-B(x', y')|(x', y')∈D b}. Preferably, a linear structural element is selected, such as a window size of W*1. If the number of times of the grid in the window exceeds a threshold T, the grid is marked as a first value, otherwise it is marked as a second value. Thus, the grid-divided line graph can be converted into a simplified graph comprising the first value and the second value. In some embodiments, the simplified graph is a binary graph. For example, the first value can be taken as 1, and the second value can be taken as 0.

[0140] In one example, please refer to FIG. 6A. Figure 19 The length of one grid is L1, W = 2*L1, and T = 2, Figure 19 As shown in FIG. 6B, five grids are arranged along the length direction, and the numbers in the grids represent the number of times. When performing line erosion, the window is aligned with the grid. After line erosion, the five grids are marked as 0, 1, 0, 0, and 0, respectively.

[0141] In another example, please refer to FIG. 6C. Figure 20 The length of one grid is L1, W = 2*L1, and T = 2, Figure 20 As shown in FIG. 6D, five grids are arranged along the length direction, and the numbers in the grids represent the number of times. When performing line erosion, the window is misaligned with the grid. After line erosion, the five grids are marked as 0, 1, 0, 0, and 0, respectively.

[0142] It should be noted that the value of W is greater than or equal to the length of one grid, and preferably, W is an integer multiple of the length of one grid. In the example, W >= L1, and preferably, W is an integer multiple of L1.

[0143] In other examples, the threshold T is in the range of [6, 8], which is selected in relation to the fluctuation of the waveform of the fold line graph. The smaller the fluctuation, the larger the value of the threshold T.

[0144] For the convenience of understanding, the following description is given by taking 1 and 0 in the binary graph as an example when describing the run-length coding. It can be understood that other types of simplified graphs and other values of the first value and the second value can be changed by those skilled in the art according to the following description.

[0145] In the run-length coding, the 8-connected mode can be used. According to the grid, the respective connected regions are recursively connected according to the principle of 8-connection, and then the connected regions are identified by using the run-length coding. Specifically, by 8-connection (for example, using a 3*3 window as shown in FIG. 8), starting from a non-0 grid Q, if the grids in the 8 directions of the grid Q are all non-0, the grids in the 8 directions of the grid Q are identified as the same value as the grid Q, and so on. After the entire simplified graph is completed, the identification graph as shown in FIG. 9 can be obtained. Figure 21 Figure 18

[0146] In Figure 22 , different connected regions are identified by different values, and when calculating the area of each connected region, the number of times the same number appears is recorded as the area of the connected region, as shown in Figure 22 , the number 9 appears 9 times, so the area of the connected region corresponding to the number 9 is 9, and the number 7 appears 20 times, so the area of the connected region corresponding to the number 7 is 20.

[0147] The above example uses a recursive algorithm, and in other examples, a traversal algorithm can also be used to find the connected regions.

[0148] If the area of the connected region is greater than a third set threshold P, one such connected region corresponds to one single molecule. The value of P is related to the decay time of the single molecule fluorescence. In one example, the third set threshold P is in the range of [5, 10].

[0149] Please refer to Figure 23 ​​This invention discloses a method for counting single molecules, comprising the following steps: S81, inputting a time series of image bright spot intensities; S82, forming a line graph of image bright spot time versus intensity based on the time series, the line graph consisting of multiple line segments; S83, dividing the line graph into a grid to form multiple arrayed grids, and counting the number of times the line segments and / or the endpoints of the line segments fall within each grid; S84, grouping based on intensity magnitude, and performing frequency statistics to obtain a histogram; S85, finding the maximum point of the histogram, and determining that a peak containing a maximum point corresponds to a single molecule if the value of the maximum point is greater than a first set threshold and the width of the peak containing the maximum point is greater than a second set threshold; S86, calculating the number of single molecules S1. This single molecule counting method, by converting the time series of bright spot intensities into an image for histogram processing, can quickly count single molecules with high accuracy. It should be noted that the descriptions of the technical features and advantages of the single-molecule identification and / or counting methods in any of the above embodiments and examples, including the explanations and descriptions of the steps, parameter settings, and image preprocessing bright spot detection, are also applicable to the single-molecule counting method of this embodiment. To avoid redundancy, they will not be elaborated in detail here.

[0150] For example, in some embodiments, before dividing the line graph into a grid in step S83, the single-molecule counting method further includes a step of filtering the line graph. As another example, in some embodiments, please refer to... Figure 24 The single-molecule counting method further includes the following steps: S91, performing line erosion on the gridded line graph according to the number of times each grid corresponds to, to convert the gridded line graph into a simplified graph; S92, performing run-length encoding on the simplified graph to identify connected regions; S93, calculating the area of ​​each connected region, and determining that a connected region corresponds to a single molecule if the area of ​​the connected region is greater than a third preset threshold; S94, calculating the number of single molecules S2; S95, taking the smaller of S1 and S2 as the final number of single molecules. This single-molecule counting method based on histogram statistics is particularly suitable for accurately finding bright spots containing more than 3 single molecules, while the single-molecule counting method based on run-length encoding is particularly suitable for accurately finding bright spots containing less than or equal to 3 single molecules. In this embodiment, combining the two methods can accurately find and count single molecules in line graphs of various waveforms. In some embodiments, the simplified graph is a binary graph.

[0151] Please refer to Figure 25The single-molecule counting method of the embodiment includes the steps of: S61, inputting a time sequence of image bright spot intensity; S62, forming a time-intensity polyline of the image bright spot according to the time sequence, the polyline being composed of a plurality of line segments; S63, performing grid division on the polyline to form a plurality of grids arranged in an array, and counting the number of line segments and / or end points of the line segments falling in each grid; S64, grouping based on the size of the intensity, and frequency counting the number to obtain a histogram; S65, finding a maximum value point of the histogram, and determining to add 1 to the count of the single molecule when the value of the maximum value point is greater than a first set threshold and the width of the peak where the maximum value point is located is greater than a second set threshold. The single-molecule counting method described above can quickly count the single molecules by converting the polyline of the time sequence of the bright spot intensity into image processing to obtain the histogram, and the counting accuracy is also high.

[0152] It should be noted that the descriptions of the technical features and advantages of the single-molecule recognition and / or counting method in any of the above embodiments and examples, including the explanations and descriptions of the steps, parameter settings, and image preprocessing bright spot detection, are also applicable to the single-molecule counting method of the present embodiment. To avoid redundancy, they will not be described in detail here.

[0153] For example, in some embodiments, before the grid division on the polyline in step S63, the single-molecule counting method further includes a step of filtering the polyline. For another example, in some embodiments, please refer to Figure 26 The single-molecule counting method further includes the steps of: S71, performing line erosion on the grid-divided polyline to convert the grid-divided polyline into a simplified graph according to the number corresponding to each grid; S72, performing run-length encoding on the simplified graph to identify connected regions; S73, calculating the area of each connected region, and determining to add 1 to the count of the single molecule when the area of the connected region is greater than a third set threshold; and S74, taking the smaller one of the count of the single molecule obtained based on the histogram and the count of the single molecule obtained based on the run-length encoding as the final single molecule count. In this way, the single-molecule counting method can be applied to a wider range of applications, and a more accurate single molecule count can be obtained.

[0154] The single-molecule counting method based on histogram statistics is particularly suitable for accurately finding the number of single molecules contained in a bright spot when the number is greater than 3, and the single-molecule counting method based on run-length encoding is particularly suitable for accurately finding the number of single molecules contained in a bright spot when the number is less than or equal to 3. In this embodiment, the two methods are combined to accurately find and count single molecules in polylines of various waveforms. For example, the number of single molecules obtained based on the histogram is S1, the number of single molecules obtained based on the run-length encoding is S2, the smaller one of S1 and S2 is taken as the final single molecule count.

[0155] Please refer to Figure 27 The single molecule recognition device 200 of the embodiment of the present application is used to implement all or part of the steps of the single molecule recognition method in any of the above embodiments or examples. The single molecule recognition device 200 comprises: a first input unit 202 configured to input a time sequence of image bright spot intensity; a first conversion unit 204 configured to form a time-intensity broken line graph of image bright spots according to the time sequence in the first input unit 202, the broken line graph being composed of a plurality of line segments; a first grid statistical unit 206 configured to perform grid division on the broken line graph from the first conversion unit 204 to form a plurality of grids arranged in an array, and count the number of line segments and / or end points of line segments falling in each grid; a first histogram statistical unit 208 configured to group based on the size of intensity, and count the frequency of the number from the first grid statistical unit 206 to obtain a histogram; and a first determination unit 210 configured to find a maximum value point of the histogram from the first histogram statistical unit 208, and determine that a peak where one maximum value point satisfying the following conditions corresponds to one single molecule: the value of the maximum value point is greater than a first set threshold value, and the width of the peak where the maximum value point is located is greater than a second set threshold value. The single molecule recognition device 200 described above can quickly recognize single molecules by converting the broken line graph of the time sequence of bright spot intensity into image processing to obtain a histogram, and the recognition accuracy is also high.

[0156] It should be noted that the explanations and descriptions of the technical features and beneficial effects of the single molecule recognition method in any of the above embodiments and examples also apply to the single molecule recognition device 200 of the present embodiment. To avoid redundancy, they will not be described in detail here.

[0157] For example, in some embodiments, please refer to Figure 28 The single molecule recognition device 200 further comprises a first filtering unit 212 connected to the first grid statistical unit 206, configured to filter the broken line graph from the first conversion unit 204 before grid division.

[0158] In some embodiments, in the first grid statistical unit 206, the grid division of the broken line graph is performed according to the number of time frames of intensity and the size of intensity.

[0159] In some embodiments, in the first histogram statistical unit 208, the grouping based on the size of intensity and the frequency counting of the number to obtain a histogram comprises: dividing according to the size of intensity into N groups, and counting the frequency of the number falling in the N groups: Wherein, ni represents the sum of the frequency of the number of times falling in the ith row of the grid, j represents the number of time frames, gij represents the frequency of the number of times falling in the grid (i, j), and M represents the number of time frames.

[0160] In some embodiments, in the first histogram statistics unit 208, the grouping based on the size of the intensity and the frequency statistics of the number of times are performed to obtain the histogram, including performing histogram equalization of the L window: wherein nprepresents the equalization of ni, n i represents the equalization result of ni, and p is an integer related to the size of the window L and the i-th row. i

[0161] In some embodiments, referring to Figure 29 The single molecule recognition device 200 further comprises: a first simplification unit 214, configured to perform line erosion on the grid-divided step graph according to the number of times corresponding to each grid to convert the grid-divided step graph into a simplified graph; a first marking unit 216, configured to perform run-length encoding on the simplified graph to mark a connected region; and the first determination unit 210, configured to calculate the area of each connected region and determine that one connected region corresponding to one single molecule satisfies the following condition: the area of the connected region is greater than a third set threshold. In some embodiments, the simplified graph is a binary graph.

[0162] In some embodiments, referring to Figure 30 The single molecule recognition device 200 further comprises: a first image preprocessing unit 218, configured to analyze an input to-be-processed image to obtain a first image, the to-be-processed image containing at least one image bright spot, and the image bright spot having at least one pixel point; and a first bright spot detection unit 220, configured to: analyze the first image to calculate a bright spot determination threshold, analyze the first image to obtain a candidate bright spot, determine whether the candidate bright spot is an image bright spot according to the bright spot determination threshold, if the determination result is yes, obtain a time sequence of image bright spot intensity, and if the determination result is no, discard the candidate bright spot.

[0163] In some embodiments, referring to Figure 31 The first image preprocessing unit 218 comprises a first background reduction unit 226, configured to perform background reduction processing on the to-be-processed image to obtain the first image.

[0164] In some embodiments, referring to Figure 32 The first image preprocessing unit 218 comprises a first image simplification unit 222, configured to perform simplification processing on the to-be-processed image after the background reduction processing to obtain the first image.

[0165] In some embodiments, referring to Figure 33 ​The first image preprocessing unit 218 comprises a first image filtering unit 224, which is configured to perform filtering processing on the to-be-processed image to obtain the first image.

[0166] In some embodiments, the first image preprocessing unit 218 comprises a first image simplification unit 222, which is configured to perform simplification processing on the to-be-processed image to obtain the first image. Figure 34 The first image preprocessing unit 218 comprises a first image filtering unit 224, which is configured to perform filtering processing on the to-be-processed image to obtain the first image.

[0167] In some embodiments, the first image preprocessing unit 218 comprises a first image simplification unit 222, which is configured to perform simplification processing on the to-be-processed image to obtain the first image. Figure 35 The first image preprocessing unit 218 comprises a first image filtering unit 224, which is configured to perform filtering processing on the to-be-processed image to obtain the first image.

[0168] In some embodiments, the first image preprocessing unit 218 comprises a first image simplification unit 222, which is configured to perform simplification processing on the to-be-processed image to obtain the first image. Figure 36 The first image preprocessing unit 218 comprises a first image filtering unit 224, which is configured to perform filtering processing on the to-be-processed image to obtain the first image.

[0169] In some embodiments, the background reduction processing on the to-be-processed image in the first background reduction unit 226 comprises: determining the background of the to-be-processed image by using an opening operation, and performing the background reduction processing on the to-be-processed image according to the background. In some embodiments, the filtering processing is a Mexican hat filtering processing. In some embodiments, the simplification processing is a binarization processing.

[0170] In some embodiments, the first image simplification unit 222 is configured to obtain a signal-to-noise ratio matrix according to the to-be-processed image before the simplification processing, and to simplify the to-be-processed image before the simplification processing according to the signal-to-noise ratio matrix to obtain the first image.

[0171] In some embodiments, the analysis of the first image to calculate the bright spot determination threshold in the first bright spot detection unit 220 comprises: processing the first image by using the Otsu method to calculate the bright spot determination threshold.

[0172] In some embodiments, the judgment of whether the candidate bright spot is an image bright spot according to the bright spot determination threshold in the first bright spot detection unit 220 comprises: searching for a pixel point connected greater than (h*h-1) in the first image, and taking the searched pixel point as the center of the candidate bright spot, h being a natural number and an odd number greater than 1; judging whether the center of the candidate bright spot satisfies the condition: I max *A BI *ceof guass >T, wherein I maxA is the center strongest intensity of the h*h window BI ceof is the ratio of the h*h window in which the first image is set to the value guass T is the correlation coefficient of the h*h window and the two-dimensional Gaussian distribution, and the candidate bright spot center corresponding to the bright spot is determined as the image bright spot if the above conditions are met, and the bright spot corresponding to the candidate bright spot center is discarded if the above conditions are not met.

[0173] Please refer to Figure 37 The single molecule counting device 400 of the embodiment of the present application is used to implement all or part of the steps of the single molecule counting method in any of the embodiments of the present application described above, and the single molecule counting device 400 comprises: a second input unit 402 configured to input a time sequence of image bright spot intensity; a second conversion unit 404 configured to form a time and intensity line graph of the image bright spot according to the time sequence in the second input unit 402, the line graph being composed of a plurality of line segments; a second grid statistical unit 406 configured to perform grid division on the line graph from the second conversion unit 404 to form a plurality of grids arranged in an array, and count the number of line segments and / or endpoints of the line segments falling in each grid; a second histogram statistical unit 408 configured to group based on the size of the intensity, and count the frequency of the number from the second grid statistical unit 406 to obtain a histogram; a second determination unit 410 configured to find a maximum value point of the histogram from the second histogram statistical unit 408, and determine that the peak where one maximum value point satisfying the following conditions corresponds to one single molecule: the value of the maximum value point is greater than a first set threshold, and the width of the peak where the maximum value point is located is greater than a second set threshold; and a calculation unit 412 configured to calculate the number of single molecules S1. The single molecule counting device 400 described above can quickly count single molecules by converting the line graph of the time sequence of bright spot intensity into image processing to obtain a histogram, and the counting accuracy is also high.

[0174] It should be noted that the explanations and descriptions of the technical features and advantages of the single molecule counting method in any of the embodiments and examples described above also apply to the single molecule counting device 400 of the present embodiment, and to avoid redundancy, they will not be described in detail here.

[0175] For example, in some embodiments, please refer to Figure 38 The single molecule counting device 400 further comprises a second filtering unit 414 connected to the second grid statistical unit 406, configured to filter the line graph from the second conversion unit 404 before grid division.

[0176] In some embodiments, please refer to Figure 39The single-molecule counting device 400 further comprises a second simplifying unit 416 configured to perform line erosion on the grid-divided line graph according to the number of times corresponding to each grid, so as to convert the grid-divided line graph into a simplified graph; a second marking unit 418 configured to perform run-length encoding on the simplified graph to mark a connected region; in the second judging unit 410, the area of each connected region is calculated, and it is determined that one connected region corresponds to one single molecule when the area of the connected region is greater than a third set threshold; in the calculating unit 412, the number S2 of single molecules is calculated, and the smaller one of S1 and S2 is taken as the final number of single molecules.

[0177] Please refer to Figure 40 The single-molecule counting device 600 according to an embodiment of the present application comprises: a third input unit 602 configured to input a time sequence of image bright spot intensity; a third conversion unit 604 configured to form a line graph of time and intensity of image bright spots according to the time sequence in the third input unit 602, the line graph being composed of a plurality of line segments; a third grid counting unit 606 configured to perform grid division on the line graph from the third conversion unit 604 to form a plurality of grids arranged in an array, and count the number of line segments and / or end points of line segments falling in each grid; a third histogram counting unit 608 configured to group based on the size of intensity, and count the frequency of the number from the third grid counting unit 606 to obtain a histogram; and a third judging unit 610 configured to find a maximum value point of the histogram from the third histogram counting unit 608, and determine to add 1 to the count of single molecules when the following conditions are met: the value of the maximum value point is greater than a first set threshold, and the width of the peak where the maximum value point is located is greater than a second set threshold. The single-molecule counting device 600 described above can quickly count single molecules by converting the line graph of the time sequence of bright spot intensity into image processing to obtain a histogram, and the counting accuracy is also high.

[0178] It should be noted that the explanations and descriptions of the technical features and advantages of the single-molecule counting method in any of the above embodiments and examples are also applicable to the single-molecule counting device 600 according to the present embodiment, and to avoid redundancy, they will not be described in detail here.

[0179] For example, in some embodiments, please refer to Figure 41 The single-molecule counting device 600 further comprises a third filtering unit 612 connected to the third grid counting unit 606, configured to filter the line graph from the third conversion unit 604 before performing grid division on the line graph.

[0180] In some embodiments, please refer to Figure 42The single-molecule counting device 600 further comprises a third simplifying unit 614 configured to perform line erosion on the grid-divided line graph according to the number of times corresponding to each grid to convert the grid-divided line graph into a simplified graph; a third marking unit 616 configured to perform run-length encoding on the simplified graph to mark a connected region; in the third determining unit 610, the area of each connected region is calculated, and when the following condition is met, the count of single molecules is increased by 1: the area of the connected region is greater than a third set threshold; and the smaller one of the count of single molecules obtained based on the histogram and the count of single molecules obtained based on the run-length encoding is taken as the final single-molecule count.

[0181] Please refer to Figure 43 The single-molecule processing system 300 according to an embodiment of the present application comprises: a data input device 302 configured to input data; a data output device 304 configured to output data; a storage device 306 configured to store data, the data comprising computer executable programs; and a processor 308 configured to execute the computer executable programs, the execution of the computer executable programs comprising the completion of the method according to any one of the above embodiments.

[0182] The computer readable storage medium according to an embodiment of the present application is configured to store programs for execution by a computer, the execution of the programs comprising the completion of the method according to any one of the above embodiments. The computer readable storage medium can comprise a read-only memory, a random access memory, a magnetic disk or an optical disk, etc.

[0183] In the description of the present application, the description of the terms "one embodiment", "certain embodiments", "exemplary embodiment", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the exemplary description of the above terms does not necessarily mean the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0184] In addition, each functional unit in each embodiment of the present application can be integrated in one processing module, or each unit can exist physically independently, or two or more units can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software functional module. When the integrated module is realized in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer readable storage medium.

[0185] Although the embodiments of the present application have been shown and described above, it is understood that the above-described embodiments are exemplary and are not to be construed as limiting the present application, and that changes, modifications, substitutions and variations can be made by those skilled in the art without departing from the scope of the present application.

Claims

1. A method of recognizing a single molecule, characterized by, The method comprises the steps of: inputting a time sequence of image bright spot intensity, the image being from a single molecule sequencing platform; forming a time-intensity polygon of the image bright spot according to the time sequence, the polygon being composed of a plurality of line segments; grid dividing the polygon to form a plurality of grids arranged in an array, and counting the number of line segments and / or endpoints of the line segments falling in each grid; grouping according to the magnitude of the intensity, and frequency counting the number to obtain a histogram; finding a maximum point of the histogram, and determining that a peak where the maximum point is located corresponds to a single molecule when the value of the maximum point is greater than a first set threshold and the width of the peak where the maximum point is located is greater than a second set threshold.

2. A method of counting single molecules, characterized in that The method comprises the steps of: inputting a time sequence of image bright spot intensity, the image being from a single molecule sequencing platform; forming a time-intensity polygon of the image bright spot according to the time sequence, the polygon being composed of a plurality of line segments; grid dividing the polygon to form a plurality of grids arranged in an array, and counting the number of line segments and / or endpoints of the line segments falling in each grid; grouping according to the magnitude of the intensity, and frequency counting the number to obtain a histogram; finding a maximum point of the histogram, and determining that a peak where the maximum point is located corresponds to a single molecule when the value of the maximum point is greater than a first set threshold and the width of the peak where the maximum point is located is greater than a second set threshold. calculating the number of single molecules S1.

3. A method of counting single molecules, characterized by, The method comprises the steps of: inputting a time sequence of image bright spot intensity, the image being from a single molecule sequencing platform; forming a time-intensity polygon of the image bright spot according to the time sequence, the polygon being composed of a plurality of line segments; grid dividing the polygon to form a plurality of grids arranged in an array, and counting the number of line segments and / or endpoints of the line segments falling in each grid; grouping according to the magnitude of the intensity, and frequency counting the number to obtain a histogram; finding a maximum point of the histogram, and determining that a peak where the maximum point is located corresponds to a single molecule when the value of the maximum point is greater than a first set threshold and the width of the peak where the maximum point is located is greater than a second set threshold.

4. The method according to any one of claims 1 to 3, characterized in that, Before the grid dividing of the polygon, the method further comprises the step of: filtering the polygon.

5. The method according to any one of claims 1 to 3, characterized in that, The grid dividing of the polygon is according to the number of time frames of collecting the intensity and the magnitude of the intensity.

6. The method according to any one of claims 1 to 3, characterized in that, The step of grouping according to the magnitude of the intensity, and frequency counting the number to obtain a histogram comprises the steps of: dividing into N groups according to the magnitude of the intensity, and counting the frequency of the number falling in the N groups: where n i represents the sum of the frequencies of the number of times falling on the i-th row of the grid, j represents the number of time frames, g ij represents the frequency of the number of times falling on the grid (i, j), and M represents the number of time frames.

7. The method of claim 6, wherein, performing L-window histogram equalization. where n p represents the sum of the equalization results of n i , n i represents the equalization result of n i , and p is an integer related to the size of the window L and the i-th line where the window L is located.

8. The method of claim 1, wherein, The method further comprises the steps of: According to the number corresponding to each of the grids, line erosion is performed on the line graph after grid division to convert the line graph after grid division into a simplified graph, and optionally, the simplified graph is a binary graph; Run-length encoding is performed on the simplified graph to identify connected regions; The area of each of the connected regions is calculated, and it is determined that one of the connected regions corresponding to one single molecule satisfies the following condition: the area of the connected region is greater than a third set threshold.

9. The method of claim 2, wherein, Further comprising steps: According to the number corresponding to each of the grids, line erosion is performed on the line graph after grid division to convert the line graph after grid division into a simplified graph, and optionally, the simplified graph is a binary graph; Run-length encoding is performed on the simplified graph to identify connected regions; The area of each of the connected regions is calculated, and it is determined that one of the connected regions corresponding to one single molecule satisfies the following condition: the area of the connected region is greater than a third set threshold. The number of single molecules S2 is calculated; The smaller one of S1 and S2 is taken as the final number of single molecules.

10. The method of claim 3, wherein, Further comprising steps: According to the number corresponding to each of the grids, line erosion is performed on the line graph after grid division to convert the line graph after grid division into a simplified graph, and optionally, the simplified graph is a binary graph; Run-length encoding is performed on the simplified graph to identify connected regions; The area of each of the connected regions is calculated, and it is determined that one of the connected regions corresponding to one single molecule satisfies the following condition: the area of the connected region is greater than a third set threshold. The smaller one of the number of single molecules based on the histogram and the number of single molecules based on the run-length encoding is taken as the final number of single molecules.

11. The method according to any one of claims 1 to 3, 7 to 10, characterized in that, Further comprising: An image preprocessing step, which analyzes an input image to be processed to obtain a first image, the image to be processed containing at least one image bright spot, the image bright spot having at least one pixel point; A bright spot detection step, which comprises the steps of: analyzing the first image to calculate a bright spot determination threshold, analyzing the first image to obtain a candidate bright spot, determining whether the candidate bright spot is the image bright spot according to the bright spot determination threshold, if the determination result is yes, obtaining a time sequence of the image bright spot intensity, if the determination result is no, discarding the candidate bright spot.

12. The method of claim 11, wherein, The image preprocessing step comprises: performing background subtraction on the image to be processed to obtain the first image.

13. The method of claim 12, wherein, The image preprocessing step comprises: performing simplification processing on the image to be processed after background subtraction to obtain the first image.

14. The method of claim 11, wherein, The image preprocessing step comprises: performing filtering processing on the image to be processed to obtain the first image.

15. The method of claim 11, wherein, The image preprocessing step comprises: performing background subtraction on the image to be processed and then performing filtering processing to obtain the first image.

16. The method of claim 15, wherein, The image preprocessing step comprises: performing simplification processing on the image to be processed after background subtraction and filtering processing to obtain the first image.

17. The method of claim 11, wherein, The image preprocessing step comprises: performing simplification processing on the image to be processed to obtain the first image.

18. The method of claim 12, 13, 15, or 16, wherein, The background reduction processing is performed on the image to be processed, comprising: determining the background of the image to be processed by using an opening operation, performing the background reduction processing on the image to be processed according to the background.

19. The method according to any one of claims 14-16, characterized by, The filtering processing is a Mexican hat filtering processing.

20. The method of claim 13, 16 or 17, wherein, The simplification processing is a binarization processing.

21. The method of claim 13, 16 or 17, wherein, In the simplification processing, a signal-to-noise ratio matrix is obtained according to the image to be processed before the simplification processing, and the image to be processed before the simplification processing is simplified according to the signal-to-noise ratio matrix to obtain the first image.

22. The method of claim 11, wherein, The step of analyzing the first image to calculate a bright spot determination threshold comprises: processing the first image by Otsu method to calculate the bright spot determination threshold.

23. The method of claim 13, 16 or 17, wherein, The step of determining whether the candidate bright spot is the image bright spot according to the bright spot determination threshold comprises: finding a pixel point greater than (h*h-1) connected in the first image, and taking the pixel point as the center of the candidate bright spot, h is a natural number and an odd number greater than 1; determining whether the center of the candidate bright spot satisfies the condition: I max A BI ceof guass > T, wherein I max is the center strongest intensity of the h*h window, A BI is the ratio of the first image in the h*h window that is a set value, ceof guass is the correlation coefficient of the h*h window and the two-dimensional Gaussian distribution, and T is the bright spot determination threshold. if the above condition is met, determining that the bright spot corresponding to the center of the candidate bright spot is the image bright spot, if the above condition is not met, discarding the bright spot corresponding to the center of the candidate bright spot.

24. A single molecule recognition device, characterized by comprise: an input unit configured to input a time series of image bright spot intensity, the image being from a single molecule sequencing platform; a conversion unit configured to form a time-intensity line graph of the image bright spot according to the time series in the input unit, the line graph being composed of a plurality of line segments; a grid counting unit configured to divide the line graph from the conversion unit into a plurality of grids arranged in an array to count the number of line segments and / or endpoints of the line segments falling in each grid; a histogram counting unit configured to group based on the magnitude of the intensity and count the frequency of the number from the grid counting unit to obtain a histogram; a determination unit configured to find a maximum value point of the histogram from the histogram counting unit, and determine that a peak where one maximum value point is located corresponds to one single molecule if the value of the maximum value point is greater than a first set threshold and the width of the peak where the maximum value point is located is greater than a second set threshold.

25. A single molecule counting device, characterized by comprise: an input unit configured to input a time series of image bright spot intensity, the image being from a single molecule sequencing platform; a conversion unit configured to form a time-intensity line graph of the image bright spot according to the time series in the input unit, the line graph being composed of a plurality of line segments; a grid counting unit configured to divide the line graph from the conversion unit into a plurality of grids arranged in an array to count the number of line segments and / or endpoints of the line segments falling in each grid; a histogram counting unit configured to group based on the magnitude of the intensity and count the frequency of the number from the grid counting unit to obtain a histogram; a determination unit configured to find a maximum value point of the histogram from the histogram counting unit, and determine that a peak where one maximum value point is located corresponds to one single molecule if the value of the maximum value point is greater than a first set threshold and the width of the peak where the maximum value point is located is greater than a second set threshold. A calculating unit configured to calculate a number of single molecules S1.

26. A single molecule counting device, characterized by The method comprises: An inputting unit configured to input a time series of image spot intensity, the image being from a single molecule sequencing platform; A transforming unit configured to form a time-intensity plot of the image spot according to the time series in the inputting unit, the plot being composed of a plurality of line segments; A grid counting unit configured to grid the plot from the transforming unit to form a plurality of grids arranged in an array, count the number of line segments and / or endpoints of the line segments falling in each grid; A histogram counting unit configured to group the number from the grid counting unit based on the magnitude of the intensity, and count the frequency of the number to obtain a histogram; A judging unit configured to find a maximum point of the histogram from the histogram counting unit, and determine to add 1 to the count of single molecules when the value of the maximum point is greater than a first set threshold and the width of the peak where the maximum point is located is greater than a second set threshold.

27. The apparatus of any of claims 24-26, wherein, The device further comprises a filtering unit connected to the grid counting unit, configured to filter the plot from the transforming unit before gridding the plot.

28. The apparatus of any one of claims 24-26, wherein, In the grid counting unit, the gridding of the plot is performed according to the number of time frames of collecting the intensity and the magnitude of the intensity.

29. The apparatus of any one of claims 24-26, wherein, In the histogram counting unit, the grouping based on the magnitude of the intensity and the counting of the frequency of the number to obtain the histogram comprises: dividing the magnitude of the intensity into N groups, and counting the frequency of the number falling in the N groups: where n i represents the sum of the frequencies of the number of times falling on the i-th row of the grid, j represents the number of time frames, g ij represents the frequency of the number of times falling on the grid (i, j), and M represents the number of time frames.

30. The apparatus of any one of claims 24-26, wherein, In the histogram counting unit, the grouping based on the magnitude of the intensity and the counting of the frequency of the number to obtain the histogram comprises: performing histogram equalization by L windows: where n p represents the sum of the equalization results of n i , and p is an integer related to the size of the window L and the i-th row where the window L is located. i i where n i represents the sum of the equalization results of n i , and p is an integer related to the size of the window L and the i-th row where the window L is located.

31. The apparatus of claim 24, wherein, Further comprising: A simplifying unit configured to perform line erosion on the gridded plot according to the number corresponding to each grid to convert the gridded plot into a simplified plot, and optionally, the simplified plot is a binary plot; An identifying unit configured to perform run-length encoding on the simplified plot to identify connected regions; In the judging unit, the area of each connected region is calculated, and one connected region corresponding to one single molecule is determined when the area of the connected region is greater than a third set threshold.

32. The apparatus of claim 25, wherein, Further comprising: A simplifying unit configured to perform line erosion on the gridded plot according to the number corresponding to each grid to convert the gridded plot into a simplified plot, and optionally, the simplified plot is a binary plot; An identifying unit configured to perform run-length encoding on the simplified plot to identify connected regions; In the judging unit, the area of each connected region is calculated, and one connected region corresponding to one single molecule is determined when the area of the connected region is greater than a third set threshold; In the calculating unit, a number of single molecules S2 is calculated, and the smaller one of S1 and S2 is taken as the final number of single molecules.

33. The apparatus of claim 26, wherein, Further comprising: Simplifying unit, configured to perform line erosion on the meshed line graph according to the number of times corresponding to each mesh to convert the meshed line graph into a simplified graph, and optionally, the simplified graph is a binary graph; Identifying unit, configured to perform run-length encoding on the simplified graph to identify connected regions; In the determining unit, the area of each connected region is calculated, and when the area of the connected region is greater than a third preset threshold, the count of single molecules is increased by 1; And The smaller one of the count of single molecules obtained based on the histogram and the count of single molecules obtained based on the run-length encoding is taken as the final count of single molecules.

34. The apparatus of any one of claims 24-26, 31-33, wherein, Further comprising: An image preprocessing unit, configured to analyze an input image to be processed to obtain a first image, the image to be processed containing at least one image bright spot, the image bright spot having at least one pixel point; A bright spot detection unit, configured to: analyze the first image to calculate a bright spot determination threshold, analyze the first image to obtain a candidate bright spot, determine whether the candidate bright spot is the image bright spot according to the bright spot determination threshold, if the determination result is yes, obtain a time sequence of the image bright spot intensity, if the determination result is no, discard the candidate bright spot.

35. The apparatus of claim 34, wherein, The image preprocessing unit comprises a background subtraction unit, configured to perform background subtraction on the image to be processed to obtain the first image.

36. The device of claim 35, wherein, The image preprocessing unit comprises an image simplifying unit, configured to perform simplification on the image to be processed after the background subtraction to obtain the first image.

37. The apparatus of claim 34, wherein, The image preprocessing unit comprises an image filtering unit, configured to perform filtering on the image to be processed to obtain the first image.

38. The apparatus of claim 34, wherein, The image preprocessing unit comprises a background subtraction unit and an image filtering unit, the background subtraction unit is configured to perform background subtraction on the image to be processed, and the image filtering unit is configured to perform filtering on the image to be processed after the background subtraction to obtain the first image.

39. The device of claim 38, wherein, The image preprocessing unit comprises a simplifying unit, configured to perform simplification on the image to be processed after the background subtraction and the filtering to obtain the first image.

40. The apparatus of claim 34, wherein, The image preprocessing unit comprises an image simplifying unit, configured to perform simplification on the image to be processed to obtain the first image.

41. The apparatus of claim 35 or 38, wherein, In the background subtraction unit, the background subtraction on the image to be processed comprises: determining the background of the image to be processed by using an opening operation, performing background subtraction on the image to be processed according to the background.

42. The device of claim 36, 38, or 39, wherein, In the background subtraction unit, the background subtraction on the image to be processed comprises: determining the background of the image to be processed by using an opening operation, performing background subtraction on the image to be processed according to the background.

43. The device of any one of claims 37-39, wherein, The filtering is a Mexican hat filtering.

44. The device of claim 36, 39, or 40, wherein, The simplification is a binary processing.

45. The device of claim 36, 39, or 40, wherein, The image simplification unit is configured to obtain a signal-to-noise ratio matrix according to the image to be processed before simplification, and simplify the image to be processed before simplification according to the signal-to-noise ratio matrix to obtain the first image.

46. The device of claim 34, wherein, In the bright spot detection unit, the analyzing the first image to calculate a bright spot determination threshold comprises: processing the first image by Otsu method to calculate the bright spot determination threshold.

47. The device of claim 36, 39, or 40, wherein, In the bright spot detection unit, the judging whether the candidate bright spot is the image bright spot according to the bright spot determination threshold comprises: determining whether a center of the candidate bright spot satisfies a condition: I max *A BI *ceof guass > T, wherein I max is a strongest intensity of a h*h window, A BI is a ratio of the first image in which a set value is occupied in the h*h window, ceof guass is a correlation coefficient of a pixel of the h*h window and a two-dimensional Gaussian distribution, and T is the bright spot determination threshold, finding a pixel point greater than (h*h-1) connected pixels in the first image and taking the found pixel point as the center of the candidate bright spot, h being a natural number and an odd number greater than 1; if the above condition is satisfied, judging that the bright spot corresponding to the center of the candidate bright spot is the image bright spot, 48. A single molecule processing system, comprising: if the above condition is not satisfied, discarding the bright spot corresponding to the center of the candidate bright spot. comprise: a data input device configured to input data; a data output device configured to output data; a storage device configured to store data, the data comprising computer executable programs; a processor configured to execute the computer executable programs, the execution of the computer executable programs comprising completing the method according to any one of claims 1-23.

Citation Information

Patent Citations

  • Automatic cell localization method based on minimized model L1

    CN103236052A

  • Real-time single-molecule positioning method guaranteeing precision and system thereof

    CN105243677A

  • Single-molecule image correction device

    CN105303540A

  • Single-molecule positioning method

    CN105303551A

  • Optical detection method of membrane protein molecule mutual action

    CN1793862A