Single molecule recognition, counting method and device and processing system

By processing the time series of image bright spot intensity into a line graph, and utilizing grid partitioning, line erosion, and run-length encoding, the problems of slow speed and low accuracy in single-molecule recognition and counting in existing technologies are solved, achieving fast and efficient single-molecule recognition and counting.

CN108229098BActive Publication Date: 2026-01-16GENEMIND BIOSCIENCES CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN201710607586.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-12-09
Filing Date
2017-07-24
Publication Date
2026-01-16
Estimated Expiration
2037-07-24

AI Technical Summary

Technical Problem

Existing single-molecule recognition and counting methods rely on human eye recognition, which is slow and labor-intensive. Methods based on HMM and machine learning require a large number of samples for training and are not efficient. They are also subject to fluorescence recognition errors and background noise interference.

Method used

By processing the time series of image bright spot intensity into a line graph, performing mesh generation, line erosion, and run-length encoding, and calculating the area of ​​connected regions, single molecules are identified and counted. This image processing technique enables fast and high-precision single molecule identification and counting.

Benefits of technology

It achieves rapid and high-precision single-molecule identification and counting, reduces the need for manual intervention and sample training, and improves the efficiency and accuracy of identification and counting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN108229098B_ABST
    Figure CN108229098B_ABST
Patent Text Reader

Abstract

The application discloses a single molecule recognition and counting method and device. The single molecule recognition method comprises the following steps: inputting a time sequence of image bright spot intensity; forming a time-intensity broken line graph of the image bright spot according to the time sequence, wherein the broken line graph is composed of multiple line segments; performing grid division on the broken line graph to form multiple grids arranged in an array, and counting the number of line segments and / or end points of the line segments falling in each grid; performing line corrosion on the broken line graph after the grid division according to the number corresponding to each grid to convert the broken line graph after the grid division into a simplified graph; performing run-length encoding on the simplified graph to identify connected regions; and calculating the area of each connected region to determine that one connected region corresponding to one single molecule satisfies the following condition: the area of the connected region is greater than a first set threshold. The single molecule recognition method can quickly recognize single molecules by converting the time sequence of bright spot intensity into a histogram through image processing, and the recognition accuracy is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of gene sequencing, in particular to a single molecule recognition and counting method, a recognition and counting device and a processing system. BACKGROUND

[0002] In the related art, the third generation sequencing technology is single molecule sequencing, and the single molecule sequencing technology based on imaging optical detection is a base recognition technology relying on optical signals and electrical signals. Among them, the base group relying on fluorescence recognition emits light intensity from the excited state to the ground state under the irradiation of a specific power laser. However, the different lengths of time of different fluorescent groups, the different light intensities emitted, and the existence of background noise will all cause errors in single molecule recognition. At the same time, uneven distribution of DNA chains and base group clustering will also cause a decrease in effective single molecules.

[0003] The existing method mainly relies on the human eye to recognize and count single molecules on the collected fluorescence image, but such a method consumes manpower and is also slow. Referring to voice recognition, the method based on HMM and machine learning not only needs a large number of sample training, but also has low running efficiency. SUMMARY

[0004] The embodiments of the present application aim to at least solve one of the technical problems existing in the prior art. To this end, the embodiments of the present application need to provide a single molecule recognition and counting method, a recognition and counting device and a processing system.

[0005] The single molecule recognition method of the embodiments of the present application comprises the steps of: inputting a time sequence of image bright spot intensity; forming a time and intensity broken line graph of the image bright spot according to the time sequence, the broken line graph being composed of a plurality of line segments; performing grid division on the broken line graph to form a plurality of grids arranged in an array, and counting the number of line segments and / or end points of the line segments falling in each grid; performing line erosion on the broken line graph after grid division to convert the broken line graph after grid division into a simplified graph according to the number corresponding to each grid; performing run-length encoding on the simplified graph to identify connected regions; and calculating the area of each connected region to determine that one connected region meeting the following condition corresponds to one single molecule: the area of the connected region is greater than a first set threshold. The above single molecule recognition method can quickly recognize single molecules by converting the broken line graph of the time sequence of bright spot intensity into image processing to obtain a histogram, and the recognition accuracy is also high.

[0006] The single molecule counting method of the embodiment of the present application comprises the following steps: inputting a time sequence of image bright spot intensity; forming a time-intensity line graph of the image bright spot according to the time sequence, the line graph being composed of a plurality of line segments; performing grid division on the line graph to form a plurality of grids arranged in an array, and counting the number of line segments and / or end points of the line segments falling in each grid; performing line erosion on the line graph after grid division according to the number corresponding to each grid to convert the line graph after grid division into a simplified graph; performing run-length encoding on the simplified graph to identify connected regions; calculating the area of each connected region, and determining that one connected region corresponding to one single molecule satisfies the following condition: the area of the connected region is greater than a first set threshold; and calculating the number of single molecules S2. The single molecule counting method can quickly count single molecules with high accuracy by converting the time sequence of bright spot intensity into a histogram through image processing.

[0007] The single molecule counting method of the embodiment of the present application comprises the following steps: inputting a time sequence of image bright spot intensity; forming a time-intensity line graph of the image bright spot according to the time sequence, the line graph being composed of a plurality of line segments; performing grid division on the line graph to form a plurality of grids arranged in an array, and counting the number of line segments and / or end points of the line segments falling in each grid; performing line erosion on the line graph after grid division according to the number corresponding to each grid to convert the line graph after grid division into a simplified graph; performing run-length encoding on the simplified graph to identify connected regions; calculating the area of each connected region, and determining that one connected region corresponding to one single molecule satisfies the following condition: the area of the connected region is greater than a first set threshold; and calculating the number of single molecules S2. The single molecule counting method can quickly count single molecules with high accuracy by converting the time sequence of bright spot intensity into a histogram through image processing.

[0008] The single molecule recognition device of the embodiment of the present application is used to implement part or all steps of the single molecule recognition method of the above-mentioned aspect of the present application, and comprises: an input unit configured to input a time sequence of image bright spot intensity; a conversion unit configured to form a time-intensity polyline of the image bright spot according to the time sequence in the input unit, the polyline being composed of a plurality of line segments; a grid counting unit configured to perform grid division on the polyline from the conversion unit to form a plurality of grids arranged in an array, and count the number of the line segments and / or the end points of the line segments falling in each grid; a simplification unit configured to perform line erosion on the grid-divided polyline according to the number corresponding to each grid, so as to convert the grid-divided polyline into a simplified graph; a marking unit configured to perform run-length encoding on the simplified graph to mark a connected region; and a judging unit configured to calculate the area of each connected region, and determine that one connected region corresponding to one single molecule satisfies the following condition: the area of the connected region is greater than a first set threshold. The single molecule recognition device can quickly recognize single molecules by converting the time sequence of bright spot intensity into a polyline and then into a histogram through image processing, and the recognition accuracy is relatively high.

[0009] The single molecule counting device of the embodiment of the present application is used to implement part or all steps of the single molecule counting method of the above-mentioned aspect of the present application, and comprises: an input unit configured to input a time sequence of image bright spot intensity; a conversion unit configured to form a time-intensity polyline of the image bright spot according to the time sequence in the input unit, the polyline being composed of a plurality of line segments; a grid counting unit configured to perform grid division on the polyline from the conversion unit to form a plurality of grids arranged in an array, and count the number of the line segments and / or the end points of the line segments falling in each grid; a simplification unit configured to perform line erosion on the grid-divided polyline according to the number corresponding to each grid, so as to convert the grid-divided polyline into a simplified graph; a marking unit configured to perform run-length encoding on the simplified graph to mark a connected region; a judging unit configured to calculate the area of each connected region, and determine that one connected region corresponding to one single molecule satisfies the following condition: the area of the connected region is greater than a first set threshold; and a calculation unit configured to calculate the number S2 of obtained single molecules. The single molecule counting device can quickly count single molecules by converting the time sequence of bright spot intensity into a polyline and then into a histogram through image processing, and the counting accuracy is relatively high.

[0010] The single-molecule counting device of the embodiment of the present application is used to implement part or all steps of the single-molecule counting method of the above-mentioned aspect of the present application, and comprises: an input unit configured to input a time sequence of image bright spot intensity; a conversion unit configured to form a time-intensity polyline of the image bright spot according to the time sequence in the input unit, the polyline being composed of a plurality of line segments; a grid counting unit configured to perform grid division on the polyline from the conversion unit to form a plurality of grids arranged in an array, and count the number of the line segments and / or the end points of the line segments falling in each grid; a simplification unit configured to perform line erosion on the grid-divided polyline according to the number corresponding to each grid to convert the grid-divided polyline into a simplified graph; an identification unit configured to perform run-length encoding on the simplified graph to identify connected regions; and a determination unit configured to calculate the area of each connected region and determine to add 1 to the count of single molecules when the area of the connected region is greater than a first set threshold. The single-molecule counting device described above can quickly count single molecules by converting the time-intensity polyline of the bright spot intensity into an image processing to obtain a histogram, and the accuracy of the counting is also high.

[0011] The single-molecule processing system of the embodiment of the present application comprises: a data input device configured to input data; a data output device configured to output data; a storage device configured to store data, the data comprising a computer executable program; and a processor configured to execute the computer executable program, the execution of the computer executable program comprising the method of any of the above-mentioned embodiments. The single-molecule processing system can realize single-molecule identification and / or single-molecule counting.

[0012] The computer readable storage medium of the embodiment of the present application is used to store a program for execution by a computer, and the execution of the program comprises the method of any of the above-mentioned embodiments. The computer readable storage medium can comprise a read-only memory, a random access memory, a magnetic disk or an optical disk, etc.

[0013] Additional aspects and advantages of the embodiments of the present application will be in part apparent and in part pointed out hereinafter in the description of the embodiments of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0014] The above-mentioned and / or additional aspects and advantages of the embodiments of the present application will become more apparent and easily understood from the description of the embodiments of the present application in conjunction with the following drawings, in which:

[0015] Figure 1 is a flowchart of the single-molecule identification method of the embodiment of the present application.

[0016] Figure 2is another flowchart of the single molecule recognition method of the embodiments of the present application.

[0017] Figure 3 is yet another flowchart of the single molecule recognition method of the embodiments of the present application.

[0018] Figure 4 is yet another flowchart of the single molecule recognition method of the embodiments of the present application.

[0019] Figure 5 is yet another flowchart of the single molecule recognition method of the embodiments of the present application.

[0020] Figure 6 is yet another flowchart of the single molecule recognition method of the embodiments of the present application.

[0021] Figure 7 is yet another flowchart of the single molecule recognition method of the embodiments of the present application.

[0022] Figure 8 is yet another flowchart of the single molecule recognition method of the embodiments of the present application.

[0023] Figure 9 is a plot of the Mexican hat filter of the single molecule recognition method of the embodiments of the present application.

[0024] Figure 10 is still another flowchart of the single molecule recognition method of the embodiments of the present application.

[0025] Figure 11 is a plot of the 8-connected pixels of the single molecule recognition method of the embodiments of the present application.

[0026] Figure 12 is a plot of the single molecule recognition method of the embodiments of the present application.

[0027] Figure 13 is a plot of the grid division of the single molecule recognition method of the embodiments of the present application.

[0028] Figure 14 is a plot of the single molecule recognition method of the embodiments of the present application.

[0029] Figure 15 is a plot of the single molecule recognition method of the embodiments of the present application.

[0030] Figure 16 is another plot of the single molecule recognition method of the embodiments of the present application.

[0031] Figure 17is a schematic diagram of a histogram after equalization in a single molecule recognition method of an embodiment of the present invention.

[0032] Figure 18 is another schematic diagram of a flowchart of a single molecule recognition method of an embodiment of the present invention.

[0033] Figure 19 is a schematic diagram of a line erosion process in a single molecule recognition method of an embodiment of the present invention.

[0034] Figure 20 is another schematic diagram of a line erosion process in a single molecule recognition method of an embodiment of the present invention.

[0035] Figure 21 is a schematic diagram of 8-connected windows in a single molecule recognition method of an embodiment of the present invention.

[0036] Figure 22 is a schematic diagram of identifying connected regions in a single molecule recognition method of an embodiment of the present invention.

[0037] Figure 23 is a schematic diagram of a flowchart of a single molecule counting method of an embodiment of the present invention.

[0038] Figure 24 is another schematic diagram of a flowchart of a single molecule counting method of an embodiment of the present invention.

[0039] Figure 25 is yet another schematic diagram of a flowchart of a single molecule counting method of an embodiment of the present invention.

[0040] Figure 26 is still another schematic diagram of a flowchart of a single molecule counting method of an embodiment of the present invention.

[0041] Figure 27 is a schematic diagram of a module of a single molecule recognition apparatus of an embodiment of the present invention.

[0042] Figure 28 is another schematic diagram of a module of a single molecule recognition apparatus of an embodiment of the present invention.

[0043] Figure 29 is yet another schematic diagram of a module of a single molecule recognition apparatus of an embodiment of the present invention.

[0044] Figure 30 is another schematic diagram of a module of a single molecule recognition apparatus of an embodiment of the present invention.

[0045] Figure 31 is yet another schematic diagram of a module of a single molecule recognition apparatus of an embodiment of the present invention.

[0046] Figure 32is another module schematic diagram of the single molecule recognition device of the embodiment of the present application.

[0047] Figure 33 is another module schematic diagram of the single molecule recognition device of the embodiment of the present application.

[0048] Figure 34 is another module schematic diagram of the single molecule recognition device of the embodiment of the present application.

[0049] Figure 35 is another module schematic diagram of the single molecule recognition device of the embodiment of the present application.

[0050] Figure 36 is another module schematic diagram of the single molecule recognition device of the embodiment of the present application.

[0051] Figure 37 is a module schematic diagram of the single molecule counting device of the embodiment of the present application.

[0052] Figure 38 is another module schematic diagram of the single molecule counting device of the embodiment of the present application.

[0053] Figure 39 is another module schematic diagram of the single molecule counting device of the embodiment of the present application.

[0054] Figure 40 is yet another module schematic diagram of the single molecule counting device of the embodiment of the present application.

[0055] Figure 41 is still another module schematic diagram of the single molecule counting device of the embodiment of the present application.

[0056] Figure 42 is still another module schematic diagram of the single molecule counting device of the embodiment of the present application.

[0057] Figure 43 is a module schematic diagram of the single molecule processing system of the embodiment of the present application. DETAILED DESCRIPTION

[0058] Embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals are used throughout the drawings and the same or similar components are denoted by the same or similar reference numerals to indicate the same or similar components or components having the same or similar functions. The embodiments described below by reference to the accompanying drawings are exemplary and are for the purpose of explanation only, and are not to be understood as limiting the present application.

[0059] In the description of the application, it needs to be understood that the terms "first", "second" are only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the application, the meaning of "a plurality of" is two or more, unless otherwise explicitly specified and limited.

[0060] In the description of the application, it needs to be noted that, unless otherwise explicitly specified and limited, "connection" should be understood broadly, for example, it can be fixed connection, or detachable connection, or integrally connected; it can be mechanical connection, or electrical connection or can communicate with each other; it can be directly connected, or indirectly connected through intermediate medium, it can be the internal communication of two elements or the interaction relationship between two elements. For those skilled in the art, the specific meaning of the above terms in the application can be understood according to the specific circumstances.

[0061] The following disclosure provides many different embodiments or examples for implementing different structures of the application. In order to simplify the disclosure of the application, the components and settings of specific examples are described below. In addition, the present application can repeatedly refer to numbers and / or letters in different examples, and such repetition is for the purpose of simplification and clarity, which itself does not indicate the relationship between the various embodiments and / or settings discussed.

[0062] The single molecule recognition method and counting method of the embodiments of the application can be applied to gene sequencing. The "gene sequencing" referred to in the embodiments of the application is the same as nucleic acid sequencing, including DNA sequencing and / or RNA sequencing, including long fragment sequencing and / or short fragment sequencing.

[0063] Please refer to Figure 1The single molecule recognition method of the embodiment of the present application comprises the following steps: S01, inputting a time sequence of image bright spot intensity; S02, forming a time-intensity broken line graph of the image bright spot according to the time sequence, the broken line graph being composed of a plurality of line segments; S03, performing grid division on the broken line graph to form a plurality of grids arranged in an array, and counting the number of line segments and / or end points of the line segments falling in each grid; S04, performing line erosion on the broken line graph after the grid division according to the number corresponding to each grid, so as to convert the broken line graph after the grid division into a simplified graph; S05, performing run-length encoding on the simplified graph to identify connected regions; and S06, calculating the area of each connected region, and determining that one connected region corresponding to one single molecule satisfies the following condition: the area of the connected region is greater than a first set threshold. The single molecule recognition method described above can quickly recognize single molecules by converting the time sequence of the bright spot intensity into a connected region of run-length encoding through image processing, and the recognition accuracy is also high. The single molecule recognition method based on run-length encoding can accurately recognize single molecules according to the time sequence data of the bright spot intensity, and is especially suitable for the case that the number of single molecules contained in one bright spot is not greater than 3.

[0064] Specifically, in step S01, when the image bright spot is formed, a specific wavelength laser is used to irradiate the test sample to make the test sample emit fluorescence, and then a camera is used to collect the fluorescence to form an image. In the image, there are image bright spots corresponding to the part (nucleic acid molecules) of the test sample emitting fluorescence. The so-called "bright spot" refers to a light-emitting point on the image, and one light-emitting point occupies at least one pixel point. The so-called "pixel point" is the same as "pixel".

[0065] In an embodiment of the present application, the image comes from a single molecule sequencing platform, such as the sequencing platform of Helicos or Pacific Biosciences (PacBio) company. The input raw data is the parameter of the pixel point of the image, and the detection of the so-called "bright spot" is the detection of the single molecule optical signal.

[0066] In some embodiments, please refer to Figure 2 The single molecule recognition method further comprises: an image preprocessing step S31, the image preprocessing step S31 analyzes the input image to be processed to obtain a first image, the image to be processed contains at least one image bright spot, and the image bright spot has at least one pixel point; a bright spot detection step S32, the bright spot detection step S32 comprises the following steps: S321, analyzing the first image to calculate a bright spot determination threshold, S322, analyzing the first image to obtain a candidate bright spot, S323, judging whether the candidate bright spot is an image bright spot according to the bright spot determination threshold, if the judgment result is yes, S324, then obtaining a time sequence of image bright spot intensity, if the judgment result is no, S325, discarding the candidate bright spot.

[0067] Therefore, the denoising processing of the to-be-processed image by the image preprocessing step can reduce the calculation amount of the bright spot detection step, and the bright spot judgment threshold can be used to judge whether the candidate bright spot is an image bright spot, thereby improving the accuracy of judging the image bright spot.

[0068] Specifically, in one example, the input to-be-processed image can be a 512*512 or 2048*2048 16-bit tiff format image, and the tiff format image can be a grayscale image. In this way, the processing process of the single molecule recognition method can be simplified.

[0069] In some embodiments, referring to Figure 3 The image preprocessing step S31 includes: performing background subtraction processing on the to-be-processed image to obtain a first image. In this way, the noise of the to-be-processed image can be further reduced, and the accuracy of the single molecule counting method is higher.

[0070] In some embodiments, referring to Figure 4 The image preprocessing step S31 includes: performing simplification processing on the to-be-processed image after the background subtraction processing to obtain a first image. In this way, the calculation amount of the subsequent single molecule recognition and / or counting method can be reduced.

[0071] In some embodiments, referring to Figure 5 The image preprocessing step S31 includes: performing filtering processing on the to-be-processed image to obtain a first image. In this way, the filtering processing on the to-be-processed image can obtain the first image while preserving the image detail features as much as possible, thereby improving the accuracy of the single molecule recognition and / or counting method.

[0072] In some embodiments, referring to Figure 6 The image preprocessing step S31 includes: performing filtering processing on the to-be-processed image after the background subtraction processing to obtain a first image. In this way, the filtering processing on the to-be-processed image after the background subtraction processing can further reduce the noise of the to-be-processed image, and the accuracy of the single molecule recognition and / or counting method is higher.

[0073] In some embodiments, referring to Figure 7 The image preprocessing step S31 includes: performing simplification processing on the to-be-processed image after the background subtraction processing and the filtering processing to obtain a first image. In this way, the calculation amount of the subsequent image processing method can be reduced.

[0074] In some embodiments, referring to Figure 8 The image preprocessing step S31 includes: performing simplification processing on the to-be-processed image to obtain a first image. In this way, the calculation amount of the subsequent single molecule recognition and / or counting method can be reduced.

[0075] In some embodiments, the background subtraction processing on the image to be processed comprises: determining the background of the image to be processed by using the open operation, and performing the background subtraction processing on the image to be processed according to the background. In this way, the open operation is used to eliminate small objects, separate objects at fine points, and smooth the boundaries of large objects without significantly changing the area of the image, so that the image after the background subtraction processing can be more accurately obtained.

[0076] Specifically, in the embodiments of the present application, a window (for example, a 15*15 window) is moved in the image to be processed f(x, y) (for example, a gray image), and the background of the image to be processed is estimated by using the open operation (first erosion and then dilation), as shown in the following formula 1 and formula 2:

[0077] g(x, y) = erode [f(x, y), B] = min{f(x+x', y+y')-B(x', y')|(x', y')∈D b} Formula 1,

[0078] wherein g(x, y) is the eroded gray image, f(x, y) is the original gray image, and B is the structural element.

[0079] g(x, y) = dilate [f(x, y), B] = max{f(x-x', y-y')-B(x', y')|(x', y')∈D b} Formula 2.

[0080] wherein g(x, y) is the dilated gray image, f(x, y) is the original gray image, and B is the structural element.

[0081] Therefore, the background noise g = imopen (f(x, y), B) = dilate [erode (f(x, y), B)] can be obtained. Formula 3.

[0082] The original image is subjected to background subtraction:

[0083] f = f-g = {f(x, y)-g(x, y)|(x, y)∈D} Formula 4.

[0084] It can be understood that the specific method of the present embodiment for performing the background subtraction processing on the image to be processed can be applied to the step of performing the background subtraction processing on the image to be processed mentioned in any of the above embodiments.

[0085] In some embodiments, the filtering processing is a Mexican hat filtering processing. The Mexican hat filtering is easy to implement, which reduces the cost of the single molecule recognition and / or counting method, and at the same time, the Mexican hat filtering can improve the contrast between the foreground and the background, so that the foreground is brighter and the background is darker.

[0086] In performing the Mexican hat filtering, a m*m window is used to perform Gaussian filtering on the to-be-processed image before the filtering, and two-dimensional Laplacian sharpening is performed on the to-be-processed image after the Gaussian filtering, where m is a natural number and an odd number greater than 1. In this way, the Mexican hat filtering is realized through two steps.

[0087] Specifically, please refer to Figure 9 The Mexican hat kernel can be expressed as:

[0088] where x and y represent the coordinates of a pixel.

[0089] First, a m*m window is used to perform Gaussian filtering on the to-be-processed image, as shown in the following formula 6:

[0090]

[0091] where t1 and t2 represent the position of the filtering window, and wt1,t2 represents the weight of the Gaussian filtering.

[0092] Then, two-dimensional Laplacian sharpening is performed on the to-be-processed image, as shown in the following formula 7:

[0093]

[0094] where K and k both represent the Laplacian operator, which is related to the sharpening target. If it is necessary to strengthen the sharpening or weaken the sharpening, K and k are modified.

[0095] In one example, m = 3, so m*m = 3*3. When performing Gaussian filtering, formula 6 becomes:

[0096]

[0097] It can be understood that the specific method of the Mexican hat filtering of the present embodiment can be applied to the step of performing filtering processing on the to-be-processed image mentioned in any of the above embodiments.

[0098] In some embodiments, the simplified image is a binary image. In this way, the binary image is easy to process and has a wide range of applications.

[0099] Specifically, in one example, the binary image can contain two numerical values of 0 and 1 representing different attributes of a pixel. The binary image can be expressed as:

[0100] In some embodiments, when performing the simplification processing, a signal-to-noise ratio matrix is obtained according to the to-be-processed image before the simplification processing, and the to-be-processed image before the simplification processing is simplified according to the signal-to-noise ratio matrix to obtain a first image.

[0101] In one specific example, the background subtraction processing can be performed on the to-be-processed image first, and then the SNR matrix can be obtained according to the to-be-processed image after the background subtraction processing. In this way, information can be obtained from the image with less noise, and the accuracy of the processing result of the single molecule recognition and / or counting method can be improved.

[0102] Specifically, in one example, the SNR matrix can be represented as: wherein x and y represent the coordinates of the pixel points, h represents the height of the image, w represents the width of the image, i∈w, j∈h.

[0103] In one example, the simplified image is a binary image, and the binary image can be obtained according to the SNR matrix, as shown in formula 9:

[0104]

[0105] In the calculation of the SNR matrix, the background subtraction processing and / or the filtering processing can be performed on the to-be-processed image first, such as the background subtraction processing step and the filtering processing step in the above embodiment, and then the ratio matrix of the to-be-processed image after the background subtraction processing and the background can be obtained according to formula 4:

[0106] R = f / g = {f(x, y) / g(x, y) | (x, y) ∈ D} formula 10, wherein D represents the dimension (height * width) of the image f.

[0107] Thus, the SNR matrix can be obtained:

[0108] In some embodiments, the step of analyzing the first image to calculate the bright spot determination threshold value includes: processing the first image by the Otsu method to calculate the bright spot determination threshold value. In this way, the bright spot determination threshold value is found by a more mature and simple method, thereby improving the accuracy of the single molecule recognition and / or counting method and reducing the cost of the single molecule recognition and / or counting method. At the same time, the finding of the bright spot determination threshold value by using the first image can improve the efficiency and accuracy of the single molecule recognition and / or counting method.

[0109] Specifically, the Otsu method (OTSU algorithm) can also be referred to as the maximum inter-class variance method. The Otsu method uses the maximum inter-class variance to segment the image, which means the minimum misclassification probability and high accuracy. Assuming that the segmentation threshold value of the foreground and the background of the to-be-processed image is T, the proportion of the number of pixel points belonging to the foreground in the entire image is ω0, and the average gray value is μ0; the proportion of the number of pixel points belonging to the background in the entire image is ω1, and the average gray value is μ1. The total average gray value of the to-be-processed image is denoted as μ, and the inter-class variance is denoted as var, and then:

[0110] μ = ω0*μ0 + ω1*μ1 formula 11;

[0111] var = ω0(μ0 - μ) 2 + ω1(μ1 - μ) 2 Equation 12.

[0112] Substitute Equation 11 into Equation 12 to obtain equivalent Equation 13:

[0113] var = ω0ω1(μ1 - μ0) 2 Equation 13.

[0114] The traversal method is used to obtain the segmentation threshold T that maximizes the inter-class variance, i.e., the obtained highlight determination threshold T.

[0115] In some embodiments, referring to Figure 10 The step of determining whether the candidate highlight is an image highlight according to the highlight determination threshold includes:

[0116] S41, find the pixel points greater than (h*h - 1) connected in the first image and take the found pixel points as the centers of the candidate highlights, h*h is in one-to-one correspondence with the highlights, each value in h*h corresponds to a pixel point, h is a natural number and an odd number greater than 1;

[0117] S42, determine whether the center of the candidate highlight satisfies the condition: I max *A BI *ceof guass > T, wherein I max is the strongest intensity of the center of the h*h window, A BI is the ratio of the set value in the h*h window in the first image, ceof guass is the correlation coefficient of the pixel and two-dimensional Gaussian distribution of the h*h window, and T is the highlight determination threshold.

[0118] If the above condition is satisfied, S43, determine that the highlight corresponding to the center of the candidate highlight is an image highlight contained in the image to be processed.

[0119] If the above condition is not satisfied, S44, discard the highlight corresponding to the center of the candidate highlight. In this way, the detection of the image highlight is realized.

[0120] Specifically, I max can be understood as the strongest intensity of the center of the candidate highlight. In an example, h = 3, find the pixel points greater than 8 connected, as shown in Figure 11 Take the found pixel points as the pixel points of the candidate highlights. I max is the strongest intensity of the center of the 3*3 window, A BI is the ratio of the set value in the 3*3 window in the first image, and ceof guassThe correlation coefficient is the pixel value of a 3x3 window and the two-dimensional Gaussian distribution.

[0121] The first image is a simplified image; for example, it could be a binarized image, meaning that the set values ​​in the binarized image are the values ​​corresponding to pixels that meet set conditions. In another example, the binarized image could contain two values, 0 and 1, representing different attributes of pixels, with a set value of 1, A. BI This represents the percentage of 1s in the binarized image within the h*h window. For example, refer to Formula 9; when SNR <= mean(SNR), BI = 1.

[0122] In some implementations, the value of h can be equal to the value of m selected when performing the Mexican hat filter, i.e., h = m.

[0123] In some implementations, when acquiring the images described above, the camera sequentially acquires fluorescence from multiple fields of view (FOV) in a time sequence. Therefore, when image data is obtained, the intensity of the bright spots in the image data corresponds to the time sequence acquired by the camera.

[0124] In step S02, after obtaining the desired image bright spots, the intensity of the image bright spots corresponding to adjacent acquisition times is connected to form a line graph of image bright spot time versus intensity, as shown below. Figure 12 As shown. In Figure 12 In the diagram, the horizontal axis represents the fluorescence acquisition time in milliseconds (ms), and the vertical axis represents the intensity of bright spots in the image. In one example, the time interval between two consecutive fluorescence acquisitions is 20 ms.

[0125] The vertical axis represents the corresponding bright spot intensity value. In this embodiment of the invention, the bright spot intensity value is the bright spot pixel value. For a 16-bit TIFF image, the bright spot pixel value is in the range of 0-65535; for an 8-bit grayscale image, the bright spot pixel value is in the range of 0-255. This embodiment of the invention uses a 16-bit TIFF image.

[0126] In step S03, the waveform of the line graph is converted into an image for subsequent run-length encoding. Image processing of the line graph includes dividing the line graph into a grid.

[0127] In some embodiments, the grid division of the line graph is performed according to the number of time frames of the collected intensity and the size of the intensity. In this way, the line graph can be processed more simply to obtain the grid division, and the cost of the single molecule recognition method is reduced. Specifically, the number of time frames can be divided into M and the size of the intensity can be divided into N, i.e., M*N grids are formed. The number of time frames of the collected intensity is the time interval between two adjacent collected fluorescence. In one embodiment, one grid can be referred to as the length direction along the horizontal axis direction and the height direction along the vertical axis direction. The length of one grid can be set to be a multiple of the number of time frames, such as 1 times, 2 times, 2.5 times, etc. The height of one grid can be flexibly set, for example, for a 16-bit tiff image, the value of the vertical axis is 0-65535, and when the grid is divided, the value of the vertical axis can be normalized and divided into 50 parts, and then the height of one grid is set to 0.02, i.e., N=50.

[0128] In one example, the time interval between two adjacent collected fluorescence is 20 ms, the length of one grid is equal to one time interval, and the height is 0.02. Please refer to FIG. 6A. Figure 16 In such an example, the number of line segments falling into one grid can be 0, 1 or 2. Figure 16 The black dots in FIG. 6B represent the time sequence of the intensity of the image bright spot.

[0129] In one example, please refer to FIG. 6A. Figure 13 The line graph is divided into an 8*6 grid, and the number of line segments and / or the number of endpoints of the line segments falling into each grid is counted. In Figure 13 In FIG. 6B, the number of line segments falling into each grid (i.e., the number of times each grid is passed through by the line segment) is counted, and the number in the grid represents the number of line segments falling into each grid. Figure 13 The black dots in FIG. 6B represent the time sequence of the intensity of the image bright spot.

[0130] In line erosion, the following formula can be used for morphological erosion operation: g(x, y) = erode[f(x, y), B] = min{f(x+x', y+y')-B(x', y') | (x', y') ∈ D b}. Preferably, a straight line structure element is selected, such as a window size of W*1, and if the number of grids in the window exceeds a threshold T, the grid is marked as a first value, otherwise it is marked as a second value. In this way, the line graph after grid division can be converted into a simplified graph including the first value and the second value. In some embodiments, the simplified graph is a binary graph. For example, the first value can be taken as 1, and the second value can be taken as 0.

[0131] In one example, please refer to FIG. 6A. Figure 19 The length of one grid is L1, W=2*L1, T=2, Figure 19The display shows five grids arranged along the length direction. The numbers in the grids represent the number of times. When performing line erosion, the window is aligned with the grid. After line erosion, the five grids are marked as 0, 1, 0, 0, and 0, respectively.

[0132] In another example, please refer to Figure 20 A grid has a length of L1, W = 2 * L1, and T = 2. Figure 20 The display shows five grids arranged along the length direction. The numbers in the grids represent the number of times. When performing line erosion, the window is offset from the grid. After line erosion, the five grids are marked as 0, 1, 0, 0, and 0, respectively.

[0133] It should be noted that the value of W must be greater than or equal to the length of a grid cell; preferably, W is an integer multiple of the length of a grid cell. In the example, W >= L1, and preferably, W is an integer multiple of L1.

[0134] In other examples, the threshold T ranges from [6, 8], and its selection is related to the fluctuation of the waveform of the line graph. The smaller the fluctuation, the larger the threshold T value.

[0135] For ease of understanding, the following explanation uses 1 and 0 in a binary graph as an example when illustrating run-length encoding. It is understood that those skilled in the art can modify other types of simplified graphs and other values ​​of the first and second values ​​based on the following explanation.

[0136] In run-length encoding, an 8-connectivity approach can be used. Based on the mesh, connected regions are recursively connected according to the 8-connectivity principle, and then run-length encoding is used to identify the connected regions. Specifically, through 8-connectivity (such as using...) Figure 21 (As shown in the 3x3 window), starting from a non-zero grid Q, if all eight directions of grid Q are non-zero, then the grids in those eight directions are assigned the same value as grid Q, and so on. After completing the entire simplified diagram, we can obtain the following: Figure 22 The diagram shown is an illustration.

[0137] exist Figure 22 In this approach, different connected regions are identified by different numerical values. When calculating the area of ​​each connected region, the frequency of occurrence of the same number is recorded as the area of ​​the connected region. For example... Figure 22 In the problem, the number 9 appears 9 times, so the area of ​​the connected region corresponding to the number 9 is 9. The number 7 appears 20 times, so the area of ​​the connected region corresponding to the number 7 is 20.

[0138] The above example uses a recursive algorithm. In other examples, a traversal algorithm can also be used to find connected components.

[0139] If the area of ​​a connected region is greater than a first set threshold P, then one such connected region corresponds to one single molecule. The value of P is related to the decay time of the single molecule fluorescence. In one example, the first set threshold P ranges from [5, 10].

[0140] In some implementations, please refer to Figure 18 The method for identifying single molecules also includes the following steps: S51, grouping based on the intensity and performing frequency statistics on the number of times to obtain a histogram; S52, finding the maximum point of the histogram and determining that a peak containing a maximum point corresponds to a single molecule if the value of the maximum point is greater than a second set threshold and the width of the peak containing the maximum point is greater than a third set threshold.

[0141] This broadens the application range of single-molecule identification methods. This histogram-based single-molecule identification method can accurately identify single molecules based on the time-series data of bright spot intensity, and is particularly suitable for cases where the number of single molecules contained in a bright spot is greater than 3. In this embodiment, combining histogram-based and run-length encoding methods for single-molecule identification enables accurate identification of single molecules in various waveform line graphs (time-series of bright spot intensity).

[0142] In some implementations, step S51 includes the step of: dividing the intensity into N groups and counting the frequency of the occurrences falling into the N groups. Where, n i The sum of frequencies of occurrences falling in the i-th row of the grid, j represents the time frame number, and g represents the sum of frequencies of occurrences in the i-th row of the grid. ij This represents the frequency of occurrences falling within grid (i,j), and M represents the number of time frames. This allows the frequency to be converted into a histogram of intensity versus frequency, simplifying subsequent single-molecule identification calculations.

[0143] Specifically, in one example, the horizontal axis of the histogram represents the number of groups, and the vertical axis represents the frequency of occurrences falling within the corresponding number of groups. It should be noted that the value of N is equal to the value of N in the M*N grid above. The value of M is equal to the value of M in the M*N grid above.

[0144] In some implementations, the step of grouping based on intensity and performing frequency statistics to obtain a histogram includes the step of performing histogram equalization by L-window: Where, n p Indicates that for n i Equalization, n i ' represents n iwhere p is an integer related to the size of the window L and the i-th row. In this way, the distribution of the histogram can be made more uniform and easier to identify. The L window is used for histogram equalization, and the value of L is related to the size of the decay rate of the single-molecule fluorescence. Generally, if the single-molecule fluorescence light emission is quenched quickly, the value of L should not be too large. The precision of the histogram statistics is affected by the size of the L window, and the value of L can be flexibly set to select the appropriate precision of the histogram statistics. In one example, the value of L is in the range of [5, 15].

[0145] Please refer to Figure 17 , Figure 17 is the equalized histogram, and the horizontal axis represents the group number, and the vertical axis represents the frequency of the number of times falling in the corresponding group number.

[0146] In step S52, all the maximum points of the histogram can be obtained by derivation. The second set threshold Q and the third set threshold H are related to the shape of the peak of the line graph. The sharper the peak, the larger the second set threshold Q and the smaller the third set threshold H; the fatter the peak, the smaller the second set threshold Q and the larger the third set threshold H. In one example, the value of the second set threshold Q is in the range of [2, 6], and the value of the third set threshold H is in the range of [4, 10].

[0147] In the embodiments of the present application, the maximum point is the peak point, and the maximum point is the vertex (inflection point) of the peak, that is, the peak corresponding to a single molecule is determined by a maximum point satisfying the condition.

[0148] In some embodiments, before the grid division of the line graph, the single-molecule identification method further comprises the step of filtering the line graph. In this way, the sudden error caused by light intensity flicker and camera sampling can be eliminated, and the waveform of the line graph is smoother. Specifically, the waveform modification can use the median filtering based on the L2 size window: R = medium(Z i ). In one example, please refer to Figure 14 and Figure 15 , Figure 14 is the line graph before filtering, Figure 15 is the line graph after filtering, and it can be seen from the figure that the waveform of the line graph after filtering is smoother, which is beneficial to improve the accuracy and efficiency of single-molecule identification.

[0149] Please refer to Figure 23The single-molecule counting method of the embodiment includes the following steps: S81, inputting a time sequence of image bright spot intensity; S82, forming a time-intensity broken line graph of the image bright spot according to the time sequence, the broken line graph being composed of a plurality of line segments; S83, performing grid division on the broken line graph to form a plurality of grids arranged in an array, and counting the number of line segments and / or end points of the line segments falling in each grid; S84, performing line erosion on the broken line graph after the grid division according to the number corresponding to each grid, so as to convert the broken line graph after the grid division into a simplified graph; S85, performing run-length encoding on the simplified graph to identify connected regions; S86, calculating the area of each connected region, and determining that one connected region corresponding to one single molecule meets the following condition: the area of the connected region is greater than a first set threshold; and S87, calculating the number of single molecules S2. The single-molecule counting method can quickly count single molecules with high accuracy, by converting the time sequence of bright spot intensity into a connected region of run-length encoding through image processing. It should be noted that the description of the technical features and advantages of the single-molecule recognition and / or counting method in any of the above embodiments and examples, including the steps, parameter settings, and image preprocessing and bright spot detection, is also applicable to the single-molecule counting method of the embodiment. To avoid redundancy, the description will not be repeated here.

[0150] For example, in some embodiments, before the grid division on the broken line graph in step S83, the single-molecule counting method further includes a step of filtering the broken line graph. For another example, in some embodiments, please refer to Figure 24 The single-molecule counting method further includes the following steps: S91, grouping based on the size of the intensity, and frequency counting the number to obtain a histogram; S92, finding a maximum value point of the histogram, and determining that one peak where the maximum value point meets the following condition corresponds to one single molecule: the value of the maximum value point is greater than a second set threshold, and the width of the peak where the maximum value point is located is greater than a third set threshold; S93, calculating the number of single molecules S1; and S94, taking the smaller one of S1 and S2 as the final number of single molecules. The single-molecule counting method based on histogram statistics is particularly suitable for accurately finding the number of single molecules contained in a bright spot > 3, and the single-molecule counting method based on run-length encoding is particularly suitable for accurately finding the number of single molecules contained in a bright spot <= 3. In this embodiment, the two methods are combined to accurately find and count single molecules in broken line graphs of various waveforms. In some embodiments, the simplified graph is a binary graph.

[0151] For example, in some embodiments, before the grid division on the broken line graph in step S83, the single-molecule counting method further includes a step of filtering the broken line graph. For another example, in some embodiments, please refer to Figure 25The single molecule counting method of the embodiment includes the steps of: S61, inputting a time sequence of image bright spot intensity; S62, forming a time-intensity broken line graph of the image bright spot according to the time sequence, the broken line graph being composed of a plurality of line segments; S63, performing grid division on the broken line graph to form a plurality of grids arranged in an array, and counting the number of line segments and / or end points of the line segments falling in each grid; S64, performing line erosion on the broken line graph after the grid division according to the number corresponding to each grid, so as to convert the broken line graph after the grid division into a simplified graph; S65, performing run-length encoding on the simplified graph to identify a connected region; and S66, calculating the area of each connected region, and determining to add 1 to the count of single molecules when the area of the connected region is greater than a first set threshold.

[0152] The single molecule counting method can quickly count single molecules and has high counting accuracy by converting the broken line graph of the time sequence of bright spot intensity into an image processing to obtain a run-length encoded connected region.

[0153] It should be noted that the descriptions of the technical features and advantages of the single molecule recognition / counting method in any of the above embodiments and examples, including the explanations and descriptions of the steps, parameter settings, and image preprocessing bright spot detection, are also applicable to the single molecule counting method of the present embodiment. To avoid redundancy, the descriptions will not be repeated in detail.

[0154] For example, in some embodiments, before the grid division on the broken line graph in step S63, the single molecule counting method further includes a step of filtering the broken line graph. For another example, in some embodiments, please refer to Figure 26 The single molecule counting method further includes the steps of: S71, grouping based on the size of the intensity, and frequency counting the number to obtain a histogram; S72, finding a maximum value point of the histogram, and determining to add 1 to the count of single molecules when the value of the maximum value point is greater than a second set threshold and the width of the peak where the maximum value point is located is greater than a third set threshold; and S73, taking the smaller one of the count of single molecules obtained based on the histogram and the count of single molecules obtained based on the run-length encoding as the final single molecule count. In this way, the single molecule counting method can be applied to a wider range of applications, and a more accurate single molecule count can be obtained.

[0155] The histogram-based molecule counting method is particularly suitable for accurately finding bright spots containing more than 3 molecules, while the run-length encoding-based molecule counting method is particularly suitable for accurately finding bright spots containing less than or equal to 3 molecules. In this embodiment, combining the two methods allows for accurate finding and counting of molecules in line graphs of various waveforms. For example, the number of molecules obtained from the histogram is S1, and the number of molecules obtained from the run-length encoding is S2. Comparing the magnitudes of S1 and S2, the smaller of S1 and S2 is taken as the final number of molecules.

[0156] Please refer to Figure 27 This invention discloses a single-molecule identification device 200, which is used to implement all or part of the steps of the single-molecule identification method in any of the above embodiments or examples. The single-molecule identification device 200 includes: a first input unit 202 for inputting a time series of image bright spot intensity; a first conversion unit 204 for generating a line graph of image bright spot time versus intensity based on the time series in the first input unit 202, the line graph consisting of multiple line segments; and a first grid statistics unit 206 for processing the line graph from the first conversion unit 204. The device performs grid division to form multiple grids arranged in an array, and counts the number of line segments and / or endpoints falling in each grid. A first simplification unit 208 performs line erosion on the gridded line graph based on the counts corresponding to each grid to convert the gridded line graph into a simplified graph. A first identification unit 209 performs run-length encoding on the simplified graph to identify connected regions. A first determination unit 210 calculates the area of ​​each connected region and determines that a connected region corresponds to a single molecule if the area of ​​the connected region is greater than a first set threshold. The single-molecule identification device 200, by converting a time-series line graph of bright spot intensity into an image for image processing to obtain run-length encoded connected regions, can quickly identify single molecules with high accuracy.

[0157] In some implementations, the simplified diagram is a binary diagram.

[0158] It should be noted that the explanations and descriptions of the technical features and beneficial effects of the single-molecule recognition method in any of the above embodiments and examples also apply to the single-molecule recognition device 200 of this embodiment. To avoid redundancy, they will not be elaborated further here. For example, in some embodiments, please refer to... Figure 28 The single-molecule identification device 200 also includes a first filtering unit 212, which is connected to the first grid statistics unit 206, and is used to filter the line graph from the first conversion unit 204 before dividing the line graph into grids.

[0159] In some embodiments, in the first grid statistics unit 206, the grid division of the histogram is performed according to the number of time frames of the collected intensity and the size of the intensity.

[0160] In some embodiments, referring to Figure 29 , the single-molecule recognition device 200 further comprises: a first histogram statistics unit 214, configured to group the intensity according to the size of the intensity, and to count the frequency of the number of times from the grid statistics unit to obtain a histogram; and the first determination unit 210 is configured to find the maximum value point of the histogram from the histogram statistics unit, and to determine that the peak where one maximum value point satisfying the following conditions corresponds to one single molecule: the value of the maximum value point is greater than a second set threshold value, and the width of the peak where the maximum value point is located is greater than a third set threshold value.

[0161] In some embodiments, in the first histogram statistics unit 214, the grouping according to the size of the intensity and the frequency counting of the number of times to obtain the histogram comprises: dividing the intensity according to the size of the intensity into N groups, and counting the frequency of the number of times falling in the N groups: wherein, n i represents the sum of the frequency of the number of times falling in the i-th row of the grid, j represents the number of time frames, and g ij represents the frequency of the number of times falling in the grid (i, j), and M represents the number of time frames.

[0162] In some embodiments, in the first histogram statistics unit 214, the grouping according to the size of the intensity and the frequency counting of the number of times to obtain the histogram comprises: performing histogram equalization by L window: wherein, n p represents the equalization of n i , n i ' represents the sum of the equalization results of n i , and p is an integer related to the size of the window L and the i-th row where the window L is located.

[0163] In some embodiments, referring to Figure 30 , the single-molecule recognition device 200 further comprises: a first image preprocessing unit 218, configured to analyze an input image to be processed to obtain a first image, the image to be processed containing at least one image bright spot, and the image bright spot having at least one pixel point; and a first bright spot detection unit 220, configured to: analyze the first image to calculate a bright spot determination threshold, analyze the first image to obtain a candidate bright spot, judge whether the candidate bright spot is an image bright spot according to the bright spot determination threshold, if the judgment result is yes, obtain a time sequence of the intensity of the image bright spot, and if the judgment result is no, discard the candidate bright spot.

[0164] In some embodiments, referring to Figure 31The first image preprocessing unit 218 comprises a first background reduction unit 226, which is configured to perform background reduction on the to-be-processed image to obtain the first image.

[0165] In some embodiments, the first image preprocessing unit 218 comprises a first image simplification unit 222, which is configured to perform simplification on the to-be-processed image after the background reduction to obtain the first image. Figure 32

[0166] In some embodiments, the first image preprocessing unit 218 comprises a first image filtering unit 224, which is configured to perform filtering on the to-be-processed image to obtain the first image. Figure 33

[0167] In some embodiments, the first image preprocessing unit 218 comprises a first image simplification unit 222, which is configured to perform simplification on the to-be-processed image after the background reduction and the filtering to obtain the first image. Figure 34

[0168] In some embodiments, the first image preprocessing unit 218 comprises a first image filtering unit 224 and a first image simplification unit 222, the first image filtering unit 224 is configured to perform filtering on the to-be-processed image after the background reduction, and the first image simplification unit 222 is configured to perform simplification on the to-be-processed image after the background reduction and the filtering to obtain the first image. Figure 35

[0169] In some embodiments, the first image preprocessing unit 218 comprises a first image simplification unit 222, which is configured to perform simplification on the to-be-processed image to obtain the first image. Figure 36

[0170] In some embodiments, the background reduction on the to-be-processed image in the first background reduction unit 226 comprises: determining the background of the to-be-processed image by using an opening operation, and performing background reduction on the to-be-processed image according to the background. In some embodiments, the filtering is a Mexican hat filtering. In some embodiments, the simplification is a binarization.

[0171] In some embodiments, the first image simplification unit 222 is configured to obtain a signal-to-noise ratio matrix according to the to-be-processed image before the simplification, and to simplify the to-be-processed image before the simplification according to the signal-to-noise ratio matrix to obtain the first image.

[0172] ​​​​​In some embodiments, in the first bright spot detection unit 220, the analyzing the first image to calculate the bright spot determination threshold comprises: processing the first image by Otsu method to calculate the bright spot determination threshold.

[0173] In some embodiments, in the first bright spot detection unit 220, judging whether the candidate bright spot is an image bright spot according to the bright spot determination threshold comprises: finding a pixel point connected by (h*h-1) in the first image and taking the found pixel point as the center of the candidate bright spot, h is a natural number and an odd number greater than 1; judging whether the center of the candidate bright spot satisfies the condition: I max *A BI *ceof guass >T, wherein, I max is the strongest intensity of the center of the h*h window, A BI is the ratio of the set value in the first image in the h*h window, ceof guass is the correlation coefficient of the pixel and the two-dimensional Gaussian distribution of the h*h window, and T is the bright spot determination threshold. If the above condition is satisfied, the bright spot corresponding to the center of the candidate bright spot is judged as an image bright spot. If the above condition is not satisfied, the bright spot corresponding to the center of the candidate bright spot is discarded.

[0174] Please refer to Figure 37 , a single molecule counting device 400 is used to implement all or part of the steps of the single molecule counting method in any of the embodiments of the present application. The single molecule counting device 400 comprises: a second input unit 402 for inputting a time sequence of image bright spot intensity; a second conversion unit 404 for forming a time and intensity line graph of the image bright spot according to the time sequence in the second input unit 402, the line graph being composed of a plurality of line segments; a second grid statistical unit 406 for grid dividing the line graph from the second conversion unit 404 to form a plurality of grids arranged in an array, and counting the number of line segments and / or endpoints of line segments falling in each grid; a second simplification unit 408 for line erosion of the grid-divided line graph according to the number corresponding to each grid to convert the grid-divided line graph into a simplified graph; a second identification unit 409 for run-length encoding of the simplified graph to identify connected regions; a second determination unit 410 for calculating the area of each connected region, and determining that one connected region corresponding to one single molecule satisfies the following condition: the area of the connected region is greater than a first set threshold; and a calculation unit 412 for calculating the number S2 of obtained single molecules. The single molecule counting device 400 described above can quickly count single molecules by converting the line graph of the time sequence of bright spot intensity into image processing to obtain run-length encoded connected regions, and the counting accuracy is also high.

[0175] It should be noted that the technical features and advantages of the counting method of the single molecule in any of the above embodiments are also applicable to the single molecule counting device 400 of the present embodiment, and thus will not be described in detail herein.

[0176] For example, in some embodiments, the time sequence of the intensity of the image bright spot is obtained by Figure 38 The single molecule counting device 400 further comprises a second filtering unit 414 connected to the second grid counting unit 406, configured to filter the line graph from the second conversion unit 404 before grid division.

[0177] In some embodiments, the time sequence of the intensity of the image bright spot is obtained by Figure 39 The single molecule counting device 400 further comprises a second histogram counting unit 416 configured to group the intensity values and count the number of times of the line segments falling into each group to obtain a histogram; in the second determination unit 410, the maximum value point of the histogram from the second histogram counting unit 416 is found, and it is determined that the peak where the maximum value point is located corresponds to a single molecule when the value of the maximum value point is greater than a second set threshold value and the width of the peak where the maximum value point is located is greater than a third set threshold value; in the calculation unit 412, the number of single molecules S1 is calculated, and the smaller one of S1 and S2 is taken as the final number of single molecules.

[0178] For example, in some embodiments, the time sequence of the intensity of the image bright spot is obtained by Figure 40 The single molecule counting device 600 of the present embodiment comprises: a third input unit 602 configured to input a time sequence of the intensity of the image bright spot; a third conversion unit 604 configured to form a line graph of the time and intensity of the image bright spot according to the time sequence in the third input unit 602, the line graph being composed of a plurality of line segments; a third grid counting unit 606 configured to grid divide the line graph from the third conversion unit 604 to form a plurality of grids arranged in an array, and count the number of line segments and / or end points of the line segments falling into each grid; a third simplification unit 608 configured to convert the grid divided line graph into a simplified graph by line erosion of the grid divided line graph according to the number of times corresponding to each grid; a third identification unit 609 configured to identify connected regions by run-length encoding of the simplified graph; and a third determination unit 610 configured to calculate the area of each connected region, and determine to count one single molecule when the area of the connected region is greater than a first set threshold value.

[0179] The single molecule counting device 600 described above can quickly count single molecules by converting the line graph of the time sequence of the intensity of the bright spot into image processing to obtain run-length encoded connected regions, and the counting accuracy is also high.

[0180] It should be noted that the technical features and advantages of the counting method of the single molecule in any of the above embodiments and examples are also applicable to the single molecule counting device 600 of the present embodiment, and thus will not be described in detail herein.

[0181] For example, in some embodiments, the counting method of the single molecule comprises the following steps. Figure 41 The single molecule counting device 600 further comprises a third filtering unit 612 connected to the third grid counting unit 606, configured to filter the line graph from the third conversion unit 604 before grid division.

[0182] In some embodiments, the counting method of the single molecule comprises the following steps. Figure 42 The single molecule counting device 600 further comprises a third histogram counting unit 614 configured to group the intensity values and count the number of times of the line graph from the third grid counting unit 606 to obtain a histogram; the third determination unit 610 is configured to find the maximum value point of the histogram from the third histogram counting unit 614, and determine to add 1 to the count of the single molecule when the value of the maximum value point is greater than the second set threshold value and the width of the peak where the maximum value point is located is greater than the third set threshold value; and the smaller one of the count of the single molecule based on the histogram and the count of the single molecule based on the run-length coding is taken as the final count of the single molecule.

[0183] For example, in some embodiments, the counting method of the single molecule comprises the following steps. Figure 43 The single molecule processing system 300 of the present embodiment comprises a data input device 302 configured to input data, a data output device 304 configured to output data, a storage device 306 configured to store data, wherein the data comprises computer executable programs; and a processor 308 configured to execute the computer executable programs, wherein the execution of the computer executable programs comprises the method of any of the above embodiments.

[0184] The computer readable storage medium of the present embodiment is configured to store programs for execution by a computer, wherein the execution of the programs comprises the method of any of the above embodiments. The computer readable storage medium can comprise a read-only memory, a random access memory, a magnetic disk or an optical disk, etc.

[0185] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "exemplary embodiment", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the exemplary description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0186] In addition, each functional unit in each embodiment of the present application can be integrated in one processing module, or each unit can exist physically separately, or two or more units can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software functional module. When the integrated module is realized in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer readable storage medium.

[0187] Although the embodiments of the present application have been shown and described above, it should be understood by those skilled in the art that the above embodiments are exemplary and cannot be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.

Claims

1. A method of recognizing a single molecule, characterized by, The method comprises the steps of: inputting a time series of image bright spot intensity, the image being from a single molecule sequencing platform; forming a time-intensity line graph of the image bright spot according to the time series, the line graph being composed of a plurality of line segments; grid dividing the line graph to form a plurality of grids arranged in an array, counting the number of line segments and / or endpoints of the line segments falling in each grid; line eroding the grid-divided line graph according to the number corresponding to each grid to convert the grid-divided line graph into a simplified graph; run-length encoding the simplified graph to identify connected regions; calculating the area of each connected region and determining that one connected region corresponding to one single molecule satisfies the condition that the area of the connected region is greater than a first set threshold.

2. A method of counting single molecules, characterized in that The method comprises the steps of: inputting a time series of image bright spot intensity, the image being from a single molecule sequencing platform; forming a time-intensity line graph of the image bright spot according to the time series, the line graph being composed of a plurality of line segments; grid dividing the line graph to form a plurality of grids arranged in an array, counting the number of line segments and / or endpoints of the line segments falling in each grid; line eroding the grid-divided line graph according to the number corresponding to each grid to convert the grid-divided line graph into a simplified graph; run-length encoding the simplified graph to identify connected regions; calculating the area of each connected region and determining that one connected region corresponding to one single molecule satisfies the condition that the area of the connected region is greater than a first set threshold. calculating the number of single molecules S2.

3. A method of counting single molecules, characterized by, The method comprises the steps of: inputting a time series of image bright spot intensity, the image being from a single molecule sequencing platform; forming a time-intensity line graph of the image bright spot according to the time series, the line graph being composed of a plurality of line segments; grid dividing the line graph to form a plurality of grids arranged in an array, counting the number of line segments and / or endpoints of the line segments falling in each grid; line eroding the grid-divided line graph according to the number corresponding to each grid to convert the grid-divided line graph into a simplified graph; run-length encoding the simplified graph to identify connected regions; calculating the area of each connected region and determining that one connected region corresponding to one single molecule satisfies the condition that the area of the connected region is greater than a first set threshold.

4. The method of claim 1, wherein, The method further comprises the steps of: grouping based on the size of the intensity, frequency counting the number to obtain a histogram; finding a maximum value point of the histogram and determining that one peak where the maximum value point is located corresponding to one single molecule satisfies the condition that the value of the maximum value point is greater than a second set threshold and the width of the peak where the maximum value point is located is greater than a third set threshold.

5. The method of claim 2, wherein, The method further comprises the steps of: grouping based on the size of the intensity, frequency counting the number to obtain a histogram; Finding maximum points of the histogram, and determining that a peak where a maximum point is located corresponds to a single molecule when the value of the maximum point is greater than a second set threshold and the width of the peak where the maximum point is located is greater than a third set threshold; Calculating the number of single molecules S1, and taking the smaller one of S1 and S2 as the final single molecule number.

6. The method of claim 2, wherein, Further comprising steps of: Grouping based on the magnitude of the intensity, and frequency counting the number of times to obtain a histogram; Finding maximum points of the histogram, and determining that a peak where a maximum point is located corresponds to a single molecule when the value of the maximum point is greater than a second set threshold and the width of the peak where the maximum point is located is greater than a third set threshold; Taking the smaller one of the single molecule number obtained based on the histogram and the single molecule number obtained based on the run-length coding as the final single molecule number.

7. The method according to any one of claims 1 to 6, characterized in that, Before grid division of the line graph, further comprising a step of filtering the line graph.

8. The method according to any one of claims 1 to 6, characterized in that, The grid division of the line graph is according to the time frame number of collecting the intensity and the magnitude of the intensity.

9. The method according to any one of claims 1 to 6, characterized in that, The simplified graph is a binary graph.

10. The method according to any one of claims 1 to 6, characterized in that, The step of grouping based on the magnitude of the intensity, and frequency counting the number of times to obtain a histogram comprises steps of: Dividing according to the magnitude of the intensity into N groups, and counting the frequency of the number of times falling in the N groups: where n i represents the sum of the frequencies of the number of times falling on the i-th row of the grid, j represents the number of time frames, g ij represents the frequency of the number of times falling on the grid (i, j), and M represents the number of time frames.

11. The method of claim 10, wherein, The step of grouping based on the magnitude of the intensity, and frequency counting the number of times to obtain a histogram comprises a step of histogram equalization by L window: where n p represents the sum of the equalization results of n i , n i represents the equalization result of n i , and p is an integer related to the size of the window L and the i-th line where the window L is located.

12. The method according to any one of claims 1 to 6, characterized in that, Further comprising: An image preprocessing step, which analyzes an input image to be processed to obtain a first image, the image to be processed containing at least one image bright spot, the image bright spot having at least one pixel point; A bright spot detection step, which comprises steps of: Analyzing the first image to calculate a bright spot determination threshold, Analyzing the first image to obtain a candidate bright spot, Determining whether the candidate bright spot is the image bright spot according to the bright spot determination threshold, If the determination result is yes, obtaining a time sequence of the image bright spot intensity, If the determination result is no, discarding the candidate bright spot.

13. The method of claim 12, wherein, The image preprocessing step comprises: performing background subtraction on the image to be processed to obtain the first image.

14. The method of claim 13, wherein, The image preprocessing step comprises: performing simplification processing on the image to be processed after background subtraction to obtain the first image.

15. The method of claim 12, wherein, The image preprocessing step comprises: performing filtering processing on the image to be processed to obtain the first image.

16. The method of claim 12, wherein, The image preprocessing step comprises: performing filtering processing on the image to be processed after background subtraction to obtain the first image.

17. The method of claim 16, wherein, The image preprocessing step comprises: performing simplification processing on the image to be processed after background subtraction and filtering processing to obtain the first image.

18. The method of claim 12, wherein, The image preprocessing step comprises: performing simplification processing on the image to be processed to obtain the first image.

19. The method of claim 13, wherein, The background subtraction on the image to be processed comprises: Determining the background of the image to be processed by using an opening operation, Performing background subtraction on the image to be processed according to the background.

20. The method of claim 15, wherein, The filtering processing is a Mexican hat filtering processing.

21. The method of claim 14, wherein, The simplification processing is a binarization processing.

22. The method of claim 14, wherein, In the simplification processing, a signal-to-noise ratio matrix is obtained according to the image to be processed before the simplification processing, and the image to be processed before the simplification processing is simplified according to the signal-to-noise ratio matrix to obtain the first image.

23. The method of claim 12, wherein, The step of analyzing the first image to calculate a bright spot determination threshold comprises: The first image is processed by the Otsu method to calculate the bright spot determination threshold.

24. The method of claim 12, wherein, The step of determining whether the candidate bright spot is the image bright spot according to the bright spot determination threshold comprises: finding a pixel point greater than (h*h-1) connected in the first image, and taking the found pixel point as the center of the candidate bright spot, h being a natural number and an odd number greater than 1; determining whether the center of the candidate bright spot satisfies the condition: I max *A BI *ceof guass >T, wherein I max is the strongest intensity of the h*h window, A BI is the ratio of the first image in the h*h window that is a set value, ceof guass is the correlation coefficient of the h*h window and the two-dimensional Gaussian distribution, and T is the bright spot determination threshold. If the above condition is met, the bright spot corresponding to the center of the candidate bright spot is determined as the image bright spot, If the above condition is not met, the bright spot corresponding to the center of the candidate bright spot is discarded.

25. A single molecule recognition device, characterized by Comprise: An input unit configured to input a time sequence of image bright spot intensity, the image being from a single molecule sequencing platform; A conversion unit configured to form a time-intensity polyline of the image bright spot according to the time sequence in the input unit, the polyline being composed of a plurality of line segments; A grid counting unit configured to perform grid division on the polyline from the conversion unit to form a plurality of grids arranged in an array, and count the number of the line segments and / or endpoints of the line segments falling in each grid; A simplification unit configured to perform line erosion on the grid-divided polyline according to the number corresponding to each grid to convert the grid-divided polyline into a simplified graph; An identification unit configured to perform run-length encoding on the simplified graph to identify connected regions; A determination unit configured to calculate the area of each connected region, and determine that one connected region satisfying the following condition corresponds to one single molecule: the area of the connected region is greater than a first set threshold.

26. A single molecule counting device, comprising: Comprise: An input unit configured to input a time sequence of image bright spot intensity, the image being from a single molecule sequencing platform; A conversion unit configured to form a time-intensity polyline of the image bright spot according to the time sequence in the input unit, the polyline being composed of a plurality of line segments; A grid counting unit configured to perform grid division on the polyline from the conversion unit to form a plurality of grids arranged in an array, and count the number of the line segments and / or endpoints of the line segments falling in each grid; A simplification unit configured to perform line erosion on the grid-divided polyline according to the number corresponding to each grid to convert the grid-divided polyline into a simplified graph; An identification unit configured to perform run-length encoding on the simplified graph to identify connected regions; A determination unit configured to calculate the area of each connected region, and determine that one connected region satisfying the following condition corresponds to one single molecule: the area of the connected region is greater than a first set threshold. A calculation unit configured to calculate the number S2 of single molecules.

27. A single molecule counting device, characterized by Comprise: inputting a time series of image spot intensity, the image being from a single molecule sequencing platform; transforming, according to the time series in the inputting unit, the time and intensity of the image spot into a time-intensity polyline, the polyline being composed of a plurality of line segments; grid counting, grid dividing the polyline from the transforming unit to form a plurality of grids arranged in an array, counting the number of line segments and / or endpoints of the line segments falling in each of the grids; simplifying, according to the number corresponding to each of the grids, line eroding the grid-divided polyline to convert the grid-divided polyline into a simplified graph; identifying, run-length encoding the simplified graph to identify connected regions; determining, calculating the area of each of the connected regions, and determining to add 1 to the count of single molecules when the area of the connected region is greater than a first set threshold.

28. The apparatus of claim 25, wherein, Further comprising: histogram counting, grouping based on the magnitude of the intensity, and frequency counting the number from the grid counting unit to obtain a histogram; in the determining unit, finding the maximum points of the histogram from the histogram counting unit, and determining that the peak where one of the maximum points is located corresponds to one single molecule when the value of the maximum point is greater than a second set threshold and the width of the peak where the maximum point is located is greater than a third set threshold.

29. The apparatus of claim 26, wherein, Further comprising: histogram counting, grouping based on the magnitude of the intensity, and frequency counting the number from the grid counting unit to obtain a histogram; in the determining unit, finding the maximum points of the histogram from the histogram counting unit, and determining that the peak where one of the maximum points is located corresponds to one single molecule when the value of the maximum point is greater than a second set threshold and the width of the peak where the maximum point is located is greater than a third set threshold; in the calculating unit, calculating the number of single molecules S1, and taking the smaller one of S1 and S2 as the final number of single molecules.

30. The apparatus of claim 27, wherein, Further comprising: histogram counting, grouping based on the magnitude of the intensity, and frequency counting the number from the grid counting unit to obtain a histogram; in the determining unit, finding the maximum points of the histogram from the histogram counting unit, and determining to add 1 to the count of single molecules when the value of the maximum point is greater than a second set threshold and the width of the peak where the maximum point is located is greater than a third set threshold; and taking the smaller one of the count of single molecules obtained based on the histogram and the count of single molecules obtained based on the run-length encoding as the final number of single molecules.

31. The apparatus of any of claims 25-30, wherein, The device further comprises a filtering unit connected to the grid counting unit, for filtering the polyline from the transforming unit before grid dividing the polyline.

32. The device of any one of claims 25-30, wherein, In the grid counting unit, the grid dividing of the polyline is according to the number of time frames of collecting the intensity and the magnitude of the intensity.

33. The apparatus of any one of claims 25-30, wherein, The simplified graph is a binary graph.

34. The apparatus of any one of claims 25-30, wherein, In the histogram statistics unit, the times are frequency counted based on the magnitude of the intensity to obtain a histogram, including: dividing into N groups according to the magnitude of the intensity, and counting the frequency of the times falling in the N groups: where n i represents the sum of the frequencies of the number of times falling on the i-th row of the grid, j represents the number of time frames, g ij represents the frequency of the number of times falling on the grid (i, j), and M represents the number of time frames.

35. The apparatus of claim 33, wherein, In the histogram statistics unit, the times are frequency counted based on the magnitude of the intensity to obtain a histogram, including: histogram equalization by L windows: where n p represents the sum of the equalization results of n i , and p is an integer related to the size of the window L and the i-th row where the window L is located. i i where n i represents the sum of the equalization results of n i , and p is an integer related to the size of the window L and the i-th row where the window L is located.

36. The device of any one of claims 25-30, wherein, Further comprising: An image preprocessing unit, configured to analyze an input image to be processed to obtain a first image, the image to be processed containing at least one image highlight, the image highlight having at least one pixel point; A highlight detection unit, configured to: analyze the first image to calculate a highlight determination threshold, analyze the first image to obtain a candidate highlight, determine whether the candidate highlight is the image highlight according to the highlight determination threshold, if the determination result is yes, obtain a time sequence of the image highlight intensity, if the determination result is no, discard the candidate highlight.

37. The device of claim 36, wherein, The image preprocessing unit comprises a background subtraction unit, configured to perform background subtraction on the image to be processed to obtain the first image.

38. The device of claim 37, wherein, The image preprocessing unit comprises an image simplification unit, configured to perform simplification processing on the image to be processed after the background subtraction to obtain the first image.

39. The device of claim 36, wherein, The image preprocessing unit comprises an image filtering unit, configured to perform filtering processing on the image to be processed to obtain the first image.

40. The device of claim 36, wherein, The image preprocessing unit comprises a background subtraction unit and an image filtering unit, the background subtraction unit being configured to perform background subtraction on the image to be processed, and the image filtering unit being configured to perform filtering processing on the image to be processed after the background subtraction to obtain the first image.

41. The device of claim 40, wherein, The image preprocessing unit comprises a simplification unit, configured to perform simplification processing on the image to be processed after the background subtraction and the filtering processing to obtain the first image.

42. The device of claim 36, wherein, The image preprocessing unit comprises an image simplification unit, configured to perform simplification processing on the image to be processed to obtain the first image.

43. The device of claim 37, wherein, In the background subtraction unit, the background subtraction on the image to be processed comprises: determining the background of the image to be processed by using an opening operation, performing background subtraction on the image to be processed according to the background.

44. The device of any one of claims 39, wherein, The filtering processing is Mexican hat filtering processing.

45. The device of claim 38, wherein, The simplification processing is binarization processing.

46. The device of claim 38, wherein, The image simplification unit is configured to obtain a signal-to-noise ratio matrix according to the image to be processed before the simplification processing, and simplify the image to be processed before the simplification processing according to the signal-to-noise ratio matrix to obtain the first image.

47. The device of claim 36, wherein, In the highlight detection unit, analyzing the first image to calculate a highlight determination threshold comprises: processing the first image by Otsu method to calculate the highlight determination threshold.

48. The device of claim 36, wherein, In the bright spot detection unit, the judging whether the candidate bright spot is the image bright spot according to the bright spot judging threshold comprises: finding the pixel points greater than (h*h-1) connected in the first image and taking the found pixel points as the center of the candidate bright spot, h is a natural number and an odd number greater than 1; determining whether the center of the candidate bright spot satisfies the condition: I max *A BI *ceof guass > T, wherein I max is the strongest intensity of the h*h window, A BI is the ratio of the first image in the h*h window that is a set value, ceof guass is the correlation coefficient of the h*h window and the two-dimensional Gaussian distribution, T is the bright spot determination threshold, If the above condition is met, judging that the bright spot corresponding to the center of the candidate bright spot is the image bright spot, If the above condition is not met, discarding the bright spot corresponding to the center of the candidate bright spot.

49. A single molecule processing system, comprising: Comprise: Data input device for inputting data; Data output device for outputting data; Storage device for storing data, the data comprising computer executable programs; Processor for executing the computer executable programs, executing the computer executable programs comprising completing the method according to any one of claims 1-24.

Citation Information

Patent Citations

  • Reproducible quantification of biomarker expression

    AU2009293380A1

  • Automatic cell localization method based on minimized model L1

    CN103236052A

  • Real-time single-molecule positioning method guaranteeing precision and system thereof

    CN105243677A

  • Single-molecule image correction device

    CN105303540A

  • Methods and apparatus for optical segmentation of biological samples

    US20110235875A1