Single-molecule immune array analysis model based on machine learning and image recognition algorithm
By employing a single-molecule immune array analysis model based on machine learning and image recognition algorithms, the problems of noise and interference in fluorescence images were solved, enabling accurate detection and concentration conversion of single-molecule fluorescent markers, thus improving the accuracy and reliability of detection.
Patent Information
- Application Number
- CN202510957007.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-10-21
AI Technical Summary
The fluorescence images in existing single-molecule array technology contain background noise and interference factors, which affect the accurate detection and quantitative analysis of target molecules.
A single-molecule immune array analysis model based on machine learning and image recognition algorithms is adopted, including modules for image preprocessing, denoising, reaction well identification, fluorescence signal conversion, and concentration conversion. Through image processing and analysis algorithms, interference factors are removed, fluorescence signals are accurately detected, and converted into concentrations.
It improves the accuracy and reliability of single-molecule immune array analysis, enhances the clarity and visualization of molecular signals, and realizes an automated process from fluorescence signal to concentration.
Smart Images

Figure CN120823403A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of biomolecule detection, and in particular relates to a single molecule immune array analysis model based on machine learning and image recognition algorithms. Background Art
[0002] Single Molecular Array (SIMOA) technology is a method for ultra-sensitive detection of biomarkers such as proteins and nucleic acids in serum and plasma. The key to Simoa technology is the capture and sealing of immune complexes in a chip containing over 200,000 femtoliter-level pores for reaction. This allows for ultra-sensitive detection of low-concentration biomarkers such as proteins and nucleic acids in serum and plasma that are undetectable by traditional methods, making it a leading technology in the field of femtogram-level ultra-low abundance protein detection. Simoa technology is 1,000 times more sensitive than traditional enzyme-linked immunosorbent assays (ELISAs), enhancing the ability to detect central nervous system biomarkers and opening up new research prospects in areas such as life science research, in vitro diagnostics, companion diagnostics, and blood screening.
[0003] In the single-molecule array Simoa, fluorescent markers are used as indicators to detect target molecules. However, fluorescence images often contain background noise, impurities, and other interfering factors that can interfere with the accurate detection and quantification of target molecules.
[0004] In view of this, the present invention proposes an optimized image-based single-molecule immune array analysis method, which can eliminate interference factors in fluorescence images, accurately quantify the fluorescence signal intensity, and convert the fluorescence signal intensity into an exact concentration to achieve accurate quantification. Summary of the Invention
[0005] The present invention aims to provide an image-based single-molecule immune array analysis method. During the single-molecule immune array analysis process, after fluorescently labeling the molecules bound to the antibody, the method uses image processing and analysis algorithms to accurately detect and analyze the single-molecule fluorescent markers, extracting the region of interest from the background noise and removing other interfering factors, thereby enhancing the clarity and visualization of the molecular signal. The fluorescence signal intensity of the detection image before and after detection is compared to obtain the change in fluorescence intensity corresponding to each detected molecule. Based on the change in fluorescence intensity before and after and the standard curve information, the concentration of each detected molecule is converted. This helps to improve the accuracy and reliability of single-molecule immune array analysis.
[0006] In order to achieve the above objectives, the technical solutions adopted are:
[0007] A single-molecule immune array analysis model based on machine learning and image recognition algorithms, including:
[0008] An image preprocessing module, used to convert the single-molecule fluorescent marker detection image into an image format that can be recognized by a computer program;
[0009] An image denoising module, used for pre-processing the identifiable single-molecule fluorescent marker detection image to remove noise, bubbles, and repeated proteins;
[0010] An image analysis module, comprising: a reaction well identification submodule and a reaction molecule landing identification submodule, for generating an optimized single molecule fluorescent marker detection image based on the single molecule fluorescent marker detection characteristics, identifying the reaction wells of the titration plate, and the number, position, and nearby slice images of the proteins landing in the reaction wells;
[0011] The fluorescence signal conversion module compares the fluorescence signal intensity of the detection image before and after the detection after the image analysis module identifies the molecules that have fallen into the image, and obtains the fluorescence intensity change before and after for each detected molecule;
[0012] The detection concentration conversion module calculates the corresponding concentration of each detection molecule based on the changes in fluorescence intensity before and after and the standard curve information.
[0013] Furthermore, in the image preprocessing module, the single-molecule fluorescent marker detection image is converted from the ipl format to the tiff format.
[0014] Furthermore, the reaction well identification module adaptively applies threshold processing according to the grayscale value of the local area of the image to obtain a binarized image, thereby clearly distinguishing between bright pixels and dark pixels and identifying the location of the reaction well.
[0015] Furthermore, in the fluorescence signal conversion module, the average enzyme number AEB per magnetic bead is calculated using the following formula:
[0016]
[0017] Among them, F on I is the ratio of the number of reaction wells with increased fluorescence intensity of a single sample to the total number of reaction wells. bead I is the average increase in pixel value of the reaction wells where the fluorescence intensity of a single sample increases. single It is the average pixel value corresponding to each molecule to be detected for all samples in the same batch of experiments.
[0018] Furthermore, the F on The formula is as follows:
[0019]
[0020] Wherein, the Bead total In order to identify the location and number of the total reaction wells, Bead on is the number of all reaction wells with increased fluorescence intensity;
[0021] The I bead The formula is as follows: bead =(∑ i∈onbeads Δpix i ) / Bead on ;
[0022] Among them, the Δpix i is the pixel difference of the reaction well with increased fluorescence intensity before and after the reaction;
[0023] The I single The formula is as follows:
[0024]
[0025] Among them, the ∑Δpix i It is the total pixel difference of the reaction wells with increased fluorescence intensity of all samples in the same batch.
[0026] Furthermore, in the detection concentration conversion module, the AEB value and the concentration value are converted by fitting the regression equation of the standard curve; the standard curve is fitted by a nonlinear regression method, and the fitting formula is as follows:
[0027]
[0028] The second object of the present invention is to provide a single molecule immune array analysis method based on machine learning and image recognition algorithms, using the above-mentioned single molecule immune array analysis model.
[0029] The third invention object of the present invention is to provide a single molecule immune array analysis device based on machine learning and image recognition algorithms, including the above-mentioned single molecule immune array analysis model.
[0030] The fourth inventive object of the present invention is to provide a computer device comprising a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the operation of the above-mentioned single molecule immune array analysis model.
[0031] A fifth inventive object of the present invention is to provide a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the operation of the above-mentioned single molecule immune array analysis model.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] 1. The technical solution of the present invention uses image processing and analysis algorithms to accurately detect and analyze single-molecule fluorescent markers after fluorescently labeling the molecules bound to the antibody. The single-molecule fluorescent marker detection image is processed and segmented to extract the region of interest from the background noise and remove other interfering factors, thereby enhancing the clarity and visualization of the molecular signal, which helps to improve the accuracy and reliability of single-molecule immune array analysis.
[0034] 2. Unlike other technologies that exploit the black-box trap, which results in poor interpretability of the algorithm model, the present invention's algorithm integrates feature extraction and machine learning algorithms, balancing practical relevance and classification success rate. This improves the accuracy of the algorithm while ensuring interpretability. Combined with the subsequent concentration conversion formula, this method achieves a complete automated process from raw images to precise detection results. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 To identify the location of the reaction well;
[0036] Figure 2 The reaction well position is calculated based on the reaction well contour;
[0037] Figure 3 The final identified image of the reaction well;
[0038] Figure 4 is an image of the reaction well into which the molecule to be detected falls;
[0039] Figure 5 Summary of Poisson regression. DETAILED DESCRIPTION
[0040] To further illustrate the single-molecule immune array analysis model based on machine learning and image recognition algorithms and achieve the intended purpose of the present invention, the following describes in detail the specific implementation, structure, features, and efficacy of the single-molecule immune array analysis model based on machine learning and image recognition algorithms proposed in accordance with the present invention, in conjunction with preferred embodiments. In the following description, different "one embodiment" or "embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.
[0041] The following is a detailed introduction to a single molecule immune array analysis model based on machine learning and image recognition algorithms in conjunction with specific embodiments:
[0042] Example 1.
[0043] The single molecule immune array analysis model based on machine learning and image recognition algorithms described in the present invention comprises:
[0044] The image preprocessing module is used to convert single-molecule fluorescent marker detection images into an image format that can be recognized by computer programs. Single-molecule immunoassay data is in ipl format, which needs to be converted to tiff format using ImageJ software for subsequent analysis.
[0045] The image denoising module is used to pre-process the identifiable single-molecule fluorescent marker detection images by removing noise, bubbles, and duplicate proteins. The raw image data contains a large amount of noise, such as bubbles, dust, pixel failure, camera offset, and protein binding, which can cause deviations in the final fluorescence intensity index. The image optimization module calculates and compares the pixel count of the image to locate the noise and remove it.
[0046] An image analysis module, comprising: a reaction well identification submodule and a reaction molecule landing identification submodule, for generating an optimized single molecule fluorescent marker detection image based on the single molecule fluorescent marker detection characteristics, identifying the reaction wells of the titration plate, and the number, position, and nearby slice images of the proteins landing in the reaction wells;
[0047] The fluorescence signal conversion module compares the fluorescence signal intensity of the detection image before and after the detection after the image analysis module identifies the molecules that have fallen into the image, and obtains the fluorescence intensity change before and after for each detected molecule;
[0048] The detection concentration conversion module calculates the corresponding concentration of each detection molecule based on the changes in fluorescence intensity before and after and the standard curve information.
[0049] The specific modules are:
[0050] 1. Image preprocessing module
[0051] A module used to convert the format of offline data. Since the subsequent processing module uses a self-written Python program, it is necessary to process the offline data in ipl format into an image format that can be read by the program to obtain a single-molecule fluorescent marker detection image that can be recognized by the computer program.
[0052] This module uses ImageJ software to read images in ipl format, and writes macro scripts to convert all images into tiff format. A series of operations are recorded and automatically executed to improve work efficiency, and then imported into subsequent Python programs for analysis.
[0053] 2. Image denoising module
[0054] It is used to perform denoising, de-bubble, and de-repeating protein processing on the single-molecule fluorescent marker detection images obtained by the image preprocessing module.
[0055] Single-molecule fluorescent marker detection images that can be identified by computer programs may contain a large amount of noise. The images may contain bubbles, dust noise, tiny pixel loss, camera offset, and even multi-protein connections, which can lead to deviations in the final fluorescence intensity index. The image denoising module calculates and compares the pixel count of the image to locate the location of the noise and remove it. Different targeted noise treatment methods are required to address this problem.
[0056] a. Pixel loss noise will appear noticeably brighter against a dark background, and its pixel coordinates will remain unchanged between images. If a bright, abrupt white spot appears at the same pixel coordinate in both fluorescence images before and after the experiment, this pixel is considered lost and requires correction. After identifying the coordinates of the pixel loss noise point, we perform a fitting process based on the mean pixel values of nearby points to complete the pixel value.
[0057] b. Noise caused by camera misalignment can result in missing positions for some reaction wells. This example uses a method to identify valid reaction well locations, eliminating invalid reaction wells caused by black edges due to camera misalignment. All off-machine images from the same batch are calibrated to determine the final number of reaction wells for subsequent performance analysis.
[0058] c. Dust, bubbles, and multi-protein binding can cause a large number of bright spots to become connected, making it difficult to accurately identify the reaction well locations. Therefore, these bright spots need to be identified and removed. This type of noise is handled primarily based on the difference in pixel count between a connected white bright spot and a normal white bright spot. If the number of pixels is greater than 50,000, it indicates the presence of large bubbles, and the sample is invalid. If the number of pixels is greater than 40, it indicates protein binding or dust, and the reaction wells in this area are removed.
[0059] 3. Image Analysis Module
[0060] The image analysis module includes: a reaction hole recognition submodule and a reaction molecule falling recognition submodule.
[0061] Specifically:
[0062] (1) Reaction well recognition submodule
[0063] This embodiment uses the adaptiveThreshold function in OpenCV to identify the reaction wells. The threshold processing is adaptively applied according to the grayscale value of the local area of the image to obtain a binary image, thereby clearly distinguishing between bright pixels and dark pixels, thereby identifying the location of the reaction wells. Figure 1 shown.
[0064] To identify the contour of the reaction well, use the findContours function to identify its contour, and calculate the center point of the reaction well based on the contour, and then draw the standardized reaction well position based on the center point, such as Figure 2 The coordinates of the reaction well are obtained by intercepting the 5*5 pixel coordinates around the center point of the reaction well. These coordinates are then used for subsequent analysis such as the detection molecule falling into the well and the detection molecule reaction identification.
[0065] (2) The reaction molecules fall into the recognition submodule:
[0066] In this embodiment, based on the principle of the Canny method, the fluorescence image of the sample after the reaction is subjected to Gaussian filtering, convolution, and non-maximum suppression processing, thereby obtaining a function that can be used to classify light and dark. Then, the function is binarized to obtain the final recognition image of the reaction well where the molecule to be detected falls, as shown in FIG. Figure 3 As shown. Using the above reaction well recognition module, the reaction well where the molecule to be detected falls is cut, and an image of the reaction well where the molecule to be detected falls can be obtained, as shown Figure 4As shown in the figure, the overall classification is performed based on the brightness of the interior of the cut reaction wells. The specific method is: the pixel value of each reaction well is summed, the values of all reaction wells are then counted, and then all reaction wells are binary classified using the GMM algorithm to screen out brighter reaction wells and darker reaction wells. The brighter reaction wells are the reaction wells that fall into the molecule to be detected and are marked as "1", and the darker reaction wells are marked as "0". The coordinate position of the reaction well marked as "1" is generated from this.
[0067] 4. Fluorescence signal conversion module:
[0068] This embodiment quantifies the change in fluorescence intensity before and after the reaction of the reaction well that falls into the molecule to be detected, thereby converting the light signal into a specific concentration value of the molecule to be detected. According to the principle of the single-molecule immune array platform, high-resolution fluorescence imaging is used to determine the proportion of magnetic beads associated with at least one enzyme and the fluorescence intensity of each well. The measurement unit of the single-molecule immune array is the average number of enzymes per bead (AEB Average enzymes per bead). It is calculated using the Poisson distribution at ultra-low concentration (digital mode) or the average fluorescence intensity at higher concentration (analog mode). To calculate AEB, it is necessary to calculate three indicators separately, namely Fon, Ibead, and Isingle. The specific definitions and calculation formulas of the indicators are as follows:
[0069] Define F on The ratio of the number of reaction wells with increased fluorescence intensity of a single sample to the total number of reaction wells is as follows:
[0070]
[0071] According to the above reaction well identification module, the position and number of the total reaction wells can be identified, which are recorded as Bead total According to the above reaction molecules falling into the recognition module, the position and number of the molecules to be detected can be identified. According to the calculation of the brightness difference of the reaction wells where the molecules to be detected fall before and after the reaction, if it is higher than a certain threshold, it is marked as a reaction well with increased fluorescence intensity. The number of reaction wells with increased fluorescence intensity is counted and recorded as Bead on .
[0072] Definition I bead The average increase in pixel value of the reaction wells with increased fluorescence intensity of a single sample is expressed as follows:
[0073] I bead =(∑ i∈onbeads Δpix i ) / Bead on (2)
[0074] Bead onThat is, the number of all the reaction wells with increased fluorescence intensity mentioned above. The pixel difference of the reaction well with increased fluorescence intensity before and after the reaction is calculated and recorded as Δpix i . Count the total pixel difference of all reaction wells with increased fluorescence intensity, that is, the difference between Δpix i Sum, denoted as ∑ i∈onbeads Δpix i .
[0075] Definition I single It is the average pixel value corresponding to each molecule to be detected in all samples in the same batch of experiments. The formula is as follows:
[0076]
[0077] Count the total pixel difference of the reaction wells with increased fluorescence intensity of all samples in the same batch, recorded as ∑Δpix i , count the total number of reaction wells with increased fluorescence intensity of all samples, and record it as ∑Bead on .
[0078] Based on the above three indicators F on , I bead , I single The formula for calculating the AEB value is as follows:
[0079]
[0080] When F on When the value of is ≤0.7, it indicates that the concentration of the molecule to be detected is low. The digital mode is used to calculate AEB. Specifically, the Poisson distribution is used to divide F on Directly converted to AEB, that is AEB Digital Based on the existing data, the present invention on Poisson regression fitting is performed for the case where ≤0.7, and the results are as follows Figure 5 As shown, the final fitting function is shown in the above formula (4).
[0081] When F on When the value is greater than 0.7, it indicates that the concentration of the molecules to be detected is high and accumulation of the molecules to be detected may occur. Therefore, it is necessary to adjust the concentration of the molecules to be detected according to F. on , I bead , I single Calculate AEB using simulation mode, which is AEB Analog , as shown in the above formula (5).
[0082] The present invention adopts R 2 The model performance is evaluated by two evaluation indicators, α and APE. The calculation formulas are as follows:
[0083]
[0084]
[0085] R 2 It is an indicator that measures the similarity between two variables. Its value ranges from 0 to 1. The closer it is to 1, the higher the similarity between the two variables. APE is the average percentage error, which can express the error between the overall and actual situation and indicate the deviation of the fit.
[0086] The AEB values obtained by simulating and calculating 1348 samples and the actual test results AEB are compared to obtain R 2 The APE results are shown in Table 1:
[0087] Table 1
[0088] <![CDATA[R 2 ]]> APE Fon≤0.7 0.91 0.65% Fon>0.7 0.937 0.57%
[0089] The results show that the fitting effect is very close to the true value, proving that the algorithm performs well.
[0090] 5. Detection concentration calculation module
[0091] After obtaining the AEB of the detection molecule, the AEB value needs to be converted into the final concentration value. The single-molecule immune array platform adopts the standard curve fitting method. At the same time as each batch of detection, the calibrator, standard and sample are loaded on the machine together. The calibrator takes different concentration gradients, and the AEB value and exact concentration are known to construct the standard curve. By fitting the regression equation of the standard curve, the conversion formula between the AEB value and the concentration value is obtained, and the AEB values detected by the sample and standard can be converted into concentration values, and compared with the known concentration value of the standard to evaluate the accuracy of the experiment. The function of this module is to fit the standard curve and obtain the conversion formula. The present invention fits the standard curve by the nonlinear regression method, and the fitting formula is as follows:
[0092]
[0093] The Levenberg-Marquardt (LM) algorithm is then used to calculate the A, B, C, and D values in Equation (6). The LM algorithm is a popular method for solving nonlinear least squares problems, which are common in nonlinear regression. The LM algorithm combines the characteristics of the Gauss-Newton algorithm and gradient descent to provide a balanced method that is robust in many practical situations. The following is a brief overview of its working principle:
[0094] (1) Initialization: Start with an initial guess of the model parameters.
[0095] (2) Jacobian matrix calculation: Calculate the Jacobian matrix at the current parameter value, which contains the first-order derivatives of the model prediction with respect to the parameters.
[0096] (3) Adjustment step calculation: Use the following formula to calculate the adjustment amount of the parameter:
[0097] Δp=(J T J+λD) -1 j T r (7)
[0098] where J is the Jacobian matrix, r is the residual vector between the observed and predicted values, λ is a damping factor that helps ensure convergence, and D is usually taken as J T The diagonal matrix consisting of the diagonal elements of J.
[0099] (4) Update parameters:
[0100] p new =p old +Δp (8)
[0101] (5) Check convergence: Evaluate the residual sum of squares using the new parameters. If the residual decreases sufficiently or the maximum number of iterations is reached, stop; otherwise, adjust λ and repeat step 2.
[0102] The above nonlinear regression and LM algorithm were implemented in R (version 4.3.2) using the onls package and the minpack.lm package. The final A, B, C, and D values were obtained by writing a computer program for calculation.
[0103] For example, in a standard curve calibration experiment, the AEB values corresponding to 7 gradient concentrations were measured and shown in Table 2 below.
[0104] Table 2
[0105] AEB value Concentration value (pg / ml) NaN 0 0.033 0.37 0.063 1.19 0.174 4.18 0.506 12.5 1.661 40.1 4.261 120
[0106] After fitting the formula, the A, B, C, and D values are shown in Table 3 below.
[0107] Table 3
[0108]
[0109]
[0110] The above is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the embodiments of the present invention. Any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the embodiments of the present invention are still within the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A single molecule immune array analysis model based on machine learning and image recognition algorithms, characterized in that: include: An image preprocessing module, used to convert the single-molecule fluorescent marker detection image into an image format that can be recognized by a computer program; An image denoising module, used for pre-processing the identifiable single-molecule fluorescent marker detection image to remove noise, bubbles, and repeated proteins; An image analysis module, comprising: a reaction well identification submodule and a reaction molecule landing identification submodule, for generating an optimized single molecule fluorescent marker detection image based on the single molecule fluorescent marker detection characteristics, identifying the reaction wells of the titration plate, and the number, position, and nearby slice images of the proteins landing in the reaction wells; The fluorescence signal conversion module compares the fluorescence signal intensity of the detection image before and after the detection after the image analysis module identifies the molecules that have fallen into the image, and obtains the fluorescence intensity change before and after for each detected molecule; The detection concentration conversion module calculates the corresponding concentration of each detection molecule based on the changes in fluorescence intensity before and after and the standard curve information.
2. The single molecule immune array analysis model according to claim 1, characterized in that In the image preprocessing module, the single-molecule fluorescent marker detection image is converted from the ipl format to the tiff format.
3. The single molecule immune array analysis model according to claim 1, characterized in that The reaction well identification module adaptively applies threshold processing according to the grayscale value of the local area of the image to obtain a binarized image, thereby clearly distinguishing bright pixels from dark pixels and identifying the position of the reaction well.
4. The single molecule immune array analysis model according to claim 1, characterized in that In the fluorescence signal conversion module, the average enzyme number AEB per magnetic bead is calculated using the following formula: Among them, F on I is the ratio of the number of reaction wells with increased fluorescence intensity of a single sample to the total number of reaction wells. bead I is the average increase in pixel value of the reaction wells where the fluorescence intensity of a single sample increases. single It is the average pixel value corresponding to each molecule to be detected for all samples in the same batch of experiments.
5. The single molecule immune array analysis model according to claim 4, characterized in that The F on The formula is as follows: Wherein, the Bead total In order to identify the location and number of the total reaction wells, Bead on is the number of all reaction wells with increased fluorescence intensity; The I bead The formula is as follows: bead =(∑ i∈onbeads Δpix i ) / Bead on ; Among them, the Δpix i is the pixel difference of the reaction well with increased fluorescence intensity before and after the reaction; The I single The formula is as follows: Among them, the ∑Δpix i It is the total pixel difference of the reaction wells with increased fluorescence intensity of all samples in the same batch.
6. The single molecule immune array analysis model according to claim 1, characterized in that In the detection concentration conversion module, the AEB value and the concentration value are converted by fitting the regression equation of the standard curve; the standard curve is fitted by the nonlinear regression method, and the fitting formula is as follows:
7. A single molecule immune array analysis method based on machine learning and image recognition algorithm, characterized in that: The single molecule immune array analysis method adopts the single molecule immune array analysis model described in any one of claims 1-6.
8. A single molecule immune array analysis device based on machine learning and image recognition algorithm, characterized in that: The single molecule immune array analysis device comprises the single molecule immune array analysis model according to any one of claims 1 to 6.
9. A computer device, characterized in that: The invention comprises a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the operation of the single-molecule immune array analysis model according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to implement the operation of the single-molecule immune array analysis model according to any one of claims 1 to 6.