Target molecule identification and quantification method based on solid nanopore detection technology, electronic equipment and computer readable storage medium
By filtering and multi-threaded parallel processing of the current signal data from nanopore sensors, the problem of identifying and quantifying massive signal data in solid-state nanopore detection technology has been solved, enabling rapid and accurate identification and quantitative analysis of target molecules, thus improving analytical efficiency and accuracy.
Patent Information
- Application Number
- CN202510956636.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-31
AI Technical Summary
Existing solid-state nanopore detection technologies struggle to quickly and accurately identify and quantitatively analyze massive amounts of signal data, thus limiting their widespread application in large-scale scientific research and commercialization.
By filtering the current signal data generated by the nanopore sensor, dividing the overlapping window to calculate the baseline current, using multi-threaded parallel processing to screen target through-pore events, and combining saliency and morphological filtering, the average interval time is calculated to determine the relative concentration of target molecules.
This technology enables automated identification and quantitative analysis of solid-state nanopores, improving analytical efficiency and accuracy. It can quickly and accurately confirm whether target molecules are detected and their relative content.
Smart Images

Figure HDA0005496806650000011 
Figure HDA0005496806650000012
Abstract
Description
Technical Field
[0001] This invention relates to the field of solid-state nanopore detection technology, specifically to a method for identifying and quantifying target molecules based on solid-state nanopore detection technology, an electronic device, and a computer-readable storage medium. Background Technology
[0002] Solid-state nanopore molecular detection utilizes channels ranging from 1 nm to 50 nm fabricated on materials such as silicon nitride films or monolayer graphene films. By applying an electric field and monitoring the blocking signal of the ion current in the picoampere to nanoampere range generated when a single molecule passes through the channel, molecular size, sequence, or conformational information can be directly read without fluorescent labeling or amplification. It offers advantages such as single-molecule detection, real-time processing, ease of arraying, and compatibility with CMOS. Solid-state nanopore detection technology can generate massive amounts of electrical signal data in a short time; a single channel can produce over 1 Gbyte of data within 15 minutes, while arrayed nanopores can achieve terabyte levels in the same timeframe.
[0003] However, existing methods are still insufficient for the identification and quantitative analysis of such massive amounts of signal data. Currently, this field still relies mainly on manual observation for data identification, making it difficult to quickly and accurately confirm whether target molecules have been detected and their relative abundance. This greatly limits the promotion and development of solid-state nanopore detection technology in large-scale scientific research applications and the commercial application of molecular detection. Summary of the Invention
[0004] The purpose of this invention is to provide a method, electronic device, and computer-readable storage medium for the identification and quantification of target molecules based on solid-state nanopore detection technology, which can quickly and accurately confirm whether the target molecule has been detected and its relative content.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0006] This invention provides a method for identifying and quantifying target molecules based on solid-state nanopore detection technology, comprising the following steps:
[0007] (1) Acquire the current signal data generated by the target molecule through the nanopore sensor over a period of time;
[0008] (2) Filter the current signal data;
[0009] (3) Identify the filtered data to obtain the target via event set, including:
[0010] Calculate the global baseline current of the filtered data;
[0011] The filtered data is divided into multiple overlapping windows. The baseline current of each window is calculated and compared with the global baseline current. If the ratio of the difference between the two to the global baseline current is greater than a set threshold, the window is removed.
[0012] The remaining windows are divided into multiple data blocks. The data blocks are processed in parallel using multi-threading. The processing of each thread includes filtering all extreme points below the event judgment threshold as potential event candidates based on the event judgment threshold, merging adjacent candidate points to form a continuous event interval, and extending to the complete signal decline or recovery stage to obtain candidate target via events.
[0013] Summarize the candidate target via events from the remaining multiple windows, remove overlapping target via events and linearize them to obtain the target via event set;
[0014] (4) Filter the target via event set, extract signal data from the filtered events, and obtain target molecule recognition events;
[0015] (5) Calculate the average interval time of the target molecule recognition event, and obtain the frequency of the target molecule based on the average interval time, which is the relative concentration of the target molecule.
[0016] In some implementations, the peak and valley values of the current signal data of the data block are obtained by scanning window data. Based on these peak and valley values, the local baseline and local baseline standard deviation of the data block are calculated. The local baseline is called the baseline, the local baseline standard deviation is called std, and the event determination threshold is h0. I The event determination threshold, the local baseline, and the standard deviation of the local baseline satisfy the following relationship: h0 I = 2 × baseline - (baseline + N × std), where N takes the value of 3 to 4.
[0017] In some implementations, the filtering includes filtering the target via event based on one or more of the blockage rate and duration of the target via event, and also includes salience filtering.
[0018] In some specific implementations, filtering the target via event based on the blockage rate of the target via event includes removing events whose blockage rate is less than or equal to a set blockage rate threshold.
[0019] Furthermore, the set congestion rate threshold is not less than 0.25, such as 0.25, 0.3, etc.
[0020] In some specific implementations, filtering the target via event based on the duration of the target via event includes removing events whose duration is less than or equal to a set time threshold.
[0021] Furthermore, the set event threshold is not less than 0.0003.
[0022] In some specific implementations, the saliency filtering includes setting a saliency level, setting a preset threshold based on the blockage rate or duration, filtering out target via events that meet the conditions based on the preset threshold, evaluating the saliency of the target via event based on a statistical model, and filtering out significant events.
[0023] In some implementations, the filtering also includes morphological filtering.
[0024] In some implementations, step (4) further includes filtering the extracted signal data based on the blockage rate of the signal data to obtain the target molecule recognition event.
[0025] In some implementations, sparse sampling and histograms are used to approximate the median of the global current value, and the median of the global current value is used as the global baseline current.
[0026] In some specific implementations, binary division is used to accelerate the approximate calculation of the histogram.
[0027] In some embodiments, the method further includes saving the filtered data to a recording medium in stream and / or binary form, reading the data from the recording medium, and performing the identification on the data.
[0028] In some embodiments, the filtering process includes digitally filtering the current signal data, and the digital filtering method includes, but is not limited to, Bessel digital filtering, and includes, but is not limited to, forward and reverse filtering.
[0029] In some embodiments, the method further includes performing double buffering on the current signal data before the filtering process.
[0030] In some implementations, calculating the average interval between the occurrence of the target molecule recognition events includes statistically analyzing the interval between each target molecule recognition event and another adjacent target molecule recognition event, fitting a log-normal distribution, and obtaining the average interval event.
[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the identification and quantification method as described above.
[0032] The present invention also provides a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the identification and quantification method as described above.
[0033] Due to the application of the above technical solution, the present invention has the following advantages compared with the prior art:
[0034] This invention enables automated identification and quantitative analysis of the massive signal data generated by solid-state nanopore detection technology. By optimizing the data identification method, the analysis efficiency and accuracy are improved. Attached Figure Description
[0035] Figure 1 An overview diagram of sequencing data provided for this invention;
[0036] Figure 2 The present invention provides a baseline and signal identification diagram for a target via event. Detailed Implementation
[0037] Currently, the massive signal data generated by solid-state nanopore molecule detection still mainly relies on manual observation for identification, making it difficult to quickly and accurately confirm whether target molecules have been detected and their relative content. This invention provides a method for the identification and quantification of target molecules based on solid-state nanopore detection technology. This method effectively removes noise and significantly improves signal quality by filtering the current signal generated by the target molecule passing through the nanopore sensor, providing a solid foundation for accurate target molecule identification. Furthermore, this invention efficiently removes abnormal data, enhances data consistency, and optimizes analysis efficiency by dividing overlapping windows, calculating the baseline current of each window, and comparing it with the global baseline current. Abnormal windows with differences exceeding a set threshold are eliminated. In addition, multi-threaded processing technology significantly improves the efficiency of acquiring the target via event set. Finally, by filtering the target via event set, target molecule identification events are accurately selected, and the relative concentration of the target molecule is accurately determined based on the average interval of these events. This invention automates target molecule identification and quantitative analysis, offering significant advantages in high analysis efficiency and accuracy.
[0038] The present invention will be further described below with reference to embodiments. However, the present invention is not limited to the following embodiments. The implementation conditions used in the embodiments can be further adjusted according to different requirements of specific applications, and the implementation conditions not specified are conventional conditions in the industry. The technical features involved in the various embodiments of the present invention can be combined with each other as long as they do not conflict with each other.
[0039] Example 1: A method for identifying and quantifying target molecules based on solid-state nanopore detection technology, comprising the following steps:
[0040] (1) Acquire the current signal data generated by the target molecule through the nanopore sensor over a period of time;
[0041] (2) Filter the current signal data;
[0042] (3) Identify the filtered data to obtain the target via event set, including:
[0043] Calculate the global baseline current of the filtered data;
[0044] The filtered data is divided into multiple overlapping windows. The baseline current of each window is calculated, and the baseline current of the window is compared with the global baseline current. If the ratio of the difference between the two to the global baseline current is greater than a set threshold, the window is removed. Unless otherwise specified, the difference refers to the absolute value of the difference between the two. The set threshold can be designed according to actual needs, and can be 0.15, 0.2, 0.25, etc. For example, if the threshold is set to 0.2, the baseline current of the window is 700, and the global baseline current is 1000, then the ratio of the difference between the two to the global baseline current is 0.3, which is greater than the set threshold, and the window needs to be discarded.
[0045] The remaining windows are divided into multiple data blocks. The data blocks are processed in parallel using multi-threading. The processing of each thread includes filtering all extreme points below the event judgment threshold as potential event candidates based on the event judgment threshold, merging adjacent candidate points to form a continuous event interval, and extending to the complete signal decline or recovery stage to obtain candidate target via events.
[0046] Summarize the candidate target via events from the remaining multiple windows, remove overlapping target via events and linearize them to obtain the target via event set;
[0047] (4) Filter the target via event set, extract the signal data of the filtered events, and obtain the target molecule recognition events;
[0048] (5) Calculate the average interval between the occurrence of target molecule recognition events. Based on this average interval, obtain the frequency of occurrence of target molecules, which is the relative concentration of target molecules.
[0049] The method further includes preprocessing the current signal data before filtering, including double buffering for each batch of data (current signal data generated by the target molecules through solid nanopores is uploaded in batches). For example, the last 20% of the data from the previous batch is used as a front buffer, and the last 20% of the data from the current batch is used as a back buffer. The front buffer, back buffer, and the next batch of data (100% of the data) are combined to form a 140% length data set. Then, bidirectional filtering is performed on the preprocessed data, i.e., IIR filtering is performed once in the forward direction on the 140% length data set, followed by IIR filtering in the reverse direction. Bidirectional filtering can effectively suppress phase distortion during the filtering process. After processing, each initial value will have a corresponding filtered data set, and the finally bidirectional filtered and denoised data will be recorded in a high-performance, high-precision storage format. The filtering process includes obtaining Bessel digital filter parameters and performing bidirectional filtering. Specifically, based on the characteristics of the current signal data (e.g., sampling frequency, signal frequency range, noise characteristics, etc.), the parameters of the Bessel filter (e.g., filter order, cutoff frequency) are determined, and the Bessel filter coefficients are calculated based on these parameters. The Bessel filter coefficients can be generated using digital signal processing tools (e.g., MATLAB, Python, etc.). As an example, the equation coefficient variables are a = {a0, a1, ... an}, b = {b0, b1, ... bn}, and the equation state variables are zi = {zi0, zi1, ... zin}, where n is the filter order plus 1.
[0050] The method also includes saving the filtered data to a recording medium in stream and / or binary form for easy subsequent identification. Streaming supports fast random access and allows quick jumps to a specific time point without pre-loading into memory, reducing memory consumption. Binary format also supports fast reading into data structures. As an example, it includes the following parts:
[0051] 1) File identifier: Used to identify the file type, such as a fixed string or a specific identifier code.
[0052] 2) Format version: Record the version number of the data format for easy expansion and compatibility checks later.
[0053] 3) Equipment range: Record the range of the equipment, such as the maximum and minimum values of current or voltage.
[0054] 4) Device sampling rate: Record the sampling frequency of the device, in Hz.
[0055] 5) Current bias: Record the current bias value for calibration.
[0056] 6) Current information: Current value at each sampling point.
[0057] 7) Voltage information: Voltage time point and voltage value are recorded in pairs.
[0058] In step (3), sparse sampling (e.g., 1 / 1000, i.e., taking one data point for processing every 1000 data points) and histogram approximation are used to calculate the median of the global current value, which is then used as the global baseline current. To ensure the timeliness of processing large amounts of data, binary division is used to accelerate the histogram calculation. For example, in some implementations, 8 is used as the grouping interval, with values 0-7 in one group, 8-15 in another, 16-23 in yet another, etc. This is achieved by shifting the binary value of the integer value in the computer to the right by 3 bits, which is 50 times faster. Of course, in other implementations, the grouping interval can also be 6, 10, etc.
[0059] In step (3), the filtered data is divided into multiple overlapping windows to optimize memory, avoid loading all data at once, and accelerate the independent processing of different windows by multiple processes in parallel.
[0060] Furthermore, the peak and valley values of the current signal data of the data block are obtained by scanning the window data, and the local baseline and local baseline standard deviation of the data block are calculated based on these peak and valley values. Here, the local baseline is denoted as `baseline`, the local baseline standard deviation is denoted as `std`, and the event determination threshold is `h0`. I The event determination threshold, local baseline, and local baseline standard deviation satisfy the following relationship: h0 I = 2 × baseline - (baseline + N × std), where N takes the value of 3 to 4.
[0061] The filtering in step (4) includes filtering the target via events based on one or more of the blockage rate and duration, and also includes saliency filtering. Filtering the target via events based on the blockage rate includes removing events whose blockage rate is less than or equal to a set blockage rate threshold. The set blockage rate threshold is not less than 0.25, such as 0.25, 0.3, etc. Filtering the target via events based on the duration includes removing events whose duration is less than or equal to a set time threshold. The set time threshold is not less than 0.0003. The saliency filtering includes setting a saliency level, setting a preset threshold based on the blockage rate or duration, filtering target via events that meet the conditions based on the preset threshold, evaluating the saliency of the target via event based on a statistical model, and filtering out saliency events. The filtering also includes morphological filtering. For details on saliency filtering and morphological filtering, refer to existing technologies.
[0062] Furthermore, step (4) also includes filtering the extracted signal data based on the pore-blocking rate of the signal data to obtain the target molecule recognition event. For example, single- and double-stranded DNA structures will form a circular structure, resulting in two signal drops. The filtering is performed based on the pore-blocking rate of the two signal levels. For example, if the pore-blocking rate of the first step current is >15% and the pore-blocking rate of the second step current is >45%, and the difference between the pore-blocking rates of the first and second step currents is >15%, then the molecule is finally listed as the identified target molecule. Target molecules include, but are not limited to, biomolecules such as DNA and proteins.
[0063] In step (5), calculating the average interval time of the target molecule recognition event includes statistically analyzing the interval time between each target molecule recognition event and its adjacent target molecule recognition event, fitting a log-normal distribution to obtain the average interval event. Multiple measurements under the same conditions on the same nanopore can yield the relative concentrations of multiple molecules. This can be used to determine whether a target molecule is present in the sample and the relationship between the concentrations of different molecules. This method is more accurate than directly calculating the number of events divided by the total duration and can reduce the impact of pore-clogging events on density estimation.
[0064] The accuracy of the above method was also verified in the following details. Unless otherwise specified, all reagents and instruments mentioned below are commercially available products.
[0065] Single-stranded nucleic acid templates and probes were synthesized from Sangon Biotech (Shanghai) Co., Ltd. The sequences were derived from the novel coronavirus sequence SARS-CoV-2 and a complementary strand sequence containing a guide sequence plus a complementary pairing sequence, designated M and P, respectively. M and P were dissolved and diluted to 10 μM with nuclease-free water. Equal volumes of M and P were then incubated in a PCR instrument (Biorad T100) to allow for complete binding and form the analyte. Electrolyte solution (1 M KCl) was added to the analyte, and the current signal generated by DNA molecules passing through nanopores was detected using a single-probe ultra-low noise patch-clamp amplifier (Axon Axopatch 200B Capacitor Feedback Patch Clamp Amplifier). Signal collection lasted 10 minutes, generating approximately 7 minutes of data totaling 645 Mb. This data was acquired at a frequency of 500 kHz and filtered using a Bessel digital filter with a 10 kHz cutoff frequency for noise reduction. Figure 1 An overview of the sequencing data is generated, and the data is encoded to produce a .pbf file.
[0066] The data was analyzed using the above method, with a window size of 0.5s and a step size of 25%, allowing for a baseline deviation of 25% and a minimum hole blockage rate of 18%, to obtain a set of candidate target via events. Figure 2 For baseline and signal identification, here is an example of a signal from one of the via events.
[0067] Further filtering of the candidate target via event set requires a minimum via blockage rate of 30% and a minimum via duration of 0.0003s to obtain the target via event set.
[0068] Based on the fact that the first-step current blockage rate is >15%, the second-step current blockage rate is >45%, and the difference between the first-step and second-step current blockage rates is >15%, signal data is extracted from the event to obtain the target molecule recognition event.
[0069] The above method was further used to fit the data of the target identification event to obtain the relative content of the target molecule, which was consistent with the actual content.
[0070] The present invention has been described in detail above, with the aim of enabling those skilled in the art to understand and implement the invention. However, this description should not be construed as limiting the scope of protection of the invention. All equivalent changes or modifications made in accordance with the spirit and essence of the invention should be included within the scope of protection of the invention.
Claims
1. A method for identifying and quantifying target molecules based on solid-state nanopore detection technology, characterized in that, Includes the following steps: (1) Acquire the current signal data generated by the target molecule through the nanopore sensor over a period of time; (2) Filter the current signal data; (3) Identify the filtered data to obtain the target via event set, including: Calculate the global baseline current of the filtered data; The filtered data is divided into multiple overlapping windows. The baseline current of each window is calculated and compared with the global baseline current. If the ratio of the difference between the two to the global baseline current is greater than a set threshold, the window is removed. The remaining windows are divided into multiple data blocks. The data blocks are processed in parallel using multi-threading. The processing of each thread includes filtering all extreme points below the event judgment threshold as potential event candidates based on the event judgment threshold, merging adjacent candidate points to form a continuous event interval, and extending to the complete signal decline or recovery stage to obtain candidate target via events. Summarize the candidate target via events from the remaining multiple windows, remove overlapping target via events and linearize them to obtain the target via event set; (4) Filter the target via event set, extract signal data from the filtered events, and obtain target molecule recognition events; (5) Calculate the average interval time of the target molecule recognition event, and obtain the frequency of the target molecule based on the average interval time, which is the relative concentration of the target molecule.
2. The method for identifying and quantifying target molecules based on solid-state nanopore detection technology according to claim 1, characterized in that, The peak and valley values of the current signal data of the data block are obtained by scanning the window data. Based on these peak and valley values, the local baseline and local baseline standard deviation of the data block are calculated. The local baseline is defined as the baseline, and the local baseline standard deviation is defined as std. The event determination threshold is h0. I The event determination threshold, the local baseline, and the standard deviation of the local baseline satisfy the following relationship: h0 I = 2 × baseline - (baseline + N × std), where N takes the value of 3 to 4.
3. The method for identifying and quantifying target molecules based on solid-state nanopore detection technology according to claim 1, characterized in that, The filtering includes filtering the target via event based on one or more of the blockage rate and duration of the target via event, and also includes salience filtering.
4. The method for identifying and quantifying target molecules based on solid-state nanopore detection technology according to claim 3, characterized in that, Filtering the target via events based on their blockage rate includes removing events whose blockage rate is less than or equal to a set blockage rate threshold. And / or, filtering the target via event based on the duration of the target via event includes removing events whose duration is less than or equal to a set time threshold; And / or, the filtering also includes morphological filtering.
5. The method for identifying and quantifying target molecules based on solid-state nanopore detection technology according to claim 1, characterized in that, Step (4) further includes filtering the extracted signal data based on the blockage rate of the signal data to obtain the target molecule recognition event.
6. The method for identifying and quantifying target molecules based on solid-state nanopore detection technology according to claim 1, characterized in that, The median of the global current value is calculated by sparse sampling and histogram approximation, and the median of the global current value is used as the global baseline current.
7. The method for identifying and quantifying target molecules based on solid-state nanopore detection technology according to claim 6, characterized in that, The approximate calculation of the histogram is accelerated by using binary division.
8. The method for identifying and quantifying target molecules based on solid-state nanopore detection technology according to claim 1, characterized in that, The method further includes saving the filtered data to a recording medium in stream and / or binary form, reading the data from the recording medium, and performing the identification on the data.
9. The method for identifying and quantifying target molecules based on solid-state nanopore detection technology according to claim 1, characterized in that, The filtering process includes digital filtering of the current signal data, and the digital filtering method includes Bessel digital filtering, and performs forward and reverse filtering.
10. The method for identifying and quantifying target molecules based on solid-state nanopore detection technology according to claim 1, characterized in that, The method further includes performing double buffering on the current signal data before the filtering process.
11. The method for identifying and quantifying target molecules based on solid-state nanopore detection technology according to claim 1, characterized in that, Calculating the average interval between the occurrence of the target molecule recognition events includes statistically analyzing the interval between each target molecule recognition event and another adjacent target molecule recognition event, fitting the results to a log-normal distribution, and obtaining the average interval event.
12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the identification and quantification method as described in any one of claims 1 to 11.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the identification and quantification method as described in any one of claims 1 to 11.
Citation Information
Cited By
Single molecule detection system based on solid nanopores
CN121577719A
A nanopore signal event analysis method and system
CN122432735A