Method and apparatus for processing polypeptide nanopore sequencing signal

By constructing a reference density matrix model, the distance fraction of peptide sequencing signals is calculated, and target signals are screened out. This solves the problem of low data quality in high-throughput peptide nanopore sequencing signals and improves data quality and sequencing efficiency.

WO2026020391A1PCT designated stage Publication Date: 2026-01-29SHENZHEN HUADA GENE INST +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/107379
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

In existing technologies, high-throughput peptide nanopore sequencing signals cannot be effectively screened, resulting in a large amount of interfering data, low data quality, and difficulty in handling secondary structure interference in complex protein primary structures.

Method used

By constructing a reference density matrix model and utilizing the characteristic parameters of known peptide sequencing signals, the distance fraction between the sequencing signal and the reference matrix is ​​calculated, the target peptide sequencing signal is selected, and interfering data is filtered out to obtain effective data of uniform quality.

Benefits of technology

This technology enables the selection of effective signals from sequencing signals of unknown samples, the removal of redundant and noisy signals, the improvement of data quality, the reduction of computational requirements for subsequent analysis, and the enhancement of sequencing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024107379_29012026_PF_FP_ABST
    Figure CN2024107379_29012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of biological sequencing or other related technical fields. Disclosed are a method and apparatus for processing a polypeptide nanopore sequencing signal. The method comprises: upon receiving N polypeptide sequencing signals, acquiring a distance score between each polypeptide sequencing signal and a reference density matrix model; and on the basis of the distance scores, screening the N polypeptide sequencing signals for a target polypeptide sequencing signal. The reference density matrix model contains a frequency-density reference matrix generated after signal filtering of known sequencing samples. The present invention solves the technical problem in the related art that unknown high-throughput polypeptide nanopore sequencing signals cannot be effectively screened, resulting in a large amount of interference data and low data quality.
Need to check novelty before this filing date? Find Prior Art

Description

Methods and apparatus for processing peptide nanopore sequencing signals Technical Field

[0001] This invention relates to the field of biological sequencing technology or other related fields, and more specifically, to a method and apparatus for processing polypeptide nanopore sequencing signals. Background Technology

[0002] Nanopore sequencing, as an emerging single-molecule sequencing technology, has been effectively applied in nucleic acid sequencing. Different bases generate different electrical signals as they pass through the sensing region of the nanopore, and the primary structure of nucleic acids can be sequenced by analyzing these electrical signals. However, relying solely on a single nanopore sensing device to acquire sequencing data results in extremely low throughput and detection efficiency.

[0003] Inspired by nucleic acid nanopore sequencing, the determination of protein primary structure can also employ a similar detection principle. By coupling peptides to nucleic acid chains, different amino acids generate different electrical signals as they pass through nanopores, thus enabling the resolution of the amino acid sequence. High-throughput sequencing is essential for efficiently obtaining sufficient sequencing data. However, array-based high-throughput sequencing devices face several challenges, such as crosstalk between electrical signals from different nanopore channels and reduced signal-to-noise ratios due to chip-related issues. Furthermore, heterogeneity between nanopores can lead to inconsistent sequencing data quality. Moreover, proteins, with their primary structure containing 20 amino acids and various post-translational modifications, are far more complex than nucleic acids, and secondary structures may exist during sequencing, interfering with the sequencing signal. Therefore, after obtaining a large amount of sequencing data, a method is needed to filter out interfering data caused by the above factors, obtaining high-quality, effective data for subsequent research and analysis. This ensures data quality while removing redundant noise, thereby reducing the computational demands of subsequent analysis.

[0004] Summary of the Invention

[0005] This invention provides a method and apparatus for processing peptide nanopore sequencing signals, which at least solves the technical problem in related technologies that the sequencing signals of unknown high-throughput peptide nanopore sequencing signals cannot be effectively screened, resulting in a large amount of interference data and low data quality.

[0006] According to one aspect of the present invention, a method for processing peptide nanopore sequencing signals is provided, comprising: after receiving N peptide sequencing signals, obtaining a distance score between each peptide sequencing signal and a reference density matrix model, wherein the reference density matrix model includes a frequency density reference matrix generated after signal filtering of known sequencing samples, and N is a positive integer greater than 1; and selecting a target peptide sequencing signal from the N peptide sequencing signals based on the distance score.

[0007] Optionally, constructing the reference density matrix model includes: obtaining a set of sequencing signals of the target peptide in a known sequencing sample; filtering each sequencing signal in the set of sequencing signals to retain sequencing signals whose signal characteristics conform to a normal distribution; downsampling all the retained sequencing signals; merging all the downsampled sequencing signals; and calculating the frequency values ​​of all units in the I*J matrix, where the frequency values ​​are the frequency values ​​of the sampling points in each unit relative to the total number of points in its column, and the J value is the same as the number of extended sampling points during the downsampling process; and normalizing the frequency values ​​of each column in the I*J matrix to generate the reference density matrix model.

[0008] Optionally, the step of obtaining the sequencing signal set of the target peptide in a known sequencing sample includes: obtaining a nucleic acid-peptide-nucleic acid complex of the known sequencing sample; passing the complex through a nanopore and performing nanopore sequencing to obtain the sequencing signal set of the target peptide.

[0009] Optionally, when filtering each sequencing signal in the sequencing signal set, the process includes: statistically analyzing the signal characteristic parameters of each of the filtered sequencing signals, wherein the signal characteristic parameters include at least one of the following: pore relaxation time, current hysteresis percentage, and standard deviation of current fluctuation.

[0010] Optionally, when constructing the reference density matrix model, the current blocking percentage is extended from 0-1 to 0-J.

[0011] Optionally, the step of obtaining the distance score between each peptide sequencing signal and the reference density matrix model includes: downsampling each peptide sequencing signal to J points; filtering each peptide sequencing signal through the reference density matrix model; calculating the frequency values ​​of the J units in the reference density matrix model traversed by the peptide sequencing signal; and performing logarithmic and negative sum processing on the frequency values ​​to obtain the distance score between each peptide sequencing signal and the reference density matrix model.

[0012] Optionally, when selecting the target peptide sequencing signal from the N peptide sequencing signals, a preset score threshold is used to filter the N peptide sequencing signals. The preset score threshold is the filtering threshold of the reference density matrix model. Obtaining the preset score threshold includes: constructing a density distribution map based on the distance scores of all sequencing signals in the target peptide sequencing signal set; determining the peak distance scores in the density distribution map; for all score points whose distance scores are less than the peak distance scores, constructing a target score symmetrical to the peak distance scores, merging all score points and the target score, and calculating the standard deviation; and using the sum of the peak distance scores and the standard deviation as the preset score threshold.

[0013] According to another aspect of the present invention, a processing apparatus for peptide nanopore sequencing signals is also provided, comprising: a distance score acquisition unit, configured to acquire a distance score between each peptide sequencing signal and a reference density matrix model after receiving N peptide sequencing signals, wherein the reference density matrix model includes a frequency density reference matrix generated after signal filtering of known sequencing samples, and N is a positive integer greater than 1; and a signal filtering unit, configured to filter out target peptide sequencing signals from the N peptide sequencing signals based on the distance scores.

[0014] Optionally, the peptide nanopore sequencing signal processing device, when constructing the reference density matrix model, includes: a signal acquisition unit, used to acquire a set of sequencing signals of the target peptide in a known sequencing sample, filter each sequencing signal in the set of sequencing signals, and retain sequencing signals whose signal characteristics conform to a normal distribution; a signal downsampling unit, used to downsample all the retained sequencing signals, merge all the downsampled sequencing signals, and calculate the frequency values ​​of all units in the I*J matrix, wherein the frequency values ​​are the frequency values ​​of the sampling points in each unit relative to the total number of points in its column, and the J value is the same as the number of extended sampling points during the downsampling process; and a normalization unit, used to normalize the frequency values ​​of each column in the I*J matrix to generate the reference density matrix model.

[0015] Optionally, the signal acquisition unit includes: a complex acquisition module for acquiring nucleic acid-peptide-nucleic acid complexes of known sequencing samples; and a sequencing module for passing the complexes through a nanopore for nanopore sequencing to obtain a set of sequencing signals for the target peptide.

[0016] Optionally, the peptide nanopore sequencing signal processing device further includes: a feature parameter statistics unit, used to count the signal feature parameters of each sequence signal after filtering when filtering each sequence signal in the sequence signal set, wherein the signal feature parameters include at least one of the following: pore relaxation time, current hysteresis percentage, and standard deviation of current fluctuation.

[0017] Optionally, the processing device for the peptide nanopore sequencing signal further includes an expansion unit for expanding the current hindrance percentage from 0-1 to 0-J when constructing the reference density matrix model.

[0018] Optionally, the distance score acquisition unit includes: a downsampling module, used to downsample each peptide sequencing signal to J points; and a distance score acquisition module, used to filter each peptide sequencing signal through the reference density matrix model, calculate the frequency values ​​of the J units in the reference density matrix model traversed by the peptide sequencing signal, and perform logarithmic and negative sum processing on the frequency values ​​to obtain the distance score between each peptide sequencing signal and the reference density matrix model.

[0019] Optionally, when selecting the target peptide sequencing signal from the N peptide sequencing signals, a preset score threshold is used to filter the N peptide sequencing signals. The preset score threshold is the filtering threshold of the reference density matrix model. When obtaining the preset score threshold, the peptide nanopore sequencing signal processing device includes: a density distribution map construction unit, used to construct a density distribution map based on the distance scores of all sequencing signals in the target peptide sequencing signal set; a peak distance score determination unit, used to determine the peak distance scores in the density distribution map; a standard deviation calculation unit, used to construct a target score symmetrical to the peak distance score for all score points whose distance scores are less than the peak distance scores, and to merge all score points and the target score to calculate the standard deviation; and a score threshold determination unit, used to use the sum of the peak distance scores and the standard deviation as the preset score threshold.

[0020] According to another aspect of the present invention, an electronic device is also provided, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the peptide nanopore sequencing signal processing method of any one of the above.

[0021] According to another aspect of the present invention, a computer program product is also provided, comprising a computer program that, when executed by a processor, implements the steps of the polypeptide nanopore sequencing signal processing method described in any one of the above embodiments.

[0022] In this disclosure, after receiving N peptide sequencing signals, the distance score between each peptide sequencing signal and the reference density matrix model is obtained. Based on the distance score, the target peptide sequencing signal is selected from the N peptide sequencing signals. The reference density matrix model includes a frequency density reference matrix generated after filtering known sequencing samples.

[0023] In this disclosure, when new sequencing signals are obtained, the distance score between each peptide sequencing signal and the reference density matrix model can be calculated. The peptide sequencing signals can be filtered using a known score filtering threshold. This allows for the selection of effective sequencing signals from unknown sample sequencing signals, filtering out interference data caused by various factors, obtaining effective data of uniform quality, and removing a large amount of redundant noise. This solves the technical problem in related technologies where it is impossible to effectively filter sequencing signals for unknown high-throughput peptide nanopore sequencing signals, resulting in a large amount of interference data and low data quality. Attached Figure Description

[0024] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0025] Figure 1 is a flowchart of an optional method for processing peptide nanopore sequencing signals according to Embodiment 1 of the present invention;

[0026] Figure 2 is a schematic diagram of the complete sequencing signal of an optional single "nucleic acid-peptide-nucleic acid" complex passing through a nanopore according to Embodiment 1 of the present invention;

[0027] Figure 3 is a schematic diagram of the sequencing signal of an optional single polypeptide according to Embodiment 1 of the present invention;

[0028] Figure 4 is a statistical graph of the characteristics of various parameters of an optional polypeptide sequencing signal according to Embodiment 1 of the present invention;

[0029] Figure 5 is a schematic diagram of the density matrix of the remaining polypeptide sequencing signal after parameter feature filtering according to an optional embodiment of the present invention.

[0030] Figure 6 is a schematic diagram of an optional reference density matrix model according to Embodiment 1 of the present invention;

[0031] Figure 7 is an optional distance fraction density distribution diagram according to Embodiment 1 of the present invention;

[0032] Figure 8 is a schematic diagram of the filtering threshold according to Embodiment 1 of the present invention;

[0033] Figure 9 is a superimposed diagram of all unfiltered polypeptide signals according to Embodiment 1 of the present invention;

[0034] Figure 10 is a superimposed diagram of the remaining polypeptide signals after the polypeptide signals are filtered by the reference density matrix model DM Ref. in Embodiment 1 of the present invention;

[0035] Figure 11 is a flowchart of an optional sequencing signal filtering based on a reference density matrix according to an embodiment of the present invention;

[0036] Figure 12 is a flowchart of an optional method for filtering new sequencing signals based on a reference density matrix model according to an embodiment of the present invention;

[0037] Figure 13 is a schematic diagram of an optional peptide nanopore sequencing signal processing device according to Embodiment 2 of the present invention;

[0038] Figure 14 is a hardware structure block diagram of an electronic device (or mobile device) for processing peptide nanopore sequencing signals according to an embodiment of the present invention. Detailed Implementation

[0039] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0040] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0041] To facilitate understanding of the present invention by those skilled in the art, some terms or nouns involved in the various embodiments of the present invention are explained below:

[0042] A polypeptide is a molecule composed of two or more amino acid residues linked together by peptide bonds.

[0043] The perforation time (t) in peptide sequencing signals refers to the time it takes for a single strand of DNA or RNA to pass through a nanopore, usually measured in milliseconds (ms).

[0044] Current hindrance percentage (I / I0) in peptide sequencing signals: refers to the percentage change in current when passing through a nanopore relative to the baseline current when not passing through a nanopore, used to measure the degree of current hindrance when molecules pass through a nanopore.

[0045] Standard deviation (STD) of current fluctuations in peptide sequencing signals: refers to the standard deviation of the fluctuation of current signals, used to measure the stability and noise level of current signals.

[0046] It should be noted that the method and apparatus for processing peptide nanopore sequencing signals disclosed herein can be used in the field of biological sequencing technology to achieve peptide nanopore sequencing signal processing, and can also be used in any field other than the field of biological sequencing technology to achieve peptide nanopore sequencing signal processing. The application field of the method and apparatus for processing peptide nanopore sequencing signals disclosed herein is not limited.

[0047] It should be noted that the information collected in this public disclosure (including but not limited to user device information such as collected peptide sequencing signals and user personal information) and data (including but not limited to data used for analysis, stored data, and displayed data) are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with the relevant laws, regulations, and standards of the relevant regions, and necessary confidentiality measures have been taken. This does not violate public order and good morals, and corresponding operation entry points are provided for users to choose to authorize or refuse. For example, this system has interfaces with relevant users or institutions. Before obtaining relevant information, a request to obtain the information needs to be sent to the aforementioned user or institution through the interface, and the relevant information is obtained only after receiving consent from the aforementioned user or institution.

[0048] It should be noted that in this disclosure, customer information is collected and analyzed, and users are provided with corresponding operation entry points to choose whether to agree to or reject the automated decision results; if the user chooses to reject, the process will proceed to the expert decision-making process.

[0049] The peptide sequencing signals mentioned in the following embodiments of the present invention can refer to the signals obtained after peptide rate-controlled sequencing through nanopores. High-throughput nanopore sequencing equipment is used, and the number of nanopore sequencing channels is large (for example, a nanopore array formed by 3000 nanopores, which can effectively improve the sequencing data throughput and efficiently obtain a large amount of sequencing data).

[0050] The basic principles for obtaining sequencing signals include, but are not limited to: breaking the nucleic acid to be tested (such as DNA) into segments, constructing single-stranded or double-stranded loops using specific methods, and then using a nanopore sequencing device. Under the influence of an electric field, the nucleic acid template to be tested will pass through the nanopore in a single-stranded mode and generate electrical signals of different bases. By judging the differences in electrical signals caused by different bases, different sequence template nucleic acids can be identified.

[0051] The nanopore sequencing device mentioned in this invention can be a device based on biological nanopores or solid nanopores. Biological nanopores can be protein nanopores, and solid nanopores can be silicon-based material nanopores. The narrowest part of the nanopore satisfies the property of allowing single-stranded nucleic acids to pass through. The support layer supporting the nanopore can be a phospholipid bilayer (phospholipid membrane), a polymer membrane, or other surface-modified (e.g., biotin-modified) support.

[0052] In this invention, the following drawbacks were found in the prior art regarding the processing of peptide sequencing signals after achieving high-throughput peptide sequencing:

[0053] 1. There is no effective screening scheme for high-throughput sequencing signals of peptides, and no relevant analytical method for the correspondence between peptides and signal features.

[0054] 2. Manual judgment is difficult to implement in high-throughput sequencing, is inefficient, and the judgment criteria cannot be quantified and lack universality.

[0055] 3. In rate-controlled sequencing of peptides, the sequencing signal contains information on a time scale. If the sequencing data is clustered and filtered only based on peptide pore time, current retardation percentage, and standard deviation of current fluctuation, only simple filtering of the sequencing signal can be performed. For unknown samples, it is impossible to effectively predict the signal.

[0056] The following embodiments of the present invention can be applied to various systems / applications / devices for processing peptide nanopore sequencing signals. To address the aforementioned drawbacks, the present invention utilizes multiple basic sequencing characteristics of known peptide sequencing signals (e.g., pore relaxation time (defined as t in this embodiment), current hysteresis percentage (defined as I / I0 in this embodiment), and standard deviation of current fluctuation (defined as STD in this embodiment)). Sequencing signals with high aggregation levels are selected as reference signals, and a dataset named the reference density matrix (DM) is constructed. This dataset serves as a reference feature for the peptide, i.e., a reference density matrix model (defined as DM Ref. in this embodiment), for signal alignment. A similarity filtering threshold is obtained based on known peptide sequencing signals. When a new sequencing signal is obtained, the similarity between the new sequencing signal and the reference density matrix model (DM Ref.) is calculated, and the similarity to the DM Ref. is obtained from the new sequencing signal using the known similarity filtering threshold. The most similar sequencing signal can be selected from the sequencing signals of unknown samples to identify the sequencing signal with the highest similarity to the target peptide. This filters out interference data caused by various factors (such as crosstalk between electrical signals of different nanopore channels, reduced signal-to-noise ratio of some data due to chip equipment problems, the presence of 20 amino acids in the primary structure of the protein and various post-translational modifications, and the possible existence of secondary structures during sequencing). This results in obtaining effective data of uniform quality. On the one hand, it ensures data quality, and on the other hand, it removes a large number of redundant noise signals, thereby reducing the computational requirements for subsequent analysis. For example, it can improve the efficiency of learning and classification by using artificial intelligence systems.

[0057] The present invention will now be described in detail with reference to various embodiments.

[0058] Example 1

[0059] According to an embodiment of the present invention, an embodiment of a method for processing peptide nanopore sequencing signals is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0060] Figure 1 is a flowchart of an optional method for processing peptide nanopore sequencing signals according to an embodiment of the present invention. As shown in Figure 1, the method includes the following steps:

[0061] Step S101: After receiving N peptide sequencing signals, obtain the distance score between each peptide sequencing signal and the reference density matrix model, wherein the reference density matrix model contains a frequency density reference matrix generated after filtering the known sequencing samples, and N is a positive integer greater than 1.

[0062] Step S102: Based on the distance score, select the target peptide sequencing signal from the N peptide sequencing signals.

[0063] Through the above steps, after receiving N peptide sequencing signals, the distance score between each peptide sequencing signal and the reference density matrix model can be obtained. Based on the distance score, the target peptide sequencing signal can be selected from the N peptide sequencing signals. The reference density matrix model includes a frequency density reference matrix generated after filtering signals from known sequencing samples. In this embodiment, when a new sequencing signal is obtained, the distance score between each peptide sequencing signal and the reference density matrix model can be calculated. By filtering peptide sequencing signals using a known score filtering threshold, effective sequencing signals can be selected from sequencing signals of unknown samples. This filters out interference data caused by various factors, obtains high-quality, uniform effective data, and removes a large amount of redundant noise. This solves the technical problem in related technologies where effective filtering of unknown high-throughput peptide nanopore sequencing signals is impossible, resulting in a large amount of interference data and low data quality.

[0064] The following embodiments of the present invention can be applied to a system / tool ​​for processing peptide nanopore sequencing signals. This system / tool ​​can call a high-throughput nanopore sequencing device to perform sequencing, obtain N peptide sequencing signals, and can screen out effective sequencing signals from the sequencing signals of unknown samples. It can filter out interference data caused by various factors (e.g., crosstalk between electrical signals of different nanopore channels, chip equipment problems leading to a decrease in the signal-to-noise ratio of some data, the presence of 20 amino acids in the primary structure of proteins and various post-translational modifications, and the possible existence of secondary structures during sequencing, etc.), and obtain effective data of uniform quality.

[0065] The embodiments of the present invention will now be described in detail with reference to the steps described above.

[0066] It should be noted that, in this embodiment, a reference density matrix model needs to be pre-constructed before screening effective peptide sequencing signals. Optionally, the construction of the reference density matrix model includes: obtaining a set of sequencing signals of the target peptide in a known sequencing sample; filtering each sequencing signal in the set and retaining sequencing signals whose signal characteristics conform to a normal distribution; downsampling all retained sequencing signals; merging all downsampled sequencing signals; calculating the frequency values ​​of all units in the I*J matrix, where the frequency value is the frequency of the sampling point in each unit relative to the total number of points in its column, and the J value is the same as the number of expanded sampling points during the downsampling process; normalizing the frequency values ​​of each column in the I*J matrix to generate the reference density matrix model.

[0067] The process of obtaining the sequencing signal set of the target peptide in a known sequencing sample includes: obtaining the nucleic acid-peptide-nucleic acid complex of the known sequencing sample; passing the complex through a nanopore and performing nanopore sequencing to obtain the sequencing signal set of the target peptide.

[0068] In one optional embodiment, the DNA can be arbitrarily selected, empty deoxynucleotides without bases are represented by X, and all polypeptide sequences are represented as N-terminus → C-terminus, wherein the amino group at the N-terminus of the polypeptide sequence is modified with azide, and the amino group of the lysine side chain at the C-terminus is modified with azide, denoted as LYS(N3).

[0069] For example, the sequence information is as follows:

[0070] DNA1 (SEQ ID NO:1): phosphorylated (i.e., pho) / GCTTCTCGTGXXTTTTTTTTCTCTC / dibenzocyclooctylene (i.e., DBCO), where X is an empty deoxynucleotide without a base.

[0071] DNA2 (SEQ ID NO:2): DBCO / CCCXXTTTTTTTTTTGCTGTCTTCTGTCGTCGTTTC, where X is an empty deoxynucleotide without a base.

[0072] DNA3 (SEQ ID NO:3):

[0073] AAACGACGACAGAAGACAGCAAAAAAAAAATTGGGXGAGAGAAAAAAAAAACACGAGAAGCA, where X is an empty deoxyribonucleotide without a base.

[0074] Obtain peptide powder, Peptide (SEQ ID NO:4): N3-QMNDE{Lys(N3)}.

[0075] Then, a peptide library is constructed and nanopore sequencing data is obtained. The specific procedure includes:

[0076] (1) Dissolve the powders of DNA1, DNA2 and DNA3 in 1×PBS (optional) to prepare a 100μM stock solution; mix them evenly in a 1:1:1 ratio and anneal them using a PCR instrument.

[0077] (2) Dissolve the peptide powder in 1×PBS (choose your own) to make a 100μM stock solution, and add it to the annealing solution in (1) at a ratio of DNA:peptide = 1:1.

[0078] (3) Incubate overnight at room temperature to allow the azide on the polypeptide to undergo a click chemical reaction with the DBCO on the nucleic acid, thereby obtaining a nucleic acid-polypeptide-nucleic acid complex.

[0079] (4) Connect the complex obtained in (3) to the adapter and perform nanopore sequencing to obtain sequencing signals.

[0080] Results analysis: Figure 2 is a schematic diagram of the complete sequencing signal of an optional single "nucleic acid-peptide-nucleic acid" complex passing through a nanopore according to an embodiment of the present invention. As shown in Figure 2, the signal between the two obvious peaks ("*") is the sequencing signal of the peptide, and the signals before and after the peaks are the sequencing signals of the nucleic acid.

[0081] Figure 3 is a schematic diagram of the sequencing signal of an optional single peptide according to an embodiment of the present invention. As shown in Figure 3, by using time as the X-axis and current as the Y-axis to determine the time for peptide pore passage, the sequencing signal set of the target peptide can be obtained, including the current fluctuation during peptide pore passage.

[0082] After obtaining a set of sequencing signals for the target peptide from known sequencing samples (e.g., more than 10,000 known target peptide sequencing signals), each sequencing signal in the set is filtered, and then multiple signal characteristic parameters are statistically analyzed for each signal. These signal characteristic parameters include at least one of the following: pore relaxation time, current hysteresis percentage, and standard deviation of current fluctuation. After analyzing all signal characteristic parameters, sequencing signals that meet the sequencing standard range for all signal characteristic parameters can be retained. Here, the sequencing standard range refers to the range of a normal distribution determined by the mean and standard deviation of the signal characteristic parameters.

[0083] Here, the sequencing standard range of signal characteristic parameters can be determined in a variety of ways, for example, by calculating the mean and standard deviation of each signal characteristic parameter, and determining the sequencing standard range of each signal characteristic parameter based on the mean and standard deviation. For example, calculate the mean and standard deviation of the three parameters, and finally retain the sequencing data that simultaneously meet the range of "mean ± standard deviation" of the three (conforming to a normal distribution).

[0084] After obtaining the preserved sequencing signals, all the preserved sequencing signals can be downsampled, and all the downsampled sequencing signals can be merged. The frequency values ​​of all units in the I*J matrix are calculated. The frequency value refers to the frequency value of the sampling point in each unit relative to the total number of points in its column. Here, the J value is the same as the number of expanded sampling points during the downsampling process. Finally, the frequency values ​​of each column in the I*J matrix are normalized to generate a reference density matrix model.

[0085] In this embodiment, the number of downsampled data points can be selected by the user. For example, if J is selected as 1000, the sequencing signal will be downsampled to 1000 points. All the obtained sequencing signals will be merged, and the frequency of the data points in each column of all cells in the I*J matrix (e.g., a 1000×1000 matrix) will be calculated. Finally, the sum of each column of the matrix will be normalized to 1 to obtain the final reference density matrix model.

[0086] It should be noted that after downsampling all the retained sequencing signals, the process also includes expanding the current hysteresis percentage from 0-1 to 0-J.

[0087] For example, after downsampling all the retained sequencing signals, the current hysteresis percentage I / I0 value expanded from 0-1 to 0-1000.

[0088] As shown in Figure 3, sequencing signals of peptides can be obtained through nanopore sequencing. The horizontal axis represents the relaxation time of the peptide through the pore (defined as t in this embodiment), and the vertical axis represents the current fluctuation during peptide passage. The average current of all collected points is calculated and then divided by the opening current (defined as I0 in this embodiment) to obtain the current hindrance percentage during peptide passage (defined as I / I0 in this embodiment). The standard deviation of the current of all collected points is calculated, and the resulting value is the standard deviation of the current fluctuation during peptide passage (defined as STD in this embodiment).

[0089] Sequencing data from over 10,000 pattern peptides were collected, and the three parameters mentioned above were statistically analyzed, as shown in Figure 4. The left side represents the parameter variation of pore relaxation time (horizontal axis: time t, vertical axis: frequency (%)), the middle side represents the current hysteresis percentage I / I0, and the right side represents the standard deviation of current fluctuation during peptide pore passage. Small peaks can be seen in addition to the main peak, indicating that some data have sequencing quality deviating from the center point of the dataset. After calculating the mean and standard deviation of the three parameters, sequencing data that simultaneously meet the range of "mean ± standard deviation" were retained. After preliminary filtering, the statistical results of the remaining peptides for the three parameters are shown in Figure 5.

[0090] After downsampling all peptide sequencing data to 1000 points, the frequency of each cell in the 1000×1000 cell array is calculated to obtain the DM Ref. as shown in Figure 6. (Note: This is illustrated using a 10×10 matrix. The sum of the frequency values ​​in each column of the matrix shown in Figure 6 is normalized, for example, it can be normalized to 1).

[0091] In addition, in order to filter invalid data in peptide sequencing signals, this embodiment requires the use of a filtering threshold. Optionally, the preset score threshold is the filtering threshold of the reference density matrix model. When obtaining the preset score threshold, the following steps are taken: constructing a density distribution map based on the distance scores of all sequencing signals in the target peptide sequencing signal set; determining the peak distance scores in the density distribution map; constructing a target score symmetrical to the peak distance scores for all score points whose distance scores are less than the peak distance scores, merging all score points and the target scores, and calculating the standard deviation; and using the sum of the peak distance scores and the standard deviation as the preset score threshold.

[0092] Here, when determining the score threshold, the distance scores of the sequencing signals can be obtained first. When filtering each sequencing signal using the reference density matrix model, each sequencing data can be compared with the reference density matrix model, and the frequency values ​​of the J traversed units can be calculated. The logarithms of the frequency values ​​and the negative summation are then performed to obtain the distance scores of all sequencing signals. Based on the distance scores of all sequencing signals, a density distribution map is calculated, the peak value of the density distribution map is obtained, and the distance score corresponding to the peak value is calculated to obtain the peak distance score p. For all score points s whose distance scores are less than the peak distance score p, a score sp symmetric to p is constructed. All score points s and scores sp are merged, and the standard deviation σ is calculated. The sum of the peak distance score p and the standard deviation σ is then calculated to obtain the preset score threshold.

[0093] When obtaining the score threshold, the sequencing signal data of known peptides needs to be downsampled first, for example, downsampled to 1000 points, and I / I0 expanded to 0-1000. After downsampling, the distance score is calculated. Specifically, the calculation of the distance score includes: comparing each sequencing data with the obtained reference density matrix model DM Ref, calculating the frequency values ​​of the J traversed units, taking the logarithm of the frequency values ​​and performing negative sum processing to obtain the distance score of all sequencing signals. For example, calculating the logarithm of the frequency values ​​of the 1000 traversed units and calculating their negative sum. Then, based on the obtained distance scores of all peptide signals, a density distribution map is obtained. Then, the distance score p corresponding to the peak of the obtained density distribution map is calculated. For all score points s with a distance score less than p, a score sp symmetric to p is constructed, i.e., sp = p + (ps) = 2p - s. All scores s and sp are combined, and the standard deviation σ is calculated. The preset score threshold is determined as p + σ.

[0094] By calculating the distance between each peptide and the reference density matrix model DM Ref, peptide signals that are more similar receive lower scores. The resulting density distribution map of the distance scores is shown in Figure 7 (where the horizontal axis represents the distance score and the vertical axis represents density). This density distribution map can be used to determine the filtering threshold. As shown in Figure 8, the filtering threshold for the peptides in this embodiment is 5500.

[0095] In Figure 7, ① is a schematic diagram comparing the peptide sequencing signal with the reference density matrix model (DM Ref.), and ② is a schematic diagram of the distance fraction density distribution after calculating the similarity between all peptides and DM Ref. P is the peak value of the fraction density distribution, and p+σ is the filtering threshold for the peptide.

[0096] After completing the construction of the reference density matrix model and determining the filtering score threshold, the newly acquired peptide sequencing signals can be screened using the reference density matrix model and the score threshold, retaining peptide sequencing signals corresponding to distance scores greater than or equal to the preset score threshold.

[0097] Step S101: After receiving N peptide sequencing signals, obtain the distance score between each peptide sequencing signal and the reference density matrix model, wherein the reference density matrix model contains a frequency density reference matrix generated after filtering the known sequencing samples, and N is a positive integer greater than 1.

[0098] Here, when obtaining the distance score between each peptide sequencing signal and the reference density matrix model, it may include: downsampling each peptide sequencing signal to J points, obtaining the frequency value of each peptide sequencing signal traversing J units in the reference density matrix model, and then performing logarithmic and negative sum processing on the frequency values ​​to obtain the distance score between each peptide sequencing signal and the reference density matrix model. Here, logarithmic processing refers to calculating the logarithmic value of the frequency value, and negative sum processing refers to obtaining the opposite value of the logarithmic value and accumulating all the opposite values.

[0099] For example, in this embodiment, after obtaining new peptide sequencing signal data using nanopore sequencing, the sequencing signals can be downsampled to 1000 points, I / I0 can be expanded to 0-1000, and then each sequencing signal can be compared with the reference density matrix model DM Ref. of the peptide. The logarithm of the frequency values ​​of the 1000 traversed units can be calculated, and the negative sum can be calculated to obtain the distance score.

[0100] It should be noted that, in addition to using the above-mentioned technical features to obtain the distance score, this embodiment can also use other methods, that is, obtain the distance score through other similarity calculation methods, such as cosine similarity, Euclidean distance, etc., to calculate the distance score between the newly obtained sequencing signal and the reference density matrix model.

[0101] Step S102: Based on the distance score, select the target peptide sequencing signal from the N peptide sequencing signals.

[0102] It should be noted that, in this embodiment, during the screening process, for distance scores less than or equal to a preset score threshold, the peptide sequencing signal corresponding to that distance score can be retained, and the retained peptide sequencing signal can be used as the target peptide sequencing signal.

[0103] After determining the distance fraction of each sequencing signal, invalid data is removed based on the previously determined fraction threshold. Specifically, the numerical comparison between the fraction threshold and the distance fraction can be performed, and peptide sequencing signals corresponding to distance fractions less than the preset fraction threshold can be retained. For example, data with a distance fraction less than 5500 can be deleted, and the remaining sequencing signals can be retained.

[0104] Figure 9 shows a stacked image of all peptides, revealing obvious erroneous data, such as the portion circled in the dashed box. After filtering using the reference density matrix model DM Ref., the remaining sequencing signals are stacked together, resulting in the signal stacked image shown in Figure 10. The data is significantly more concentrated and cleaner, indicating that this filtering method can be used for peptide signal filtering, improving data homogeneity for subsequent analyses.

[0105] The invention will now be described in conjunction with another alternative embodiment.

[0106] Figure 11 is a flowchart of an optional sequencing signal filtering method based on a reference density matrix according to an embodiment of the present invention. As shown in Figure 11, the sequencing signal processing method includes:

[0107] Step 1: Obtain nanopore sequencing signal data of the known target peptide.

[0108] Step 2, basic filtering: Sequencing signals whose via relaxation time (t), current hysteresis percentage (I / I0), and standard deviation of current fluctuation are all within the mean ± STD range.

[0109] Step 3: Obtain the reference density matrix model (DM Ref.) of the known target peptide.

[0110] Step 4: Calculate the distance score to obtain the filtering threshold p+σ.

[0111] Step 5: Use the density matrix model DM.Ref. and the filtering threshold p+σ as the filtering criteria for peptide sequencing signals.

[0112] In this embodiment, the basic sequencing features of known peptide sequencing signals (pore relaxation time (t), current hysteresis percentage (I / I0), and standard deviation of current fluctuation (STD)) are used to select sequencing signals with high aggregation from the above sequencing signals as reference signals. A dataset called density matrix (DM) is constructed as a reference feature of this peptide, i.e., the reference density matrix (DM Ref.) for signal alignment, and a similarity filtering threshold is obtained based on the known peptide sequencing signals.

[0113] Figure 12 is a flowchart of an optional new sequencing signal filtering based on a reference density matrix model according to an embodiment of the present invention. As shown in Figure 12, the filtering of new sequencing signals includes: when a new sequencing signal is obtained, calculating the distance score between the new sequencing signal and the reference density matrix (DM Ref.), and then determining whether it is less than a preset score threshold p+σ. If not, discarding the newly obtained sequencing signal data; if so, retaining the sequencing signal data.

[0114] It should be noted that in this embodiment, sequencing signals with the highest similarity to the target peptide can be screened from the sequencing signals of unknown samples for subsequent algorithm analysis, such as using the learning and classification of artificial intelligence systems to improve the efficiency of learning and classification.

[0115] The following provides a complete process for constructing a reference density matrix model from known target peptides and determining the filtering fraction threshold, and then using the reference density matrix model and the fraction threshold to screen newly obtained sequencing signals.

[0116] (1) Obtain the sequencing signal of a known target peptide by nanopore sequencing technology.

[0117] (2) Perform preliminary filtering on known target peptide sequencing signals. ① Statistically analyze the distribution of characteristic parameters (pore time (t), current hysteresis percentage (I / I0), and standard deviation of current fluctuation (STD)) of all known target peptide sequencing reads. These parameters are usually approximately normal or log-normal. ② Retain sequencing signals that fall within the range of the mean ± standard deviation of the three parameters.

[0118] (3) Normalize the horizontal axis (time) and vertical axis (current) of the sequencing signals obtained above to 1000, and output a 1000×1000 density matrix (DM). Specifically, perform the following operations on the sequencing signals obtained after preliminary filtering: ① Downsample each sequencing signal read length to 1000 points; ② Expand the I / I0 value from 0-1 to 0-1000; ③ Superimpose all sequencing data; ④ Calculate the frequency of data points in each column of the 1000×1000 matrix; ⑤ Normalize the sum of the frequencies in each column of the matrix to 1; ⑥ Generate a reference density matrix model (i.e., DM Ref.).

[0119] (4) Determine the filtering score threshold. ① Downsample all known target peptide sequencing signals obtained in (1) to 1000 points, and expand I / I0 to 0-1000; ② Obtain the distance score: Calculate the frequency values ​​of each sequencing signal in the 1000 units traversed in DM Ref., take the logarithm of each value, and sum the negative values ​​as the distance score of the peptide signal (obviously, when the sequencing signal sampling point passes through the unit with a higher frequency in DM Ref., the negative logarithm obtained is smaller, indicating that the sequencing signal is more similar to DM Ref., and vice versa); ③ Determine the filtering score threshold: Obtain the distance scores of all peptides through ②, and establish a distance score density distribution map; calculate the distance score p corresponding to the peak in the distribution map, and construct a score sp symmetric to p for all score points s with a distance score less than p, i.e., sp = p + (ps) = 2p - s; combine all scores s and sp, and calculate the standard deviation σ; determine the threshold as p + σ.

[0120] (5) Filtering of newly acquired peptide sequencing signals using DM Ref. The newly acquired peptide sequencing signals are filtered by calculating the similarity distribution between the newly acquired peptide sequencing signals and the aforementioned DM Ref. Specifically, ① all newly acquired peptide sequencing signals are downsampled to 1000 points, and I / I0 is expanded to 0-1000; ② distance score is obtained: the frequency values ​​of each sequencing signal sampling point in the DM Ref. are calculated, the logarithm is taken first, and then the negative is summed to obtain the distance score of the peptide signal; ③ if the distance score is less than the threshold (p+σ), the sequencing signal data is retained; otherwise, the signal data is discarded.

[0121] The following is a detailed description with reference to another embodiment.

[0122] Example 2

[0123] The polypeptide nanopore sequencing signal processing device provided in this embodiment includes multiple implementation units, each of which corresponds to a specific implementation step in Embodiment 1 above.

[0124] Figure 13 is a schematic diagram of an optional peptide nanopore sequencing signal processing device according to Embodiment 2 of the present invention. As shown in Figure 13, the peptide nanopore sequencing signal processing device may include: a distance fraction acquisition unit 1301 and a signal filtering unit 1302.

[0125] The distance score acquisition unit 1301 is used to acquire the distance score between each peptide sequencing signal and the reference density matrix model after receiving N peptide sequencing signals. The reference density matrix model contains a frequency density reference matrix generated after filtering the signals of known sequencing samples, and N is a positive integer greater than 1.

[0126] The signal screening unit 1302 is used to screen out target peptide sequencing signals from N peptide sequencing signals based on distance scores.

[0127] The aforementioned peptide nanopore sequencing signal processing device includes a distance score acquisition unit, which, upon receiving N peptide sequencing signals, acquires the distance score between each peptide sequencing signal and a reference density matrix model. The reference density matrix model contains a frequency density reference matrix generated after filtering known sequencing samples, where N is a positive integer greater than 1. A signal filtering unit is used to filter target peptide sequencing signals from the N peptide sequencing signals based on the distance scores. In this embodiment, when a new sequencing signal is obtained, the distance score between each peptide sequencing signal and the reference density matrix model can be calculated. By filtering peptide sequencing signals using a known score filtering threshold, effective sequencing signals can be filtered from sequencing signals of unknown samples. This filters out interference data caused by various factors, obtains uniform and effective data, and removes a large amount of redundant noise. This solves the technical problem in related technologies where effective filtering of unknown high-throughput peptide nanopore sequencing signals is impossible, resulting in a large amount of interference data and low data quality.

[0128] Optionally, the peptide nanopore sequencing signal processing device, when constructing the reference density matrix model, includes: a signal acquisition unit, used to acquire the sequencing signal set of the target peptide in a known sequencing sample, filter each sequencing signal in the sequencing signal set, and retain sequencing signals whose signal characteristics conform to a normal distribution; a signal downsampling unit, used to downsample all the retained sequencing signals, merge all the downsampled sequencing signals, and calculate the frequency values ​​of all units in the I*J matrix, where the frequency value is the frequency value of the sampling point in each unit relative to the total number of points in its column, and the J value is the same as the number of extended sampling points during the downsampling process; and a normalization unit, used to normalize the frequency values ​​of each column in the I*J matrix to generate the reference density matrix model.

[0129] Optionally, the signal acquisition unit includes: a complex acquisition module for acquiring nucleic acid-peptide-nucleic acid complexes of known sequencing samples; and a sequencing module for passing the complexes through nanopores for nanopore sequencing to obtain a set of sequencing signals for the target peptide.

[0130] Optionally, the peptide nanopore sequencing signal processing device further includes: a feature parameter statistics unit, used to statistically analyze the signal feature parameters of each sequence signal after filtering when filtering each sequence signal in the sequence signal set, wherein the signal feature parameters include at least one of the following: pore relaxation time, current hysteresis percentage, and standard deviation of current fluctuation.

[0131] Optionally, the processing device for peptide nanopore sequencing signals further includes an extension unit for extending the current hindrance percentage from 0-1 to 0-J when constructing a reference density matrix model.

[0132] Optionally, the distance score acquisition unit includes: a downsampling module for downsampling each peptide sequencing signal to J points; and a distance score acquisition module for filtering each peptide sequencing signal through a reference density matrix model, calculating the frequency values ​​of J units in the reference density matrix model traversed by the peptide sequencing signal, and performing logarithmic and negative sum processing on the frequency values ​​to obtain the distance score between each peptide sequencing signal and the reference density matrix model.

[0133] Optionally, when selecting the target peptide sequencing signal from N peptide sequencing signals, a preset score threshold is used to filter the N peptide sequencing signals. The preset score threshold is the filtering threshold of the reference density matrix model. The peptide nanopore sequencing signal processing device includes the following components when obtaining the preset score threshold: a density distribution map construction unit, used to construct a density distribution map based on the distance scores of all sequencing signals in the target peptide sequencing signal set; a peak distance score determination unit, used to determine the peak distance scores in the density distribution map; a standard deviation calculation unit, used to construct a target score symmetrical to the peak distance score for all score points whose distance scores are less than the peak distance scores, and to merge all score points and the target score to calculate the standard deviation; and a score threshold determination unit, used to use the sum of the peak distance scores and the standard deviation as the preset score threshold.

[0134] The aforementioned peptide nanopore sequencing signal processing device may further include a processor and a memory. The aforementioned distance fraction acquisition unit 1301, signal filtering unit 1302, etc., are all stored in the memory as program units, and the processor executes the aforementioned program units stored in the memory to realize the corresponding functions.

[0135] The aforementioned processor contains a kernel that retrieves the corresponding program units from memory. One or more kernels can be configured. By adjusting kernel parameters, when new sequencing signals are obtained, the distance score between each peptide sequencing signal and the reference density matrix model is calculated. Peptide sequencing signals are then filtered using a known score filtering threshold, enabling the selection of valid sequencing signals from unknown samples.

[0136] The aforementioned memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0137] According to another aspect of the present invention, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored computer program, wherein, when the computer program is running, it controls the device where the computer-readable storage medium is located to perform the peptide nanopore sequencing signal processing method of any one of the above embodiments.

[0138] According to another aspect of the present invention, an electronic device is also provided, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the peptide nanopore sequencing signal processing method of any one of the above embodiments.

[0139] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the peptide nanopore sequencing signal processing method in various embodiments of this application, including: after receiving N peptide sequencing signals, obtaining the distance score between each peptide sequencing signal and a reference density matrix model, wherein the reference density matrix model contains a frequency density reference matrix generated after signal filtering of known sequencing samples, and N is a positive integer greater than 1; and selecting the target peptide sequencing signal from the N peptide sequencing signals based on the distance score.

[0140] This application also provides a computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the polypeptide nanopore sequencing signal processing method described in various embodiments of this application.

[0141] Figure 14 is a hardware structure block diagram of an electronic device (or mobile device) for processing peptide nanopore sequencing signals according to an embodiment of the present invention. As shown in Figure 14, the electronic device may include one or more processors 1402 (shown as 1402a, 1402b, ..., 1402n in Figure 14) (processor 1402 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 1404 for storing data. In addition, it may include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a keyboard, a power supply, and / or a camera. Those skilled in the art will understand that the structure shown in Figure 14 is merely illustrative and does not limit the structure of the electronic device described above. For example, the electronic device may also include more or fewer components than shown in Figure 14, or have a different configuration than that shown in Figure 14.

[0142] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0143] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0144] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0145] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0146] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0147] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0148] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for processing polypeptide nanopore sequencing signals, comprising: obtaining a distance score of each of the polypeptide sequencing signals from a reference density matrix model after receiving N polypeptide sequencing signals, wherein the reference density matrix model comprises a frequency density reference matrix generated by filtering signals of known sequencing samples, and N is a positive integer greater than 1; and selecting a target polypeptide sequencing signal from the N polypeptide sequencing signals based on the distance score. In constructing the reference density matrix model, comprising: obtaining a set of sequencing signals of a target polypeptide in known sequencing samples, filtering each of the set of sequencing signals, and retaining a sequencing signal with a signal feature conforming to a normal distribution; down-sampling all the retained sequencing signals, merging all the down-sampled sequencing signals, and calculating frequency values of all cells in an I*J matrix, wherein the frequency value is a frequency value of a sampling point in each cell accounting for a total number of column points, and J is equal to a number of extended sampling points in the down-sampling process; and normalizing frequency values of each column in the I*J matrix to generate the reference density matrix model. In obtaining the set of sequencing signals of the target polypeptide in the known sequencing samples, comprising: obtaining a nucleic acid-polypeptide-nucleic acid complex of the known sequencing samples; and passing the complex through a nanopore to perform nanopore sequencing and obtain the set of sequencing signals of the target polypeptide.

2. The method of polypeptide nanopore sequencing signal processing according to claim 1, wherein, In filtering each of the set of sequencing signals, comprising: counting signal feature parameters of each of the set of sequencing signals, wherein the signal feature parameters comprise at least one of the following: a translocation relaxation time, a current blockage percentage, and a standard deviation of current fluctuation. In constructing the reference density matrix model, comprising: extending the current blockage percentage from 0-1 to 0-J. In obtaining the distance score of each of the polypeptide sequencing signals from the reference density matrix model, comprising: down-sampling each of the polypeptide sequencing signals to J points; filtering each of the polypeptide sequencing signals by the reference density matrix model, calculating frequency values of J cells in the reference density matrix model traversed by the polypeptide sequencing signal, and performing logarithmic processing and negative sum processing on the frequency values to obtain the distance score of each of the polypeptide sequencing signals from the reference density matrix model. In selecting the target polypeptide sequencing signal from the N polypeptide sequencing signals, a preset score threshold is used to select the N polypeptide sequencing signals, the preset score threshold is a filtering threshold of the reference density matrix model, and in obtaining the preset score threshold, comprising: constructing a density distribution map based on distance scores of all the set of sequencing signals of the target polypeptide; determining a peak distance score in the density distribution map; constructing a target score symmetric to the peak distance score for all score points with a distance score less than the peak distance score, merging all the score points and the target score, and calculating a standard deviation; and taking a sum of the peak distance score and the standard deviation as the preset score threshold.

3. The method of polypeptide nanopore sequencing signal processing according to claim 2, wherein, 8. A device for processing polypeptide nanopore sequencing signals, comprising: ​ ​ 4. The method of polypeptide nanopore sequencing signal processing according to claim 2, wherein, ​ 5. The method of polypeptide nanopore sequencing signal processing according to claim 2, wherein, ​ ​ 6. The method of polypeptide nanopore sequencing signal processing according to claim 1, wherein, ​ ​ ​ 7. The method of polypeptide nanopore sequencing signal processing according to claim 1, wherein, ​ ​ ​ ​ ​ ​ The distance score acquisition unit is configured to acquire a distance score of each of the N polypeptide sequencing signals with respect to a reference density matrix model after receiving the N polypeptide sequencing signals, wherein the reference density matrix model comprises a frequency density reference matrix generated after signal filtering of known sequencing samples, and N is a positive integer greater than 1. The signal screening unit is configured to screen a target polypeptide sequencing signal from the N polypeptide sequencing signals based on the distance scores.

9. An electronic device, comprising one or more processors and memory storing one or more programs, wherein: When the one or more programs are executed by the one or more processors, the one or more processors implement the polypeptide nanopore sequencing signal processing method of any one of claims 1 to 7.

10. A computer program product comprising a computer program which, when executed by a processor, implements the steps of the polypeptide nanopore sequencing signal processing method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method, apparatus, device, medium and program product for single cell sequencing

    CN114171117A

  • Electroencephalogram signal abnormity elimination method and device and computer readable storage medium

    CN115062661A

  • Nanopore sequencing original signal data compression method, device, equipment and medium

    CN115798605A

  • Design method for bar codes in nanopore sequencing based on TDFPS algorithm

    CN118116462A

  • System and method for direct subsequence searching and mapping in nanopore raw signal

    US20210350876A1