Polypeptide sequencing data screening method, neural network model training method, and related device
By extracting depth features from peptide sequencing data and estimating point cloud density using a neural network model, the problem of noisy data in nanopore sequencing technology is solved, enabling high-quality screening of peptide sequencing data.
Patent Information
- Application Number
- PCT/CN2024/105535
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2026-01-22
AI Technical Summary
Nanopore sequencing technology contains noisy data in peptide sequences, which affects the quality of peptide sequences. Furthermore, existing technologies cannot accurately extract the characteristics of peptide sequencing data, resulting in poor data quality.
A neural network model was used to extract depth features from peptide sequencing data, and point cloud density estimates were used to filter out noisy data, thereby improving the quality of peptide sequencing data.
By extracting deep features and estimating point cloud density using a neural network model, noise data in peptide sequencing data can be accurately filtered out, thus improving the quality of peptide sequencing data.
Smart Images

Figure CN2024105535_22012026_PF_FP_ABST
Abstract
Description
Methods and related equipment for peptide sequencing data screening and neural network model training Technical Field
[0001] This application relates to the field of biotechnology, and in particular to a method and related equipment for screening peptide sequencing data and training neural network models. Background Technology
[0002] Currently, nanopore sequencing is an advanced technique for determining protein sequences. It determines peptide sequences by analyzing the current signal of single molecules passing through nanopores. Peptides are linked to the phosphoribosyl backbone of amino acids to form peptide-nucleic acid complexes. Under the pull of motor proteins, these complexes stably pass through nanopores, and the amino acid sequence is determined by analyzing the signal. However, in determining peptide sequences using nanopore sequencing, the spatial resolution of nanopores is limited. The molecular diameter of amino acids is much smaller than the diameter of the nanopore, resulting in only a weak current blocking signal as the peptide passes through. Furthermore, because the pulling speed of motor proteins is not constant, noise interference exists in the electric field, leading to noisy data in the peptide sequences obtained by nanopore sequencing, thus affecting the quality of the peptide sequences.
[0003] Summary of the Invention
[0004] The main objective of this application is to propose a method and related equipment for screening peptide sequencing data and training neural network models, which aims to reduce noise data in peptide sequences and improve the quality of peptide sequences.
[0005] To achieve the above objectives, a first aspect of this application provides a method for screening peptide sequencing data, the method comprising:
[0006] Multiple raw peptide sequencing data were obtained; wherein, the raw peptide sequencing data were obtained through nanopore sequencing technology;
[0007] Based on a preset neural network model, feature extraction is performed on the multiple raw peptide sequencing data to obtain multiple deep features; wherein, the neural network model is a model trained based on preset sample data, and the preset sample data includes multiple sample peptide sequencing data and their corresponding classification labels;
[0008] Point cloud density estimation is performed on the multiple depth features to obtain the point cloud density estimate value corresponding to each depth feature;
[0009] The original polypeptide sequencing data were filtered based on the point cloud density estimate to obtain the filtering results.
[0010] In some embodiments, the step of estimating the point cloud density of the plurality of depth features to obtain a point cloud density estimate corresponding to each depth feature includes:
[0011] The multiple deep features are subjected to dimensionality reduction processing to obtain multiple low-dimensional features;
[0012] The distribution density of the multiple low-dimensional features is estimated to obtain the point cloud density estimate corresponding to each depth feature.
[0013] In some embodiments, the step of filtering the raw peptide sequencing data based on the point cloud density estimate to obtain the filtering results includes:
[0014] The raw peptide sequencing data is globally filtered based on a preset global threshold and the estimated point cloud density to obtain preliminary peptide sequencing data.
[0015] Based on the point cloud density estimate, the anomaly probability of the preliminary peptide sequencing data is estimated to obtain the anomaly probability estimate.
[0016] The preliminary peptide sequencing data are filtered based on a preset local threshold and the estimated anomaly probability to obtain the filtering results.
[0017] In some embodiments, before filtering the preliminary peptide sequencing data based on a preset local threshold and the estimated anomaly probability to obtain the filtering results, the method further includes:
[0018] Setting the global threshold and the local threshold specifically includes:
[0019] The global threshold is determined based on a preset first quantile and the estimated point cloud density.
[0020] The local threshold is determined based on a preset second quantile and the estimated anomaly probability; wherein the first quantile is less than the second quantile.
[0021] In some embodiments, acquiring multiple raw peptide sequencing data includes:
[0022] Acquire multiple initial peptide sequencing data;
[0023] The length of the multiple initial peptide sequencing data is adjusted according to a preset length to obtain multiple candidate peptide sequencing data.
[0024] The sequencing data of the multiple candidate peptides were normalized to obtain the multiple raw peptide sequencing data.
[0025] In some embodiments, before performing feature extraction on the plurality of raw peptide sequencing data based on a preset neural network model to obtain multiple deep features, the method further includes:
[0026] Training the neural network model specifically includes:
[0027] The multiple sample peptide sequencing data are input into the neural network model for classification processing to obtain the classification prediction result of each sample peptide sequencing data.
[0028] The parameters of the neural network model are adjusted based on the classification prediction results and the classification labels corresponding to the sample peptide sequencing data.
[0029] To achieve the above objectives, a second aspect of this application proposes a neural network model training method, the method comprising:
[0030] Obtain the screening results obtained by the peptide sequencing data screening method described in the first aspect;
[0031] Based on the screening results, multiple target peptide sequencing data were extracted from the multiple raw peptide sequencing data;
[0032] The neural network model is trained based on the sequencing data of the multiple target peptides and the classification labels corresponding to the sequencing data of the target peptides.
[0033] In some embodiments, after extracting multiple target peptide sequencing data from the multiple raw peptide sequencing data based on the screening results, the method further includes:
[0034] According to a preset allocation ratio, the multiple target peptide sequencing data and the corresponding classification labels of the target peptide sequencing data are divided into training datasets and test datasets.
[0035] The neural network model is trained based on the training dataset;
[0036] Based on the trained neural network model, the target peptide sequencing data in the test dataset is classified to obtain a first classification result;
[0037] Based on the first classification result, the first classification accuracy of the neural network model is determined using the classification label corresponding to the target polypeptide sequencing data.
[0038] In some embodiments, after determining the first classification accuracy of the neural network model based on the classification label corresponding to the target peptide sequencing data according to the first classification result, the method further includes:
[0039] The original polypeptide sequencing data is classified based on the trained neural network model to obtain a second classification result.
[0040] The second classification accuracy of the neural network model is determined based on the second classification result and the classification label corresponding to the original peptide sequencing data.
[0041] The screening quality results are determined based on the first classification accuracy and the second classification accuracy.
[0042] To achieve the above objectives, a third aspect of this application provides a peptide sequencing data screening device, the device comprising:
[0043] The data acquisition module is used to acquire multiple raw peptide sequencing data; wherein, the raw peptide sequencing data is obtained through nanopore sequencing technology;
[0044] The feature extraction module is used to extract features from the multiple raw peptide sequencing data based on a preset neural network model to obtain multiple deep features; wherein, the neural network model is a model trained based on preset sample data, and the preset sample data includes multiple sample peptide sequencing data and their corresponding classification labels;
[0045] The point cloud density estimation module is used to estimate the point cloud density of the plurality of depth features to obtain the point cloud density estimate value corresponding to each depth feature.
[0046] The data filtering module is used to filter the original polypeptide sequencing data based on the point cloud density estimate to obtain the filtering results.
[0047] To achieve the above objectives, a fourth aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the peptide sequencing data screening method described in the first aspect, or to implement the neural network model training method described in the second aspect.
[0048] To achieve the above objectives, a fifth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the peptide sequencing data screening method described in the first aspect or the neural network model training method described in the second aspect.
[0049] The peptide sequencing data screening and neural network model training method and related equipment proposed in this application extract multiple depth features from multiple raw peptide sequencing data through a neural network model. This allows for the acquisition of more accurate features characterizing the raw peptide sequencing data. Then, point cloud density estimation is performed on multiple depth features to obtain the point cloud density estimate value corresponding to each depth feature. This determines the feature distribution of each raw peptide sequencing data, so as to determine the screening results from multiple raw peptide sequencing data based on the point cloud density estimate value. This achieves accurate screening of multiple raw peptide sequencing data to filter out noisy data in the raw peptide sequencing data and obtain higher quality peptide sequencing data. Attached Figure Description
[0050] Figure 1 is a flowchart of the peptide sequencing data screening method provided in an embodiment of this application;
[0051] Figure 2 is a schematic diagram of the encoder structure in the polypeptide sequencing data screening method provided in the embodiments of this application;
[0052] Figure 3 is a schematic diagram of the classifier structure in the polypeptide sequencing data screening method provided in the embodiments of this application;
[0053] Figure 4 is a flowchart illustrating the peptide sequencing data screening method provided in the embodiments of this application;
[0054] Figure 5 is a schematic diagram of the distribution of point cloud density estimates in the peptide sequencing data screening method provided in the embodiments of this application;
[0055] Figure 6(a) is a schematic diagram of the point cloud density estimation value in the peptide sequencing data screening method provided in the embodiments of this application;
[0056] Figure 6(b) is a schematic diagram of the point cloud density estimate of the preliminary peptide sequencing data obtained after global screening in the peptide sequencing data screening method provided in the embodiments of this application.
[0057] Figure 6(c) is a schematic diagram of the point cloud density estimation value after local screening in the peptide sequencing data screening method provided in the embodiments of this application;
[0058] Figure 7 is a flowchart of the neural network model training method provided in an embodiment of this application;
[0059] Figure 8(a) shows a schematic diagram of the peptide classification confusion matrix of the neural network model for the original peptide sequencing data regarding the hp1 protein;
[0060] Figure 8(b) shows a schematic diagram of the peptide classification confusion matrix of the neural network model for the original peptide sequencing data regarding the hp1 protein;
[0061] Figure 9(a) shows a schematic diagram of the peptide classification confusion matrix of the neural network model for the original peptide sequencing data regarding the hp2 protein;
[0062] Figure 9(b) shows a schematic diagram of the peptide classification confusion matrix of the neural network model for the original peptide sequencing data regarding the hp2 protein;
[0063] Figure 10 is a schematic diagram of the structure of the polypeptide sequencing data screening device provided in the embodiment of this application;
[0064] Figure 11 is a schematic diagram of the structure of the electronic device provided in an embodiment of this application. Detailed Implementation
[0065] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0066] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0068] First, let's analyze some of the terms used in this application:
[0069] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.
[0070] Deep representation, also known as deep feature representation, refers to high-level abstract features extracted from raw data using deep neural networks (DNNs). These features are often more informative than the original data and are better suited for specific tasks such as classification, regression, and clustering. The key to deep representation lies in using a multi-layered structure to extract useful features from the data layer by layer, thereby improving model performance.
[0071] Deep learning networks are neural network models composed of multiple layers of neurons that solve complex tasks by extracting and learning features from data layer by layer.
[0072] Peptide sequencing data refers to data generated by determining the amino acid sequence of peptides (short-chain amino acid sequences) or proteins through experimental and computational methods. Peptide sequencing is an important technique in biomedical research, used to understand protein structure and function, discover biomarkers, and study disease mechanisms.
[0073] Nanopore peptide sequencing is an emerging biomolecular sequencing technology that uses nanopores to detect the amino acid sequence of a single peptide or protein molecule. This technology uses single-molecule conductivity sequencing, enabling direct reading of biomolecule sequences without the need for labeling or amplification. The following is a detailed introduction to nanopore peptide sequencing.
[0074] UMAP (Uniform Manifold Approximation and Projection) is a technique for dimensionality reduction and data visualization. UMAP effectively captures the local structure of data while preserving the global structure, making it particularly effective when dealing with high-dimensional data.
[0075] Nanopore sequencing is an advanced technique for determining protein sequences. It utilizes nanopore technology to determine peptide sequences by analyzing the current signal of single molecules passing through a nanopore. By linking peptides to the phosphoribosyl backbone of nucleic acids to form peptide-nucleic acid complexes, these complexes stably pass through the nanopore under the pull of motor proteins. The amino acid sequence is then determined by analyzing the signal. Therefore, nanopore sequencing technology is characterized by high sensitivity and high throughput, enabling efficient sequencing of single proteins and has broad application prospects in protein research, including the study of protein structure, function, and interactions. However, due to the limited spatial resolution of nanopores—the molecular diameter of amino acids is much smaller than the diameter of the nanopore—peptides can only generate a weak current blocking signal when passing through the nanopore. Furthermore, the non-constant pulling speed of motor proteins and the presence of noise interference in the electric field contribute to the high noise interference and poor data quality of nanopore peptide sequencing.
[0076] To filter out noisy data in peptide sequencing data, a feature extractor is needed to extract peptide sequence features from the sequencing data. Low-quality data is then filtered out based on these features. These peptide sequence features include characteristics such as the mean, standard deviation, length, and skewness of the peptide sequencing data. However, in related technologies, manually designed features cannot accurately and comprehensively cover the characteristics of the sequencing sample data.
[0077] Based on this, embodiments of this application provide a method and related equipment for screening peptide sequencing data and training neural network models. The aim is to extract accurate depth features from peptide sequencing data through neural network models and filter out noisy data in peptide sequencing data based on the point cloud density estimate of the depth features. This can achieve accurate filtering of low-noise data and improve the quality of peptide sequencing data.
[0078] The peptide sequencing data screening and neural network model training methods and related equipment provided in this application are specifically illustrated through the following embodiments. First, the peptide sequencing data screening method in this application embodiment is described.
[0079] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0080] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0081] The peptide sequencing data screening method provided in this application relates to the field of biotechnology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the peptide sequencing data screening method, but is not limited to the above forms.
[0082] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0083] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0084] Figure 1 is an optional flowchart of a peptide sequencing data screening method provided in an embodiment of this application. The method in Figure 1 may include, but is not limited to, steps 101 to 104.
[0085] Step 101: Obtain multiple raw peptide sequencing data; wherein, the raw peptide sequencing data are obtained through nanopore sequencing technology.
[0086] It should be noted that the raw peptide sequencing data is obtained by measuring the complex using nanopore sequencing technology. Specifically, the current change data output after different amino acids in the complex pass through the nanopore are used as the raw peptide sequencing data. In this embodiment, nanopore sequencing can directly read long nucleic acid sequences without amplification or labeling, offering advantages such as speed, convenience, and high throughput. Therefore, determining the raw peptide sequencing data using nanopore sequencing technology can yield high-quality peptide sequence data.
[0087] As previously disclosed, raw peptide sequencing data consists of current changes in different amino acids in the complex as they pass through nanopores. However, the number of amino acids in the complex varies, and the signal lengths of the amino acid numbers are related. Therefore, in order to output raw peptide sequencing data with consistent lengths, it is necessary to adjust the length of the current change data passing through the nanopores so that the neural network model can extract accurate features from the raw peptide sequencing data.
[0088] For example, to obtain raw peptide sequencing data for a mixed peptide sample of hNEDD8 protein using nanopore sequencing, the model protein hNEDD8 (hp1) is first selected from a protein library. hp1 is obtained through recombinant expression in *E. coli* BL21(DE3). The purified hp1 protein is then digested with the intraprotein enzyme LysC, resulting in multiple peptide fragments. Each fragment is 2 to 15 amino acids long, characterized by containing an N-terminal α-NH2 residue and a C-terminal lysine residue (with an ε-NH2 side chain residue) except for a single C-terminal peptide. The digested peptides are then treated with fluorosulfonyl azide (FSO2N3) at room temperature for 3 hours to produce peptides with N- and C-terminal azide groups, which can be coupled to DBCO-modified oligonucleotides. A DNA template-guided synthesis strategy is used to efficiently produce sequentially linked "nucleic acid-peptide-nucleic acid" complexes, which are then mixed with pre-prepared adapter complexes and incubated to form a database for nanopore protein sequencing. In this embodiment, a patch-clamp amplifier (or other electrical signal amplifiers) was used to acquire current signals to obtain raw peptide sequencing data.
[0089] Specifically, for current signal acquisition using a patch-clamp amplifier, the electrolytic cell is divided into two chambers using a planar 1,2-diphynoyl-sn-glycero-3-phosphocholine (DPhPC, Avanti Polar Lipids) phospholipid bilayer membrane: a cis chamber and a trans chamber. Each chamber contains a pair of Ag / AgCl electrodes. CsgG nanoporins are added to the phospholipid bilayer membrane, and a voltage of 180 mV is applied to promote the insertion of the nanoporins into the phospholipid bilayer membrane, forming individual nanopore channels. After the individual nanoporins are inserted into the phospholipid membrane, sequencing buffer (0.5 M KCl, 10 mM HEPES, 0.5 mM ATP, 1 mM MgCl2, pH 8) is introduced to remove excess nanoporins. Then, the above co-incubation mixture is added to the cis chamber and incubated at 25 °C for 10 min. Finally, 180 mV is applied, and nanopore current data are recorded at a frequency of 5 kHz to obtain raw peptide sequencing data.
[0090] To obtain raw peptide sequencing data for nanopore sequencing of a mixed peptide sample of the hCOMMD6 protein, another model protein, hCOMMD6 (hp2), was selected from a protein library and obtained through recombinant expression in *E. coli* BL21(DE3). The purified hp2 protein was digested with the endonuclease LysC, resulting in multiple peptide fragments, each ranging from 2 to 17 amino acids in length. These fragments were characterized by containing an N-terminal α-NH2 residue and a C-terminal lysine residue (with an ε-NH2 side chain) except for a single C-terminal peptide fragment. Subsequently, the digested peptide fragments were treated with fluorosulfonyl azide (FSO2N3) at room temperature for 3 hours to produce peptides with N- and C-terminal azide groups, which could be coupled to DBCO-modified oligonucleotides. A DNA template-guided synthesis strategy was used to efficiently produce sequentially linked "nucleic acid-peptide-nucleic acid" complexes, which were then mixed with pre-prepared adapter complexes and incubated to form a database for nanopore protein sequencing. Finally, the current signal of the complex was acquired using a patch-clamp amplifier to obtain the raw peptide sequencing data.
[0091] That is, in some embodiments, multiple raw peptide sequencing data are obtained, including:
[0092] Acquire multiple initial peptide sequencing data;
[0093] The length of multiple initial peptide sequencing data is adjusted according to a preset length to obtain multiple candidate peptide sequencing data.
[0094] Multiple candidate peptide sequencing data were normalized to obtain multiple raw peptide sequencing data.
[0095] It should be noted that the initial peptide sequencing data consists of current changes in different amino acids within the complex after passing through nanopores, and this initial peptide sequencing data has not undergone length adjustment. The preset length is a pre-set length and also represents the uniform length of multiple original peptide sequencing data sets. It should be noted that adjusting the length of the initial peptide sequencing data facilitates its loading into neural network models and also facilitates the training of neural network models. Therefore, by adjusting the length of multiple initial peptide sequencing data sets according to the preset length to obtain multiple candidate peptide sequencing data sets, the lengths of each candidate peptide sequencing data set are identical. Specifically, if the preset length is set to L, and the length of the initial peptide sequencing data is less than or greater than L, an interpolation algorithm is used to adjust the length of the initial peptide sequencing data to L to obtain the candidate peptide sequencing data.
[0096] After unifying the lengths, candidate peptide sequencing data were obtained. These candidate peptide sequencing data were then normalized to obtain multiple raw peptide sequencing data sets. The normalization process converts data of different dimensions and magnitudes into the same standard range, facilitating feature extraction from the raw peptide sequencing data by the subsequent neural network model.
[0097] It should be noted that before inputting the raw peptide sequencing data into the neural network model, the neural network model needs to be trained first so that it can extract more accurate peptide sequence features.
[0098] In some embodiments, prior to step 102, the peptide sequencing data screening method further includes:
[0099] Train the neural network model.
[0100] In this embodiment, the neural network model is a deep neural network model, capable of calculating the deep features of the raw peptide sequencing data and achieving accurate feature extraction. The neural network model includes an encoder, which comprises convolutional layers, batch normalization layers, ReLU activation functions, max pooling layers, dropout layers, and fully connected layers. It should be noted that the structure of the encoder is shown in Figure 2. In Figure 2, the convolutional layer (Conv) is a convolutional layer, the batch normalization layer (Bn) is a batch normalization layer, the pooling layer is a pooling layer, and the fully connected layer (FC) is a fully connected layer. This is achieved by setting multiple convolutional layers, batch normalization layers, and ReLU activation functions as one layer, followed by a group of pooling layers, and then setting multiple dropout layers and fully connected layers as another group of layers.
[0101] Therefore, by training a neural network model with the structure shown in Figure 2, it is possible to extract more accurate deep features representing the original peptide sequencing signal from the original peptide sequencing data.
[0102] In other words, in some embodiments, training a neural network model specifically includes:
[0103] Multiple sample peptide sequencing data are input into a neural network model for classification processing to obtain the classification prediction result of each sample peptide sequencing data.
[0104] The parameters of the neural network model are adjusted based on the classification prediction results and the classification labels corresponding to the sample peptide sequencing data.
[0105] It should be noted that the method for acquiring sample peptide sequencing data is the same as that for raw peptide sequencing data, and will not be repeated here. It should also be noted that the neural network model is a classifier. Classifiers have both encoding and classification functions, but when performing feature extraction, the neural network model only uses the encoding function, essentially acting as an encoder, omitting the classification function. However, to verify the accuracy of feature extraction by the neural network model, it is necessary to determine the classification based on the classification results and the corresponding classification labels of the sample peptide sequencing data. Therefore, the training process uses a neural network model with a classifier structure. After training the neural network model, the classification function is omitted, and it is then applied to the screening of peptide sequencing data.
[0106] Specifically, in the neural network model training process, the neural network model is a classifier, and the structure of the classifier is shown in Figure 3. As shown in Figure 3, the classifier adds a Dropout layer and a fully connected layer at the tail of the encoder. After the encoder extracts the deep features corresponding to the sample peptide sequencing data, the classification prediction result corresponding to each sample peptide sequencing data is determined. After the neural network model outputs the classification prediction result, the classification prediction result and the corresponding classification label are used to construct a loss function, and the parameters of the neural network model are adjusted according to the loss function until the parameters of the loss function meet the preset parameters. For example, if CrossEntropyLoss is used as the loss function, AdamW is used as the optimizer, and the learning rate is set to 5*10-4, the training of the neural network model is completed when the learning rate of the loss function reaches 5*10-4, so as to obtain deep features that can extract accurate representations of peptide sequencing data.
[0107] In this embodiment, a neural network model with a classifier structure is trained. The training process uses the classification prediction results and corresponding classification labels obtained from sample peptide sequencing data to train the neural network model, making the training operation simple and achieving high training accuracy. After training is complete, only the classification function needs to be removed, leaving the encoder structure, which can then be applied to feature extraction from peptide sequencing data. This allows the extracted deep features to more accurately represent the peptide sequencing data.
[0108] Step 102: Based on a preset neural network model, feature extraction is performed on multiple raw peptide sequencing data to obtain multiple deep features; wherein, the neural network model is a model trained based on preset sample data, which includes multiple sample peptide sequencing data and their corresponding classification labels.
[0109] After the neural network model is trained, deep features are extracted from the raw peptide sequencing data using the neural network model. Specifically, during feature extraction, the neural network model acts as an encoder, mapping the raw peptide sequencing data to a K-dimensional latent space to obtain K-dimensional abstract features as deep features. If there are N raw peptide sequencing data points of length L, the encoder outputs deep features of dimension N*K.
[0110] For example, multiple raw peptide sequencing data consist of 590,000 peptide sequencing signals. By adjusting the length of the peptide sequencing signals to 1,000, normalizing the peptide sequencing signals, and then using an encoder to calculate the depth representation of the peptide sequencing signals, the peptide signals are mapped to a 64-dimensional latent space.
[0111] Step 103: Perform point cloud density estimation on multiple depth features to obtain the point cloud density estimate corresponding to each depth feature;
[0112] After mapping multiple raw peptide sequencing data into a latent space to obtain multiple depth features, it is necessary to calculate the point cloud density estimate for each depth feature. It should be noted that the point cloud density estimate characterizes the distribution of the depth features. These depth features include characteristics such as the mean, standard deviation, length, and skewness of the raw peptide sequencing data. By calculating the point cloud density estimate for each depth feature, the differences between different raw peptide sequencing data under characteristics such as mean, standard deviation, length, and skewness can be clearly identified. This allows for the filtering out of low-quality raw peptide sequencing data, i.e., filtering out raw peptide sequencing data whose mean, length, standard deviation, and other characteristics do not meet the requirements.
[0113] In some embodiments, point cloud density estimation is performed on multiple depth features to obtain a point cloud density estimate corresponding to each depth feature, including:
[0114] Multiple deep features are dimensionality reduced to obtain multiple low-dimensional features;
[0115] The distribution density of multiple low-dimensional features is estimated to obtain the point cloud density estimate corresponding to each depth feature.
[0116] It should be noted that directly estimating point cloud density using multiple depth features would be time-consuming due to the multidimensional nature of these features. Therefore, it is recommended to first reduce the dimensionality of multiple depth features to obtain low-dimensional features, and then perform point cloud density estimation on these low-dimensional features. This reduces the time required for point cloud density estimation and improves the efficiency of point cloud density estimation calculation.
[0117] Specifically, in this embodiment, the low-dimensional feature is a 2-dimensional feature. The low-dimensional feature is obtained by reducing the multi-dimensional depth features to 2 dimensions, and then calculating the point cloud density estimate of the low-dimensional feature in 2-dimensional space. This embodiment uses the UMAP dimensionality reduction algorithm to reduce the depth features into low-dimensional features. The UMAP dimensionality reduction algorithm is fast and effective; therefore, using it to reduce the dimensionality of the depth features can also improve the computational efficiency of the point cloud density estimate.
[0118] For example, as shown in Figure 4, after preprocessing, N peptide sequencing signals of length L are encoded to extract depth features of dimension N*K. Using the UMAP dimensionality reduction algorithm, these depth features are reduced to a low-dimensional feature of dimension N*2. Then, the point cloud density is estimated in 2D space to obtain the point cloud density estimate. The distribution of the point cloud density estimate is shown in Figure 5. Figure 5 shows that there are 9 peptide categories in the multiple original peptide sequencing data. The color of the scatter points represents the peptide category. Therefore, the point cloud density estimate reveals the distribution of peptide categories in the multiple original peptide sequencing signals, facilitating the filtering of noise data from these signals.
[0119] Step 104: The raw peptide sequencing data is screened based on the point cloud density estimate to obtain the screening results.
[0120] After calculating the point cloud density estimate, the raw peptide sequencing data needs to be screened based on the estimate. Specifically, low-quality raw peptide sequencing data is selected as noise data from multiple raw peptide sequencing datasets based on the point cloud density estimate, and the noise data in the multiple raw peptide sequencing datasets is then filtered out to obtain the screening results.
[0121] It should be noted that in the process of screening multiple raw peptide sequencing data, global density information or local density information can be used for screening, so as to select different screening methods according to different needs and achieve efficient and accurate peptide sequencing data screening.
[0122] In some embodiments, the raw peptide sequencing data is screened based on point cloud density estimates to obtain screening results, including:
[0123] Based on a preset global threshold and point cloud density estimate, the raw peptide sequencing data is globally filtered to obtain preliminary peptide sequencing data.
[0124] Anomaly probability estimates are performed on the preliminary peptide sequencing data based on the point cloud density estimates, resulting in anomaly probability estimates.
[0125] The preliminary peptide sequencing data were screened based on preset local thresholds and anomaly probability estimates to obtain the screening results.
[0126] It should be noted that the global threshold is preset, and the setting of the global threshold is not arbitrary. It needs to be determined based on the point cloud density estimates of all depth features. Similarly, the local threshold also needs to be determined based on the point cloud density estimates of all depth features.
[0127] Specifically, in the filtering process of multiple raw peptide sequencing data, a global screening followed by a local screening is performed first to retain high-quality peptide sequencing data. Global screening is conducted on each raw peptide sequencing data point by using a global threshold and a point cloud density estimate. Data with a point cloud density estimate below the global threshold is filtered out, resulting in preliminary peptide sequencing data. This achieves the goal of filtering out noisy data from a global perspective. Subsequently, outlier detection technology in the HDBSCAN algorithm and the point cloud density estimate are used to calculate the anomaly probability of the preliminary peptide sequencing data. This anomaly probability estimate is a local estimate of the probability that the preliminary peptide sequencing data is an outlier. Preliminary peptide sequencing data with an anomaly probability estimate greater than the local threshold are filtered out, leaving high-quality preliminary peptide sequencing data as the screening result. Therefore, combining global and local screening methods to filter raw peptide sequencing data reduces the risk of filtering out noisy data and retains high-quality peptide sequencing data.
[0128] Specifically, as shown in Figures 6(a), 6(b), and 6(c), Figure 6(a) is a schematic diagram of the point cloud density estimation, Figure 6(b) is a schematic diagram of the point cloud density estimation of the preliminary peptide sequencing data after global screening, and Figure 6(c) is a schematic diagram of the point cloud density estimation after local screening. Therefore, it can be seen from Figures 6(a), 6(b), and 6(c) that using a global screening set with local screening can not only improve the screening efficiency but also improve the accuracy of peptide sequencing data screening.
[0129] In some embodiments, the peptide sequencing data screening method may further include:
[0130] Set global and local thresholds.
[0131] As previously disclosed, the global threshold and local threshold determine the screening accuracy. Therefore, the global threshold and local threshold cannot be set arbitrarily and need to be determined by combining the point cloud density estimates of multiple depth features.
[0132] That is, in some embodiments, global thresholds and local thresholds are set, specifically including:
[0133] The global threshold is determined based on the preset first quantile and the estimated point cloud density.
[0134] A local threshold is determined based on a preset second quantile and an anomaly probability estimate; wherein the first quantile is less than the second quantile.
[0135] It should be noted that the first quantile of the point cloud density estimate is used as the global threshold, and the second quantile of the anomaly probability estimate is used as the local threshold. This allows for the filtering out of some raw peptide sequencing data based on the global threshold, and then filtering out a small portion of raw peptide sequencing data locally, thus achieving efficient and accurate screening of multiple raw peptide sequencing data.
[0136] Specifically, in this embodiment, the first quantile is less than the second quantile, with the first quantile being 20% and the second quantile being 60%. The 20% quantile of the point cloud density estimate is used as the global threshold T1. Then, raw peptide sequencing data below the global threshold T1 are filtered out to obtain the initial peptide sequencing data. The 60% quantile of the anomaly probability estimate is used as the local threshold T2, and initial peptide sequencing data above the local threshold T2 are filtered out. The remaining peptide sequencing data is then used as the screening result. For example, if there are 59 peptide sequencing signals, after global and local screening, more than 280,000 peptide sequencing signals remain.
[0137] In summary, steps 101 to 104 of the embodiments of this application involve obtaining raw peptide sequencing data using nanopore sequencing technology, and then extracting multiple deep features from the raw peptide sequencing data using a neural network model. This neural network model is pre-trained using multiple sample peptide sequencing data and their corresponding classification labels. Therefore, by extracting multiple deep features through the trained neural network model, more accurate features representing the raw peptide sequencing data can be obtained. Point cloud density estimation is then performed on these multiple deep features to obtain the estimated point cloud density value for each deep feature. This point cloud density estimation is used to filter the raw peptide sequencing data and obtain the filtering results. Thus, the trained neural network model extracts deep features from the peptide sequencing data to obtain accurate deep features, calculates the estimated point cloud density value for each deep feature, and determines the filtering results from the raw peptide sequencing data based on the estimated point cloud density value, thereby accurately filtering out noisy data in the raw peptide sequencing data and improving the data quality of the peptide sequencing data.
[0138] It should be noted that the above-mentioned peptide sequencing data screening method can be used for screening peptide sequencing data, and can also be applied to medical data analysis such as electrocardiogram and electroencephalogram. In addition, it can also be applied to protein, biomarker, small molecule research, single molecule sensing detection, and biological data analysis with peptide sequencing data as the main component.
[0139] Referring to Figure 7, this application embodiment also provides a neural network model training method, including:
[0140] Step 701: Obtain the screening results obtained by the peptide sequencing data screening method described above.
[0141] It should be noted that the screening results obtained by the above peptide sequencing data screening method indicate which original peptide sequencing data were retained, and the retained original peptide sequencing data can be directly determined through the screening results.
[0142] Step 702: Based on the screening results, extract multiple target peptide sequencing data from multiple raw peptide sequencing data.
[0143] Specifically, raw peptide sequencing data that matches the screening results from multiple raw peptide sequencing datasets are selected as target peptide sequencing data. More specifically, if the screening result is a data number retained from the raw peptide sequencing data, the raw peptide sequencing data matching the data number is selected as the target peptide sequencing data. It should be noted that the target peptide sequencing data is high-quality peptide sequencing data, capable of accurate peptide analysis.
[0144] Step 703: Train the neural network model based on multiple target peptide sequencing data and the corresponding classification labels of the target peptide sequencing data.
[0145] It should be noted that high-quality sequencing data of multiple target peptides and their corresponding classification labels are used to train the neural network model to optimize it. Specifically, during neural network model training, a classifier structure is used to improve the classification accuracy of the classifier, thereby optimizing the precision of feature extraction.
[0146] As previously disclosed, the higher the quality of the target peptide sequencing data after screening, the more accurate the training classifier can be, because judging the accuracy of the classifier can determine the quality of the target peptide sequencing data.
[0147] In some embodiments, the neural network model is applied in the biological field, primarily using peptide testing data for biological data analysis to study protein structure, function, and interactions. This embodiment does not impose specific limitations on the application of the neural network model. For example, the trained neural network model can perform peptide classification on peptide sequencing data to determine peptide categories.
[0148] That is, in some embodiments, after step 702, the neural network model training method further includes:
[0149] Based on a preset allocation ratio, multiple target peptide sequencing data and their corresponding classification labels are divided into training datasets and test datasets.
[0150] It should be noted that, in order to determine the classification accuracy of the trained classifier, the sequencing data of multiple target peptides need to be divided according to the allocation ratio, and the corresponding classification labels of the target peptide sequencing data also need to be divided accordingly, forming a training dataset and a test dataset. The training dataset is used to train the classifier, and the test dataset is used to test the accuracy of the classifier.
[0151] The neural network model is trained based on the training dataset.
[0152] Specifically, the neural network model is first trained based on the training dataset, and the specific content of training the neural network model on the training dataset is the same as that of training the neural network model with the preset sample data mentioned above, so it will not be repeated here.
[0153] The target peptide sequencing data in the test dataset is classified based on the trained neural network model to obtain the first classification result.
[0154] After training a neural network model using the selected target peptide sequencing data, the target peptide sequencing data from the test dataset is input into the neural network model for classification to obtain the first classification result. It should be noted that the first classification result is the predicted peptide category corresponding to the target peptide sequencing data.
[0155] The first classification accuracy of the neural network model is determined based on the classification label corresponding to the target peptide sequencing data, according to the first classification result.
[0156] As disclosed above, the first classification result is the peptide prediction category corresponding to the target peptide sequencing data, and the classification label corresponding to the target peptide sequencing data is the peptide validation category. The first classification accuracy is obtained by calculating the loss between the peptide prediction category and the peptide validation category. The first classification accuracy characterizes the classification accuracy of the trained neural network model and also characterizes the quality of the target peptide sequencing data.
[0157] In some embodiments, the filtered data is used to train a classifier. This improves classification accuracy, thereby enhancing feature extraction precision. Furthermore, it reflects the quality of the target peptide sequencing data, thus determining the screening precision of the peptide sequencing data. It should be noted that the less noise data contained in multiple target peptide sequencing datasets, the higher the quality, and the better the classification performance of the trained classifier. Therefore, comparing the classification precision of the original peptide sequencing data and the target peptide sequencing data before and after screening is necessary to determine the quality of the screened target peptide sequencing data.
[0158] That is, in some embodiments, after determining the first classification accuracy of the neural network model based on the classification label corresponding to the target peptide sequencing data according to the first classification result, the neural network model training method further includes:
[0159] The raw peptide sequencing data is classified based on the trained neural network model to obtain a second classification result.
[0160] The second classification accuracy of the neural network model is determined based on the classification labels corresponding to the second classification results and the original peptide sequencing data.
[0161] The screening quality results are determined based on the accuracy of the first classification and the accuracy of the second classification.
[0162] The trained neural network model classifies the raw peptide sequencing data to obtain a second classification result, which is the predicted peptide category corresponding to the raw peptide sequencing data. The second classification accuracy is calculated based on the predicted peptide category and the classification label of the raw peptide sequencing data, and is the classification accuracy obtained using unscreened raw peptide sequencing data. Therefore, by comparing the first and second classification accuracies, the accuracy comparison result of the classifier using the two types of data can be determined, thereby determining the screening quality of the aforementioned peptide sequencing data screening method.
[0163] For example, in this embodiment, multiple raw peptide sequencing data for the hp1 protein were filtered to obtain multiple target peptide sequencing data. These multiple target peptide sequencing data were used to train a neural network model with a classifier structure, improving the accuracy of the neural network model from 95.5% to 96.7%. As shown in Figure 8(a) and Figure 8(b), Figure 8(a) shows the peptide classification confusion matrix of the neural network model for the raw peptide sequencing data before data filtering, with an average accuracy of 95.5%. Figure 8(b) shows the peptide classification confusion matrix of the neural network after training with multiple target peptide sequencing data, with an average accuracy of 96.7%. This demonstrates that training the classifier with target peptide sequencing data after filtering out noisy data can improve classification accuracy.
[0164] After screening multiple raw peptide sequencing data for the hp2 protein, multiple target peptide sequencing data were obtained. Originally, these raw peptide sequencing data contained 256,000 entries. After data filtering, 123,000 target peptide sequencing data remained. A neural network model for classifying the target peptide sequencing data of the hp2 protein was trained, improving the accuracy from 93% to 95%. Specifically, as shown in Figure 9(a) and Figure 9(b), Figure 9(a) shows the peptide classification confusion matrix of the neural network model before data filtering, with an average accuracy of 93%. Figure 9(b) shows the peptide classification confusion matrix after data filtering, using the target peptide sequencing data to train the neural network model, with an average accuracy of 95%.
[0165] In summary, further optimization of the classifier using the selected target peptide sequencing data can improve classification accuracy.
[0166] Please refer to Figure 10. This application embodiment also provides a peptide sequencing data screening device, which can implement the above-described peptide sequencing data screening method. The device includes:
[0167] The data acquisition module 1001 is used to acquire multiple raw peptide sequencing data; the raw peptide sequencing data is obtained through nanopore sequencing technology.
[0168] The feature extraction module 1002 is used to extract features from multiple raw peptide sequencing data based on a preset neural network model to obtain multiple deep features; wherein, the neural network model is a model trained based on preset sample data, and the preset sample data includes multiple sample peptide sequencing data and their corresponding classification labels.
[0169] The point cloud density estimation module 1003 is used to perform point cloud density estimation on multiple depth features to obtain the point cloud density estimation value corresponding to each depth feature.
[0170] The data filtering module 1004 is used to filter the raw peptide sequencing data based on the point cloud density estimate to obtain the filtering results.
[0171] The specific implementation of this peptide sequencing data screening device is basically the same as the specific embodiment of the peptide sequencing data screening method described above, and will not be repeated here.
[0172] Please refer to Figure 11, which illustrates the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0173] The processor 1101 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0174] The memory 1102 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1102 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1102 and is called and executed by the processor 1101 to execute the peptide sequencing data screening method or the neural network model training method of this embodiment.
[0175] Input / output interface 1103 is used to implement information input and output;
[0176] The communication interface 1104 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0177] Bus 1105 transmits information between various components of the device (e.g., processor 1101, memory 1102, input / output interface 1103, and communication interface 1104);
[0178] The processor 1101, memory 1102, input / output interface 1103 and communication interface 1104 are connected to each other within the device via bus 1105.
[0179] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described peptide sequencing data screening method or the neural network model training method of this embodiment.
[0180] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0181] The peptide sequencing data screening and neural network model training method and related equipment provided in this application extract multiple depth features from multiple raw peptide sequencing data through a neural network model. This allows for the acquisition of more accurate features characterizing the raw peptide sequencing data. Then, point cloud density estimation is performed on the multiple depth features to obtain the point cloud density estimate value corresponding to each depth feature. This determines the feature distribution of each raw peptide sequencing data, so as to determine the screening results from multiple raw peptide sequencing data based on the point cloud density estimate value. This achieves accurate screening of multiple raw peptide sequencing data, filtering out noisy data in the raw peptide sequencing data and obtaining higher quality peptide sequencing data.
[0182] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0183] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0184] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0185] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0186] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0187] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0188] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0189] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0190] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0191] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0192] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for screening peptide sequencing data, characterized in that, The method comprises: obtaining a plurality of original polypeptide sequencing data; wherein the original polypeptide sequencing data is obtained by nanopore sequencing technology; performing feature extraction on the plurality of original polypeptide sequencing data based on a preset neural network model to obtain a plurality of deep features; wherein the neural network model is a model trained based on preset sample data, and the preset sample data comprises a plurality of sample polypeptide sequencing data and corresponding classification labels thereof; performing point cloud density estimation on the plurality of deep features to obtain a point cloud density estimation value corresponding to each deep feature; performing screening on the original polypeptide sequencing data based on the point cloud density estimation value to obtain a screening result.
2. The method of claim 1, wherein, The point cloud density estimation on the plurality of deep features to obtain a point cloud density estimation value corresponding to each deep feature comprises: performing dimension reduction processing on the plurality of deep features to obtain a plurality of low-dimensional features; performing distribution density estimation on the plurality of low-dimensional features to obtain the point cloud density estimation value corresponding to each deep feature.
3. The method of claim 1, wherein, The screening on the original polypeptide sequencing data based on the point cloud density estimation value to obtain a screening result comprises: performing global screening on the original polypeptide sequencing data based on a preset global threshold and the point cloud density estimation value to obtain preliminary polypeptide sequencing data; performing abnormal probability estimation on the preliminary polypeptide sequencing data according to the point cloud density estimation value to obtain an abnormal probability estimation value; performing screening on the preliminary polypeptide sequencing data based on a preset local threshold and the abnormal probability estimation value to obtain the screening result.
4. The method of claim 3, wherein, Before the screening on the preliminary polypeptide sequencing data based on the preset local threshold and the abnormal probability estimation value to obtain the screening result, the method further comprises: setting the global threshold and the local threshold, specifically comprising: determining the global threshold based on a preset first quantile and the point cloud density estimation value; determining the local threshold based on a preset second quantile and the abnormal probability estimation value; wherein the first quantile is smaller than the second quantile.
5. The method according to any one of claims 1 to 4, characterized in that, The obtaining of the plurality of original polypeptide sequencing data comprises: obtaining a plurality of initial polypeptide sequencing data; performing length adjustment on the plurality of initial polypeptide sequencing data according to a preset length to obtain a plurality of candidate polypeptide sequencing data; performing normalization processing on the plurality of candidate polypeptide sequencing data to obtain the plurality of original polypeptide sequencing data.
6. The method according to any one of claims 1 to 4, characterized in that, Before the feature extraction on the plurality of original polypeptide sequencing data based on the preset neural network model to obtain a plurality of deep features, the method further comprises: training the neural network model, specifically comprising: inputting the plurality of sample polypeptide sequencing data into the neural network model for classification processing to obtain a classification prediction result of each sample polypeptide sequencing data; performing parameter adjustment on the neural network model based on the classification prediction result and the classification label corresponding to the sample polypeptide sequencing data.
7. A neural network model training method characterized by comprising: The method comprises: obtaining a screening result obtained by the polypeptide sequencing data screening method according to any one of claims 1 to 6. extract target polypeptide sequencing data from the plurality of original polypeptide sequencing data based on the screening result; train a neural network model based on the plurality of target polypeptide sequencing data and the classification labels corresponding to the target polypeptide sequencing data.
8. The method of claim 7, wherein, After the plurality of target polypeptide sequencing data is extracted from the plurality of original polypeptide sequencing data based on the screening result, the method further comprises: dividing the plurality of target polypeptide sequencing data and the classification labels corresponding to the target polypeptide sequencing data into a training data set and a test data set according to a preset allocation ratio; training the neural network model based on the training data set; performing classification processing on the target polypeptide sequencing data in the test data set based on the trained neural network model to obtain a first classification result; determining a first classification accuracy of the neural network model based on the first classification result and the classification labels corresponding to the target polypeptide sequencing data.
9. The method of claim 8, wherein, After the first classification accuracy of the neural network model is determined based on the first classification result and the classification labels corresponding to the target polypeptide sequencing data, the method further comprises: performing classification processing on the original polypeptide sequencing data based on the trained neural network model to obtain a second classification result; determining a second classification accuracy of the neural network model based on the second classification result and the classification labels corresponding to the original polypeptide sequencing data; determining a screening quality result based on the first classification accuracy and the second classification accuracy.
10. A polypeptide sequencing data screening device, characterized in that, The device comprises: a data acquisition module configured to acquire a plurality of original polypeptide sequencing data, wherein the original polypeptide sequencing data is obtained by nanopore sequencing technology; a feature extraction module configured to perform feature extraction on the plurality of original polypeptide sequencing data based on a preset neural network model to obtain a plurality of deep features, wherein the neural network model is a model trained based on preset sample data, and the preset sample data comprises a plurality of sample polypeptide sequencing data and classification labels corresponding thereto; a point cloud density estimation module configured to perform point cloud density estimation on the plurality of deep features to obtain a point cloud density estimation value corresponding to each deep feature; a data screening module configured to perform screening on the original polypeptide sequencing data based on the point cloud density estimation value to obtain a screening result.
11. An electronic device, comprising: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the polypeptide sequencing data screening method of any one of claims 1 to 6 or the neural network model training method of any one of claims 7 to 9 when executing the computer program.
12. A computer-readable storage medium, the computer-readable storage medium storing a computer program, characterized in that, The computer program is executed by the processor to implement the polypeptide sequencing data screening method of any one of claims 1 to 6 or the neural network model training method of any one of claims 7 to 9.
Citation Information
Patent Citations
Solid nanopore sequencing electric signal noise reduction processing method based on residual autoencoder convolutional neural network
CN113743301A
Nanopore protein sequencing data processing method based on machine learning and application thereof
CN116741265A
Systems and methods for improved nanopore-based analysis of nucleic acids
US20220366313A1
Method and apparatus for classification model training and classification, computer device, and storage medium
US20230084638A1
Method and device for classifying nanopore sequencing time series electrical signal
WO2024124521A1