A method and system for population stratification based on protein data

By employing a protein-based population differentiation method, utilizing OpenMS and maxquant tools to extract feature information, and combining a filtering joint embedding method for feature selection, the problem of differentiation difficulties caused by DNA loss was solved, achieving stable population classification results.

CN116230084BActive Publication Date: 2025-12-30CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310207004.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-03
Publication Date
2025-12-30
Estimated Expiration
2043-03-03

AI Technical Summary

Technical Problem

Existing DNA typing techniques are ineffective in distinguishing populations in bioarchaeology and forensic medicine due to DNA loss, and a stable and abundant alternative method is urgently needed.

Method used

A population discrimination method based on protein data is adopted, including feature extraction, preprocessing, screening and model training. Protein data is used to distinguish populations. Openms and maxquant tools are used to extract mass spectrometry feature peaks and peptide quantitative information. Feature screening is carried out by filtering and joint embedding method, and support vector machine is used for classification.

Benefits of technology

It achieves effective population differentiation in cases of DNA loss, improves classification accuracy and stability, and is applicable to machine learning models in the life sciences field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116230084B_ABST
    Figure CN116230084B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of life science, and particularly relates to a population distinguishing method and system based on protein data; the method comprises the following steps: obtaining protein data with race information, performing feature extraction, preprocessing and screening processing on the protein data respectively, inputting the processed protein data into a trained population distinguishing model based on protein data for processing, and obtaining a population distinguishing result of the protein data; the present application uses protein level data to distinguish populations, which is helpful to solve the problem that DNA typing cannot be used due to DNA loss in special sites, and the model has high classification accuracy and good application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of life sciences, specifically relating to a method and system for population differentiation based on protein data. Background Technology

[0002] The problem of population differentiation involves multiple fields, including biomedicine, population evolution, forensics, and archaeology. DNA typing is a commonly used method for identifying individuals. It can statistically place individuals in specific locations, correlate them with physical evidence, and determine biometric and biogeographical genetic information. DNA typing methods rely on the presence of a sufficient quantity and quality of DNA template to amplify and generate genotypic information for short tandem repeat loci (STRs), single nucleotide polymorphisms (SNPs), or mitochondrial DNA haplotypes via PCR. However, a major limitation of these techniques is the sensitivity of DNA to biological, environmental, and chemical processes that reduce template length and modify base structure. These processes can lead to the loss of template DNA in the sample, sometimes exceeding the compensatory capacity of PCR and sequencing strategies. If DNA typing produces partial or invalid results, researchers have few quantifiable genetic alternatives; for example, in archaeological or crime scenes, DNA typing may be unusable due to DNA loss.

[0003] Beyond the development of identification techniques that rely solely on DNA typing, new methods for population differentiation are urgently needed in forensic and bioarchaeological fields. With advancements in mass spectrometry, proteins, which are chemically more stable, abundant, and durable than DNA, have gradually come into focus. The condition of proteins in bioarchaeological samples is often used as an indicator of biomolecular integrity. For example, protein yield and carbon-to-nitrogen ratio are considered necessary, but not sufficient, indicators of the presence of residual endogenous DNA templates. Hair keratin, bone collagen, and tooth collagen are now routinely used for 14C dating and for stable optical isotope analysis of ancient diets and related information.

[0004] In conclusion, there is an urgent need for a method that uses protein data to differentiate populations, in order to overcome the shortcomings of using DNA for population differentiation. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention proposes a method and system for population discrimination based on protein data. The method includes: acquiring protein data with population information; performing feature extraction, preprocessing, and screening on the protein data; inputting the processed protein data into a trained population discrimination model based on protein data for further processing; and obtaining the population discrimination result of the protein data.

[0006] The process of training a population discrimination model based on protein data includes:

[0007] S1: Obtain a protein dataset with population information;

[0008] S2: Perform feature extraction on the protein dataset to obtain the first feature information; the first feature information includes mass spectrometry characteristic peak information, peptide quantification information, and protein quantification information.

[0009] S3: Preprocess the feature information to obtain the second feature information; sample the intermediate feature information to obtain the third feature information;

[0010] S4: Perform feature filtering on the third feature information to obtain the filtered feature information;

[0011] S5: Use the filtered feature information as training data to train the protein data-based population discrimination model, and obtain the trained protein data-based population discrimination model.

[0012] Preferably, the process of feature extraction for protein data includes: using Openms to extract mass spectrometry characteristic peak information, and using maxquant library search to extract peptide quantification information and protein quantification information.

[0013] Preferably, the preprocessing of feature information includes: data alignment, complementation, and normalization of the feature information.

[0014] Preferably, the process of sampling and processing intermediate feature information includes: selecting the sample number with the largest actual population as N; if the sample data of other populations is less than N, artificially simulating new samples of other populations to fill the sample number of other populations to N.

[0015] Preferably, the methods for feature selection of the third feature information include: filtering methods, embedding methods, and filtering combined with embedding methods.

[0016] Furthermore, the filtering joint embedding method specifically includes: using a filtering method to select the top 250, top 500, top 1000, top 1500, and top 2000 features as feature subsets for model training; and using an embedding method to further filter these feature subsets to obtain the filtered feature information.

[0017] Preferably, the population discrimination model based on protein data is the support vector machine.

[0018] A population differentiation system based on protein data includes: a feature extraction module, a preprocessing module, a screening module, and a classification module;

[0019] The feature extraction module is used to extract features from protein data to obtain feature information;

[0020] The preprocessing module is used to preprocess the feature information to obtain preprocessed feature information;

[0021] The filtering module is used to filter the preprocessed feature information to obtain filtered feature information.

[0022] The classification module is used to classify protein data based on the filtered feature information and output the population distinction results of the protein data.

[0023] The beneficial effects of this invention are as follows:

[0024] 1. This invention makes full use of protein-level data and innovatively utilizes protein characteristic peak data for population differentiation, which helps to solve the problem that DNA typing in special situations cannot be used due to DNA loss.

[0025] 2. The machine learning model of this invention can be applied to the field of life sciences to attempt to distinguish populations using protein-level data, and it has good population classification results;

[0026] 3. This invention innovatively proposes a feature selection method using filtering joint embedding, which improves the accuracy of the model. Attached Figure Description

[0027] Figure 1 This is a flowchart of the training process for the population discrimination model based on protein data in this invention. Detailed Implementation

[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] This invention proposes a method and system for population differentiation based on protein data, such as... Figure 1 As shown, the method includes the following:

[0030] Obtaining protein data with population information can be used in practical applications to extract protein data with population information from special occasions such as archaeological sites and crime scenes where DNA is easily lost or damaged.

[0031] The protein data undergoes feature extraction, preprocessing, and screening. The processed protein data is then input into a pre-trained protein-based population discrimination model to obtain the population discrimination results. The training process for the protein-based population discrimination model includes:

[0032] S1: Obtain a protein dataset with population information.

[0033] A protein dataset with population information is obtained for training; preferably, this protein dataset can be obtained from the CPCAC public dataset.

[0034] S2: Extract features from the protein dataset to obtain the first feature information; the first feature information includes mass spectrometry characteristic peak information, peptide quantification information, and protein quantification information.

[0035] Peptide quantification and protein quantification are commonly used features for protein analysis. Protein mass spectrometry characteristic peaks are the numerous mass spectrometry peaks generated during the protein mass spectrometry process. When comparing the primary and secondary mass spectrometry peaks generated after trypsin digestion to identify proteins, some meaningful characteristic peaks may be discarded. Directly using protein mass spectrometry characteristic peaks as features can avoid the rejection of meaningful protein characteristic peaks. By selecting the above three types of feature information as features of the protein dataset, the model can achieve better classification results.

[0036] This invention employs OpenMS to extract mass spectrometry characteristic peak information and maxquant to search the library for peptide and protein quantification information. OpenMS is a software tool for analyzing proteomics datasets. Starting with input files in vendor-specific raw formats (which can be converted using the integrated ProteoWizard) or many HUPO-compliant open formats (mzML, mzXML), data can be processed by combining a set of nearly 200 off-the-shelf tools from the "OpenMS Proteomics Pipeline (TOPP)". These tools can be used from workflow engines (such as KNIME and Galaxy) on the command line. maxquant is one of the most commonly used platforms for mass spectrometry (MS)-based proteomics data analysis, supporting various quantitative labeling techniques and non-standard quantification for high-resolution MS data, and currently enjoys relatively high acceptance.

[0037] S3: Preprocess the feature information to obtain the second feature information; sample the intermediate feature information to obtain the third feature information.

[0038] Preprocessing of feature information includes: data alignment, imputation, and normalization; specifically: checking the original data to ensure no missing or duplicate values; imputing blank values ​​using the mean; and normalizing the data using the min-max scaling method, as shown below:

[0039]

[0040] in, This represents the normalized sample data value, where x represents the sample data, max represents the maximum value of sample data x, and min represents the minimum value of sample data x.

[0041] The intermediate feature information is sampled and processed, specifically:

[0042] Many models output categories based on thresholds; for example, in logistic regression, values ​​less than 0.5 are considered negative examples, while values ​​greater than 0.5 are considered positive examples. When data is imbalanced, the default threshold can cause the model output to favor the category with more data. To address this imbalance, this invention employs a sampling method, which includes undersampling and oversampling. Undersampling is used to sample from three population categories, with the number of samples for each category equal to the size of the smallest sample category. However, a relatively small dataset can negatively impact classifier accuracy. Preferably, to create a balanced and reasonably sized dataset, this invention uses oversampling to address the imbalance based on the original data. Specifically, the largest sample size N is selected. To ensure a reasonable sample size, N=70 is chosen; 70 samples from each class are selected from the total intermediate feature information samples to create a total of 210 sample datasets. Classes with 70 samples are directly retained; for classes with fewer than 70 samples, oversampling is performed by randomly selecting from their sample pool to fill each class with 70 samples. Because the simple replication resulting from sampling with replacement can lead to overfitting, the SMOTE algorithm, an improved version of the random oversampling algorithm, can perform oversampling on minority class samples. Preferably, this invention uses the SMOTE algorithm for oversampling. The basic idea of ​​the SMOTE algorithm is to analyze and simulate minority class samples and add artificially simulated new samples to the dataset, thereby correcting the severe class imbalance in the original data. The simulation process of this algorithm employs the KNN technique.

[0043] S4: Perform feature filtering on the third feature information to obtain the filtered feature information.

[0044] Whether in the learning or prediction phase, using such a large feature vector would be computationally expensive without feature filtering. Furthermore, not all features may always be effective in a learning model. Therefore, this invention selects a set of relevant features that contribute to improving the accuracy of the learning model—that is, performs feature filtering on the third feature information. Methods for feature filtering include filtering, wrapping, embedding, and a combined filtering and embedding approach. Filtering methods rank features according to certain criteria. A subset of the top-ranked features is then used to train the classifier. Therefore, these methods are independent of the choice of classifier. On the other hand, wrapping methods search the feature space to find the optimal feature subset. The quality of the feature subset is measured by training and testing a specific classification model. Therefore, these methods are related to a specific classification algorithm. Embedding methods are similar to wrapping methods. However, in this approach, the search for the optimal feature subset is inherently built into the classification algorithm.

[0045] Preferably, the present invention uses a filtering joint embedding method; the specific method is as follows: the filtering method is used to select the top 250, top 500, top 1000, top 1500 and top 2000 features as feature subsets for model training; these feature subsets are further filtered using the embedding method to obtain the filtered feature information.

[0046] The filtering method employs the chi-square test, a relevance filtering technique specifically designed for discrete labels (i.e., classification problems). The chi-square test calculates the chi-square statistic between each non-negative feature and the label, using this statistic as a feature score. Features are then ranked from highest to lowest chi-square statistic. The top K features are selected from this ranking to form a feature subset; ideally, K is set to 250, 500, 1000, 1500, or 2000. This filtering method removes features most likely independent of the label and irrelevant to the classification objective.

[0047] Embedding is a method that allows the algorithm to decide which features to use, meaning feature selection and model training occur simultaneously. When using embedding, a subset of features is input, and a machine learning algorithm or model is used for training, resulting in weight coefficients (between 0 and 1) for each feature. These weight coefficients often represent the feature's contribution to the model (its importance). An optimal threshold is then found using the learning curve, and features are selected based on this threshold to obtain the final set of features.

[0048] S5: Use the filtered feature information as training data to train the protein data-based population discrimination model, and obtain the trained protein data-based population discrimination model.

[0049] Population differentiation models based on protein data can include Support Vector Machine (SVM), Random Forest, etc.

[0050] Evaluation of this invention: Several performance metrics can be used to evaluate the predictive performance of the trained model. The metrics selected in this invention are accuracy, precision, and recall. Through multiple experiments, using support vector machines and an oversampling sampling method, a combination of filtering and embedding methods was employed for feature selection. When the number of features selected by the filtering method was 1000, all the model's metrics reached their optimal levels.

[0051] This invention also proposes a population discrimination system based on protein data, which can implement the above-mentioned population discrimination method based on protein data, including: a feature extraction module, a preprocessing module, a screening module, and a classification module;

[0052] The feature extraction module is used to extract features from protein data to obtain feature information;

[0053] The preprocessing module is used to preprocess the feature information to obtain preprocessed feature information;

[0054] The filtering module is used to filter the preprocessed feature information to obtain filtered feature information.

[0055] The classification module is used to classify protein data based on the filtered feature information and output the population distinction results of the protein data.

[0056] The above-described embodiments further illustrate the purpose, technical solution, and advantages of the present invention. It should be understood that the above-described embodiments are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made to the present invention within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for population differentiation based on protein data, characterized in that, The method comprises the following steps: The process of training the population differentiation model based on protein data comprises the following steps: S1: obtaining a protein data set with population information; S2: performing feature extraction on the protein data set to obtain first feature information; the first feature information comprises mass spectrometry feature peak information, peptide segment quantitative information, and protein quantitative information; S3: performing preprocessing on the feature information to obtain second feature information; performing sampling processing on the intermediate feature information to obtain third feature information; S4: performing feature screening on the third feature information to obtain screened feature information; S5: using the screened feature information as training data to train the population differentiation model based on protein data, to obtain the trained population differentiation model based on protein data. The process of performing feature extraction on the protein data comprises the following steps: using Openms to extract mass spectrometry feature peak information, and using maxquant to extract peptide segment quantitative information and protein quantitative information.

2. The method of claim 1, wherein the method comprises: The preprocessing of the feature information comprises the following steps: performing data alignment, value filling, and normalization processing on the feature information.

3. The method of claim 1, wherein the method comprises: The process of performing sampling processing on the intermediate feature information comprises the following steps: selecting the sample number of the population with the largest number of actual populations as N; if the sample data of other populations is less than N, artificially simulating new samples of other populations to fill the sample number of other populations to N.

4. The method of claim 1, wherein the method comprises: The feature screening method for the third feature information comprises a filtering method, an embedding method, and a filtering-embedding method.

5. The method of claim 1, wherein the method is based on protein data.

6. The population differentiation method based on protein data according to claim 5, wherein the filtering-embedding method comprises the following steps: The population differentiation model based on protein data is a support vector machine. The method comprises the following steps:

7. The method of claim 1, wherein the method is based on protein data. The feature extraction module is configured to perform feature extraction on protein data to obtain feature information; 8. A population differentiation system based on protein data, the system being configured to perform the population differentiation method based on protein data according to any one of claims 1 to 7, characterized in that, The preprocessing module is configured to perform preprocessing on the feature information to obtain preprocessed feature information; The screening module is configured to perform feature screening on the preprocessed feature information to obtain screened feature information; The classification module is configured to perform protein data classification according to the screened feature information, and output a population differentiation result of protein data. ​ ​ ​

Citation Information

Patent Citations

  • Heterogeneous Feature Fusion Based Risk Prediction Method, Model and System for Coronary Heart Disease

    CN109117864A

  • Protein model quality evaluation method based on deep learning

    CN114530195A