Method and device for predicting ancestral category of unknown biological sample and electronic equipment

By combining the machine learning model of PCA and XGBoost algorithms, the problem of low accuracy of ancestral category prediction caused by single dimensional characteristics in the prior art is solved, and high accuracy prediction of ancestral category of unknown biological samples is achieved.

CN120196978APending Publication Date: 2025-06-24INST OF FORENSIC SCI OF MIN OF PUBLIC SECURITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311774111.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The ancestral category prediction methods for existing unknown biological samples involve relatively single dimensional characteristics, resulting in a low accuracy in predicting ancestral category of unknown biological samples.

Method used

A machine learning model combining unsupervised learning PCA algorithm and supervised learning XGBoost algorithm was established. Based on high-density SNP classification data of unknown biological samples, the dimensionality reduction of principal component data is obtained in order to improve the accuracy of ancestral category prediction.

Benefits of technology

Through this method, the ancestral source corresponding to unknown biological samples can be accurately determined, effectively improving the accuracy of prediction of ancestral source categories of unknown biological samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196978A_ABST
    Figure CN120196978A_ABST
Patent Text Reader

Abstract

The invention provides an ancestral category prediction method and device for an unknown biological sample and electronic equipment, and the method comprises the steps: determining principal component data of the unknown biological sample, the principal component data of the unknown biological sample being used for representing the ethnic group characteristics of the unknown biological sample, and the data dimension of the principal component data of the unknown biological sample being greater than or equal to a preset dimension threshold; and inputting the principal component data of the unknown biological sample into an ancestral category prediction model to obtain an ancestral prediction result output by the ancestral category prediction model, the ancestral category prediction model being constructed by the principal component data after dimension reduction of the genetic data of the reference population sample and the real ancestral information corresponding to the reference population sample. According to the method, the ancestral source category prediction model combining the unsupervised learning PCA algorithm and the supervised learning XGBoost algorithm is established, the ancestral source corresponding to the unknown biological sample can be accurately determined by adopting the ancestral source category prediction model, and the ancestral source category prediction accuracy of the unknown biological sample is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ancestral origin prediction, and particularly to a method, device and electronic device for predicting the ancestral origin category of an unknown biological sample. Background Art

[0002] BiogeoGraphical Ancestry (BGA) prediction is widely applied in multiple fields such as population genetics, forensic genetics, epidemiology and pharmacogenomics. With the decline in the cost of high-density Single Nucleotide Polymorphism (SNP) detection technologies (such as sequencing and SNP chips, etc.) and the rapid growth of the scale of population genomic datasets, the concept of spatial population genetics has attracted attention again.

[0003] The existing methods for predicting the ancestral origin category of an unknown biological sample predict the ancestral origin category of the unknown biological sample through unsupervised learning methods (such as Principal Component Analysis (PCA) method) and supervised learning methods (such as linear regression method). However, the dimensional features involved in this method are relatively single, resulting in a low accuracy rate for predicting the ancestral origin category of the unknown biological sample. Summary of the Invention

[0004] The present invention provides a method, device and electronic device for predicting the ancestral origin category of an unknown biological sample, so as to solve the defect that the dimensional features involved in the existing methods for predicting the ancestral origin category of an unknown biological sample are relatively single, resulting in a low accuracy rate for predicting the ancestral origin category of the unknown biological sample. This method establishes a machine learning model (ancestral origin category prediction model) that combines the PCA algorithm of unsupervised learning and the eXtreme Gradient Boosting (XGBoost) algorithm of supervised learning. Based on the high-density SNP genotyping data of the unknown biological sample, which contains rich genetic information, the principal component data of a higher dimension is obtained by dimensionality reduction, which can reduce the amount of calculation while retaining more information. Furthermore, by adopting the ancestral origin category prediction model subsequently, the ancestral origin corresponding to the unknown biological sample can be accurately determined, effectively improving the accuracy rate of predicting the ancestral origin category of the unknown biological sample.

[0005] The present invention provides a method for predicting the ancestral origin category of an unknown biological sample, including:

[0006] Determine the principal component data of the unknown biological sample, where the principal component data of the unknown biological sample is used to characterize the ethnic group characteristics of the unknown biological sample, and the data dimension of the principal component data of the unknown biological sample is greater than or equal to a preset dimension threshold;

[0007] Input the principal component data of the unknown biological sample into the ancestral origin category prediction model to obtain the ancestral origin prediction result output by the ancestral origin category prediction model. The ancestral origin category prediction model is constructed based on the principal component data after dimensionality reduction of the genetic data of the reference population sample and the corresponding true ancestral origin information of the reference population sample.

[0008] The present invention also provides an ancestral origin category prediction device for an unknown biological sample, including:

[0009] A principal component data determination module, configured to determine the principal component data of the unknown biological sample. The principal component data of the unknown biological sample is used to characterize the ethnic group characteristics of the unknown biological sample, and the data dimension of the principal component data of the unknown biological sample is greater than or equal to a preset dimension threshold;

[0010] An ancestral origin prediction result determination module, configured to input the principal component data of the unknown biological sample into the ancestral origin category prediction model to obtain the ancestral origin prediction result output by the ancestral origin category prediction model. The ancestral origin category prediction model is constructed based on the principal component data after dimensionality reduction of the genetic data of the reference population sample and the corresponding true ancestral origin information of the reference population sample.

[0011] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the ancestral origin category prediction method for the unknown biological sample as described in any one of the above.

[0012] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the ancestral origin category prediction method for the unknown biological sample as described in any one of the above.

[0013] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the ancestral origin category prediction method for the unknown biological sample as described in any one of the above.

[0014] The method, device and electronic device for predicting the ancestral origin category of an unknown biological sample provided by the present invention determine the principal component data of the unknown biological sample, which is used to characterize the ethnic characteristics of the unknown biological sample, and the data dimension of the principal component data of the unknown biological sample is greater than or equal to a preset dimension threshold; input the principal component data of the unknown biological sample into the ancestral origin category prediction model to obtain the ancestral origin prediction result output by the ancestral origin category prediction model. The ancestral origin category prediction model is constructed based on the principal component data after dimensionality reduction of the genetic data of the reference population sample and the corresponding true ancestral origin information of the reference population sample. This method establishes a machine learning model (ancestral origin category prediction model) that combines the PCA algorithm of unsupervised learning and the XGBoost algorithm of supervised learning. Based on the high-density SNP genotyping data of the unknown biological sample, which contains rich genetic information, dimensionality reduction is performed to obtain the principal component data of a relatively high dimension, which can reduce the amount of calculation while retaining more information. Then, by using the ancestral origin category prediction model subsequently, the ancestral origin corresponding to the unknown biological sample can be accurately determined, effectively improving the prediction accuracy of the ancestral origin category of the unknown biological sample. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a flowchart showing the method for predicting the ancestral origin category of an unknown biological sample provided by the present invention;

[0017] Figure 2 It is a schematic diagram showing the construction and application process of the ancestral origin category prediction model provided by the present invention;

[0018] Figure 3 It is a schematic diagram showing the 10-fold cross-validation provided by the present invention;

[0019] Figure 4 It is a schematic diagram showing the principal component analysis result provided by the present invention;

[0020] Figure 5 It is a broken line diagram showing the energy lithotripsy provided by the present invention;

[0021] Figure 6 It is a schematic diagram showing the population and individual ethnic components provided by the present invention;

[0022] Figure 7 It is a schematic diagram showing the prediction accuracy of the ancestral origin category prediction model under different principal component dimensions provided by the present invention;

[0023] Figure 8 It is a schematic diagram of the prediction accuracy of the ancestral origin category prediction model under different training rounds provided by the present invention;

[0024] Figure 9 It is a schematic structural diagram of the device for predicting the ancestral origin category of an unknown biological sample provided by the present invention;

[0025] Figure 10 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] To better understand the embodiments of the present invention, the background technology will be elaborated in detail first:

[0028] Optionally, the biological sample may be a single biological sample or multiple biological samples (such as a population sample).

[0029] The application of biogeographical ancestry prediction is very extensive. For example: identifying disease susceptibility markers, correcting population stratification, and explaining ethnic expansion, migration and mixing, etc. At present, in the field of deoxyribonucleic acid (DNA) ethnic inference, a system construction and inference method based on a small number of screened ancestry informative markers (AIMs) is used for the differentiation of intercontinental populations. However, AIMs need to be repeatedly screened and verified for specific populations, and there are also certain limitations in identifying fine substructures of populations with low-density genotyping data.

[0030] With the development of spatial population genetics, ancestry prediction often uses probability and statistical inference methods based on models to estimate the parameters of the assumed parametric model using known genotype data, and obtain the spatio-temporal distribution changes of the biological attributes of individuals or populations. Among them, a representative one is the spatial distribution model of gene frequencies, which is estimated by maximizing the likelihood function value or the approximate Bayesian method to obtain the spatial probability distribution of individuals with specified genotypes. However, the fitting of this spatial distribution model of gene frequencies takes a long time for the analysis of high-dimensional biological characteristics, and the complex model assumptions make the fitting more difficult and it is difficult to improve the prediction accuracy.

[0031] Machine Learning (ML)-based methods can avoid using hypothetical idealized parametric models, optimize the model through data learning, and do not require excessive biological prior knowledge. Machine learning methods have good adaptability to high-dimensional genomic data. In addition, as a new generation of machine learning methods, deep learning has been applied in fields such as population genetics research. However, due to the black box characteristics of deep learning, the interpretability of deep learning models is insufficient.

[0032] In existing methods for predicting the ancestral category of unknown biological samples, based on the experience explored by unsupervised learning methods (such as the PCA method) and combined with supervised learning methods (such as the linear regression method), the ancestral category of unknown biological samples is predicted. However, the dimensional features involved in this method are relatively single, resulting in a low accuracy rate for predicting the ancestral category of unknown biological samples.

[0033] To solve the above problems, the embodiments of the present invention provide a method, device, and electronic device for predicting the ancestral category of unknown biological samples, and establish a machine learning model (ancestral category prediction model) that combines the PCA algorithm of unsupervised learning and the XGBoost algorithm of supervised learning. Based on the high-density SNP genotyping data of unknown biological samples, which contains rich genetic information, the data is reduced in dimension to obtain principal component data of a higher dimension, which can reduce the amount of calculation while retaining more information. Then, in subsequent use of the ancestral category prediction model, the corresponding ancestry of the unknown biological sample can be accurately determined, effectively improving the accuracy rate of predicting the ancestral category of unknown biological samples.

[0034] It should be noted that the execution subject involved in the embodiments of the present invention can be a device for predicting the ancestral category of unknown biological samples, or an electronic device. Optionally, the electronic device may include: a computer, a mobile terminal, a wearable device, etc.

[0035] The following further illustrates the embodiments of the present invention by taking an electronic device as an example.

[0036] As Figure 1 shown, it is a schematic flowchart of the method for predicting the ancestral category of unknown biological samples provided by the present invention, which may include:

[0037] 101. Determine the principal component data of the unknown biological sample. The principal component data of the unknown biological sample is used to characterize the ethnic group characteristics of the unknown biological sample, and the data dimension of the principal component data of the unknown biological sample is greater than or equal to a preset dimension threshold.

[0038] Among them, the principal component (PC) data refers to the principal component data generated after the SNP genotyping data is reduced in dimension.

[0039] Ethnic characteristics refer to the common traits exhibited by an ethnic group due to its common history, sociocultural heritage, traditions, race, religion, language, and national identity, etc.

[0040] Optionally, the preset dimension threshold can be set before the electronic device leaves the factory or can be user-defined. For example, the preset dimension threshold is set to 40.

[0041] The electronic device can determine the principal component data of the unknown biological sample according to the SNP genotyping data of the unknown biological sample, so as to determine the ancestral origin prediction result subsequently.

[0042] 102. Input the principal component data of the unknown biological sample into the ancestral origin category prediction model to obtain the ancestral origin prediction result output by the ancestral origin category prediction model. The ancestral origin category prediction model is constructed based on the principal component data after dimensionality reduction of the genetic data of the reference population sample and the corresponding true ancestral origin information of the reference population sample.

[0043] Among them, the ancestral origin category prediction model is also called the PCA-XGBoost model.

[0044] Ancestral origin information is used to characterize the ancestral origin category.

[0045] It should be noted that the genetic data is the SNP genotyping data.

[0046] It should be noted that the XGBoost algorithm is a supervised learning model based on the ensemble decision tree and gradient boosting algorithm. New decision trees are added in the iteration to fit the residuals and fuse them into the original decision tree. Finally, the prediction scores of each category can be obtained at the leaf nodes of the decision tree.

[0047] The electronic device takes the principal component data of the unknown biological sample as the input of the ancestral origin category prediction model. Then, the output of the ancestral origin category prediction model is the ancestral origin prediction result corresponding to the unknown biological sample, and this ancestral origin prediction result is relatively accurate. The whole process effectively improves the accuracy of predicting the ancestral origin category of the unknown biological sample.

[0048] In some embodiments, the construction steps of the ancestral origin category prediction model are as follows: The electronic device determines the principal component data after dimensionality reduction of the genetic data of the reference population sample according to the single nucleotide polymorphism (SNP) genotyping data of the reference population sample; the electronic device inputs the principal component data after dimensionality reduction of the genetic data of the reference population sample into the original ancestral origin category prediction model to obtain the ancestral origin prediction result corresponding to the reference population sample output by the original ancestral origin category prediction model; the electronic device updates the model parameters of the original ancestral origin category prediction model according to the ancestral origin prediction result corresponding to the reference population sample and the true ancestral origin information corresponding to the reference population sample to obtain the trained ancestral origin category prediction model.

[0049] In the process of constructing an ancestral category prediction model, an electronic device first obtains the SNP genotyping data of a reference population sample, and can determine the principal component data after dimensionality reduction of the genetic data of the reference population sample based on the SNP genotyping data of the reference population sample; then, the electronic device uses the principal component data after dimensionality reduction of the genetic data of the reference population sample as the input of the original ancestral category prediction model, and the output of the original ancestral category prediction model is the ancestral prediction result corresponding to the reference population sample; then, the electronic device updates the model parameters of the original ancestral category prediction model according to the ancestral prediction result corresponding to the reference population sample and the true ancestral information corresponding to the reference population sample to obtain a trained ancestral category prediction model.

[0050] Among them, because the genetic information contained in the SNP genotyping data of the reference population sample is rich, the prediction accuracy of the trained ancestral category prediction model is relatively high.

[0051] Optionally, the above reference population sample may include 2,504 samples from 26 populations on 5 continents, sourced from the dataset of the 1000 Genomes project Phase 3 (1KG Phase 3).

[0052] Optionally, the above model parameters may include at least one of the following: learning rate (eta), the descent value of the minimum loss function required for node splitting (gamma), the regularization term of weight L2 (lambda), and the maximum depth of the tree (max_depth), etc.

[0053] Optionally, for the electronic device to update the model parameters of the original ancestral category prediction model, it may include: the electronic device uses a preset strategy (such as a greedy strategy) to update the model parameters, and comprehensively considers the model complexity and running efficiency to determine the updated model parameters.

[0054] Exemplarily, the updated model parameters are: eta: 0.007, gamma: 0.1, lambda: 2, max_depth: 12. In addition, the number of training rounds (round_num) is set to 1000, and the early stopping condition (early_stopping) is set to 100 (1 / 10 of round_num).

[0055] In some embodiments, for the electronic device to determine the principal component data after dimensionality reduction of the genetic data of the reference population sample based on the single nucleotide polymorphism SNP genotyping data of the reference population sample, it may include: the electronic device performs SNP locus screening on the SNP genotyping data of the reference population sample to obtain target loci; the electronic device determines the principal component data after dimensionality reduction of the genetic data of the reference population sample according to the target loci and the SNP genotyping data of the reference population sample.

[0056] In the process of determining the principal component data after dimensionality reduction of the genetic data of the reference population sample, the electronic device first obtains the SNP genotyping data of the reference population sample, screens the SNP sites of the SNP genotyping data of the reference population sample, and determines the screened SNP sites as the target sites; then, the electronic device can determine the principal component data after dimensionality reduction of the genetic data of the reference population sample according to the target sites and the SNP genotyping data of the reference population sample.

[0057] In some embodiments, the electronic device screens the SNP sites of the SNP genotyping data of the reference population sample to obtain the target sites, which may include: the electronic device determines the intersection sites between the global screening array chip and the set of genotyping sites of the target population; the electronic device takes the intersection of the intersection sites and the SNP genotyping data of the reference population sample based on the SNP sites to obtain the initial target sites; the electronic device screens the initial target sites according to the relevant parameters of the initial target sites to obtain the target sites, and the relevant parameters include at least one of the following: site detection rate, minor allele frequency (MAF), Hardy-Weinberg Equilibrium (HWE) test value, and Linkage Disequilibrium (LD) test value.

[0058] Among them, the Global Screening Array (GSA) chip refers to a genotyping chip.

[0059] The set of genotyping sites of the target population refers to the set of sites where the target population is detected to have genotyping.

[0060] Optionally, the target population can be the Chinese population. The set of genotyping sites of the target population can be the sites detected by the Chinese Genotyping Array (CGA) chip.

[0061] Exemplarily, in the process of obtaining the target sites, the electronic device first determines the intersection sites (a total of 553,849 SNP sites) between the global screening array chip and the set of genotyping sites of the target population. Then, based on the SNP sites, it takes the intersection of the SNP genotyping data of these intersection sites and the reference population samples to obtain the initial target sites (480,698 SNP sites). Next, according to the relevant parameters of the initial target sites, the electronic device can use Plink v1.9 software to screen the initial target sites: site detection rate (geno parameter < 1%), minor allele frequency (maf parameter < 0.01), Hardy-Weinberg equilibrium test value (p value > 10 -10 ) and linkage disequilibrium test value (indep-pairwise parameter: sliding window 500, step size 50 and measurement index r 2 > 0.8), to obtain the target sites (307,866 SNP sites).

[0062] In some embodiments, the electronic device determines the principal component data after dimensionality reduction of the genetic data of the reference population samples based on the target sites and the SNP genotyping data of the reference population samples, which may include: the electronic device determines the SNP genotyping data of the reference population samples after quality control based on the target sites and the SNP genotyping data of the reference population samples; the electronic device reduces the dimension of the SNP genotyping data of the reference population samples after quality control to obtain the principal component data after dimensionality reduction of the genetic data of the reference population samples.

[0063] In the process of determining the principal component data after dimensionality reduction of the genetic data of the reference population samples, the electronic device can determine the SNP genotyping data of the reference population samples after quality control based on the target sites and the SNP genotyping data of the reference population samples, and then use the PCA algorithm to reduce the dimension of the SNP genotyping data of the reference population samples after quality control to obtain the principal component data after dimensionality reduction of the genetic data of the reference population samples.

[0064] Optionally, the electronic device determines the SNP genotyping data of the reference population samples after quality control based on the target sites and the SNP genotyping data of the reference population samples, which may include: the electronic device obtains the SNP genotyping data of the test samples; the electronic device determines the intersection of the SNP genotyping data of the reference population samples and the SNP genotyping data of the test samples as the SNP genotyping data of the reference population samples after quality control based on the target sites.

[0065] Optionally, the above test samples may include 700 samples from 76 populations, specifically including: individuals of modern populations in the 1240K Version v54.1.p1 dataset of the The Allen Ancient DNA Resource (AADR), excluding individuals sourced from 1KGPhase3, and selecting populations from the country or region where the initial reference ethnic group data samples are located; some Han samples published by a certain laboratory (C.C.Wang Lab).

[0066] It should be noted that the reference population samples mentioned above include the genotype data (i.e., SNP genotyping data) and ancestral origin labels (i.e., true ancestral origins) of the samples; the above test samples include the genotype data of the samples and do not include ancestral origin labels.

[0067] In some embodiments, the electronic device inputs the principal component data of an unknown biological sample into the ancestral origin category prediction model, and obtains the ancestral origin prediction result output by the ancestral origin category prediction model, which may include: the electronic device uses the ancestral origin category prediction model to determine the probability values corresponding to the respective multiple ethnic group sources of the principal component data of the unknown biological sample; the electronic device determines the ancestral origin information corresponding to the maximum probability value among the multiple probability values as the ancestral origin prediction result corresponding to the unknown biological sample.

[0068] In the process of obtaining the ancestral origin prediction result, the electronic device inputs the principal component data of the unknown biological sample into the ancestral origin category prediction model, and obtains the probability values corresponding to the respective multiple ethnic group sources of the principal component data of the unknown biological sample, that is, the probability values of the unknown biological sample being predicted to the respective ancestral origin categories of multiple reference population samples; then, the electronic device sorts these multiple probability values (for example, sorts them in descending order of probability value), determines the maximum probability value (i.e., the probability value ranked first) from the sorted multiple probability values, and determines the ancestral origin information corresponding to the maximum probability value as the ancestral origin prediction result corresponding to the unknown biological sample.

[0069] In some embodiments, the method may further include: the electronic device determines the likelihood ratio (Likelihood Ratio, LR) between the maximum probability value and each other probability value; the electronic device determines the reliability degree of the ancestral origin prediction result based on at least one likelihood ratio.

[0070] After sorting the probability values corresponding to the respective multiple ethnic group sources of the principal component data of the unknown biological sample (for example, sorting them in descending order of probability value), the electronic device determines the ratio between the maximum probability value and each other probability value, that is, the likelihood ratio; then, the electronic device can determine the reliability degree of the ancestral origin prediction result based on at least one likelihood ratio.

[0071] Exemplarily, the electronic device inputs the principal component data obtained by dimensionality reduction of the genetic data of the reference population sample into the ancestral origin category prediction model, obtains the ancestral origin prediction result corresponding to the reference population sample output by the ancestral origin category prediction model, and compares the ancestral origin prediction result with the true ancestral origin (i.e., the ancestral origin label) corresponding to the reference population sample (divided into the intercontinental level and the population level), and counts the accuracy rate of the first prediction result (i.e., the ancestral origin information corresponding to the maximum probability value) and the prediction accuracy rate based on LR. Among them, LR x The value is the ratio between the probability value (1stProb) corresponding to the category of the first prediction result (1stPred) of a certain test sample and the probability value (xthProb) corresponding to the category of the xth (x>1) prediction result (xthPred). The LR x value of each test sample increases monotonically with x, and the LR x value can reflect the certainty of the ancestral origin category prediction model for the xthPred of the test sample.

[0072] Exemplarily, assume that the category of the kth (k≥1) prediction result kthPred is consistent with the true ancestral origin, and assume that m = max x (LR x ≤10). m represents how many ethnic groups in the ancestral origin prediction result have LR values less than or equal to the preset LR threshold (such as 10), LR x ≤10; LR x >10. The evaluation results based on the LR value can include the following 3 cases:

[0073] 1. Consistent Conclusion (CC), k = m = 1, that is, LR2>10, indicating that the ancestral origin category prediction model can determine that 1stPred is the correct ancestral origin category and exclude other ancestral origin categories;

[0074] 2. Inconclusive Conclusion (IC), m≥2, and k≤m, indicating that the correct ancestral origin category is included in the top m ancestral origin categories of the ancestral origin prediction result, and the conclusion is not to exclude these m ancestral origin categories;

[0075] 3. Error Conclusion (EC), k>m, indicating that the correct ancestral origin category is not included in the m possible ancestral origin categories predicted by the ancestral origin category prediction model, and the conclusion is that the prediction is incorrect.

[0076] Taking the test sample as an example, the evaluation indicators used are as follows:

[0077] 1. The accuracy of the 1st prediction (1stAcc), which is the traditional accuracy calculation method and is equivalent to the positive predictive value (PPV) of the ancestral category prediction model. It only calculates the accuracy of 1stPred and does not consider other possible results.

[0078] 2. The consistency rate (Cons), LR ) The higher this value is, the higher the accuracy and certainty of 1stPred of the ancestral category prediction model.

[0079] 3. The inconclusive rate (Incc), LR ) This value can reflect the uncertainty degree of the ancestral category prediction model for 1stPred, and at the same time can reflect the difficulty degree of the ancestral category prediction model in distinguishing a certain category from other categories.

[0080] 4. The accuracy (Acc), LR ) That is, Acc LR = Cons LR + Incc LR ; It measures the accuracy of the top several prediction results obtained by the ancestral category prediction model based on a certain reasonable degree of suspicion (determined by the preset LR threshold).

[0081] 5. The error rate (Err), LR ) That is, Err LR = 1 - Acc LR .

[0082] In an embodiment of the present invention, the principal component data of an unknown biological sample is determined. The principal component data of the unknown biological sample is used to characterize the ethnic group characteristics of the unknown biological sample, and the data dimension of the principal component data of the unknown biological sample is greater than or equal to a preset dimension threshold; the principal component data of the unknown biological sample is input into an ancestral origin category prediction model, and an ancestral origin prediction result output by the ancestral origin category prediction model is obtained. The ancestral origin category prediction model is constructed based on the principal component data after dimensionality reduction of the genetic data of the reference population sample and the corresponding true ancestral origin information of the reference population sample. This method establishes a machine learning model (ancestral origin category prediction model) that combines the PCA algorithm of unsupervised learning and the XGBoost algorithm of supervised learning. Based on the high-density SNP genotyping data of the unknown biological sample, which contains rich genetic information, the principal component data of a higher dimension is obtained by dimensionality reduction, which can reduce the computational amount while retaining more information. Then, in the subsequent use of the ancestral origin category prediction model, the ancestral origin corresponding to the unknown biological sample can be accurately determined, effectively improving the accuracy of the ancestral origin category prediction of the unknown biological sample.

[0083] The embodiments of the present invention will be further described in conjunction with the following several examples:

[0084] Exemplarily, the construction idea of the ancestral origin category prediction model (PCA-XGBoost) is to use the XGBoost classification model to perform ancestral origin category prediction with the principal component data after PCA dimensionality reduction as features. For population genetic structure analysis, the electronic device can use the GCTA v1.94 software to perform PCA analysis on the reference population sample, and use the ADMIXTURE v1.3.0 software to perform ethnic group component analysis on the reference population sample, and run 10-fold cross-validation, where the value of K is from 2 to 10. The electronic device can use the ggplot2 package of Rv3.14 to visualize the above results.

[0085] Exemplarily, based on the Python v3.9 environment, the electronic device uses the Plink v1.9 software to take the intersection of SNP loci of the reference population sample and the test sample. Then, the electronic device takes the first 10-dimensional principal component data after PCA analysis as the feature input of the XGBoost model, and the ancestral origin information (continent or population) of the reference population sample as the ancestral origin label. Using the scikit-learn v1.2.1 package, the random sampling method is adopted to divide the reference population sample into a training set and a validation set in a ratio of 8:2, which are used for model training and testing the iterative effect of the model respectively. Use the multi:softprob option of xgboost v1.7.3 for multi-classification, use the SoftMax function to calculate the probability that the test sample is predicted to each category, and calculate the likelihood ratio of the predicted category. Arrange the predicted result categories in descending order of probability values.

[0086] Exemplarily, such asFigure 2 As shown, it is a schematic diagram of the construction and application process of the ancestral origin category prediction model provided by the present invention. Figure 2 In it, Figure (A): The genotype data and ancestral origin labels of the reference population sample including samples of k ancestral origin categories at 308,766 SNP loci; the test sample does not include the ancestral origin label; Figure (B): The reference population sample and the test sample are combined (taking the SNP locus intersection); the first 10 PC data are retained during PCA dimensionality reduction; Figure (C): The principal component data obtained after dimensionality reduction include the values on each PC and the ancestral origin labels of the reference population sample; Figure (D): The 10-dimensional principal component data is used as the feature input of the original ancestral origin category prediction model, and the training set and the validation set are used iteratively to establish the ancestral origin category prediction model; Figure (E): Predict the test sample, and the output of the ancestral origin category prediction model is the probability (Prob1 to Prob k ) values and the likelihood ratio (LR1 to LR k ) values of the test sample being predicted to the respective ancestral origin categories of multiple reference population samples. In addition,

[0087] Among them, G represents guanine deoxynucleotide, C represents cytosine deoxynucleotide; A represents adenine deoxynucleotide; AA, GG, and CC all represent homozygous genotypes; AG represents heterozygous genotype.

[0088] Exemplarily, as Figure 3 shown, it is a schematic diagram of 10-fold cross-validation provided by the present invention. The electronic device can use the pandas v 1.5.3 package and scikit-learn v1.2.1, and according to the population ratio, use 10-fold cross-validation to sequentially extract 10% of the samples from the reference population sample as the validation set for model training and prediction, and finally obtain the test results of all samples in the reference population sample after integration.

[0089] Exemplarily, the electronic device respectively uses the first 5, 10, 20, 40, 80, 160, or 1233-dimensional PC data as the input features of the ancestral origin category prediction model, constructs and evaluates the ancestral origin category prediction model based on the reference population sample, calculates the prediction accuracy of the ancestral origin category prediction model under different PC dimensions to determine the optimal PC dimension. Then, on the basis of round_num = 1000 and early_stopping = 100, the number gradients of the training rounds are increased to 100, 300, 500, 2000, 3000, and 4000 (early_stopping is set to 1 / 10 of the corresponding round_num quantity), and the impacts of different training rounds on the accuracy and running time of the ancestral origin category prediction model are compared. The above process can be called the model optimization process.

[0090] Exemplarily, an optimized ancestral category prediction model is used to conduct tests based on test samples to verify the generalization ability of the ancestral category prediction model. Genotype data at target loci (a total of 307,866 SNP loci) are extracted from the test samples, and an ancestral category prediction model is constructed using 2,504 reference population samples to predict and evaluate 700 test samples.

[0091] The reference population samples involved in the embodiments of the present invention do not include all countries and regions within a continent. To study the accuracy and generalization ability of the optimized ancestral category prediction model, the 700 selected test samples are divided into two categories: (A) 108 test samples that are the same as the reference population samples in terms of the country or region of origin and highly similar in terms of population type (such as ethnicity, region, city, etc.); (B) the remaining 592 test samples are only the same as the reference population samples in terms of the country or region of origin.

[0092] For the ancestral prediction results of the test samples, the accuracy of the ancestral prediction results corresponding to these two types of test samples is analyzed and the differences are compared according to the above two types of test samples. For the intercontinental prediction results of all test samples and the ancestral prediction results of type A test samples, the evaluation methods and evaluation indicators described above are used; for the ancestral prediction results of type B test samples, the accuracy is statistically analyzed according to the following three situations: reference population samples predicted to the same country or region, other reference population samples within the continent, and reference population samples from other continents.

[0093] The experimental results involved in the embodiments of the present invention are further elaborated below:

[0094] 1. Population genetic structure research

[0095] The reference population samples include 2,504 samples from 26 populations on 5 continents, with a total of 307,866 SNP loci. First, PCA is performed to analyze population genetic clustering and genetic relationships. Exemplarily, as Figure 4 shown, it is a schematic diagram of the principal component analysis results provided by the present invention.

[0096] Figure 4Among them, EUR represents Europe; EAS represents East Asia; AFR represents Africa; AMR represents Americas; SAS represents South Asia; PC1 represents the first-dimensional principal component data; PC2 represents the second-dimensional principal component data; PC3 represents the third-dimensional principal component data; PC4 represents the fourth-dimensional principal component data; PC5 represents the fifth-dimensional principal component data; A1 to A5 successively represent 5 sub-populations of the EUR population; B1 to B5 successively represent 5 sub-populations of the EAS population; C1 to C7 successively represent 7 sub-populations of the AFR population; D1 to D4 successively represent 4 sub-populations of the AMR population; E1 to E5 successively represent 5 sub-populations of the SAS population.

[0097] As can be seen from Figure 4 it that the PCA results can, to a certain extent, reflect the genetic distance and genetic relationship formed under the influence of various factors such as geographical distance or gene flow among populations in the PC space. For example, the distribution of populations in the PC1-PC2 space can roughly reflect the distribution of populations in the global geographical space. Populations that are geographically far apart are usually located far away in the PCA graph, while populations that are geographically close usually show similar positions in the PCA graph. The formation of modern AMR populations, some admixed populations with AFR ancestry, etc. is related to gene flow among multiple populations and shows a trend of being located among multiple populations and extending towards EUR in the PCA graph. By synthesizing the results in the PC space of the first 5 dimensions, it can be found that the D4 and D1 sub-populations of the AMR population are close to the EUR population, while the D3 sub-population is the farthest from the EUR population. The C6 and C5 sub-populations of the AFR population extend towards the cluster corresponding to the EUR population to form an elongated cluster, located between the EUR population and other major AFR populations. In addition, the boundary between the AMR population and the AFR population appears to be relatively blurred, and a small number of samples of the D4 sub-population (about 4%) deviate significantly from the main part of the clustering and are close to the C6 and C5 sub-populations.

[0098] Examining the performance of PCs in differentiating populations, it can be found that in each two-dimensional PC space, it is relatively easy to distinguish the five continental populations. However, when further divided into 26 sub-populations, it is difficult to obtain clear and effective differentiation boundaries between some populations. But in different PC spaces, the overlap and separation of populations can vary greatly, that is, different PCs have different differentiation abilities for specific populations. For example, using the dimension of PC3 to differentiate the SAS population from other continental populations is very effective, and there is no overlap in the PC1-PC2 space; in addition, compared with PC1 to PC4, PC5 has better differentiation ability for the five sub-populations of the EAS population. This indicates that different PCs have different focuses in identifying population patterns. Some dimensions mainly reflect the relative geographical positions of the overall continental populations, while some dimensions are more focused on the differentiation between certain fine sub-populations. Therefore, incorporating more PCs helps to more comprehensively reflect population information, thereby achieving more refined population differentiation.

[0099] To effectively distinguish continental sub-populations and avoid incorporating too many PCs from affecting the running efficiency of the ancestral category prediction model, the embodiments of the present invention analyzed the individual explained variance (EV) and cumulative explained variance (CEV) of the first 40 PC dimensions. Exemplarily, as Figure 5 shown, is the energy lithotripsy broken line schematic diagram provided by the present invention. From Figure 5 it can be seen that after PC5, the EV is small and the decreasing trend is gentle. Therefore, we first tried to construct an ancestral category prediction model using the first 10-dimensional PCs.

[0100] To verify the genetic difference degree and population admixture among the reference population samples and analyze the genetic sub-structure of the population, ADMIXTURE was used to perform genetic ancestry component analysis of the world population (K = 2 to 10). The optimal K value with high individual component consistency within the population and low cross-validation (CV) error was selected as 5 (CV error = 0.45299). Exemplarily, as Figure 6As shown, it is a schematic diagram of population and individual ethnic group components provided by the present invention. Five ancestral components can explain the genetic ancestral components of the reference population samples, namely the ancestral components corresponding to red, green, blue, and purple that account for the largest proportions in the EUR, EAS, AFR, and SAS populations, respectively, and the yellow ancestral component that mainly exists in the AMR population, especially in the D3 sub-population. The main components of the EUR, EAS, and AFR populations account for the vast majority of their respective proportions (usually greater than 95%), and the consistency among their respective sub-populations is relatively high. The mixing situation of the AMR population is the most obvious. Except for the D3 sub-population, the other three sub-populations have relatively large amounts of EUR components (blue, with an average proportion of about 50% to 70%). The C5 and C6 sub-populations in the AFR population have mixed more EUR components (accounting for about 10%) compared to other sub-populations in the AFR population.

[0101] 2. Construction and evaluation of the ancestral origin category prediction model

[0102] The 10PC-XGBoost model was initially constructed using the first 10 dimensions of PCs. The prediction accuracy results at the continental level are shown in Table 1. In Table 1, each number represents the number of samples predicted to the corresponding continent, and the number in parentheses represents the number of samples with an inconclusive conclusion (IC).

[0103] Table 1:

[0104]

[0105] The prediction results at the continental level show that the first prediction accuracy (1stAcc), the consistency rate (Cons LR ) and the accuracy rate (Acc LR ) all reach more than 98%. Among them, the AMR population samples have the lowest values in the three prediction indicators. Three D4 sub-population samples were misjudged as having an AFR ancestral origin, and one sample each from the D2 and D1 sub-populations was misjudged as having a EUR ancestral origin. In addition, four C6 sub-population samples with an AFR ancestral origin were misjudged as having an AMR ancestral origin. The small number of mutually predictive confusion results in the test results of AFR and AMR are consistent with Figure 4 the relatively blurred boundary between the two in

[0106] The prediction accuracy results at the population level are shown in Table 2.

[0107] Table 2:

[0108]

[0109]

[0110] In Table 2, POP represents the abbreviation of population; Accurate POP represents the reference population sample predicted to the same country or region; Other POP Within Continent represents other reference population samples within the continent; POP in Other Continents represents the reference population samples in other continents. Each number in the grid represents the number of samples predicted to the corresponding category, and the number in parentheses represents the number of samples without excluding the conclusion (IC).

[0111] The prediction results at the population level show that the average accuracy rate is relatively high (97.16%), which is about 10% higher than the first prediction accuracy rate (87.02%). At the same time, referring to Table 3, it can be found that the cumulative prediction accuracy rate (In first 2P%) of the top 2 is as high as 98.4%. Among them, the prediction accuracy rate of the second place (2ndP%) accounts for about 11%, while the accuracy rate of the subsequent prediction results is relatively low (all less than 1%). This indicates that the 10PC-XGBoost model has more confusion in the identification of the population in the first two prediction results. The accuracy rate based on LR (97.16%) is about 10% higher than the first prediction accuracy rate (87.02%), approaching 98.4% of In first 2P%. That is, by not excluding all the predicted populations with LR < 10 (mainly not excluding 2nd Pred) as the true ancestral origin, the possibility of prediction errors can be reduced. However, the 10PC-XGBoost model has a relatively low agreement rate (Cons LR <70%) and a relatively high non-exclusion rate (Incc LR > 25%) at the population level, indicating that the current 10PC-XGBoost model has insufficient ability to distinguish sub-populations. Further analysis of the confusion matrix of the test results also found that there are often cases of mutual prediction confusion between two sub-populations within the continent. For example, for sub-populations A1 and A3 of the EUR population, the 10PC-XGBoost model is difficult to distinguish between sub-populations A1 and A3, and their Incc LR is greater than 90%; about 10% of the B2 sub-population of the EAS population is misclassified into the B3 sub-population, etc. This shows that the initial constructed 10PC-XGBoost model has insufficient discrimination fineness and needs to improve its ability to distinguish sub-populations within the continent.

[0112] Table 3:

[0113]

[0114]

[0115] 3. Optimization of the Ancestral Origin Category Prediction Model

[0116] The above results indicate that there may be factors that need to be optimized in the 10PC-XGBoost model, such as too few selected features, underfitting (e.g., insufficient number of training rounds), or insufficient model complexity (e.g., insufficient depth of classification trees).

[0117] The parameter optimization of the model first adjusts the number of PCs input to the XGBoost model. The XGBoost model is trained using the first 5, 10, 20, 40, 80, 160, or 1233-dimensional PC data respectively (the CEV of the first 1233 dimensions > 60%). Exemplarily, as Figure 7 shown, it is a schematic diagram of the prediction accuracy of the ancestral category prediction model under different principal component dimensions provided by the present invention. Figure 7 In [the figure], based on the broken lines of the first prediction accuracy and the agreement rate, it can be seen that as the PC dimension increases, the prediction accuracy rises and the growth rate significantly becomes smaller, and reaches a peak near 40PC to 80PC, and decreases after 80PC. Considering the time cost of model training comprehensively, and since the first prediction accuracy and the agreement rate do not increase significantly after 40PC, the first 40 dimensions of PCs are finally selected as the model features.

[0118] The accuracy results of the 40PC-XGBoost model at the continental level and population level after increasing the feature PC dimension are shown in Table 4 and Table 5 respectively. The first prediction accuracy, agreement rate, and accuracy at the continental level change little (although the agreement rate of the AMR population increases by about 0.2%), reflecting that increasing the PC dimension has approached the upper limit of the model on the reference population sample in terms of continental population differentiation. The optimized prediction results at the population level show that the average first prediction accuracy and the average agreement rate increase by about 5.3% and 8.3% respectively, indicating that the 1stPred of the optimized ancestral category prediction model (i.e., the 40PC-XGBoost model) is more accurate and reliable. The results show that more-dimensional PCs may contain more refined population structure information, which helps to distinguish and predict more refined sub-populations.

[0119] Table 4:

[0120]

[0121]

[0122] In Table 4, each number represents the number of samples predicted to the corresponding continent, and the number in parentheses represents the number of samples with an inconclusive conclusion (IC) among them.

[0123] Table 5:

[0124]

[0125]

[0126] In Table 5, each number in the grid represents the number of samples predicted for the corresponding category, and the number in parentheses represents the number of samples that do not rule out the conclusion (IC).

[0127] After increasing the PC dimension, the feature space processed by the ancestral origin category prediction model increases, and it may become insufficient in terms of training sufficiency, easily leading to underfitting of the model. Therefore, on the basis of the original 1000 training epochs (round_num), it is extended to a quantity gradient from 100 to 4000 to compare the impact of the model training epochs on the model accuracy.

[0128] Exemplarily, as Figure 8 shown, it is a schematic diagram of the prediction accuracy rate of the ancestral origin category prediction model under different training epochs provided by the present invention. Figure 8 Among them, the first prediction accuracy rate of the ancestral origin category prediction model slowly increases with the training epochs, and remains at a level between 92% and 93% in experiments with more than 500 epochs, while the prediction agreement rate significantly increases with the increase of the training epochs before 1000 epochs, and then remains at a level between 78% and 80%. On the contrary, AccLR slowly decreases in the range of 98% to 97%.

[0129] For the experiment with the minimum training epochs (100 epochs) designed in the embodiments of the present invention, the agreement rate is relatively low (18.8%), but the AccLR value is about 98.4%, and the 1stAcc also reaches more than 90%. That is, in the case of relatively low training epochs, the ancestral origin category prediction model already has a certain ancestral origin prediction ability, but the non-exclusion rate is relatively high (about 70%), which increases the result space and decision-making difficulty when considering LR. The increase in the training epochs will also lead to an increase in the test time. Generally, there is a positive linear relationship between the time cost and the training epochs. For example, the training time of 1000 epochs is about 10 times that of 100 epochs. Considering factors such as the time cost, accuracy rate, and non-exclusion rate of the model comprehensively, the embodiments of the present invention will adopt the ancestral origin category prediction model with 1000 training times (the early stopping condition is 100) in the model verification part based on the test samples to further verify the generalization ability.

[0130] 4. Verification of the ancestral origin category prediction model

[0131] Extract 307866 SNP locus data from the test samples. From the test samples sourced from AADR, 111025 SNP loci can be extracted. Among the partial Han samples sourced from a certain laboratory (C.C.Wang Lab): for a certain Han population, 306500 SNP loci can be extracted, and for the remaining populations, 42269 or 67366 SNP loci can be extracted.

[0132] The test results of 700 test samples at the intercontinental level are shown in Table 6. As can be seen from Table 6, the AccLR of the intercontinental populations except SAS is 100%. Among them, the ancestral category prediction model has the best prediction effect on the EAS population and the AFR population, and the consistency rates are both 100%.

[0133] In Table 6, each number in the grid represents the number of samples predicted to the corresponding continent, and the number in parentheses represents the number of samples with an inconclusive conclusion (IC).

[0134] Table 6:

[0135]

[0136] The test results of 108 Class A test samples highly similar to the reference population samples are shown in Table 7. As can be seen from Table 7, the consistency rate predicted by the ancestral category prediction model can reach 100% in more than half (8) of the test samples, and there are no test samples predicted to other continents. However, the consistency rate and accuracy rate of test sample c4 are 35% and 73.46% respectively, which are significantly lower than the test results of the C4 sub-population in the reference population samples (82.41% and 95.37%). Among the test samples of c4, 11 samples (about 42%) are predicted to be its neighboring C7 sub-population.

[0137] Table 7:

[0138]

[0139]

[0140] In Table 7, each number in the grid represents the number of samples predicted to the corresponding category, and the number in parentheses represents the number of samples with an inconclusive conclusion (IC).

[0141] Next, the ancestral category prediction device for unknown biological samples provided by the present invention will be described. The ancestral category prediction device for unknown biological samples described below can be correspondingly referred to the ancestral category prediction method for unknown biological samples described above.

[0142] As Figure 9 shown, it is a schematic structural diagram of the ancestral category prediction device for unknown biological samples provided by the present invention, which may include:

[0143] A principal component data determination module 901, configured to determine the principal component data of the unknown biological sample, where the principal component data of the unknown biological sample is used to characterize the ethnic characteristics of the unknown biological sample, and the data dimension of the principal component data of the unknown biological sample is greater than or equal to a preset dimension threshold;

[0144] The ancestral origin prediction result determination module 902 is configured to input the principal component data of the unknown biological sample into the ancestral origin category prediction model, and obtain the ancestral origin prediction result output by the ancestral origin category prediction model. The ancestral origin category prediction model is constructed based on the principal component data after dimensionality reduction of the genetic data of the reference population sample and the corresponding true ancestral origin information of the reference population sample.

[0145] Optionally, the construction steps of the ancestral origin category prediction model are as follows: Determine the principal component data after dimensionality reduction of the genetic data of the reference population sample according to the single nucleotide polymorphism (SNP) typing data of the reference population sample; Input the principal component data after dimensionality reduction of the genetic data of the reference population sample into the original ancestral origin category prediction model, and obtain the ancestral origin prediction result corresponding to the reference population sample output by the original ancestral origin category prediction model; Update the model parameters of the original ancestral origin category prediction model according to the ancestral origin prediction result corresponding to the reference population sample and the corresponding true ancestral origin information of the reference population sample, and obtain the trained ancestral origin category prediction model.

[0146] Optionally, the ancestral origin prediction result determination module 902 is specifically configured to perform SNP locus screening on the SNP typing data of the reference population sample to obtain target loci; Determine the principal component data after dimensionality reduction of the genetic data of the reference population sample according to the target loci and the SNP typing data of the reference population sample.

[0147] Optionally, the ancestral origin prediction result determination module 902 is specifically configured to determine the SNP typing data of the reference population sample after quality control according to the target loci and the SNP typing data of the reference population sample; Perform dimensionality reduction on the dimension of the SNP typing data of the reference population sample after quality control to obtain the principal component data after dimensionality reduction of the genetic data of the reference population sample.

[0148] Optionally, the ancestral origin prediction result determination module 902 is specifically configured to use the ancestral origin category prediction model to determine the probability values corresponding to the respective ethnic group origins of the principal component data of the unknown biological sample; Determine the ancestral origin information corresponding to the maximum probability value among the multiple probability values as the ancestral origin prediction result corresponding to the unknown biological sample.

[0149] Optionally, the ancestral origin prediction result determination module 902 is specifically configured to determine the likelihood ratio between the maximum probability value and each other probability value; Determine the reliability degree of the ancestral origin prediction result according to at least one likelihood ratio.

[0150] Optionally, the ancestral origin prediction result determination module 902 is specifically configured to determine the intersection sites between the global screening array chip and the set of genotyping sites of the target population; based on the SNP sites, take the intersection of the intersection sites and the SNP genotyping data of the reference population samples to obtain the initial target sites; screen the initial target sites according to the relevant parameters of the initial target sites, and obtain the target sites, where the relevant parameters include at least one of the following: site detection rate, minor allele frequency, Hardy-Weinberg equilibrium test value, and linkage disequilibrium test value.

[0151] As Figure 10 shown, it is a schematic structural diagram of an electronic device provided by the present invention. The electronic device may include: a processor 1010, a communication interface 1020, a memory 1030, and a communication bus 1040. Among them, the processor 1010, the communication interface 1020, and the memory 1030 complete mutual communication through the communication bus 1040. The processor 1010 may call the logical instructions in the memory 1030 to execute the method for predicting the ancestral origin category of an unknown biological sample.

[0152] In addition, when the logical instructions in the above-mentioned memory 1030 are implemented in the form of a software functional unit and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0153] On the other hand, the present invention also provides a computer program product, where the computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for predicting the ancestral origin category of an unknown biological sample provided by the above-mentioned methods.

[0154] On yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the method for predicting the ancestral origin category of an unknown biological sample provided by the above-mentioned methods.

[0155] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0156] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A method for predicting the ancestral category of an unknown biological sample, characterized in that, Including: Determine the principal component data of an unknown biological sample, where the principal component data of the unknown biological sample is used to characterize the ethnic group characteristics of the unknown biological sample, and the data dimension of the principal component data of the unknown biological sample is greater than or equal to a preset dimension threshold; Input the principal component data of the unknown biological sample into an ancestral origin category prediction model to obtain an ancestral origin prediction result output by the ancestral origin category prediction model. The ancestral origin category prediction model is constructed based on the principal component data after dimensionality reduction of the genetic data of a reference population sample and the corresponding true ancestral origin information of the reference population sample.

2. The method according to claim 1, characterized in that, The construction steps of the ancestral origin category prediction model are as follows: Determine the principal component data after dimensionality reduction of the genetic data of the reference population sample according to the single nucleotide polymorphism (SNP) genotyping data of the reference population sample; Input the principal component data after dimensionality reduction of the genetic data of the reference population sample into an original ancestral origin category prediction model to obtain an ancestral origin prediction result corresponding to the reference population sample output by the original ancestral origin category prediction model; Update the model parameters of the original ancestral origin category prediction model according to the ancestral origin prediction result corresponding to the reference population sample and the corresponding true ancestral origin information of the reference population sample to obtain a trained ancestral origin category prediction model.

3. The method according to claim 2, wherein The determining the principal component data after dimensionality reduction of the genetic data of the reference population sample according to the single nucleotide polymorphism (SNP) genotyping data of the reference population sample includes: Perform SNP locus screening on the SNP genotyping data of the reference population sample to obtain target loci; Determine the principal component data after dimensionality reduction of the genetic data of the reference population sample according to the target loci and the SNP genotyping data of the reference population sample.

4. The method according to claim 3, characterized in that The determining the principal component data after dimensionality reduction of the genetic data of the reference population sample according to the target loci and the SNP genotyping data of the reference population sample includes: Determine the SNP genotyping data of the reference population sample after quality control according to the target loci and the SNP genotyping data of the reference population sample; Reduce the dimension of the SNP genotyping data of the reference population sample after quality control to obtain the principal component data after dimensionality reduction of the genetic data of the reference population sample.

5. The method according to any one of claims 1-4, characterized in that, The inputting the principal component data of the unknown biological sample into an ancestral origin category prediction model to obtain an ancestral origin prediction result output by the ancestral origin category prediction model includes: Use the ancestral origin category prediction model to determine the probability values corresponding to each of the multiple ethnic group origins corresponding to the principal component data of the unknown biological sample; Determine the ancestral origin information corresponding to the maximum probability value among the multiple probability values as the ancestral origin prediction result corresponding to the unknown biological sample.

6. The method according to claim 5, wherein The method further includes: Determine the likelihood ratio between the maximum probability value and each other probability value; Determine the reliability degree of the ancestral origin prediction result according to at least one likelihood ratio.

7. The method according to claim 3, characterized in that The performing SNP locus screening on the SNP genotyping data of the reference population sample to obtain target loci includes: Determine the intersection loci between the global screening array chip and the target population genotyping locus set; Based on the SNP locus, taking the intersection of the SNP genotyping data of the intersection locus and the reference population sample to obtain an initial target locus; According to the relevant parameters of the initial target locus, screening the initial target locus to obtain the target locus, where the relevant parameters include at least one of the following: locus detection rate, minor allele frequency, Hardy-Weinberg equilibrium test value, and linkage disequilibrium test value.

8. An ancestral category prediction device for an unknown biological sample, characterized in that, Including: A principal component data determination module, configured to determine the principal component data of an unknown biological sample, where the principal component data of the unknown biological sample is used to characterize the ethnic characteristics of the unknown biological sample, and the data dimension of the principal component data of the unknown biological sample is greater than or equal to a preset dimension threshold; An ancestral origin prediction result determination module, configured to input the principal component data of the unknown biological sample into an ancestral origin category prediction model to obtain an ancestral origin prediction result output by the ancestral origin category prediction model, where the ancestral origin category prediction model is constructed based on the principal component data after dimensionality reduction of the genetic data of the reference population sample and the corresponding true ancestral origin information of the reference population sample.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for predicting the ancestral origin category of an unknown biological sample according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for predicting the ancestral origin category of an unknown biological sample according to any one of claims 1 to 7.