A DNA database retrieval method based on machine learning algorithms

By applying machine learning algorithms in the DNA database search method, the DNA typing and calculation probability ratios of different contributors were simulated, and the problem of insufficient retrieval efficiency and accuracy of large-scale DNA databases in the existing technology was solved, and efficient and accurate target individual screening was achieved.

CN119517181BActive Publication Date: 2025-06-17SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411499345.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-25
Publication Date
2025-06-17
Estimated Expiration
2044-10-25

AI Technical Summary

Technical Problem

When faced with a huge database, the computing speed and accuracy of existing DNA database search methods are difficult to meet the efficient and accurate screening needs of front-line investigation work.

Method used

Using a DNA database search method based on machine learning algorithm, we can improve retrieval efficiency and accuracy by collecting large-scale STR-DNA map datasets, calculating LR prior parameters, simulating DNA typing of different contributors, setting mutex assumptions, calculating probability ratios, obtaining eigenvalues ​​and normalizing processing, and using regression models for training and prediction.

Benefits of technology

It has achieved efficient and accurate screening of target individuals in large-scale DNA databases, improved search speed and accuracy, and met the needs of front-line investigation work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119517181B_ABST
    Figure CN119517181B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of database information analysis, and specifically relates to a DNA database retrieval method based on machine learning algorithms; collect information on the map dataset and calculate the LR parameter θ; simulate individuals with parent-child relationships, full-sibling relationships, and unrelated individuals for each known contributor respectively; compare each mixed DNA map with the genotypes of each candidate individual to obtain eigenvalue; obtain the training dataset and test dataset of features and labels; use the training dataset to train the regression model, and perform hyperparameter optimization and feature selection during the training process; use the test dataset as the input of all trained models to obtain the predicted values of all models; obtain the mixed DNA map from the scene, traverse each candidate individual in the DNA database in turn to calculate the eigenvalue corresponding to each candidate individual, and perform LR prediction to obtain the target individual; through the above method, the requirements of efficient and accurate screening are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of database information analysis, and in particular to a DNA database retrieval method based on a machine learning algorithm. Background Art

[0002] At present, my country's DNA database mainly uses STR (short tandem repeats) as genetic markers, which is an important tool for modern criminal investigation science. It aims to compare the DNA profile of biological samples obtained at the crime scene with the STR typing of the profile in the database, find the target individual, and provide clues for solving the case. At present, my country's forensic DNA database has a three-level network architecture of ministry-province-city, which is large in scale and provides strong data support for criminal investigation. With the continuous advancement of DNA testing technology and information technology, the amount of data in the DNA database continues to grow, which poses a challenge to retrieval efficiency. Efficient and accurate retrieval schemes have become a key issue to be urgently solved in the current application and development of DNA databases. The DNA profile at the crime scene may be a complex mixed profile, containing the STR typing of multiple contributors. Its retrieval strategy is to regard each individual in the DNA database as a candidate for potential contributor, and set a set of mutually exclusive hypotheses. By calculating the probability ratio under the two hypotheses, the likelihood ratio (LR), the association between each candidate individual and the mixed profile is evaluated. Among them, the plaintiff assumes that the candidate individual is a contributor to the mixed graph, and the defendant assumes that the candidate individual is not a contributor to the mixed graph. When LR is greater than 1, the plaintiff's hypothesis is supported, indicating that the candidate individual may be the target individual.

[0003] In view of the scientificity and practicality of this strategy, in the existing search, all candidate individuals in the database can be traversed, and the LR value of each individual in the mixed map can be calculated to exclude irrelevant individuals and narrow the target screening range, thereby improving the detection efficiency. To achieve this goal, a variety of optimization methods can be used to increase the search speed, such as streamlining the number of parameters required for LR calculation, limiting the possible combination range of genotypes of contributors to the mixed map, and using a hierarchical screening mechanism to gradually eliminate unlikely candidates, thereby accelerating the search process. Although the above method has shown significant advantages in improving detection efficiency and accuracy, when faced with a huge database of millions or even tens of millions of records, its computing speed and the accuracy of avoiding the erroneous inclusion of irrelevant individuals in the target range are still facing challenges, and it is difficult to meet the urgent needs of front-line detection work for efficient and accurate screening.

[0004] In summary, it is necessary to propose a DNA database retrieval method that meets the needs of efficient and accurate screening. Summary of the invention

[0005] The object of the present invention is to provide a DNA database retrieval method based on machine learning algorithms, aiming to solve the technical problem that in the prior art, when facing a huge database with millions or even tens of millions of records, the computing speed and the accuracy of avoiding misclassifying irrelevant individuals into the target range still face challenges, and it is difficult to meet the urgent needs of front-line investigation work for efficient and accurate screening.

[0006] To achieve the above object, a DNA database retrieval method based on machine learning algorithms adopted by the present invention includes the following steps:

[0007] Collect information on a large-scale STR-DNA map dataset; wherein the dataset includes single-source maps and mixed DNA maps, each map includes allele data and peak height data, and the mixed DNA map is jointly composed of the DNA of at least two contributors;

[0008] Using the single-source maps in the dataset, calculate the LR prior parameter θ;

[0009] For each mixed DNA map in the dataset, respectively simulate individuals with parent-child relationships, full-sibling relationships, and unrelated individuals of each known contributor;

[0010] For each mixed DNA map in the dataset, the mixed DNA map is composed of K known contributors. Traverse each candidate individual and set mutually exclusive hypothesis propositions H p 、H d ; Calculate the probability ratio LR under the two hypotheses, and take Log 10 (LR) as the label value;

[0011] Where H p means that the mixed DNA map is composed of a candidate individual and K-1 unknown unrelated individuals; H d means that the mixed DNA map is composed of K unknown unrelated individuals;

[0012] Compare the genotyping of each mixed DNA map with each candidate individual to obtain feature values, and normalize the feature values;

[0013] Obtain a dataset of features and labels, and divide this dataset into a training dataset and a test dataset;

[0014] Use a regression model, train the regression model using the training dataset, and perform hyperparameter optimization and feature selection during the training process;

[0015] Use the test dataset as the input of all trained models, obtain the predicted values of all models, and judge the target model;

[0016] Obtain the mixed DNA profile from the scene, traverse each candidate individual in the DNA database in turn, calculate the eigenvalue corresponding to each candidate individual, and perform Log 10 (LR) prediction to obtain the target individual.

[0017] Among them, in the steps of simulating individuals with parent-child relationship, full-sibling relationship, and unrelated individuals of each known contributor for each mixed DNA profile in the dataset, the definitions are as follows:

[0018] The candidate individuals include: K known contributors, K individuals with parent-child relationship with the known contributors, K individuals with full-sibling relationship with the known contributors, and K unrelated individuals;

[0019] The genotype of the known contributor is g, and the genotype g at locus m m ={a1, a2}, where a1 and a2 respectively represent the two alleles at locus m.

[0020] Among them, in the steps of simulating individuals with parent-child relationship, full-sibling relationship, and unrelated individuals of each known contributor for each mixed DNA profile in the dataset, the simulation process is as follows:

[0021] Randomly select an allele a1 at this locus, and randomly select an allele a3 at this locus from the population frequency data. Then, the genotype of the simulated individual with parent-child relationship with this contributor at locus m is g m_PO ={a1, a3};

[0022] Obtain the simulated father g of this contributor m_F ={a1, a3}, and randomly select an allele a4 at this locus from the population frequency data. Then, the simulated mother g of this contributor m_M ={a2, a4}; Randomly select one allele from the genotypes of the simulated father and the simulated mother to form the genotype g of the simulated individual with full-sibling relationship with this contributor at locus m m_FS ;

[0023] Randomly select two alleles at this locus from the population frequency data to form the genotype g of the unrelated individual at locus m m_UN ;

[0024] Traverse all loci to obtain the genotype g of the simulated individual with parent-child relationship with this contributor PO , the genotype g of the simulated individual with full-sibling relationship FS , and the genotype g of the unrelated simulated individual UN .

[0025] Among them, for each mixed DNA profile in the dataset, where the mixed DNA profile consists of K known contributors, traverse each candidate individual and set mutually exclusive hypothesis propositions H p and H d in the steps of:

[0026] Calculate the probability ratio LR under two hypotheses, and take Log 10 (LR) as the label value. Using the corresponding prior parameter θ, the probability ratio LR formula is:

[0027]

[0028] Among them, in the steps of comparing the typing of each mixed DNA profile with each candidate individual to obtain eigenvalue and normalizing the eigenvalue:

[0029] The eigenvalues include: the total number of alleles in the mixed profile, the total peak height of alleles in the mixed profile, the total number of matching alleles, the proportion of non-lost alleles, the maximum peak height of matching alleles, the minimum peak height of matching alleles, the peak height ratio of matching alleles, the p-value of matching alleles, the product of frequencies, the expected peak height, the coefficient of variation of peak height, and the peak height degradation parameter.

[0030] Among them, in the steps of taking the test dataset as the input of all trained models to obtain the predicted values of all models:

[0031] Compare the predicted values with the true values to judge the fitting effect, and obtain the target model according to the fitting effect.

[0032] Among them, in the steps of obtaining the mixed DNA profile from the scene, traversing each candidate individual in the DNA database in turn to calculate the corresponding eigenvalue for each candidate individual, and performing Log 10 (LR) prediction to obtain the target individual:

[0033] When the predicted Log 10 (LR) value > 0, the corresponding candidate individual is included in the target individual range;

[0034] When the predicted Log 10 (LR) value < 0, the corresponding candidate individual is excluded from the target individual range.

[0035] A method for retrieving a DNA database based on a machine learning algorithm of the present invention collects information on a large-scale STR-DNA map dataset; wherein the dataset includes single-source maps and mixed DNA maps, each map includes allele data and peak height data, and the mixed DNA map is composed of the DNA of at least two contributors; uses the single-source maps in the dataset to calculate the LR prior parameter θ; for each mixed DNA map in the dataset, respectively simulate individuals with parent-child relationships, full-sibling relationships, and unrelated individuals of each known contributor; for each mixed DNA map in the dataset, the mixed DNA map is composed of K known contributors, traverse each candidate individual, and set mutually exclusive hypothesis propositions H p 、H d ; where H p means that the mixed DNA map is composed of a candidate individual and K-1 unknown unrelated individuals; H d means that the mixed DNA map is composed of K unknown unrelated individuals; compare the genotyping of each mixed DNA map with that of each candidate individual to obtain eigenvalue, and normalize the eigenvalue; obtain a dataset of features and labels, divide the dataset into a training dataset and a test dataset, and normalize the features of the dataset; use a regression model, use the training dataset to train the regression model, and perform hyperparameter optimization and feature selection during the training process; use the test dataset as the input of all trained models, obtain the predicted values of all models, and judge the target model; obtain a mixed DNA map from the scene, sequentially traverse each candidate individual in the DNA database to calculate the corresponding eigenvalue of each candidate individual, and perform Log 10 (LR) prediction to obtain the target individual; meeting the requirements of efficient and accurate screening. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0037] Figure 1 is a flowchart of the steps of the DNA database retrieval method based on a machine learning algorithm of the present invention.

[0038] Figure 2 is a schematic diagram of a partial dataset of features and labels of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0039] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application.

[0040] Please refer to Figure 1 and Figure 2 , the present invention provides a method for retrieving a DNA database based on a machine learning algorithm, including the following steps:

[0041] Step 1: Collect information on a large-scale STR-DNA map dataset; wherein the dataset includes single-source maps and mixed DNA maps, each map includes allele data and peak height data, and the mixed DNA map is composed of the DNA of at least two contributors.

[0042] In this embodiment, information on a large-scale STR-DNA map dataset is collected; wherein the dataset includes single-source maps and mixed DNA maps, each map includes allele data and peak height data, and the mixed DNA map is composed of the DNA of at least two contributors. The experimental conditions for generating the mixed DNA map are: using different amounts of DNA templates, applying different STR kits, performing different PCR reaction conditions, and adopting different electrophoresis platforms and electrophoresis injection conditions; the DNA templates include: raw DNA templates without any treatment, DNA templates that have undergone degradation, and DNA templates with PCR inhibition factors.

[0043] Such as using the STR-DNA map online dataset PROVEDIT. This dataset contains single-source maps and mixed DNA maps. Each map contains allele data and peak height data. And the mixed map contains contributors of at least two people. The kits for generating the mixed DNA map include the Identifiler kit, the PowerPlex16 kit, the Fusion 6C kit, and the GlobalFiler kit. The experimental conditions are shown in Table 1. Its DNA templates cover various types.

[0044] Table 1 Prior parameter θ values of single-source maps under different experimental conditions

[0045]

[0046]

[0047] Step 2: Calculate the LR prior parameter θ using the single-source maps in the dataset.

[0048] In this embodiment, a single-source map in the dataset is used to calculate the LR prior parameter θ. Among them, the single-source maps generated by different kits, different cycling parameters, different electrophoresis platforms, and different injection times have different LR prior parameters θ.

[0049] For example, use the single-source map in PROVEDIT to calculate the prior parameter θ related to LR.

[0050] For example, θ = {AT, pC, λ}, where AT represents the analysis threshold (analysis threshold, AT), pC represents the probability of allele insertion, and λ represents the parameter of the exponential distribution followed by the inserted allele.

[0051] The calculation results are shown in Table 1, and the calculation method is as follows:

[0052] 1. For the single-source map generated with the same kit, the same cycling parameters, the same electrophoresis platform, and the same injection time, remove the OL peaks, the true allele peaks and their corresponding -1stutter peaks and +1stutter peaks in the single-source map. All the remaining peaks are noise peaks, and the 95% quantile of the empirical cumulative distribution of their peak values is AT;

[0053] 2.

[0054] Where n represents the number of noise peaks, N represents the number of single-source maps, and L represents the number of alleles in each single-source map;

[0055]

[0056] x i represents the peak height of the i-th noise peak.

[0057] Step 3: For each mixed DNA map in the dataset, simulate individuals with parent-child relationships, full-sibling relationships, and unrelated individuals of each known contributor respectively.

[0058] In this embodiment, the following are defined:

[0059] The candidate individuals include: K known contributors, K individuals with a parent-child relationship with the known contributors, K individuals with a full-sibling relationship with the known contributors, and K unrelated individuals; for each mixed genetic map E in the dataset, the number of contributors participating in the mixed genetic map is K. When K = 2, one individual with a parent-child relationship with each known contributor, one individual with a full-sibling relationship with each known contributor, and one unrelated individual are simulated respectively. Then the candidate individuals include: 2 known contributors, 2 individuals with a parent-child relationship with the known contributors, 2 individuals with a full-sibling relationship with the known contributors, and 2 unrelated individuals, a total of 8. When K = 3, there are 12 candidate individuals. And so on.

[0060] The genotype of the known contributor is g, and the genotype at locus m is g m = {a1, a2}, where a1 and a2 respectively represent the two alleles at locus m; taking locus D3S1358 as an example, the genotype at this locus is g m = {13, 15}

[0061] The simulation process is as follows:

[0062] Randomly select an allele 13 at this locus, and randomly select an allele 16 at this locus from the population frequency data. Then the genotype of the simulated individual with a parent-child relationship with this contributor at locus D3S1358 is g m_PO = {13, 16};

[0063] Obtain the simulated father g of this contributor according to the above steps m_F = {13, 16}, and randomly select an allele 18 at this locus from the population frequency data. Then the simulated mother g of this contributor m_M = {15, 18}; Randomly select one allele from the genotypes of the simulated father and the simulated mother respectively to form the genotype g of the simulated individual with a full-sibling relationship with this contributor at locus m m_FS = {15, 16};

[0064] Randomly select two alleles at this locus from the population frequency data to form the genotype g of the unrelated individual at locus m m_UN = {19, 20};

[0065] Traverse all loci to obtain the genotype g of the simulated individual with a parent-child relationship with this contributor PO , the genotype g of the simulated individual with a full-sibling relationship FS , and the genotype g of the unrelated simulated individual UN .

[0066] Step 4: For each mixed DNA profile in the dataset, where the mixed DNA profile consists of K known contributors, traverse each candidate individual and set mutually exclusive hypothesis propositions H p and H d ; where H p is that the mixed DNA profile consists of the candidate individual and K - 1 unknown unrelated individuals; H d is that the mixed DNA profile consists of K unknown unrelated individuals.

[0067] In this embodiment, for each mixed profile E, where the mixed profile E consists of K known contributors, traverse each candidate individual and set mutually exclusive hypothesis propositions H p and H d ; when the number of contributors K participating in the mixed profile is 2, H p is that the mixed profile consists of the candidate individual and 1 unknown unrelated individual; H d is that the mixed profile consists of 2 unknown unrelated individuals; when K = 3, H p is that the mixed profile consists of the candidate individual and 2 unknown unrelated individuals; H d is that the mixed profile consists of 3 unknown unrelated individuals; and so on. Calculate the probability ratio LR under the two hypotheses, and take Log 10 (LR) as the label value. The probability ratio LR formula is:

[0068]

[0069] Step 5: Compare the genotyping of each mixed DNA profile with each candidate individual to obtain feature values, and perform normalization processing on the feature values.

[0070] In this embodiment, for each mixed profile E in the dataset, compare it with the genotyping of each candidate individual to obtain the following feature values: total number of alleles in the mixed profile, total peak height of alleles in the mixed profile, total number of matching alleles, proportion of non-lost alleles, maximum peak height of matching alleles, minimum peak height of matching alleles, peak height ratio of matching alleles, p-value of matching alleles, product of frequencies, expected peak height, coefficient of variation of peak height, peak height degradation parameter; and perform normalization processing on the feature values;

[0071] where the definitions of the feature data are as follows:

[0072] Total number of alleles in the mixed profile: The sum of the number of alleles at all loci in the mixed profile.

[0073] Total peak height of alleles in the mixed profile: The sum of the peak heights of alleles at all loci in the mixed profile.

[0074] Matched allele: An allele that is the same at locus m in both the mixed profile and the candidate individual's genotype.

[0075] Total number of matched alleles: The sum of the number of matched alleles at all loci.

[0076] Proportion of alleles not lost: The total number of matched alleles divided by the total number of alleles of the candidate individual.

[0077] Maximum peak height of matched alleles: The maximum peak height corresponding to the matched alleles at all loci.

[0078] Minimum peak height of matched alleles: The minimum peak height corresponding to the matched alleles at all loci.

[0079] Proportion of peak heights of matched alleles: The sum of the peak heights of the matched alleles at all loci divided by the sum of the peak heights of the alleles in the mixed profile.

[0080] p-value of matched alleles: The negative of the probability of observing a result better than the number of matched alleles at all loci. The formula is:

[0081]

[0082] where c m represents the actual number of matched alleles at locus m, and C m represents the theoretical number of matched alleles observed at locus m, which can take values of 0, 1, or 2. p(C m ≥ c m ) represents the probability that C m is greater than c m , and M represents the total number of loci on the mixed DNA profile.

[0083] Product of frequencies: The negative of the product of the frequencies of the matched alleles and the frequencies of the non-matched alleles at all loci. The formula is:

[0084]

[0085] where p m represents the frequency of the first matched allele at locus m, and q m represents the frequency of the second matched allele at locus m; if there are no matched alleles, the result for this locus is 1.

[0086] Expected peak height, coefficient of variation of peak height, peak height degradation parameter: All are parameter values of the gamma model. The gamma model is used to fit the peak height data of the mixed profile, and these values are obtained after maximum likelihood estimation. The formula of the gamma model is as follows:

[0087]

[0088] Where Y represents the mean peak height of the mixed DNA profile, and μ, ω, and σ represent the peak height expectation, the peak height coefficient of variation, and the peak height degradation parameter, respectively; represents the probability of observing the peak height of the mixed DNA profile as under the condition of given μ, ω, and σ; represents the mean peak height of the alleles at locus m on the mixed DNA profile; represents the mean molecular weight of the alleles at locus m on the mixed DNA profile; M represents the total number of loci on the mixed DNA profile.

[0089] Step Six: Obtain a dataset of features and labels, and divide this dataset into a training dataset and a test dataset.

[0090] In this embodiment, a dataset of features and labels is obtained, and this dataset is divided into an 80% training dataset and a 20% test dataset; part of the dataset is as Figure 2 shown. Perform Z-Score standardization on the feature values in each group of features. The formula is z = (x i - μ) / σ, where z represents the standardized feature value, x i represents the i-th feature value in this group, μ represents the mean of the feature values in this group, and σ represents the variance of the feature values in this group.

[0091] Step Seven: Use a regression model, train the regression model using the training dataset, and perform hyperparameter optimization and feature selection during the training process.

[0092] In this embodiment, models such as ridge regression, lasso regression, K-nearest neighbor regression, random forest regression, support vector regression, and elastic net regression in machine learning are used. The above regression models are trained respectively using the aforementioned training dataset. The grid search method is used to find the optimal hyperparameter combination during the training process. Recursive feature elimination is used to select features.

[0093] Step Eight: Use the test dataset as the input of all trained models, obtain the predicted values of all models, compare the predicted values with the true values, judge the fitting effect, and obtain the target model according to the fitting effect.

[0094] In this embodiment, the test dataset is used as the input of all trained models, and the predicted values of all models are obtained. The predicted values are compared with the true values. According to the comparison results, it is judged which model has the highest fitting effect and reaches the highest prediction accuracy. The model with the highest fitting effect and the highest prediction accuracy is selected as the target model. The model fitting effect is evaluated by the coefficient of determination. R 2The larger it is, the better the model fitting effect indicates. The prediction accuracy uses the mean square error as the evaluation index. The smaller the MSE is, the higher the model prediction accuracy indicates. The 2 , and the MSE calculation formulas are as follows:

[0095]

[0096]

[0097] Among them, y i represents the true value, represents the predicted value, represents the mean of the true values, and m is the number of data in the prediction set. The R2 and MSE of each of the above models in the training set and the test set are shown in Table 2. Among them, the random forest regression has the largest 2 R in the training set and the prediction set, and the smallest MSE. Subsequently, this model is used to predict the Log 10 (LR) values in the database. After feature selection, this model uses the proportion of non-lost alleles, the total peak height of alleles in the mixed profile, the proportion of peak heights of matching alleles, the p-value of matching alleles, the total number of alleles in the mixed profile, and the coefficient of variation of peak height as features.

[0098] Table 2 R 2 and MSE of six models in the training set and the test set

[0099] Model <![CDATA[Training set R 2 > <![CDATA[Test set R 2 > MSE Ridge regression 0.68 0.71 18.05 Lasso regression 0.68 0.71 18.04 K-nearest neighbor regression 0.43 0.19 49.51 Random forest regression 0.98 0.85 6.32 Support vector regression 0.88 0.84 16.53 Elastic net regression 0.3 0.6 24.57

[0100] Step Nine: Obtain the mixed DNA profile from the scene, sequentially traverse each candidate individual in the DNA database to calculate the corresponding feature values of each candidate individual, and perform Log 10 (LR) prediction to obtain the target individual.

[0101] In this embodiment, the simulated mixed profile E' is an actual case, consisting of 2 contributor individuals. Using the Caucasian population data, two alleles are randomly selected from the population data at each locus to simulate unrelated individuals. Four DNA databases are respectively simulated, and each database contains 10,000, 100,000, 1,000,000, and 5,000,000 unrelated candidate individuals respectively. Calculate the corresponding feature values of the 2 true contributor individuals and the corresponding feature values of each candidate individual in each database, and use the model in Step Eight to perform Log 10 (LR) prediction.

[0102] If the predicted Log 10 (LR) value > 0, then the corresponding candidate individual is included in the target individual range. If the predicted Log 10If the (LR) value < 0, the corresponding candidate individual is excluded from the target individual range. Finally, N target individuals are obtained from the simulation database. The number of target individuals and the calculation speed of each simulation database are shown in Table 2. At the same time, the Log 10 (LR) predicted values are 39.82 and 12.43 respectively, and both are included in the target individual range. The corresponding true values are 41.52 and 10.44 respectively, with a small gap from the predicted values. The cpu used in the calculation process is i7–13700k.

[0103] Table 3 Number of target individuals screened from the simulation database and calculation speed table

[0104]

[0105] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the content disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present application.

[0106] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. A DNA database retrieval method based on a machine learning algorithm, characterized in that: The steps include: Collecting large-scale STR-DNA profile data set information; wherein the data set includes single-source profiles and mixed DNA profiles, each profile includes allele data and peak height data, and the mixed DNA profile is composed of DNA from at least two contributors; The single-source atlas in the dataset is used to calculate the LR prior parameter θ; For each mixed DNA profile in the data set, simulate the individuals with parent-child relationship, individuals with full sibling relationship, and unrelated individuals of each known contributor; For each mixed DNA profile in the data set, the mixed DNA profile consists of K known contributors, traverse each candidate individual, set a mutually exclusive hypothesis proposition H p , H d ; Calculate the probability ratio LR under the two hypotheses, and convert Log 10 (LR) as the label value; where H p The mixed DNA map consists of candidate individuals and K-1 unknown unrelated individuals; H d The mixed DNA profile consists of K unknown unrelated individuals; Compare each mixed DNA profile with the typing of each candidate individual to obtain characteristic values, and perform normalization on the characteristic values; Obtain a dataset of features and labels, and divide the dataset into a training dataset and a test dataset; Use the regression model, train the regression model using the training data set, and perform hyperparameter optimization and feature selection during the training process; Use the test data set as the input of all trained models, obtain the predicted values ​​of all models, and judge the target model; Obtain the mixed DNA map from the scene, traverse each candidate individual in the DNA database in turn, calculate the corresponding feature value of each candidate individual, and perform Log 10 (LR) prediction, get the target individual.

2. A DNA database retrieval method based on a machine learning algorithm as claimed in claim 1, characterized in that: In the step of simulating the parent-child relationship, full-sibling relationship, and unrelated individuals of each known contributor for each mixed DNA profile in the data set, the following definitions are made: The candidate individuals include: K known contributors, K individuals with parent-child relationships with known contributors, K individuals with full sibling relationships with known contributors, and K unrelated individuals; The known contributor is typed as g, and the genotype at site m is g m ={a1,a2}, where a1 and a2 represent the two alleles at site m respectively.

3. A DNA database retrieval method based on a machine learning algorithm as claimed in claim 2, characterized in that: In the step of simulating the parent-child relationship, full-sibling relationship, and unrelated individuals of each known contributor for each mixed DNA profile in the data set, the simulation process is as follows: Randomly select an allele a1 at the site, and randomly select an allele a3 at the site from the population frequency data. Then the simulated individual with a parent-child relationship with the contributor at site m is typed as g m_PO ={a1,a3}; Get the simulated parent of this contributor m_F ={a1,a3}, randomly select an allele a4 at the site from the population frequency data, then the simulated maternal g of the contributor m_M ={a2,a4}; randomly select one allele from each of the simulated father and simulated mother typing to form the typing g of the simulated individual with a full sibling relationship with the contributor at site m m_FS ; Randomly select two alleles at this site from the population frequency data to form the typing g of unrelated individuals at site m. m_UN ; Traverse all sites and obtain the simulated individual typing g with parent-child relationship with the contributor PO , the simulated individual typing g with full sibling relationship FS , irrelevant simulated individual typing g UN .

4. A DNA database retrieval method based on a machine learning algorithm as claimed in claim 3, characterized in that: For each mixed DNA profile in the data set, the mixed DNA profile consists of K known contributors, traverse each candidate individual, set the mutually exclusive hypothesis proposition H p , H d In the steps: The probability ratio LR formula is:

5. A DNA database retrieval method based on a machine learning algorithm as claimed in claim 4, characterized in that: In the step of comparing each mixed DNA profile with the typing of each candidate individual, obtaining characteristic values, and normalizing the characteristic values: The characteristic values ​​include: total number of mixed map alleles, sum of mixed map allele peak heights, total number of matching alleles, proportion of alleles not lost, maximum matching allele peak height, minimum matching allele peak height, proportion of matching allele peak height, matching allele p-value, frequency product, peak height expectation, peak height coefficient of variation, and peak height degradation parameter.

6. A DNA database retrieval method based on a machine learning algorithm as claimed in claim 5, characterized in that: In the step of using the test dataset as input to all trained models and obtaining the predicted values ​​of all models: Compare the predicted value with the true value, judge the fitting effect, and obtain the target model based on the fitting effect.

7. A DNA database retrieval method based on a machine learning algorithm as claimed in claim 6, characterized in that: After obtaining the mixed DNA map from the scene, we traverse each candidate individual in the DNA database in turn, calculate the corresponding feature value of each candidate individual, and perform Log 10 (LR) prediction, the steps to obtain the target individual: When predicting Log 10 If the (LR) value is > 0, the corresponding candidate individual is included in the target individual range; When predicting Log 10 If the (LR) value is <0, the corresponding candidate individual is excluded from the target individual range.

Citation Information

Patent Citations

  • Method and system for DNA mixture analysis

    US20020152035A1

  • System and method for the deconvolution of mixed DNA profiles using a proportionately shared allele approach

    US20090270264A1