Method, system, equipment and medium for identifying DNA proportion of sperm and yin mixed spot sample
By increasing the number of SNP sites and using the SNP sites of Y chromosome and autosome for sequencing, the problems of insufficient identification of DNA mixtures and limited degradation DNA analysis in biological examination materials analysis in sexual assault cases were solved, and efficient and accurate DNA proportion identification was achieved, improving the efficiency and accuracy of case detection.
Patent Information
- Application Number
- CN202510195921.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art has problems in the analysis of biological examination materials for sexual assault cases, such as difficulty in obtaining single-source DNA, insufficient identification of DNA mixtures, limited analysis of degraded DNA, and limited number of detection sites, which seriously affect the efficiency and accuracy of case investigation.
By increasing the number of SNP sites, using the SNP sites of the Y chromosome and autosome for sequencing, the homozygous rate of Y chromosome SNP and autosomal related sites are calculated, and the sample type and DNA proportion are judged, thereby improving the discrimination of DNA mixtures and the ability to degrade DNA.
It significantly improves the discrimination of DNA mixtures, can accurately judge the type of mixed samples and the proportion of DNA, and improves the efficiency and accuracy of case detection, especially when facing the degradation of DNA.
Smart Images

Figure CN120126554A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of forensic medicine, criminal investigation and computer technology, and particularly to a method, system, device and medium for identifying the DNA proportion in a semen-vaginal fluid mixed stain sample. Background Art
[0002] In the field of criminal cases, the severe situation of sexual assault crimes has become increasingly prominent. The physical and mental harm caused by such crimes to victims is long-term and profound, not only seriously damaging the physical and mental health of victims, but also posing a great threat to the normal social order. Therefore, quickly and accurately solving sexual assault cases is crucial for deterring lawbreakers and maintaining social stability.
[0003] In the process of solving sexual assault cases, the analysis of biological samples plays a key role. Common biological samples, such as vaginal / penile swabs, underwear, bed sheets, toilet paper, condoms, etc., are mostly mixtures of sperm cells and female epithelial cells, which are called mixed stains in forensic physical evidence. Although technology has advanced and the separation technology of semen-vaginal fluid mixed stains has been continuously improved, and the separation rate of sperm cells has also increased, there are still many problems in practical applications. On the one hand, the ratio of male and female cells in the sample is extremely different, and this huge ratio makes the separation work extremely difficult; on the other hand, some of the carriers attached to the mixed stains have strong adhesiveness, which further increases the difficulty of obtaining DNA from a single source. In most cases, even with advanced separation technology, it is difficult to obtain pure and single-source DNA, which severely restricts the subsequent DNA analysis work.
[0004] Currently, analyzing the DNA typing of semen-vaginal fluid mixed DNA by detecting genetic markers is a commonly used method. The genetic markers involved include short tandem repeats (STR), single nucleotide polymorphisms (SNP), microhaplotypes (MH), etc. These methods are based on the differences between men and women in different genetic marker systems, and by analyzing the peak height and peak area in the DNA mixed map, they infer the number of male individuals and their proportion in the mixed stain, so as to provide clues for solving cases. However, the existing technology has obvious limitations. First, the existing genetic marker detection methods have deficiencies in the discrimination ability of DNA mixtures, and it is difficult to accurately distinguish and judge the DNA information of different individuals in complex mixed samples, resulting in the inability to accurately lock suspects or determine the relationship between individuals in some cases. Second, when faced with degraded DNA, the accuracy and reliability of the analysis results are greatly reduced. Degraded DNA is relatively common in actual case samples. Due to environmental factors and other influences, DNA is prone to degradation, and the existing technology is difficult to effectively process such degraded DNA and cannot fully exploit the key information in it. In addition, the existing analysis methods generally have the problem of limited number of detection sites, unable to provide rich enough genetic information and difficult to meet the increasingly complex needs of case solving.
[0005] That is, the existing biological sample analysis technology for sexual assault cases has problems such as difficulty in obtaining single-source DNA, insufficient discrimination power for DNA mixtures, limited analysis of degraded DNA, and limited number of detection sites, which seriously affect the detection efficiency and accuracy of cases.
[0006] Focusing on the pain points of the existing technology, the present invention significantly improves the discrimination power of DNA mixtures by increasing the number of SNP sites. Especially in the analysis of degraded DNA, the present invention has incomparable advantages. With the continuous maturity of detection and sequencing technologies, it is expected to bring new breakthroughs to the forensic field and effectively solve the problems of the existing technology in the analysis of biological samples in sexual assault cases. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to provide a method, system, device and medium for identifying the DNA proportion in a semen-vaginal mixed stain sample, so as to solve the problems in the existing technology that the existing biological sample analysis technology for sexual assault cases has difficulties in obtaining single-source DNA, insufficient discrimination power for DNA mixtures, limited analysis of degraded DNA, and limited number of detection sites, which seriously affect the detection efficiency and accuracy of cases.
[0008] According to the first aspect of the embodiments of the present invention, a method for identifying the DNA proportion in a semen-vaginal mixed stain sample is provided. The method includes:
[0009] Screen SNP sites in a preset population;
[0010] Obtain a sequencing data sample of the target region of the mixed stain DNA to obtain the number of Y chromosome coverage sites; wherein, the number of Y chromosome coverage sites refers to the number of specific positions on the Y chromosome covered by the sequencing data in the obtained sequencing data sample of the target region of the mixed stain DNA;
[0011] Use the number of Y chromosome coverage sites to determine whether the corresponding sequencing data sample of the target region of the mixed stain DNA contains specific individual data;
[0012] If it contains specific individual data, use the SNP sites screened in the preset population to calculate the Y chromosome SNP homozygosity rate and determine the sample type;
[0013] Use the SNP sites screened in the preset population to obtain autosome-related sites, and use the autosome-related sites to determine whether the sequencing data sample of the target region of the mixed stain DNA is a mixed sample and the DNA proportion of the sequencing data sample of the target region of the mixed stain DNA;
[0014] Determine the gender of the main signal individual in a specific mixed sample by using the homozygosity rate of Y chromosome SNPs, and the results of whether the DNA target region sequencing data sample of the mixed stain is a mixed sample and the DNA proportion of the DNA target region sequencing data sample of the mixed stain.
[0015] Further, the use of the number of Y chromosome coverage sites to determine whether the corresponding DNA target region sequencing data sample of the mixed stain contains specific individual data includes:
[0016] If the number of Y chromosome coverage sites is less than the first threshold, the corresponding DNA target region sequencing data sample of the mixed stain does not contain male data, and male-female / male-male mixed samples can be excluded;
[0017] If the number of Y chromosome coverage sites is greater than or equal to the first threshold, the corresponding DNA target region sequencing data sample of the mixed stain contains male data.
[0018] Further, if it contains specific individual data, using the SNP sites screened in the preset population, calculating the homozygosity rate of Y chromosome SNPs and judging the sample type includes:
[0019] If it contains specific individual data, using the SNP sites screened in the preset population, obtain the number of sites with a minor allele frequency less than the second threshold among the number of Y chromosome coverage sites to obtain the total number of sites of homozygous SNP points on the Y chromosome;
[0020] Use the ratio of the total number of homozygous SNP points on the Y chromosome to the number of Y chromosome coverage sites to obtain the homozygosity rate of Y chromosome SNPs;
[0021] Use the homozygosity rate of Y chromosome SNPs to judge the sample type of the DNA target region sequencing data sample of the mixed stain. If the homozygosity rate of Y chromosome SNPs is greater than or equal to the third threshold,
[0022] then the mixed stain is a specific mixed sample; the specific mixed sample includes: male-female mixed sample or non-mixed sample;
[0023] Otherwise, it is a male-male mixed sample.
[0024] Further, using the SNP sites screened in the preset population to obtain autosome-related sites, and using the autosome-related sites to judge whether the DNA target region sequencing data sample of the mixed stain is a mixed sample and the DNA proportion of the DNA target region sequencing data sample of the mixed stain includes:
[0025] Use the SNP sites screened in the preset population to obtain autosome coverage sites;
[0026] Using the autosomal coverage sites, mark the sites with a sequencing depth greater than or equal to 50X as available SNP sites;
[0027] Using the autosomal coverage sites, mark the sites with a sequencing depth greater than or equal to 50X as available SNP sites, and the total number of such sites is denoted as A;
[0028] Obtain the sites with a minor allele frequency greater than 0.01 and less than 0.25, and mark them as mixed SNP sites containing mixed DNA information. The total number of such sites is denoted as H;
[0029] The sum of the minor allele frequencies is denoted as F;
[0030] Use the ratio of the total number of the mixed SNP sites containing mixed DNA information to the total number of the sites marked as available SNP sites with a sequencing depth greater than or equal to 50X to obtain the proportion of SNP sites containing mixed DNA information;
[0031] If the ratio of the proportion of SNP sites containing mixed DNA information is greater than or equal to the fourth threshold, it indicates that the sequencing data sample of the target region of the mixed stain DNA is a mixed sample;
[0032] Otherwise, it is determined as not a mixed sample;
[0033] Use the ratio of the sum of the minor allele frequencies F to the total number of the mixed SNP sites containing mixed DNA information to obtain the mean value of the secondary signals in the sequencing data sample of the target region of the mixed stain DNA, and use the mean value of the secondary signals in the mixed sample to judge the proportion of sample DNA.
[0034] Further, determining the gender of the main signal individual in a specific mixed sample by using the homozygosity rate of Y chromosome SNPs, and the results of whether the sequencing data sample of the target region of the mixed stain DNA is a mixed sample and the proportion of DNA in the sequencing data sample of the target region of the mixed stain DNA, including:
[0035] If the homozygosity rate of Y chromosome SNPs is greater than or equal to the sixth threshold and the proportion of SNP sites containing mixed DNA information is greater than or equal to the seventh threshold, it is determined as a specific mixed sample; the specific mixed sample includes: male-female mixed sample;
[0036] Obtain the total number of all covered sites on the Y chromosome, denoted as K, and use the preset first formula to obtain the proportion of effective sites on the Y chromosome. If the proportion of effective sites on the Y chromosome is greater than or equal to the eighth threshold, then the male is the main signal;
[0037] Otherwise, the female is the main signal;
[0038] D = M / K (1)
[0039] D is the proportion of valid sites on the Y chromosome; M represents the number of covered sites on the Y chromosome; K is the total number of covered sites on the Y chromosome.
[0040] According to the second aspect of the embodiments of the present invention, there is provided a system for identifying the DNA proportion of a semen-vaginal fluid mixed stain sample, which is applied to the method for identifying the DNA proportion of a semen-vaginal fluid mixed stain sample described in any one of the above. The system includes:
[0041] An acquisition module, configured to screen SNP sites in a preset population;
[0042] A first processing module, configured to obtain a sequencing data sample of the target region of the mixed stain DNA and obtain the number of covered sites on the Y chromosome; wherein, the number of covered sites on the Y chromosome refers to the number of specific positions on the Y chromosome that are covered by the sequencing data in the obtained sequencing data sample of the target region of the mixed stain DNA;
[0043] A second processing module, configured to use the number of covered sites on the Y chromosome to determine whether the corresponding sequencing data sample of the target region of the mixed stain DNA contains specific individual data;
[0044] A third processing module, configured to, if it contains specific individual data, use the SNP sites screened in the preset population to calculate the Y chromosome SNP homozygosity rate and determine the sample type;
[0045] A fourth processing module, configured to use the SNP sites screened in the preset population to obtain autosome-related sites, and use the autosome-related sites to determine whether the sequencing data sample of the target region of the mixed stain DNA is a mixed sample and the DNA proportion of the sequencing data sample of the target region of the mixed stain DNA;
[0046] A fifth processing module, configured to use the Y chromosome SNP homozygosity rate, and the results of whether the sequencing data sample of the target region of the mixed stain DNA is a mixed sample and the DNA proportion of the sequencing data sample of the target region of the mixed stain DNA, to determine the gender of the main signal individual in a specific mixed sample.
[0047] According to the third aspect of the embodiments of the present invention, there is provided an apparatus for identifying the DNA proportion of a semen-vaginal fluid mixed stain sample, including:
[0048] A memory, on which an executable program is stored;
[0049] A processor, configured to execute the executable program in the memory to implement the steps of the method described in any one of the above.
[0050] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the steps of any one of the above methods.
[0051] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0052] 1. High discrimination ability: Using SNP markers, it is suitable for degraded DNA with short amplification fragments, has a low mutation rate, and effectively improves the discrimination ability of DNA mixtures.
[0053] 2. Accurately judge the sample type: It can accurately judge the type of mixed samples, including specific mixed samples (male-female, male-male mixed samples) or non-mixed samples.
[0054] 3. High accuracy rate: It can detect more than 20,000 SNP sites, ensuring a high accuracy rate of the identification results.
[0055] 4. High efficiency and speed: The analysis time can be controlled within half an hour, which is convenient and fast, meeting the needs of actual case investigations.
[0056] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0058] Figure 1 is a schematic diagram of the steps of a method for identifying the DNA proportion of a semen-vaginal mixed stain sample shown according to an exemplary embodiment;
[0059] Figure 2 is a schematic diagram of the implementation process of a method for identifying the DNA proportion of a semen-vaginal mixed stain sample shown according to an exemplary embodiment;
[0060] Figure 3 is a schematic diagram of the system composition of a method for identifying the DNA proportion of a semen-vaginal mixed stain sample shown according to an exemplary embodiment;
[0061] Figure 4 is a schematic diagram of the device composition of a method for identifying the DNA proportion of a semen-vaginal mixed stain sample shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0063] Embodiment
[0064] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the steps of a method for identifying the DNA proportion of a semen-vaginal mixed stain sample according to an exemplary embodiment. The method includes:
[0065] S1. Screen SNP sites in a preset population;
[0066] S2. Obtain a sequencing data sample of the target region of the mixed stain DNA to obtain the number of Y-chromosome covered sites; wherein, the number of Y-chromosome covered sites refers to the number of specific positions on the Y chromosome covered by the sequencing data in the obtained sequencing data sample of the target region of the mixed stain DNA;
[0067] S3. Use the number of Y-chromosome covered sites to determine whether the corresponding sequencing data sample of the target region of the mixed stain DNA contains specific individual data;
[0068] S4. If it contains specific individual data, use the SNP sites screened in the preset population to calculate the Y-chromosome SNP homozygosity rate and determine the sample type;
[0069] S5. Use the SNP sites screened in the preset population to obtain autosomal related sites, and use the autosomal related sites to determine whether the sequencing data sample of the target region of the mixed stain DNA is a mixed sample and the DNA proportion of the sequencing data sample of the target region of the mixed stain DNA;
[0070] S6. Use the Y-chromosome SNP homozygosity rate, and the results of whether the sequencing data sample of the target region of the mixed stain DNA is a mixed sample and the DNA proportion of the sequencing data sample of the target region of the mixed stain DNA to determine the gender of the main signal individual in a specific mixed sample.
[0071] In specific implementation, the purpose of the embodiment of the present invention is to provide a method for identifying the DNA proportion of a semen-vaginal mixed stain sample, which can be applied to criminal cases related to sexual assault crimes and the field of forensic medicine. By using a high-throughput sequencing platform to construct a library and analyze SNP sets, it can accurately identify whether the sample is a mixed sample, and if it is a mixed sample, what are the signal proportions of male-female / male-male / female-female respectively.
[0072] Specifically, the principle of identifying the DNA proportion in the sperm-vaginal fluid mixed stain sample is as follows:
[0073] 1. Screen for dimorphic SNP loci that are widely distributed across autosomes and sex chromosomes in the Chinese population. The sex chromosomes include the X chromosome and the Y chromosome;
[0074] 2. Obtain the sequencing data of the target region of the mixed stain DNA, and count the number of covered sites on the Y chromosome, denoted as M;
[0075] 3. If M is less than 2, then the mixed stain does not contain male data, and male-female / male-male mixed samples can be excluded;
[0076] 4. If M is greater than or equal to 2, then the mixed stain contains male data. Count the number of sites among the M sites where the minor allele frequency is less than 0.01, and mark them as homozygous SNP sites on the Y chromosome. The total number of such sites is denoted as M'. Let m = M' / M. m is the homozygosity rate of the SNPs on the Y chromosome. Through this homozygosity rate, it can be determined whether this sample is a male-female mixed sample or a male-male mixed sample;
[0077] 5. Count the covered sites on autosomes, mark the sites with a sequencing depth greater than or equal to 50X as available SNP sites, and the total number of such sites is denoted as A; count the sites among them where the minor allele frequency is greater than 0.01 and less than 0.25, and mark them as mixed sample SNP sites containing mixed DNA information. The total number of such sites is denoted as H, and the total sum of the minor allele frequencies is denoted as F;
[0078] 6. Let I = H / A. I is the proportion of SNP sites containing mixed DNA information. Through this proportion, it can be determined whether this sample is a mixed sample;
[0079] 7. Let J = F / H. J is the mean value of the secondary signals in the mixed sample. Through this value, it can be determined the DNA proportion of this sample. Specific implementation method
[0081] Data processing and analysis steps:
[0082] 1. Data quality control and preprocessing: Use the fastp software to perform quality control and preprocessing on the original off-machine data, filter low-quality sequences, and process adapters.
[0083] 2. Data alignment: Use the BWA software to align the clean data obtained after quality control to the human reference genome.
[0084] 3. File sorting and index building: Use the sort tool in the Samtools software to sort the aligned bam file and build an index.
[0085] 4. Generate a pileup format file: Use the mpileup command of Samtools, taking the sorted bam file as input information to generate an output file in pileup format.
[0086] 5. Obtain the VCF format result file: Use the mpileup2cns command of VarScan, taking the pileup as the input file to obtain a VCF format result file containing specific SNP information.
[0087] 6. Count the number of covered sites on the Y chromosome: Count the number of covered sites on the Y chromosome, denoted as M; if M is less than 2, it is determined that the mixed stain does not contain specific individual data, and specific mixed sample types can be excluded; if M is greater than or equal to 2, the mixed stain contains specific individual data. Count the number of sites with a minor allele frequency less than 0.01 among the M sites, marked as homozygous SNP sites on the Y chromosome, and the total number of such sites is denoted as M'. m = M' / M, where m is the homozygosity rate of SNPs on the Y chromosome. Based on this homozygosity rate, it can be determined whether this sample is a specific mixed sample (male-female mixed sample) or other mixed samples (male-male mixed sample). If m is greater than or equal to 0.99, the mixed stain is a specific mixed sample (male-female mixed sample) or an unmixed sample, otherwise it is other mixed samples (male-male mixed sample). Whether it is a mixed sample can be identified through subsequent steps.
[0088] 7. Count the covered sites on autosomes: Count the covered sites on autosomes, mark the sites with a sequencing depth greater than or equal to 30X as available SNP sites, and the total number is denoted as A; count the sites with a minor allele frequency greater than 0.01 and less than 0.25 among them, marked as mixed SNP sites containing mixed DNA information, and the total number of such sites is denoted as H, and the total sum of minor allele frequencies is denoted as F.
[0089] 8. Determine whether the sample is a mixed sample: I = H / A, where I is the proportion of SNP sites containing mixed DNA information. Based on this proportion, it can be determined whether this sample is a mixed sample. If I is greater than or equal to 0.05, it is a mixed sample, otherwise it is determined to be an unmixed sample.
[0090] 9. Determine the DNA proportion of the sample: J = F / H, where J is the mean value of secondary signals in the mixed sample. Based on this value, the DNA proportion of this sample can be determined. If J is 0.05, the proportion of the main signal in the mixed sample is approximately 95%, and the proportion of the secondary signal is approximately 5%.
[0091] 10. Determine the gender of the individual with the main signal in a specific mixed sample: If m is greater than or equal to 0.99 and I is greater than or equal to 0.05, it is determined to be a specific mixed sample (male-female mixed sample). It is necessary to determine whether the main signal is a specific individual (male) or another gender (female). Count the total number of covered sites on the Y chromosome, denoted as K. D = M / K, where D is the proportion of valid sites on the Y chromosome. If D is greater than or equal to 0.5, the specific individual (male) is the main signal, otherwise the other gender (female) is the main signal.
[0092] The analysis includes multiple SNP loci, and some of the locus information is shown in Table 1 below:
[0093] Table 1
[0094]
[0095]
[0096] In specific implementation, this information is the specific data of some loci selected from numerous SNP loci in the analysis, and they are of great significance for judging the nature of the semen-vaginal fluid mixed stain sample and the DNA proportion.
[0097] SNP locus: Each locus has a specific number, such as rs61766321, which is the identifier for distinguishing different SNP loci. Just like everyone's ID number, it is used for precise positioning and identification of specific single nucleotide polymorphism positions in genetic analysis.
[0098] Chromosome: It indicates the chromosome where the SNP locus is located. For example, chr1 represents chromosome 1. Different chromosomes carry different gene information. Sex chromosomes (chrX, chrY) are related to gender, and autosomes (chr1-chr22) contain a large number of genes related to various characteristics of an individual. By analyzing the SNP locus information on different chromosomes, multi-faceted genetic information of the sample can be obtained.
[0099] Chromosome coordinate: It precisely indicates the specific position of the SNP locus on the corresponding chromosome in digital form. For example, 1004202 represents the position of the rs61766321 locus on chromosome 1. This helps to precisely locate the target SNP locus in the vast genome and provides an accurate coordinate for subsequent analysis.
[0100] Minor allele frequency: It refers to the frequency of the allele with a lower frequency at a certain locus in the population. Taking the 0.49 of the rs61766321 locus as an example, it means that at this locus, the minor allele appears in the population at a frequency of 49%. In the identification of the DNA proportion of the semen-vaginal fluid mixed stain sample, this frequency can be used to judge whether the sample is a mixed sample and calculate the DNA proportion in the mixed sample. For example, when counting the autosomal coverage loci, SNP loci containing mixed DNA information (frequency greater than 0.01 and less than 0.25) need to be screened according to the minor allele frequency, and then relevant ratios (such as I, J values) are calculated to determine the nature of the sample and the DNA proportion situation.
[0101] Example 1
[0102] A mixed stain sample from underwear, labeled L1QZ1E, was subjected to DNA proportion identification of sperm-vagina mixed stain samples. In the SNP genotyping results of L1QZ1E obtained through sequencing analysis, no Y chromosome information was detected. The proportion of SNP sites containing mixed DNA information was 0.001, and the average value of secondary signals in the mixed sample was 0.0002. Detection conclusion: The sample labeled L1QZ1E only contains DNA information of one individual and is a female sample.
[0103] In specific implementation, it was mentioned above that whether the sample contains male data was judged by counting the number of Y chromosome coverage sites. In Example 1, no Y chromosome information was detected in the L1QZ1E sample, which is consistent with the judgment basis in the technical solution that "if M is less than 2, then this mixed stain does not contain male data, and male-female / male-male mixed samples can be excluded". At the same time, according to the method in the technical solution for judging whether the sample is a mixed sample by the proportion of SNP sites containing mixed DNA information (I value), the I value of this sample is 0.001, which is less than 0.05, and it is determined that the sample is not mixed. This series of judgment steps fully follow the identification principle set by the technical solution.
[0104] From the specific implementation steps, Example 1 carried out a complete analysis process on the sample, that is, operations such as using the fastp software for quality control and preprocessing, the BWA software for data alignment, the Samtools software for sorting and building an index and generating a pileup file, and the VarScan software for obtaining SNP information, etc. Finally, key data such as Y chromosome information, the proportion of SNP sites, and the average value of secondary signals were obtained. These data were obtained based on the specific implementation steps in the technical solution, and then the detection conclusion of "only contains DNA information of one individual and is a female sample" was obtained according to the judgment criteria in the technical solution, fully verifying the feasibility and accuracy of the technical solution in actual sample detection.
[0105] Example 2
[0106] A mixed stain sample from a condom, labeled L1QZ2E, was subjected to DNA proportion identification of sperm-vagina mixed stain samples. In the SNP genotyping results of L1QZ2E obtained through sequencing analysis, Y chromosome information was detected. The homozygosity rate m of the SNPs on the Y chromosome was 100%, the proportion of SNP sites containing mixed DNA information was 0.58, the average value of secondary signals in the mixed sample was 0.2046, and the proportion of effective sites on the Y chromosome was 0.9842. Detection conclusion: The sample labeled L1QZ2E contains DNA information of two individuals. The main signal is a male sample (about 79.54%), and the secondary signal is a female sample (about 20.46%).
[0107] More specifically, when the Y chromosome information is detected and the number of covered Y chromosome sites M ≥ 2, the sample situation needs to be further analyzed. The Y chromosome information is detected in this sample, and the SNP homozygosity rate m of the Y chromosome is 100% (i.e., m ≥ 0.99), which meets the preliminary judgment conditions for male-female mixed samples or non-mixed samples in the technical solution. At the same time, the proportion I of SNP sites containing mixed DNA information is 0.58 (I ≥ 0.05), and it is determined as a mixed sample according to the technical solution. Considering these two key indicators, it is determined that this sample is a male-female mixed sample.
[0108] DNA proportion judgment: The above technical solution judges the DNA proportion through the mean value J of the secondary signals in the mixed sample. In Example 2, J is 0.2046, which means that the proportion of secondary signals is about 20.46%, so the proportion of primary signals is about 1 - 0.2046 = 79.54%.
[0109] Primary signal gender judgment: The technical solution states that for male-female mixed samples, the gender of the primary signal needs to be judged according to the proportion D of effective sites on the Y chromosome. The proportion D of effective sites on the Y chromosome in this sample is 0.9842 (D ≥ 0.5), so it is judged that the primary signal is a male sample and the secondary signal is a female sample.
[0110] Example 2 fully demonstrates the whole process from sample data acquisition to the final identification conclusion. All judgments strictly follow the process and standards of the technical solution, indicating that this technical solution can accurately judge the type of mixed stain samples and the DNA proportion of each gender individual in practical applications, providing a reliable technical means for sample analysis in criminal cases related to sexual assault crimes and the field of forensic medicine.
[0111] Example 3
[0112] A mixed stain sample from toilet paper, labeled L1QZ3E, is subjected to DNA proportion identification of sperm-vaginal mixed stain samples. In the SNP genotyping results of L1QZ3E obtained through sequencing analysis, the Y chromosome information is detected, the SNP homozygosity rate m of the Y chromosome is 100%, the proportion of SNP sites containing mixed DNA information is 0.0002, the mean value of secondary signals in the mixed sample is 0.0001, and the proportion of effective sites on the Y chromosome is 0.9999. Detection conclusion: The sample labeled L1QZ3E only contains the DNA information of one individual and is a male sample.
[0113] In specific implementation, according to the technical solution, when the Y chromosome information is detected and the number of covered Y chromosome sites M ≥ 2, the sample needs to be further analyzed. The Y chromosome information is detected in this sample, and the SNP homozygosity rate m of the Y chromosome is 100% (m ≥ 0.99), which preliminarily indicates that the sample may be a male-female mixed sample or a non-mixed sample, and specific judgment also needs to combine other indicators.
[0114] Judgment of mixed samples: In the technical solution, the proportion I of SNP sites containing mixed DNA information is used to determine whether a sample is a mixed sample. If I≥0.05, it is a mixed sample; otherwise, it is not. The I value of this sample is 0.0002, which is less than 0.05. According to the technical solution, it is determined as not a mixed sample. This indicates that there is only DNA information of one individual in the sample.
[0115] Judgment of sample gender: Since Y chromosome information is detected in the sample and it is determined as not a mixed sample, it can be determined that this sample is a male sample. The proportion of valid sites of the Y chromosome is 0.9999. This data further corroborates the accuracy of the Y chromosome-related information in the sample and also conforms to the characteristics of male samples.
[0116] In Example 3, from the detection data of the sample to the final conclusion, it is completely carried out according to the identification principle and steps set by the technical solution. This not only demonstrates the operability of the technical solution in actual sample detection but also verifies its effectiveness in accurately identifying the DNA proportion and judging the nature of samples of mixed stains from different sources, providing strong technical support for forensic medicine and criminal investigation practice.
[0117] Example 4
[0118] A mixed stain sample from a bedsheet, labeled L1QZ4E, is subjected to DNA proportion identification of sperm-vaginal mixed stain samples. In the SNP typing results of L1QZ4E obtained by sequencing analysis, Y chromosome information is detected. The homozygosity rate m of the SNPs on the Y chromosome is 100%. The proportion of SNP sites containing mixed DNA information is 0.44. The average value of the secondary signals in the mixed sample is 0.1257. The proportion of valid sites of the Y chromosome is 0.2561. Detection conclusion: For the sample labeled L1QZ4E, there is DNA information of two individuals. The main signal is a female sample (about 87.43%), and the secondary signal is a male sample (about 12.57%).
[0119] In specific implementation, the proportion (I value) of SNP sites containing mixed DNA information is statistically analyzed to determine whether a sample is a mixed sample. If I≥0.05, it is a mixed sample. The I value of the L1QZ4E sample is 0.44, which is greater than 0.05. So it is determined that this sample is a mixed sample, that is, there is DNA information of two or more individuals.
[0120] Judgment of male-female mixed samples: When Y chromosome information is detected and the number of covered sites M of the Y chromosome ≥2, the SNP homozygosity rate m of the Y chromosome needs to be calculated. The m of this sample is 100%, that is, m≥0.99, which meets the preliminary judgment conditions for male-female mixed samples or non-mixed samples. Combining the result that it has been determined as a mixed sample, it can be judged that this sample is a male-female mixed sample.
[0121] Determination of DNA proportion: According to the technical solution, the mean value J of the secondary signal in the mixed sample can be used to judge the DNA proportion of the sample. The J value of the L1QZ4E sample is 0.1257, which means that the proportion of the secondary signal is about 12.57%, so the proportion of the main signal is about 1 - 0.1257 = 87.43%.
[0122] Determination of the gender of the main signal: For a mixed sample of male and female, it is necessary to judge the gender of the main signal according to the proportion D of the effective sites of the Y chromosome. If D≥0.5, then the male is the main signal, otherwise the female is the main signal. The D value of this sample is 0.2561, which is less than 0.5, so the main signal is a female sample and the secondary signal is a male sample.
[0123] This case fully demonstrates the process of using the technical solution for sample analysis. From judging whether the sample is mixed, to determining the mixed sample of male and female, to calculating the DNA proportion and judging the gender of the main signal, each step is strictly carried out according to the technical solution, verifying the accuracy and reliability of the technical solution in practical applications, and being able to provide accurate DNA analysis results for the detection of sexual assault cases.
[0124] Example 5
[0125] A mixed stain sample from a car seat, labeled L1QZ5E, was subjected to DNA proportion identification of a semen-vaginal fluid mixed stain sample. In the SNP typing results of L1QZ5E obtained by sequencing analysis, Y chromosome information was detected. The homozygosity rate m of the SNP of the Y chromosome was 67.51%, the proportion of SNP sites containing mixed DNA information was 0.3644, the mean value of the secondary signal in the mixed sample was 0.2499, and the proportion of effective sites of the Y chromosome was 100%. Detection conclusion: The sample labeled L1QZ5E contains DNA information of two individuals. The main signal is a male sample (about 75.01%), and the secondary signal is a male sample (about 24.99%).
[0126] In specific implementation, it is judged whether the sample is a mixed sample by calculating the proportion (I value) of SNP sites containing mixed DNA information. The I value of the L1QZ5E sample is 0.3644. Since 0.3644≥0.05, according to the standard of the technical solution, it can be determined that this sample is a mixed sample, that is, there is DNA information of two or more individuals in the sample.
[0127] Judgment of sample type: When Y chromosome information is detected and the number of covered sites M of the Y chromosome ≥ 2, it is necessary to calculate the SNP homozygosity rate m of the Y chromosome. The m of this sample is 67.51%. Since m<0.99, according to the technical solution, this sample can be judged as a male-male mixed sample.
[0128] Determination of DNA proportion: In the technical solution, the mean value J of the secondary signals in the mixed sample can be used to judge the DNA proportion of each signal in the sample. The J value of the L1QZ5E sample is 0.2499, which indicates that the proportion of the secondary signal is approximately 24.99%. Then the proportion of the main signal is approximately 1 - 0.2499 = 75.01%.
[0129] Comprehensive judgment of gender and proportion: Since it has been determined that this sample is a male-male mixed sample, combined with the calculated DNA proportion, it can be concluded that there is DNA information of two individuals in this sample. The main signal is a male sample (about 75.01%), and the secondary signal is a male sample (about 24.99%).
[0130] From the acquisition of sample data to the conclusion of the final identification in this embodiment, it strictly follows the process of the technical solution for identifying the DNA proportion of sperm-vaginal mixed stain samples. This not only demonstrates the accuracy of this technical solution in dealing with complex mixed samples (male-male mixed), but also further verifies its reliability in practical applications, and can provide effective technical support for the analysis of biological samples in sexual assault cases.
[0131] Example 6
[0132] A mixed stain sample from underwear, labeled L1QZ6E, was subjected to DNA proportion identification of sperm-vaginal mixed stain samples. In the SNP typing results of L1QZ6E obtained by sequencing analysis, no Y chromosome information was detected, the proportion of SNP sites containing mixed DNA information was 0.2978, and the mean value of the secondary signals in the mixed sample was 0.0678. The detection conclusion: The sample labeled L1QZ6E has DNA information of two individuals. The main signal is a female sample (about 93.22%), and the secondary signal is a female sample (about 6.78%).
[0133] In specific implementation, when no Y chromosome information is detected, the possibility that the sample contains male data is initially excluded, that is, the situation of male-female / male-male mixed samples is excluded, and it is speculated that the sample may be a female sample or a female-female mixed sample.
[0134] Judgment of mixed sample: The technical solution judges whether the sample is a mixed sample by calculating the proportion (I value) of SNP sites containing mixed DNA information. The I value of the L1QZ6E sample is 0.2978. Since 0.2978 ≥ 0.05, according to the technical solution, this sample is determined to be a mixed sample, which indicates that there is DNA information of two or more individuals in the sample. Combining the situation of no Y chromosome information detected, it is determined that this sample is a female-female mixed sample.
[0135] Determination of DNA proportion: In the technical solution, the mean value J of the secondary signal in the mixed sample can be used to judge the DNA proportion of the sample. The J value of the L1QZ6E sample is 0.0678, which means that the proportion of the secondary signal is about 6.78%, so the proportion of the main signal is about 1 - 0.0678 = 93.22%.
[0136] Based on the above analysis, there is DNA information of two individuals in this sample, and both are female. The main signal is from a female sample (about 93.22%), and the secondary signal is from a female sample (about 6.78%).
[0137] This embodiment fully demonstrates the identification process of the technical solution when dealing with a mixed stain sample without detected Y chromosome information. From the judgment of whether the sample is mixed, to determining the sample type as female-female mixture, and then to calculating the DNA proportion, each step is strictly carried out according to the technical solution, verifying the accuracy and comprehensiveness of the technical solution in practical applications, and being able to effectively analyze and identify various complex sperm-vaginal mixed stain samples, providing a reliable basis for forensic medicine and criminal investigation.
[0138] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the composition of a system for identifying the DNA proportion of a sperm-vaginal mixed stain sample shown according to an exemplary embodiment. The system includes:
[0139] An acquisition module 30 for screening SNP sites in a preset population;
[0140] A first processing module 31 for obtaining a sequencing data sample of the target region of the mixed stain DNA to obtain the number of Y chromosome coverage sites; wherein, the number of Y chromosome coverage sites refers to the number of specific positions on the Y chromosome covered by the sequencing data in the obtained sequencing data sample of the target region of the mixed stain DNA;
[0141] A second processing module 32 for using the number of Y chromosome coverage sites to judge whether the corresponding sequencing data sample of the target region of the mixed stain DNA contains specific individual data;
[0142] A third processing module 33 for, if it contains specific individual data, using the SNP sites screened in the preset population to calculate the Y chromosome SNP homozygosity rate and judge the sample type;
[0143] A fourth processing module 43 for using the SNP sites screened in the preset population to obtain autosome-related sites, and using the autosome-related sites to judge whether the sequencing data sample of the target region of the mixed stain DNA is a mixed sample and the DNA proportion of the sequencing data sample of the target region of the mixed stain DNA;
[0144] The fifth processing module 35 is configured to determine the gender of the main signal individual in a specific mixed sample by using the Y-chromosome SNP homozygosity rate, and the results of whether the DNA target region sequencing data sample of the mixed stain is a mixed sample and the DNA proportion of the DNA target region sequencing data sample of the mixed stain.
[0145] Please refer to Figure 4 , Figure 4 which is a schematic diagram of the composition of a device for identifying the DNA proportion of a semen and vaginal fluid mixed stain sample according to an exemplary embodiment. The device includes:
[0146] A memory 41, on which an executable program is stored;
[0147] A processor 42, configured to execute the executable program in the memory 41 to implement the steps of the method described in any one of the above.
[0148] In addition, the present application provides a computer-readable storage medium storing computer instructions for causing a computer to execute the steps of the method described in any one of the above. Wherein, the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (abbreviation: HDD) or a solid-state drive (SSD), etc.; the storage medium may also include a combination of the above types of memories.
[0149] It can be understood that the same or similar parts in the above embodiments can be referred to each other, and the content not detailed in some embodiments can be seen in the same or similar content of other embodiments.
[0150] It should be noted that in the description of the present invention, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "a plurality" refers to at least two.
[0151] Any process or method description in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process, and the scope of the preferred embodiments of the present invention includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the technical field of the embodiments of the present invention.
[0152] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0153] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program, and the said program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0154] In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0155] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.
[0156] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0157] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for identifying the proportion of DNA in a spermatozoon mixed spot sample, characterized in that: The method comprises: Screening of SNP sites in a preset population; Acquire a mixed spot DNA target region sequencing data sample to obtain the number of Y chromosome coverage sites; wherein the number of Y chromosome coverage sites refers to the number of specific positions on the Y chromosome covered by the sequencing data in the acquired mixed spot DNA target region sequencing data sample; Using the number of sites covered by the Y chromosome, determine whether the corresponding mixed spot DNA target region sequencing data sample contains specific individual data; If specific individual data is included, the SNP sites in the preset population are screened, the Y chromosome SNP homozygosity rate is calculated and the sample type is determined; Using the SNP sites in the preset screening population, autosomal related sites are obtained, and using the autosomal related sites, whether the mixed spot DNA target region sequencing data sample is a mixed sample and the DNA proportion of the mixed spot DNA target region sequencing data sample is determined; The sex of the main signal individual in a specific mixed sample is determined by using the homozygous rate of the Y chromosome SNP, whether the mixed spot DNA target area sequencing data sample is a mixed sample, and the result of the DNA proportion of the mixed spot DNA target area sequencing data sample.
2. The method according to claim 1, characterized in that The method of using the number of sites covered by the Y chromosome to determine whether the corresponding mixed spot DNA target region sequencing data sample contains specific individual data includes: If the number of Y chromosome coverage sites is less than the first threshold, the corresponding mixed spot DNA target region sequencing data sample does not contain male data, and male-female / male-male mixed samples can be excluded; If the number of sites covered by the Y chromosome is greater than or equal to the first threshold, the corresponding mixed spot DNA target region sequencing data sample contains male data.
3. The method according to claim 1, characterized in that If the specific individual data is included, the SNP sites in the screening preset population are used to calculate the Y chromosome SNP homozygosity rate and determine the sample type, including: If specific individual data is included, the SNP sites in the preset population are screened, the number of sites whose minor allele frequency is less than the second threshold in the number of sites covered by the Y chromosome is obtained, and the total number of sites of homozygous SNP points on the Y chromosome is obtained; The total number of homozygous SNP points of the Y chromosome is compared with the number of covered sites of the Y chromosome to obtain the homozygous rate of the SNP of the Y chromosome; The sample type of the mixed spot DNA target region sequencing data sample is determined by using the homozygosity rate of the SNP of the Y chromosome. If the homozygosity rate of the SNP of the Y chromosome is greater than or equal to the third threshold, The mixed spot is a specific mixed sample; the specific mixed sample includes: a male and female mixed sample or an unmixed sample; Otherwise, it is a mixed male-male sample.
4. The method according to claim 1, characterized in that: The method of using the SNP sites in the screening preset population to obtain autosomal related sites, and using the autosomal related sites to determine whether the mixed spot DNA target region sequencing data sample is a mixed sample and the DNA proportion of the mixed spot DNA target region sequencing data sample, includes: Using the SNP sites in the screening preset population to obtain autosomal coverage sites; Using the autosomal coverage sites, sites with a sequencing depth greater than or equal to 50X are marked as available SNP points; Using the autosomal coverage sites, sites with a sequencing depth greater than or equal to 50X are marked as available SNP sites, and the total number of sites is recorded as A; The sites with the minor allele frequency greater than 0.01 and less than 0.25 were obtained and marked as mixed SNP sites containing mixed DNA information, and the total number of sites was recorded as H; The sum of minor allele frequencies is recorded as F; The total number of sites of the mixed sample SNP points containing mixed DNA information is compared with the total number of sites marked as available SNP points at the sites with a sequencing depth greater than or equal to 50X to obtain the proportion of SNP sites containing mixed DNA information; If the SNP site proportion ratio containing mixed DNA information is greater than or equal to the fourth threshold, it means that the mixed spot DNA target area sequencing data sample is a mixed sample; Otherwise, it is judged as unmixed sample; The sum of the minor allele frequencies F and the total number of sites of the mixed SNP points containing mixed DNA information are used to obtain the mean of the minor signals in the sequencing data sample of the mixed spot DNA target area, and the mean of the minor signals in the mixed sample is used to determine the proportion of sample DNA.
5. The method according to claim 1, characterized in that The method of using the homozygosity rate of the Y chromosome SNP, whether the mixed spot DNA target region sequencing data sample is a mixed sample, and the DNA proportion of the mixed spot DNA target region sequencing data sample to determine the gender of the main signal individual in a specific mixed sample includes: If the Y chromosome SNP homozygosity rate is greater than or equal to the sixth threshold and the proportion of SNP sites containing mixed DNA information is greater than or equal to the seventh threshold, it is determined to be a specific mixed sample; the specific mixed sample includes: a male and female mixed sample; The number of all covered sites of the Y chromosome is obtained, recorded as K, and the effective site ratio of the Y chromosome is obtained by using the preset first formula. If the effective site ratio of the Y chromosome is greater than or equal to the eighth threshold, the male is the main signal; Otherwise, females are the dominant signal; D = M / K (1) D is the percentage of effective sites on chromosome Y; M represents the number of sites covered on chromosome Y; K is the total number of sites covered on chromosome Y.
6. A system for identifying the proportion of DNA in a sperm-yin mixed spot sample, applied to a method for identifying the proportion of DNA in a sperm-yin mixed spot sample as claimed in any one of claims 1 to 5, characterized in that: The system comprises: An acquisition module is used to screen SNP sites in a preset population; The first processing module is used to obtain a mixed spot DNA target region sequencing data sample and obtain the number of Y chromosome coverage sites; wherein the number of Y chromosome coverage sites refers to the number of specific positions on the Y chromosome covered by the sequencing data in the obtained mixed spot DNA target region sequencing data sample; The second processing module is used to use the number of sites covered by the Y chromosome to determine whether the corresponding mixed spot DNA target region sequencing data sample contains specific individual data; The third processing module is used to calculate the Y chromosome SNP homozygosity rate and determine the sample type by using the SNP sites in the preset population screened if specific individual data is included; A fourth processing module is used to obtain autosomal related sites by using the SNP sites in the preset screening population, and to determine whether the mixed spot DNA target region sequencing data sample is a mixed sample and the DNA proportion of the mixed spot DNA target region sequencing data sample by using the autosomal related sites; The fifth processing module is used to determine the gender of the main signal individual in a specific mixed sample by using the homozygosity rate of the Y chromosome SNP, whether the mixed spot DNA target area sequencing data sample is a mixed sample, and the DNA proportion of the mixed spot DNA target area sequencing data sample.
7. The device for identifying the DNA ratio of spermatozoa mixed stain samples is characterized by: include: a memory having an executable program stored therein; A processor, configured to execute the executable program in the memory to implement the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the steps of the method according to any one of claims 1 to 5.