A method and system for parentage testing based on SNP marker combinations
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 厦门赛尔吉亚医学检验所有限公司
- Filing Date
- 2026-01-16
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]例如某家庭产后因亲子关系存疑提出鉴定申请,采用传统方法对被检对象与可疑对象的样本进行检测后,由于检测所涉及的遗传位点信息有限,且缺乏对检测数据的针对性校准环节,受样本检测过程中可能出现的质量波动等因素影响,最终得出的鉴定结果未能有效化解当事人的疑虑,甚至引发后续争议
因采用以SNP标记组合为核心,结合多维度质控过滤去除碱基质量低于Q30、长度异常及含过多N碱基的低质量数据、人类参考基因组比对与重复序列去除、SNP分型数据的多维特征空间映射及形状直径函数动态校准、测序深度与等位基因频率0.3到0.7范围双重筛选,以及孟德尔遗传规律位点比例与遗传相似度双指标综合判断的技术手段,克服了传统亲子鉴定方法检测位点有限、数据质量控制不全面、SNP分型准确性易受干扰、判断依据单一导致的精准度与可靠性不足的技术问题,进而达到了提升亲子鉴定结果的精准度、稳定性和可重复性,确保鉴定结论科学严谨且具有强说服力的技术效果。
Smart Images

Figure CN121538327B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biological detection technology, and in particular to a paternity testing method and system based on SNP marker combinations. Background Technology
[0002] In the field of paternity testing, traditional paternity testing methods mostly rely on specific short tandem repeat sequence markers for detection. Although such methods have been applied to some extent in routine scenarios, their inherent limitations have gradually become apparent as the demand for accuracy and reliability of results in postnatal paternity testing continues to increase.
[0003] For example, a family may request a paternity test after childbirth due to doubts about the parentage of their child. Using traditional methods to test samples from both the tested individual and the suspected individual, the results may fail to effectively address the family's concerns and even lead to further disputes. This is because traditional methods have insufficient coverage of genetic loci, making it difficult to comprehensively capture the genetic connections between parents and children. Furthermore, the lack of precise optimization mechanisms for raw data results in inaccurate and unconvincing results in complex family scenarios or when sample quality is slightly compromised, failing to meet the demands of higher standards of paternity testing. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a paternity testing method and system based on SNP marker combinations, so as to improve the scientificity and reliability of paternity testing.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, a paternity testing method based on SNP marker combinations, the method comprising: Library construction and sequencing were performed on samples from the tested object and the suspected object respectively to obtain the raw sequencing data of the samples; By splitting the raw sequencing data, individual sample data of the tested object and the suspected object are obtained separately; by quality control filtering of the individual sample data, low-quality sequencing data is removed to obtain quality-controlled data; The quality-controlled data is compared with the human reference genome to obtain sequence alignment data; the sequence alignment data is then deduplicated to eliminate duplicate sequences, resulting in deduplicated data. By performing SNP typing analysis on the deduplicated data, SNP typing data within a preset target area can be obtained; The SNP genotyping data is mapped to feature vectors in a multidimensional feature space. A multidimensional reference structure is constructed and analysis partitions are divided. The feature vectors are classified into the corresponding analysis partitions. The dynamic calibration coefficients of the feature vector distribution are calculated based on the shape diameter function. The SNP genotyping data is calibrated using the dynamic calibration coefficients to obtain optimized SNP genotyping data. By filtering the optimized SNP genotyping data, SNP sites that meet the sequencing depth requirements and whose allele frequencies are in the middle range in the reference population database are selected to obtain the final analysis data for comparison. Based on the final analysis data, the common SNP loci between the tested subjects and the suspected subjects were counted, the proportion of loci that conform to Mendelian inheritance laws was calculated, and a comprehensive analysis was conducted in combination with genetic similarity. A preset threshold was used to determine whether there was a parent-child relationship between the two.
[0006] Furthermore, by splitting the raw sequencing data, individual sample data of the tested subjects and suspected subjects are obtained separately; by performing quality control filtering on the individual sample data to remove low-quality sequencing data, the quality-controlled data is obtained, including: Based on the sample identification information in the original sequencing dataset, index matching and data separation are performed on the original sequencing data, and sequencing fragments corresponding to the tested object and the suspected object are extracted to obtain the initial individual sample data. By performing quality assessment on the initial individual sample data, the base quality score, sequence length distribution, and sequencing depth coverage of each sequencing fragment are calculated to obtain a quality assessment report; Based on the quality assessment report, a quality filtering threshold is set to remove sequencing fragments with a base quality score below Q30, sequences with abnormal lengths, and low-quality data containing too many N bases, thus obtaining high-quality sequencing data after filtering. The filtered high-quality sequencing data is integrated and recombined to generate data quality statistics, thus obtaining quality-controlled data.
[0007] Furthermore, the quality-controlled data is compared with the human reference genome to obtain sequence alignment data; the sequence alignment data is then deduplicated to eliminate repetitive sequences, resulting in deduplicated data, including: The quality-controlled data is compared with the human reference genome, an index is built based on the human reference genome, and sequence alignment is performed to obtain sequence alignment data with alignment position information. By evaluating the alignment quality of the sequence alignment data and filtering out alignment results with alignment quality scores below a preset threshold, a high-quality sequence alignment dataset is obtained. Based on the identification of repetitive sequences in high-quality sequence alignment datasets, and a comprehensive judgment based on the sequence start position, end position, and alignment quality, repetitive sequencing fragments are marked and removed. The deduplicated sequence alignment data are then integrated to obtain the deduplicated data.
[0008] Furthermore, by performing SNP genotyping analysis on the deduplicated data, SNP genotyping data within a preset target region is obtained, including: Based on the deduplication data, genotype identification was performed on SNP loci across the entire genome to obtain preliminary genome-wide SNP genotyping results; By filtering probe capture region data from whole-genome SNP genotyping results, SNP site data located within the target capture region are selected based on preset SNP site probe coordinate information. The SNP locus data of the target region are standardized to obtain an SNP genotyping data table containing locus number, genotype and sequencing depth information; By performing quality verification on the SNP genotyping data table, calculating the genotyping confidence of each SNP locus, and removing loci with confidence scores below the threshold, the final SNP genotyping data for the preset target area is formed.
[0009] Furthermore, the SNP genotyping data is mapped to feature vectors in a multidimensional feature space, a multidimensional reference structure is constructed, and analysis partitions are defined. Feature vectors are then categorized into corresponding analysis partitions, and dynamic calibration coefficients for the feature vector distribution are calculated based on a shape diameter function. These dynamic calibration coefficients are used to calibrate the SNP genotyping data, resulting in optimized SNP genotyping data, including: Based on the SNP genotyping dataset, the genotype information of each SNP locus is converted into numerical features, and a set of feature vectors in a multidimensional feature space is constructed. Based on the feature vector set, a multidimensional reference structure is constructed, and the multidimensional reference structure is divided into multiple analysis partitions based on the feature distribution density; The feature vectors are classified into corresponding analysis partitions according to their spatial coordinate positions, forming a partition feature vector distribution set; By applying the feature vector distribution set of the partitions, the morphological feature parameters of the feature vector distribution within each analysis partition are calculated to obtain the dynamic calibration coefficients; The dynamic calibration coefficients are applied to the original SNP genotyping data to perform weighted calibration on the genotyping confidence of each SNP locus, and the calibrated SNP genotyping confidence assessment results are obtained. Based on the calibration results, the SNP genotyping data are optimized and integrated, and sites with confidence levels below the threshold after calibration are filtered out to form an optimized SNP genotyping dataset.
[0010] Furthermore, by filtering the optimized SNP genotyping data, SNP loci with sequencing depths meeting the requirements and allele frequencies in the reference population database were selected to obtain the final analysis data for alignment, including: Based on the optimized SNP genotyping dataset, the sequencing depth of each SNP site is evaluated, and SNP sites whose sequencing depth meets the preset depth standard are selected to obtain a subset of SNP sites that meet the depth requirements. A subset of SNP loci that meet the depth requirements are correlated with a reference population database to calculate the allele frequency of each SNP locus in the reference population. Based on allele frequencies, SNP sites with allele frequencies within a preset intermediate frequency range are selected. The selected SNP sites are integrated to generate final analysis data containing genotype information of the tested and suspected subjects.
[0011] Furthermore, based on the final analysis data, the common SNP loci between the tested subjects and suspected subjects are counted, the proportion of loci conforming to Mendelian inheritance laws is calculated, and a comprehensive analysis is performed in conjunction with genetic similarity. A preset threshold is then used to determine whether a parent-child relationship exists between the two individuals, including: By analyzing the dataset, SNP genotype data of the tested subjects and suspected subjects were extracted; the extracted SNP genotype data were matched by locus coordinates to identify SNP loci with the same genotype locus coordinates between the tested subjects and suspected subjects, and the total number of common SNP loci was counted. Based on shared SNP loci, a Mendelian inheritance conformity analysis was performed on the genotype of each locus, the number of SNP loci conforming to Mendelian inheritance was counted, and the proportion of loci conforming to Mendelian inheritance was calculated. Using shared SNP locus data, a genotype conversion value summary matrix and an allele frequency summary matrix were constructed between the two samples, and the genetic similarity index between the two samples was calculated. The proportion of loci conforming to Mendelian inheritance laws is compared with a first preset threshold, and the genetic similarity index is compared with a second preset threshold to obtain the comparison results. Based on the comparison results, when the proportion of loci conforming to Mendelian inheritance laws is greater than the first preset threshold and the genetic similarity index is greater than the second preset threshold, it is determined that the tested object and the suspected object have a parent-child relationship; otherwise, it is determined that there is no parent-child relationship.
[0012] Secondly, a paternity testing system based on SNP marker combinations includes: The acquisition module is used to construct libraries and sequence samples from the tested object and the suspected object, respectively, to obtain the raw sequencing data of the samples; The filtering module is used to split the raw sequencing data to obtain individual sample data of the tested object and the suspected object; and to remove low-quality sequencing data by quality control filtering of the individual sample data to obtain quality-controlled data. The alignment module is used to align the quality-controlled data with the human reference genome to obtain sequence alignment data; the sequence alignment data is then deduplicated to eliminate duplicate sequences, resulting in deduplicated data. The analysis module is used to perform SNP typing analysis on the deduplicated data to obtain SNP typing data within a preset target area; The optimization module maps SNP genotyping data into feature vectors in a multidimensional feature space, constructs a multidimensional reference structure and divides the analysis partitions; it classifies the feature vectors into the corresponding analysis partitions, calculates the dynamic calibration coefficients of the feature vector distribution based on the shape diameter function, and calibrates the SNP genotyping data using the dynamic calibration coefficients to obtain optimized SNP genotyping data. The screening module is used to filter the optimized SNP genotyping data to select SNP sites that meet the sequencing depth requirements and whose allele frequencies are in the middle range in the reference population database, so as to obtain the final analysis data for comparison. The processing module is used to count the common SNP loci between the tested object and the suspected object based on the final analysis data, calculate the proportion of loci that conform to Mendelian inheritance laws, and perform a comprehensive analysis in combination with genetic similarity to determine whether there is a parent-child relationship between the two through a preset threshold.
[0013] The above-described solution of the present invention has at least the following beneficial effects: By employing a technology centered on SNP marker combinations, combined with multi-dimensional quality control filtering to remove low-quality data with base quality below Q30, abnormal length, and excessive N bases, human reference genome alignment and repetitive sequence removal, multi-dimensional feature space mapping and dynamic calibration of SNP genotyping data using shape-diameter functions, dual screening of sequencing depth and allele frequency within the range of 0.3 to 0.7, and comprehensive judgment based on Mendelian inheritance locus proportion and genetic similarity, this technology overcomes the technical problems of traditional paternity testing methods, such as limited detection sites, incomplete data quality control, susceptibility to interference in SNP genotyping accuracy, and insufficient precision and reliability due to a single judgment basis. This results in improved accuracy, stability, and reproducibility of paternity testing results, ensuring that the identification conclusions are scientifically rigorous and highly persuasive. Attached Figure Description
[0014] Figure 1 This is a schematic flowchart of a paternity testing method based on SNP marker combinations provided by an embodiment of the present invention.
[0015] Figure 2This is a schematic diagram of a paternity testing system based on SNP marker combinations provided by an embodiment of the present invention. Detailed Implementation
[0016] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0017] like Figure 1 As shown, an embodiment of the present invention proposes a paternity testing method based on SNP marker combinations, the method comprising the following steps: Step 1: Perform library construction and sequencing on samples from the tested object and suspected object respectively to obtain the raw sequencing data of the samples; Step 2: The raw sequencing data is split to obtain individual sample data of the tested object and the suspected object; the individual sample data is filtered by quality control to remove low-quality sequencing data and obtain the quality-controlled data. Step 3: Align the quality-controlled data with the human reference genome to obtain sequence alignment data; perform deduplication on the sequence alignment data to eliminate duplicate sequences and obtain deduplicated data. Step 4: Perform SNP typing analysis on the deduplicated data to obtain SNP typing data within the preset target area; Step 5: Map the SNP genotyping data to feature vectors in a multidimensional feature space, construct a multidimensional reference structure and divide the analysis partitions; classify the feature vectors to the corresponding analysis partitions, calculate the dynamic calibration coefficients of the feature vector distribution based on the shape diameter function, and calibrate the SNP genotyping data using the dynamic calibration coefficients to obtain optimized SNP genotyping data; Step 6: By filtering the optimized SNP genotyping data, SNP sites that meet the sequencing depth requirements and whose allele frequencies are in the middle range in the reference population database are selected to obtain the final analysis data for comparison. Step 7: Based on the final analysis data, count the common SNP loci between the tested subjects and the suspected subjects, calculate the proportion of loci that conform to Mendelian inheritance laws, and combine them with genetic similarity for comprehensive analysis. Use a preset threshold to determine whether there is a parent-child relationship between the two.
[0018] In this embodiment of the invention, because a series of technical means are adopted, including SNP marker combination as the core, sample library construction and sequencing, raw data splitting and quality control filtering, human reference genome alignment and repetitive sequence removal, SNP genotyping analysis of preset target regions, multidimensional feature space mapping and dynamic calibration of SNP genotyping data and shape diameter function, dual screening of sequencing depth and allele intermediate frequency in the reference population, as well as statistics of common SNP sites, calculation of the proportion of Mendelian inheritance sites and comprehensive analysis of genetic similarity combined with preset threshold judgment, the technical problems of incomplete control of test data quality, easy interference of SNP genotyping accuracy, insufficient targeting of effective site screening, and lack of accuracy and reliability caused by single judgment basis in traditional paternity testing methods are effectively overcome. Thus, the technical effect of improving the accuracy, stability and rigor of paternity testing results is achieved, ensuring that the identification conclusion is scientific, credible and highly persuasive.
[0019] In a preferred embodiment of the present invention, step 1 above may include: Step 1.1: Obtain DNA samples from the tested and suspected individuals, and perform library construction on the DNA samples to obtain sequencing libraries with sample identification information. Specifically, this includes: collecting biological samples from the tested and suspected individuals through compliant sampling procedures. Suitable sample types for DNA extraction include peripheral venous blood, oral swabs, or hair follicles. Each collected biological sample is individually numbered to ensure unique and traceable identification. Subsequently, using a standard DNA extraction kit, DNA is extracted from each numbered biological sample according to the kit's standard operating procedures. During the process, the temperature, humidity, and reagent dosage of the extraction environment are strictly controlled to avoid DNA degradation or contamination, ultimately obtaining high-purity, high-integrity genomic DNA. Next, the extracted genomic DNA undergoes quality testing. The measurement indicators include DNA concentration, purity, and fragment integrity. Only DNA samples that meet the preset quality standards are retained for subsequent library construction. During library construction, qualified genomic DNA is first randomly fragmented to break it into fragments of suitable sequencing length. Then, the fragmented DNA undergoes end repair, filling in the sticky ends of the DNA fragments and performing phosphorylation modification. Subsequently, adapter sequences with sample-specific identifiers are ligated to both ends of the DNA fragments. Each sample's adapter sequence contains a unique base recognition tag to ensure that sequencing data from different samples can be distinguished. Finally, the DNA fragments with ligated adapters are amplified using PCR amplification technology. The number of cycles is strictly controlled during the amplification process to avoid sequence bias caused by over-amplification, ultimately yielding a sequencing library with unique sample identification information.
[0020] Step 1.2: Load the sequencing libraries into the high-throughput sequencer and perform sequencing to obtain mixed sequencing data. This includes: First, performing a comprehensive pre-power-on check of the high-throughput sequencer, including the instrument hardware status, reagent compartment sealing, and sequencing chip cleanliness, to ensure the instrument is in normal working order. Then, based on the concentration and volume of the sequencing libraries, perform serial dilutions of each library according to the high-throughput sequencer's standard requirements, adjusting the library concentration to the optimal sequencing concentration range for the instrument. Denature the diluted libraries by adding denaturing reagents to unwind the double-stranded DNA into single-stranded DNA, ensuring effective primer binding during sequencing. Finally, add all the processed sequencing libraries sequentially into the sequencer's sample pool in a pre-defined order, with each library added to the sample label. The sequencer first identifies and records the specific coordinates of the library within the sample pool. Next, it loads the sequencer with the appropriate sequencing reagents, including sequencing primers, polymerase, and fluorescently labeled dNTPs. After closing the instrument compartment, it initiates calibration, calibrating the sequencing chip, laser detection system, and signal acquisition system to ensure sequencing accuracy. Once calibration is complete, sequencing parameters are set, including read length, number of sequencing cycles, and signal acquisition frequency. These parameter settings must be determined based on the library fragment length and detection requirements. Finally, the sequencer is started, automatically performing the sequencing reaction according to preset parameters. It uses laser excitation of fluorescently labeled dNTPs to collect fluorescence signals from each reaction site in real time and converts them into electrical signals. After preliminary processing by the instrument's built-in software, it generates mixed sequencing data containing sequencing information from all samples.
[0021] Step 1.3: Based on the mixed sequencing data, the mixed sequencing data is standardized and converted into a raw sequencing data file conforming to the standard format. Specifically, this includes: First, acquiring the mixed sequencing data output from the high-throughput sequencer. This data initially uses the sequencer's proprietary raw data format, containing base sequence information, fluorescence signal intensity information, and preliminary quality assessment information. Then, initiating data format conversion, the mixed sequencing data is parsed to extract core information, including the sequence data of each sequencing read, the corresponding quality value data, and sample identification-related tag data. Following the standard data format requirements commonly used in paternity testing, the parsed information is reorganized and structured, mapping the base sequence data to the corresponding quality value data, and unifying the data arrangement order and storage structure. During the format conversion process, a preliminary integrity check is performed on the data to check for missing sequences, abnormal quality values, or missing tag data. Reads with missing or abnormal data are marked. The marked abnormal data is separated, and the normal data undergoes further format conversion, ultimately generating a raw sequencing data file conforming to the industry standard format.
[0022] Step 1.4: Based on the original sequencing data file, create a sample metadata index by combining the sample identification information to obtain the original sequencing dataset with complete sample identification information. Specifically, this includes: First, organizing the complete metadata information for each sample. The metadata information includes basic sample information and experimental information. The basic sample information includes the sample number, the identity of the tested or suspected object, the sample collection time, the collection location, the sampling personnel, etc.; the experimental information includes the DNA extraction time, the extraction kit model, the library construction time, the adapter sequence information, the PCR amplification parameters, the sequencer model, the sequencing parameter settings, etc. Subsequently, the compiled metadata information is associated with the generated standard-format raw sequencing data file, establishing a mapping relationship according to the sample number to ensure that the metadata information of each sample accurately corresponds to the corresponding sequencing data. Next, a sample metadata index is created, with an index structure including an index field and a data storage address field. The index field uses the sample number as the core index item and also includes auxiliary index items such as sample identification and adapter tag sequences. The data storage address field records the starting storage location and data length of the sequencing data corresponding to each sample in the raw sequencing data file. Finally, the completeness and accuracy of the created sample metadata index are verified to ensure that all sequencing data of each sample can be quickly located through the index, ultimately forming a raw sequencing dataset with complete sample identification information and a metadata index.
[0023] In this embodiment of the invention, by employing technical means such as obtaining DNA samples of the tested object and suspected object and constructing a sequencing library with sample identification information, obtaining mixed sequencing data through high-throughput sequencing and performing format standardization processing, and creating a metadata index in combination with sample identification information to form a complete and identified original sequencing dataset, the technical problems that may occur in the traditional sample library construction and sequencing process, such as chaotic sample identification, inconsistent data format, and ambiguous correspondence between samples and sequencing data, are effectively overcome. This ensures that the original sequencing data format is standardized, the sample traceability is clear and accurate, and the integrity and reliability of the data are guaranteed.
[0024] In a preferred embodiment of the present invention, step 2 above may include: Step 2.1: Based on the sample identification information in the original sequencing dataset, perform index matching and data separation on the original sequencing data, extracting the sequencing fragments corresponding to the tested and suspected objects to obtain the initial individual sample data. Specifically, this includes: First, obtaining the original sequencing dataset with complete sample identification information and metadata index. The dataset contains mixed sequencing data of the tested and suspected objects and the corresponding sample metadata index. The core identification information in the sample metadata index includes the sample number, the unique adapter tag sequence, and the identification identifiers for distinguishing the tested and suspected objects. Then, initiate data separation by calling the sample metadata index, using the sample number as the core matching item, and combining it with the adapter tag sequence as an auxiliary verification item to perform comprehensive index matching on the original sequencing data. During the matching process, each sequencing read in the raw sequencing data is scanned line by line, and the sequence fragments containing the sample identifier contained in the read are extracted and accurately compared with the corresponding identifier information in the metadata index. For sequencing reads that are successfully matched and whose identifier information is completely consistent, they are determined to be valid data of the corresponding sample and are classified and collected into the dedicated data storage area of the sample. For sequencing reads that fail to match, have ambiguous identifier information, or have conflicts, they are marked as abnormal data and temporarily stored. Through index matching and classification and collection operations, all sequencing fragments corresponding to the tested object and the suspected object are separated from the mixed raw sequencing data respectively. The sequencing fragments of each sample retain complete sequence information and quality value information, and finally form initial individual sample data that are not confused and clearly identified.
[0025] Step 2.2 involves quality assessment of the initial individual sample data, calculating the base quality score, sequence length distribution, and sequencing depth coverage for each sequencing fragment to obtain a quality assessment report. Specifically, this includes: for each initial individual sample data, firstly, calculating the base quality score for each sequencing fragment; for each base in each sequencing read, converting it into a corresponding quality score based on the fluorescence signal intensity collected during sequencing; the quality score reflects the accuracy of base identification, with higher scores indicating higher accuracy; simultaneously, calculating the average base quality score for each sequencing fragment; next, analyzing the sequence length distribution, counting the actual length of each sequencing fragment, and recording the length of all fragments. The sequence length is calculated by taking the maximum, minimum, average, and median length values, and simultaneously counting the number and proportion of sequencing fragments within different length intervals to form sequence length distribution characteristics. Then, the sequencing depth coverage is calculated by first identifying the target genomic region related to paternity testing, counting the number of times each base site in the target region is covered by sequencing fragments, and calculating the average number of times all base sites in the target region are covered, which is the sequencing depth coverage of the sample. At the same time, the number and distribution of sites in the target region with coverage lower than the preset reference value are recorded. The results of base quality score statistics, sequence length distribution characteristics, sequencing depth coverage data, and abnormal index annotations are integrated and a detailed quality assessment report is generated in a unified format.
[0026] Step 2.3: Based on the quality assessment report, set a quality filtering threshold to remove sequencing fragments with a base quality score below Q30, sequences with abnormal lengths, and low-quality data containing too many N bases, thus obtaining filtered high-quality sequencing data. Specifically, this involves referencing the generated quality assessment report and considering the high data quality requirements of paternity testing, setting a clear quality filtering threshold. The base quality score threshold is set to Q30. In high-throughput sequencing bases, Q30 is derived from the Phred quality score (Q...). The "Score" represents the sequencing accuracy ≥ 99.9%. Sequencing fragments with a base quality score below Q30 are removed, meaning fragments with an accuracy less than 99.9% are retained, while those with an average base quality score greater than or equal to Q30 are removed. The sequence length threshold is set to 100-500 bases, retaining fragments within this range and removing abnormal sequences shorter than 100 or longer than 500 bases. The N-base percentage threshold is set to 5%, calculating the proportion of N bases in each fragment and retaining fragments with a percentage less than or equal to 5%, removing fragments with an excessive N-base percentage exceeding 5%. Data filtering is then initiated, screening each fragment of the initial individual sample data according to the three set quality filtering thresholds. During the screening process, each sequencing fragment is first checked for base quality score, then for sequence length, and finally for N base ratio. Only sequencing fragments that meet the threshold requirements for all three indicators are retained. Fragments that do not meet the threshold requirements for any one indicator are judged as low-quality data and are directly removed. In the end, high-quality sequencing data after strict screening are obtained.
[0027] Step 2.4 involves integrating and recombining the filtered high-quality sequencing data to generate data quality statistics, resulting in quality-controlled data. This includes: classifying and organizing the filtered high-quality sequencing data according to sample affiliation; sorting the high-quality sequencing data for each sample according to the genomic location information corresponding to the sequencing fragments, following the principle of chromosome number from smallest to largest, and on the same chromosome from left to right according to base site coordinates; recombining and integrating the sorted sequencing data; and grouping sequencing fragments of the same genomic region to ensure a clear data structure and logical coherence, facilitating subsequent alignment with the reference genome. During the integration and recombination process, detailed data quality statistics are generated simultaneously. These statistics include the total number of fragments, total number of bases, average base quality score, maximum, minimum, average, and median values of sequence length distribution, specific values of sequencing depth coverage, the number and proportion of low-quality data removed, and the average level of N-base percentage. The integrated and recombined sequencing data is then linked with the generated data quality statistics to form complete quality-controlled data.
[0028] In this embodiment of the invention, because it employs technical means such as index matching and data separation based on sample identification information to extract initial individual sample data, conducting quality assessment from multiple dimensions including base quality score, sequence length distribution, and sequencing depth coverage, setting targeted thresholds to remove low-quality data, integrating and recombining high-quality data, and generating data quality statistics, it effectively overcomes the technical problems of traditional data splitting being prone to sample confusion, low-quality data residue due to a single quality assessment dimension, and lack of quality traceability basis after data integration. Thus, it achieves the goal of ensuring accurate correspondence between individual sample data of the tested object and suspected object, controllable sequencing data quality and high purity, and data quality traceability through quality statistics.
[0029] In a preferred embodiment of the present invention, step 3 above may include: Step 3.1: Align the quality-controlled data with the human reference genome, construct an index based on the human reference genome, perform sequence alignment operations, and obtain sequence alignment data with alignment position information. Specifically, this includes: First, selecting the internationally recognized standard version of the human reference genome as the alignment benchmark. This version has been fully validated and contains complete genome sequence information, which can provide a reliable alignment reference for sequencing data. Before performing the alignment operation, the human reference genome is preprocessed to remove redundant sequences and uncertain regions, ensuring the accuracy and consistency of the reference sequence. Then, a professional genome indexing tool is used to construct an index file based on the preprocessed human reference genome. The index file segments and maps the genome sequence, significantly improving the efficiency and accuracy of subsequent sequence alignment. After the index is built, sequence alignment is initiated. The obtained quality-controlled data is input into the alignment system, and alignment parameters are set, including a seed length of 20 bases, a mismatch allowance of 2, a gap opening penalty of 10, and a gap extension penalty of 2. These parameter settings balance the sensitivity and specificity of the alignment. The alignment process uses the positional information in the index file to match each sequencing fragment in the quality-controlled data with the human reference genome, determining the precise location of each sequencing fragment on the reference genome. Information such as the number of mismatches and gaps during the matching process is recorded, ultimately generating sequence alignment data containing the alignment location, matching status, and related parameters for each sequencing fragment.
[0030] Step 3.2 involves evaluating the alignment quality of the sequence alignment data, filtering out alignment results with quality scores below a preset threshold to obtain a high-quality sequence alignment dataset. Specifically, this includes: initiating an alignment quality assessment for the generated sequence alignment data, using multi-dimensional indicators to comprehensively evaluate each alignment result; firstly, calculating the alignment quality score for each alignment result, which comprehensively considers factors such as the matching rate between the sequencing fragment and the reference genome, the number of mismatches, gap length and location, with a score range of 0 to 100, where a higher score indicates a more reliable alignment result; simultaneously, evaluating the alignment quality of each sequencing fragment... Regarding uniqueness, it is determined whether the fragment can only match a unique position in the reference genome to avoid interference in subsequent analysis caused by multiple matches. Referring to the quality requirements of alignment data in the field of paternity testing, and combined with the statistical analysis results of previous experimental data, the alignment quality score threshold is set at 60. Subsequently, the alignment quality assessment screens each alignment result one by one, retaining the results with an alignment quality score greater than or equal to 60 and unique alignment, and filtering out the results with an alignment quality score lower than 60, multiple matches, or ambiguous alignment positions. All the screened reliable alignment results are collected and organized to form a high-quality sequence alignment dataset.
[0031] Step 3.3: Based on the repetitive sequence identification of the high-quality sequence alignment dataset, and a comprehensive judgment of the sequence start position, end position, and alignment quality, repetitive sequencing fragments are marked and removed. Specifically, this includes: performing repetitive sequence identification and removal operations on the obtained high-quality sequence alignment dataset. First, extract key information for each sequencing fragment in the high-quality sequence alignment dataset, including the start and end positions on the human reference genome and the corresponding alignment quality score; then, start the repetitive sequence identification algorithm, sorting and traversing all sequencing fragments according to the chromosome number and base coordinate order of the reference genome; after traversing... During the process, the start and end positions of the current sequencing fragment are compared one by one with the sequencing fragments that have been traversed. If the start and end positions of two sequencing fragments are completely identical and the difference in alignment quality score is within 5, then these two sequencing fragments are determined to be repetitive sequences. For the identified repetitive sequences, the sequencing fragment with the higher alignment quality score is retained, and the repetitive fragment with the lower quality score is marked as redundant data. After all sequencing fragments have been traversed, all marked redundant data are removed in batches. At the same time, the number of repetitive sequences identified, the number removed, and the quality distribution of the retained fragments are recorded to ensure the traceability of the deduplication operation.
[0032] Step 3.4 involves integrating the deduplicated sequence alignment data to obtain deduplicated data. Specifically, this includes: first, classifying and grouping all deduplicated sequencing fragments according to the chromosome numbering order of the human reference genome, organizing fragments belonging to the same chromosome into their corresponding data; within each chromosome data module, further sorting the sequencing fragments from left to right according to their base coordinates to ensure logical consistency of the positional information of each fragment; subsequently, standardizing the data format of the sorted sequencing fragments, uniformly recording key parameters such as alignment position, matching status, mismatch information, and alignment quality score for each fragment, forming a structured sequence alignment dataset; during the integration process, a data integration report is generated simultaneously, including statistical information such as the number of sequencing fragments corresponding to each chromosome, the average alignment quality score, the data compression ratio after deduplication, and the number of fragments without alignment results; finally, associating the structured sequence alignment dataset with the data integration report to form complete deduplicated data.
[0033] In this embodiment of the invention, by employing technical means such as first constructing an index based on the human reference genome and then performing sequence alignment to obtain data containing alignment position information, evaluating the quality of the sequence alignment data and filtering results below a preset threshold, identifying and removing duplicate sequencing fragments based on the sequence start position, end position, and alignment quality, and integrating the deduplicated data, the invention effectively overcomes the technical problems of traditional genome alignment, such as low alignment efficiency and poor position accuracy due to lack of index support, low-quality alignment data interfering with subsequent analysis due to lack of quality screening of alignment results, and the single dimension of duplicate sequence identification leading to duplicate fragment residues increasing the server's operating pressure. This achieves the technical effects of improving the accuracy and efficiency of sequence alignment, purifying the quality of alignment data, reducing the interference of redundant data on analysis, providing a high-quality, low-redundancy sequence data foundation for SNP genotyping analysis, and ensuring the stability and reliability of the analysis process.
[0034] In a preferred embodiment of the present invention, step 4 above may include: Step 4.1: Based on the deduplication data, genotyping of SNP loci across the entire genome is performed to obtain preliminary genome-wide SNP genotyping results. This includes: First, obtaining the deduplication data, which possesses low redundancy, high orderliness, and high reliability, containing the precise alignment position, matching status, and related quality information of each sequencing fragment on the human reference genome. Genotyping of the entire genome SNP is initiated by preprocessing the deduplication data to remove sequencing fragments without alignment positions or with abnormal alignment quality, ensuring the purity of the input data. Then, using a professional SNP identification tool, based on the standard sequence of the human reference genome, each base locus across the entire genome is scanned and analyzed one by one. For each base locus, the coverage of that locus is statistically analyzed. The base types of all sequencing fragments are analyzed to determine if there are any variations inconsistent with the base types of the reference genome. If a stable base variation is found at a certain site, and the number of sequencing fragments supporting the variation (i.e., the sequencing depth) reaches 10 or more, and the quality score of the varied bases is not lower than 30, then the site is identified as a potential SNP site. Further analysis of the base distribution of the sequencing fragments is conducted to determine the genotype of each potential SNP site. If the proportion of varied bases at a site exceeds 90%, it is identified as a homozygous genotype; if the proportions of the two different bases are both between 40% and 60%, it is identified as a heterozygous genotype. The location information, variation type, genotype, and corresponding sequencing depth data of all identified SNP sites are collected and organized to form preliminary whole-genome SNP genotyping results.
[0035] Step 4.2 involves filtering the probe capture region data from the genome-wide SNP genotyping results. Based on pre-defined SNP probe coordinate information, SNP locus data located within the target capture region are selected. Specifically, this includes: First, retrieving the pre-defined SNP probe coordinate information. This information is based on high-information-value SNP loci designed through extensive experimental validation in the field of paternity testing. It covers target regions in the genome that are closely related to parentage and exhibit stable polymorphism. Each probe coordinate clearly corresponds to a specific interval on the genome. Then, the probe capture region filtering is activated, and the coordinates of each SNP locus in the preliminary genome-wide SNP genotyping results are selected. Each SNP site is precisely compared with the preset probe coordinates. During the comparison, the chromosome number and base position of the SNP site are used as the core matching criteria to determine whether each SNP site falls completely within the target capture region corresponding to a certain probe coordinate. If the base position of a certain SNP site is between the start and end positions of a certain probe coordinate, the site is determined to be a valid site within the target capture region. If the SNP site coordinate does not match any probe coordinate region, or only partially overlaps, it is determined to be a non-target region site and marked. After all SNP sites are compared, all SNP data determined to be valid sites are selected to form the target region SNP site data.
[0036] Step 4.3 involves standardizing the SNP locus data in the target region to obtain an SNP genotyping data table containing locus numbers, genotypes, and sequencing depth information. Specifically, this includes: standardizing the format of the selected target region SNP locus data to ensure complete uniformity in storage structure, field definitions, and representation; assigning a unique locus number to each target region SNP locus, following a combination of chromosome number and base position. For example, a locus located at position 12345678 on chromosome 1 would be numbered 112345678, ensuring direct association with the locus's specific location on the genome; then, clearly recording the genotype information for each locus, clearly labeling it as homozygous or heterozygous based on the determination results. Homozygous genotypes are labeled as combinations of two identical bases, while heterozygous genotypes are labeled as combinations of two different bases; simultaneously, calculating the sequencing depth information for each SNP locus, i.e., the total number of effective sequencing fragments covering the locus, and accurately recording the specific value. The three core pieces of information—locus number, genotype, and sequencing depth—are arranged in a fixed order, with each field separated by a uniform delimiter to ensure data readability and parsability. After processing the above information for all SNP loci in the target region, a structured SNP genotyping data table is formed.
[0037] Step 4.4 involves quality validation of the SNP genotyping data table, calculating the genotyping confidence score for each SNP locus, and removing loci with confidence scores below a threshold to form the final SNP genotyping data within the preset target region. Specifically, this includes: first, initiating quality validation of the SNP genotyping data table; performing multi-dimensional quality assessments on each SNP locus in the generated SNP genotyping data table; and then calculating the genotyping confidence score. The calculation of genotyping confidence score comprehensively considers multiple key factors, including the sequencing depth of the locus, the consistency of base distribution during genotyping, the alignment quality score of sequencing fragments, and the quality score of variant bases. Each factor participates in the confidence score calculation with different weights, and the final confidence score ranges from 0 to 100. A higher score indicates a more reliable genotyping result for that locus. Referring to the high reliability requirements for SNP genotyping results in paternity testing, a genotyping confidence threshold of 80 was set. Subsequently, quality verification assessed the confidence score of each SNP locus individually, retaining loci with confidence scores greater than or equal to 80, and marking loci with confidence scores below 80 as low-reliability loci and removing them from the SNP genotyping data table. During the removal process, the locus number, confidence score, and main reasons for non-compliance were recorded simultaneously, generating a quality verification report. Finally, the retained high-reliability SNP locus data were integrated to form the final SNP genotyping data within the preset target area, focusing on high-value, high-reliability SNP loci.
[0038] In this embodiment of the invention, a series of technical measures are employed, including whole-genome SNP genotyping of deduplicated data, screening of SNP loci within the target capture region based on preset probe coordinates, standardizing the format of SNP data in the target region to explicitly include locus numbers, genotypes, and sequencing depth information, and calculating genotyping confidence and removing low-confidence loci. These measures effectively overcome the technical problems of traditional SNP genotyping, such as weak target locus specificity, chaotic data formats without unified standards, and lack of genotyping quality verification leading to low-confidence loci interfering with subsequent analyses. This results in the precise focusing on high-value SNP loci in the preset target region, the standardization and traceability of SNP genotyping data, and the improved reliability and effectiveness of genotyping results. Ultimately, this provides high-quality, standardized genotyping data support for SNP data calibration and parentage determination.
[0039] In a preferred embodiment of the present invention, step 5 above may include: Step 5.1: Based on the SNP genotyping dataset, convert the genotype information of each SNP locus into numerical features and construct a set of feature vectors in a multidimensional feature space. Specifically, this includes: First, obtaining the SNP genotyping data within the final preset target region. The SNP genotyping data contains core information such as high-value, high-reliability SNP locus numbers, genotypes, sequencing depth, and genotyping confidence. To achieve accurate calibration of the SNP genotyping data, initiate numerical conversion, standardizing the genotype information of each SNP locus. Specifically, homozygous genotypes are uniformly converted to fixed numerical combinations based on their corresponding base types. For example, a homozygous genotype with two identical dominant bases is converted to 2, and a homozygous genotype with two identical recessive bases is converted to 2. Genotypes are converted to 0; heterozygous genotypes are converted to 1, ensuring clear and consistent numerical distinctions between different genotypes. Simultaneously, key parameters such as sequencing depth and genotyping confidence for each SNP locus are converted to corresponding standardized values. Sequencing depth is directly incorporated based on actual values, while genotyping confidence retains the original score and is converted to a standardized value between 0 and 100. Using each SNP locus as an independent unit, genotype values, sequencing depth values, genotyping confidence values, and preset allele frequency-related values are used as feature parameters of different dimensions, combined to form a single feature vector in a multidimensional feature space. After performing the above numerical conversion and feature vector construction sequentially for all SNP loci, a complete set of feature vectors is aggregated.
[0040] Step 5.2: Based on the feature vector set, construct a multidimensional reference structure and divide the multidimensional reference structure into multiple analysis partitions based on the feature distribution density. Specifically, this includes: analyzing the distribution range, density distribution, and genetic correlation characteristics of all feature vectors in the multidimensional feature space based on the constructed feature vector set; constructing the multidimensional reference structure based on this analysis, where each dimension of the multidimensional reference structure corresponds to a feature parameter of the feature vector; and determining the structural boundary range based on the maximum and minimum values of each dimension of all feature vectors to ensure coverage of the distribution range of all feature vectors. Subsequently, initiate the analysis partitioning, setting the feature distribution density threshold to a certain value per unit area. The bit space contains 10 feature vectors. Based on this threshold, the multidimensional reference structure is adaptively partitioned. During the partitioning process, the feature vector distribution density of each region in the multidimensional space is calculated. Dense regions with density higher than the threshold are subdivided to ensure that the feature vector distribution is uniform within each subdivided region. Sparse regions with density lower than the threshold are merged to avoid computational redundancy caused by too many partitions. Finally, the multidimensional reference structure is divided into multiple analysis partitions with clear boundaries and uniform internal feature distribution. Each partition clearly records the corresponding multidimensional space coordinate range, providing a reasonable spatial partitioning basis for accurate classification and targeted calibration of feature vectors.
[0041] Step 5.3 involves classifying the feature vectors according to their spatial coordinates into the corresponding analysis partitions, forming a partition feature vector distribution set. Specifically, this includes: First, extracting the dimensional values of each constructed feature vector. Based on the dimensional definition of the multidimensional reference structure, determining the specific coordinate position of each feature vector in the multidimensional space. The coordinate values correspond one-to-one with the standardized values of the feature parameters of each dimension. Feature vector classification is then initiated, comparing the spatial coordinates of each feature vector with the coordinate range of each analysis partition defined in Step 5.2. During the comparison, it is determined whether the coordinates of each dimension of the feature vector fall within the corresponding dimensional boundary of a certain analysis partition. If all dimensional coordinates meet the range requirements of a certain partition, the feature vector is classified into that analysis partition. If the feature vector coordinates exceed the range of all partitions or are in an ambiguous area at the partition boundary, they are marked as an abnormal vector and temporarily stored separately. After all feature vectors are classified, the feature vectors within each analysis partition are statistically analyzed and organized. Information such as the total number of feature vectors in each partition, the mean and variance of the distribution of each dimension feature parameter, etc., are recorded to form a structured partition feature vector distribution set, ensuring that the feature vectors in each partition have similar distribution characteristics.
[0042] Step 5.4: By applying the partition feature vector distribution set, calculate the morphological feature parameters of the feature vector distribution in each analysis partition to obtain the dynamic calibration coefficient. Specifically, this includes: using the obtained partition feature vector distribution set as input data, and calculating the morphological feature parameters for each analysis partition. For each partition, the geometric center of the feature vector distribution area is calculated by traversing the spatial coordinates of all feature vectors within the partition. Then, the straight-line distance (radius) from each feature vector to the geometric center is calculated. The average, maximum, minimum, and standard deviation of all radii are calculated to obtain core morphological characteristic parameters such as the average diameter and diameter fluctuation range of the partition. Simultaneously, the distribution compactness parameter is calculated based on the density distribution of feature vectors within the partition; a more concentrated distribution results in a higher compactness value, while a more dispersed distribution results in a lower compactness value. Based on the obtained morphological characteristic parameters, a dynamic calibration coefficient for the partition is determined using explicit numerical judgment rules. If the average diameter of the partition is within a reasonable range of 5 to 20, and the distribution compactness is higher than 80, it indicates that the feature vector distribution within the partition is stable, and the corresponding SNP genotyping data error is small; the dynamic calibration coefficient is directly set to 1.0. If the average diameter of the partition is greater than 20, or the distribution compactness is less than or equal to 60, it indicates a significant deviation in the feature vector distribution, and the corresponding SNP genotyping data may have a large error; further adjustments are needed based on… The deviation degree is adjusted by the calibration coefficient. The deviation degree is quantified by the deviation rate: the average diameter deviation rate is only calculated when the average diameter is greater than 20, which is the difference between the average diameter and 20 divided by 20; the compactness deviation rate is only calculated when the compactness is less than or equal to 60, which is the difference between 80 and compactness divided by 80. The maximum value of the average diameter deviation rate and the compactness deviation rate is taken as the core deviation rate. Then, 0.8 is added to the result of 0.2 multiplied by 1 and the core deviation rate is subtracted to obtain the dynamic calibration coefficient. The larger the deviation rate, the closer the calibration coefficient is to 0.8. If the average diameter of the partition is within a reasonable range of 5 to 20, and the distribution compactness is between 60 and 80, it indicates that the feature vector distribution is relatively stable but there is a slight deviation. The dynamic calibration coefficient is set to 0.9 plus 0.1 multiplied by the result of the difference between compactness and 60 divided by 20, so that the coefficient is dynamically adjusted between 0.9 and 1.0. The error of the SNP genotyping data in the partition is corrected by adjusting the coefficient. After completing the morphological feature parameter calculation and dynamic calibration coefficient determination for all analysis partitions in sequence, a dynamic calibration coefficient exclusive to each partition is formed.
[0043] Step 5.5 involves applying the dynamic calibration coefficients to the original SNP genotyping data to perform weighted calibration on the genotyping confidence of each SNP locus, obtaining the calibrated SNP genotyping confidence assessment results. Specifically, this includes: First, establishing a mapping relationship between feature vectors and the original SNP genotyping data. By using locus numbers, the feature vectors within each analysis partition are precisely bound to the corresponding original SNP genotyping data, ensuring that the genotyping confidence of each SNP locus corresponds to the dynamic calibration coefficient of its respective partition. Then, weighted calibration of genotyping confidence is initiated, extracting the original genotyping confidence score of each SNP locus and multiplying it with the dynamic calibration coefficient of its respective analysis partition to obtain the calibrated... For example, if the original genotyping confidence score of a SNP locus is 85 and the dynamic calibration coefficient of its partition is 0.9, then the calibrated genotyping confidence score is 76.5. If the original genotyping confidence score of another SNP locus is 90 and the dynamic calibration coefficient of its partition is 1.0, then the calibrated genotyping confidence score remains unchanged at 90. During the calibration process, key information such as the original confidence score, partition, dynamic calibration coefficient, and post-calibration confidence score of each SNP locus are recorded simultaneously to form a complete calibration record. After weighted calibration is performed on all SNP loci in sequence, all calibrated genotyping confidence scores are collected to generate the calibrated SNP genotyping confidence assessment result.
[0044] Step 5.6: Based on the calibration results, optimize and integrate the SNP genotyping data, filtering out loci with post-calibration confidence scores below the threshold to form an optimized SNP genotyping dataset. Specifically, this includes: First, considering the high reliability requirements of SNP genotyping data in paternity testing and combining the statistical analysis results of previous calibration data, setting the post-calibration genotyping confidence threshold to 85; initiating the SNP genotyping data optimization and integration, comparing the obtained post-calibration SNP genotyping confidence assessment results with the set threshold one by one; for SNPs with a post-calibration genotyping confidence score greater than or equal to 85... NP loci were identified as high-reliability loci, and their complete genotyping data, including locus number, genotype, sequencing depth, and post-calibration confidence score, were retained. SNP loci with a post-calibration confidence score below 85 were identified as low-reliability loci and removed from the dataset to avoid interference from low-quality loci in subsequent analysis results. Subsequently, the retained high-reliability SNP locus data were systematically integrated, sorted according to chromosome number and base coordinate order, and the data format and storage structure were standardized to form a structured and optimized SNP genotyping dataset.
[0045] In this embodiment of the invention, a series of technical means are employed to convert SNP locus genotype information into numerical features and construct a multidimensional feature vector set, construct a multidimensional reference structure based on feature distribution density and divide the analysis into partitions, classify feature vectors according to spatial coordinates to form partition distribution sets, apply and calculate dynamic calibration coefficients, and use these coefficients to weight and calibrate the SNP genotyping confidence and filter low-confidence loci after calibration. Therefore, the technical problems of traditional SNP genotyping data lacking systematic numerical conversion and multidimensional classification processing, static calibration methods leading to large deviations in genotyping confidence assessment, and the inability to accurately screen low-reliability loci are effectively overcome. This achieves the technical effect of realizing dynamic and accurate calibration and in-depth optimization of SNP genotyping data, improving the accuracy, consistency and reliability of genotyping results, and providing high-precision data support for locus screening and comprehensive judgment of parentage.
[0046] In a preferred embodiment of the present invention, step 6 above may include: Step 6.1: Based on the optimized SNP genotyping dataset, evaluate the sequencing depth of each SNP locus, and select SNP loci whose sequencing depth meets the preset depth standard to obtain a subset of SNP loci that meet the depth requirements. Specifically, this includes: First, obtaining the optimized SNP genotyping dataset. The SNP genotyping dataset has removed low-reliability loci, retaining only SNP loci with calibrated confidence levels and accurate genotyping, and containing complete sequencing depth information for each locus. Initiate sequencing depth evaluation, screening each SNP locus in the dataset one by one, and accurately extracting the sequencing depth value corresponding to each locus. The sequencing depth value represents the total number of effective sequencing fragments covering that locus. To reflect the support strength of the locus data, and considering the high requirements for sequencing depth in the field of paternity testing, combined with the verification results of previous experimental data, a preset depth standard of 15 was set, meaning that only SNP loci with a sequencing depth greater than or equal to 15 were retained. During the evaluation process, the sequencing depth of each locus was compared with the preset standard one by one. If the sequencing depth of a locus reached or exceeded 15, it was determined to meet the depth requirement and was included in the candidate locus set. If the sequencing depth was less than 15, it was determined to have insufficient support strength, which may lead to unstable genotyping results, and was directly removed from the dataset. After all loci were evaluated, all SNP locus data that met the depth requirement were collected and organized to form a subset of SNP loci that met the depth requirement.
[0047] Step 6.2 involves performing an association analysis between the subset of SNP loci that meet the depth requirements and the reference population database to calculate the allele frequency of each SNP locus in the reference population. Specifically, this includes: First, adjusting the reference population database, which contains whole-genome SNP data from large-scale healthy populations across different regions and ethnicities. The data has undergone rigorous quality verification and standardization to provide a reliable population genetic background reference for allele frequency calculation. Then, initiating the association analysis, the subset of SNP loci that meet the depth requirements is used as the analysis object. Through core identifiers such as locus number and chromosome coordinates, each SNP locus in the subset is precisely associated and matched with the corresponding locus in the reference population database. After a successful match, all allele types corresponding to the locus and the frequency of each allele in the population are extracted from the reference population database. Then, the allele frequency of each SNP locus is calculated. For each allele at each locus, the frequency value of the allele is obtained by dividing the frequency in the reference population by the total number of samples in the reference population. The frequency value is accurate to four decimal places. At the same time, auxiliary information such as the frequency distribution of each allele and the sample size of the reference population are recorded to ensure the traceability and reliability of the frequency calculation results. After performing association analysis and frequency calculation on all SNP loci that meet the depth requirements, data containing the allele type and corresponding frequency of each locus is generated.
[0048] Step 6.3: Based on allele frequencies, select SNP loci whose allele frequencies fall within a preset intermediate frequency range. Specifically, this includes: considering the discriminative power of SNP loci in paternity testing, defining the preset intermediate frequency range as 0.3 to 0.7. This intermediate frequency range is determined based on population genetics principles. Alleles within this range exhibit higher polymorphism and genetic information, more accurately reflecting the genetic association characteristics between parents and offspring, and effectively avoiding insufficient discriminative power due to excessively high or low allele frequencies. Extract the allele frequency data for each SNP locus and analyze the major allele frequencies of each locus. The first step is to focus on SNP loci corresponding to alleles with frequencies in the middle range. During the screening process, if the frequency of at least one allele at a certain SNP locus falls between 0.3 and 0.7, and the allele may appear in the genotypes of both the tested and suspected individuals, then the locus is considered a highly discriminative and effective locus and is retained. If the frequencies of all alleles at a certain locus are below 0.3 or above 0.7, it is considered insufficiently discriminative and cannot effectively support the determination of parentage, so it is removed. After all loci have been screened, the data of all highly discriminative and effective loci are collected to form a set of SNP loci focusing on core genetic information.
[0049] Step 6.4 integrates the screened SNP loci to generate final analysis data containing genotype information of the tested and suspected subjects. This includes: First, organizing the screened high-discrimination SNP loci by chromosome number from smallest to largest, and by base coordinates from left to right on the same chromosome, ensuring logical consistency in the locus arrangement; then, extracting the complete genotype information of the tested and suspected subjects corresponding to each effective locus, including homozygous or heterozygous type, specific base combinations, etc., while retaining key quality parameters such as sequencing depth, calibrated genotyping confidence, and allele frequency for each locus. Ensure data integrity; initiate data integration by combining the sorted locus information, genotype data of the tested and suspected subjects, and related quality parameters in a unified format to construct a structured data table. Each row of the data table corresponds to a valid SNP locus, and each column corresponds to a core piece of information, ensuring data readability and parsability. During the integration process, perform data integrity verification to check for missing genotypes, abnormal parameters, etc. If abnormal data is found, promptly trace back to previous steps for verification and correction. Finally, generate final analysis data containing genotype information and comprehensive quality parameters of the tested and suspected subjects.
[0050] In this embodiment of the invention, by employing technical means such as evaluating sequencing depth of optimized SNP genotyping data to screen loci that meet preset standards, associating a subset of these loci with a reference population database to calculate allele frequencies, screening loci within a preset intermediate frequency range, and integrating them to generate final analysis data containing genotype information of the tested and suspected subjects, the invention effectively overcomes the technical problems of traditional SNP locus screening, which focuses only on a single dimension and does not combine the genetic background of the reference population for targeted screening, resulting in low value of effective locus information and insufficient data specificity, thus interfering with the accuracy of parentage determination. This achieves the technical effect of accurately screening SNP loci with high information value and high discriminative power, ensuring the effectiveness and specificity of the final analysis data, providing high-quality data support for comprehensive parentage determination, and further improving the accuracy and rigor of parentage identification results.
[0051] In a preferred embodiment of the present invention, step 7 above may include: Step 7.1: Extract SNP genotype data from the final analysis dataset for both the tested and suspected subjects. Perform locus coordinate matching on the extracted SNP genotype data to identify SNP loci with the same genotype locus coordinates between the tested and suspected subjects, and count the total number of common SNP loci. Specifically, this includes: First, obtaining the generated final analysis data, which focuses on core SNP loci with high information value and high discriminative power, containing complete genotype information, locus coordinates, and related quality parameters corresponding to the tested and suspected subjects; initiating genotype data extraction to separate the SNP genotype data of the tested and suspected subjects from the final analysis data, including the chromosome number, base coordinates, genotype type, and corresponding allele information for each locus, ensuring that the structures of the two types of data are completely consistent; subsequently, initiating locus coordinate matching, using chromosome number and base coordinates as the core matching dimensions, to accurately compare each SNP locus of the tested subject with all SNP loci of the suspected subject one by one. During the comparison process, if the chromosome numbers of two loci are exactly the same and the base coordinates are exactly the same, they are determined to be shared SNP loci with the same genotype locus coordinates; if the chromosome numbers are different or the base coordinates are different, they are determined to be non-shared loci. During the matching process, the numbers, coordinates and genotype information of both parties of all shared loci are recorded simultaneously to avoid omissions or misjudgments. After all loci are compared, the number of all SNP loci determined to be shared is counted to obtain the total number of shared SNP loci.
[0052] Step 7.2: Based on shared SNP loci, perform Mendelian inheritance conformity analysis on the genotype of each locus, count the number of SNP loci conforming to Mendelian inheritance, and calculate the proportion of loci conforming to Mendelian inheritance. Specifically, based on the statistically analyzed shared SNP loci, first clarify the application standard of Mendelian inheritance in paternity testing, namely, that the genotype of the parents will be passed on to the offspring through inheritance, and each allele of the offspring must come from one parent. For each shared SNP locus, extract the genotypes of the tested individual and the suspected individual, and analyze whether the genotype transmission between the two conforms to this law. Specifically, if the suspected individual is a parental candidate and the tested individual is an offspring candidate, and if the suspected individual's genotype is homozygous dominant, the tested individual's genotype will be considered as a parental candidate. The genotype of the test subject must contain the dominant allele; if the genotype of the suspected test subject is homozygous recessive, the genotype of the tested test subject must contain the recessive allele; if the genotype of the suspected test subject is heterozygous, the genotype of the tested test subject must contain one of the alleles; conversely, if the tested test subject is a parental candidate and the suspected test subject is an offspring candidate, the same transmission logic is followed. After completing the conformity judgment for each shared locus, the number of all SNP loci conforming to Mendelian inheritance is counted; then the proportion of loci conforming to Mendelian inheritance is calculated by dividing the number of conforming loci by the total number of shared SNP loci, and the result is rounded to four decimal places. The proportion directly reflects the degree of consistency in genetic transmission between the two parties, providing core genetic evidence for determining parentage.
[0053] Step 7.3: Using shared SNP locus data, construct a genotype conversion value summary matrix and an allele frequency summary matrix between the two samples, and calculate the genetic similarity index between the two samples. Specifically, this includes: First, based on the determined shared SNP locus data, construct a genotype conversion value summary matrix between the two samples. Convert the genotypes of the tested and suspected subjects at each shared SNP locus into corresponding values according to the set numerical rules: homozygous dominant is 2, homozygous recessive is 0, and heterozygous is 1. The rows of the matrix correspond to each shared SNP locus, and the columns correspond to the genotype conversion values of the tested and suspected subjects, respectively, clearly presenting the numerical genetic characteristics of both parties at each locus. Next, construct an allele frequency summary matrix, extracting the allele frequency of each shared SNP locus... The SNP loci are allele frequency data from the reference population database. The rows of the matrix correspond to each shared SNP locus, and the columns correspond to the frequency values of each allele at that locus, providing population genetic background support for genetic similarity calculation. The genetic similarity index calculation is initiated by combining two summary matrices to comprehensively consider the degree of genotype matching and the weight of allele frequencies at each shared locus. During the calculation, for each shared locus, if the genotype conversion values of both parties are the same, a higher weight is assigned; if the conversion values are different, the degree of difference is calculated based on the allele frequencies and assigned a corresponding weight. The weight values of all shared loci are summed and averaged to obtain the genetic similarity index between the two samples. The index ranges from 0 to 100, with a higher value indicating greater similarity in genetic characteristics between the two parties.
[0054] Step 7.4 compares the proportion of loci conforming to Mendelian inheritance with a first preset threshold and the genetic similarity index with a second preset threshold to obtain a comparison result. Specifically, this includes: referencing industry standards in the field of paternity testing, verification results from extensive experimental data, and the optimization goals of the technical solution of this invention, setting a first preset threshold and a second preset threshold. The first preset threshold is 95, corresponding to the criterion for determining the proportion of loci conforming to Mendelian inheritance; the second preset threshold is 90, corresponding to the criterion for determining the genetic similarity index. The threshold comparison is initiated by first comparing the proportion of loci conforming to Mendelian inheritance calculated in step 7.2 with the first preset threshold of 95 to determine if the proportion is greater than 95, and recording the comparison result as either conforming or not conforming. Then, the genetic similarity index calculated in step 7.3 is compared with the second preset threshold of 90 to determine if the index is greater than 90, and again, recording the comparison result as either conforming or not conforming. During the comparison process, the specific proportion and index values, as well as the difference from the thresholds, are retained to ensure the traceability of the comparison results. Finally, a specific result containing both comparison indicators is generated.
[0055] Step 7.5: Based on the comparison results, when the proportion of loci conforming to Mendelian inheritance laws is greater than the first preset threshold and the genetic similarity index is greater than the second preset threshold, it is determined that the tested subject and the suspected subject have a parent-child relationship; otherwise, it is determined that no parent-child relationship exists. Specifically, this includes: First, summarizing the two comparison results and establishing a dual-indicator comprehensive judgment logic. If the first comparison result shows that the proportion of loci conforming to Mendelian inheritance laws is greater than 95%, and the second comparison result shows that the genetic similarity index is greater than 90, that is, both indicators meet the preset threshold requirements, then it is comprehensively determined that the tested subject and the suspected subject have a parent-child relationship. The judgment logic is based on the dual verification of genetic transmission laws and genetic characteristic similarity to ensure the rigor of the results. This comprehensive approach effectively avoids potential misjudgments that may occur with a single indicator. If any comparison result does not meet the preset threshold requirements—that is, the proportion of loci conforming to Mendelian inheritance is less than or equal to 95%, or the genetic similarity index is less than or equal to 90—then it is determined that the tested individual and the suspected individual do not have a parent-child relationship. During the determination process, a determination report is generated simultaneously. The report clearly records the total number of shared SNP loci, the number and proportion of loci conforming to Mendelian inheritance, the genetic similarity index, the two preset thresholds, and the specific comparison results, ensuring that the determination process is traceable and verifiable throughout. This comprehensive determination method solves the problems of insufficient accuracy and persuasiveness of traditional methods, meeting the needs of complex family scenarios and high-standard parentage testing.
[0056] In this embodiment of the invention, because it employs the following technical means: extracting SNP genotype data of the tested and suspected subjects from the final analysis data; accurately identifying and counting the total number of shared SNP loci through locus coordinate matching; analyzing the conformity to Mendelian inheritance laws based on shared loci and calculating the proportion of corresponding loci; constructing a genotype conversion value summary matrix and an allele frequency summary matrix between the two samples to calculate the genetic similarity index; and comparing the locus proportion and genetic similarity index through two preset thresholds to comprehensively determine the parent-child relationship, it effectively overcomes the technical problems of traditional parentage testing, such as single basis for judgment, lack of precision in locus matching, and insufficient comprehensiveness in genetic association analysis, which lead to large deviations in the judgment results and insufficient rigor and credibility. Thus, it achieves a comprehensive determination of parentage from two dimensions: conformity to genetic laws and genetic similarity, improving the accuracy, scientificity, and rigor of the identification results, and ensuring that the identification conclusions have strong persuasiveness and high credibility.
[0057] like Figure 2 As shown, embodiments of the present invention also provide a paternity testing system based on SNP marker combinations, comprising: The acquisition module is used to construct libraries and sequence samples from the tested object and the suspected object, respectively, to obtain the raw sequencing data of the samples; The filtering module is used to split the raw sequencing data to obtain individual sample data of the tested object and the suspected object; and to remove low-quality sequencing data by quality control filtering of the individual sample data to obtain quality-controlled data. The alignment module is used to align the quality-controlled data with the human reference genome to obtain sequence alignment data; the sequence alignment data is then deduplicated to eliminate duplicate sequences, resulting in deduplicated data. The analysis module is used to perform SNP typing analysis on the deduplicated data to obtain SNP typing data within a preset target area; The optimization module maps SNP genotyping data into feature vectors in a multidimensional feature space, constructs a multidimensional reference structure and divides the analysis partitions; it classifies the feature vectors into the corresponding analysis partitions, calculates the dynamic calibration coefficients of the feature vector distribution based on the shape diameter function, and calibrates the SNP genotyping data using the dynamic calibration coefficients to obtain optimized SNP genotyping data. The screening module is used to filter the optimized SNP genotyping data to select SNP sites that meet the sequencing depth requirements and whose allele frequencies are in the middle range in the reference population database, so as to obtain the final analysis data for comparison. The processing module is used to count the common SNP loci between the tested object and the suspected object based on the final analysis data, calculate the proportion of loci that conform to Mendelian inheritance laws, and perform a comprehensive analysis in combination with genetic similarity to determine whether there is a parent-child relationship between the two through a preset threshold.
[0058] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A paternity testing method based on SNP marker combinations, characterized in that, The method includes: Library construction and sequencing were performed on samples from the tested object and the suspected object respectively to obtain the raw sequencing data of the samples; By splitting the raw sequencing data, individual sample data of the tested object and the suspected object are obtained separately; by quality control filtering of the individual sample data, low-quality sequencing data is removed to obtain quality-controlled data; The quality-controlled data is compared with the human reference genome to obtain sequence alignment data; the sequence alignment data is then deduplicated to eliminate duplicate sequences, resulting in deduplicated data. By performing SNP genotyping analysis on the deduplicated data, SNP genotyping data within a preset target region is obtained, including: Based on the deduplication data, genotype identification was performed on SNP loci across the entire genome to obtain preliminary genome-wide SNP genotyping results; Retrieve the preset SNP site probe coordinate information, start the probe capture region filtering, and accurately compare the coordinates of each SNP site in the preliminary whole genome SNP genotyping results with the preset probe coordinate information one by one to filter out all SNP data that are determined to be valid sites and form the target region SNP site data. The SNP locus data of the target region are standardized to obtain an SNP genotyping data table containing locus number, genotype and sequencing depth information; By performing quality verification on the SNP genotyping data table, calculating the genotyping confidence of each SNP locus, and removing loci with confidence scores below the threshold, the final SNP genotyping data within the preset target area is formed. The SNP genotyping data within a preset target region is mapped to feature vectors in a multidimensional feature space. A multidimensional reference structure is constructed and analysis partitions are defined. The feature vectors are then categorized into their corresponding analysis partitions. Dynamic calibration coefficients for the feature vector distribution are calculated based on a shape diameter function. These dynamic calibration coefficients are used to calibrate the SNP genotyping data, resulting in optimized SNP genotyping data, including: Based on the SNP genotyping dataset, the genotype information of each SNP locus is converted into numerical features, and a set of feature vectors in a multidimensional feature space is constructed. Based on the feature vector set, a multidimensional reference structure is constructed, and the multidimensional reference structure is divided into multiple analysis partitions based on the feature distribution density; The feature vectors are classified into corresponding analysis partitions according to their spatial coordinate positions, forming a partition feature vector distribution set; For each partition, the geometric center of the feature vector distribution area is calculated by traversing the spatial coordinates of all feature vectors within the partition. The straight-line distance (radius) from each feature vector to the geometric center is calculated, and the average, maximum, minimum, and standard deviation of all radii are statistically analyzed to obtain core morphological feature parameters such as the average diameter and diameter fluctuation range of the partition. Combined with the density distribution of feature vectors within the partition, the distribution compactness parameter is calculated. The more concentrated the distribution, the higher the compactness value; the more dispersed the distribution, the lower the compactness value. Based on the obtained morphological feature parameters, the dynamic calibration coefficient of the partition is determined through explicit numerical judgment rules. The error of the SNP genotyping data within the partition is corrected by adjusting the coefficient. After completing the morphological feature parameter calculation and dynamic calibration coefficient determination for all analysis partitions in sequence, a unique dynamic calibration coefficient for each partition is formed. The dynamic calibration coefficients are applied to the original SNP genotyping data to perform weighted calibration on the genotyping confidence of each SNP locus, and the calibrated SNP genotyping confidence assessment results are obtained. Based on the calibration results, the SNP genotyping data within the preset target area are optimized and integrated, and sites with confidence levels below the threshold after calibration are filtered out to form an optimized SNP genotyping dataset. By filtering the optimized SNP genotyping data, SNP sites that meet the sequencing depth requirements and whose allele frequencies are in the middle range in the reference population database are selected to obtain the final analysis data for comparison. Based on the final analysis data, the common SNP loci between the tested subjects and the suspected subjects were counted, the proportion of loci that conform to Mendelian inheritance laws was calculated, and a comprehensive analysis was conducted in combination with genetic similarity. A preset threshold was used to determine whether there was a parent-child relationship between the two.
2. The paternity testing method based on SNP marker combinations according to claim 1, characterized in that, By splitting the raw sequencing data, individual sample data of the tested subjects and suspected subjects were obtained separately. Low-quality sequencing data was then removed from the individual sample data through quality control filtering, resulting in quality-controlled data, including: Based on the sample identification information in the original sequencing dataset, index matching and data separation are performed on the original sequencing data, and sequencing fragments corresponding to the tested object and the suspected object are extracted to obtain the initial individual sample data. By performing quality assessment on the initial individual sample data, the base quality score, sequence length distribution, and sequencing depth coverage of each sequencing fragment are calculated to obtain a quality assessment report; Based on the quality assessment report, a quality filtering threshold is set to remove sequencing fragments with a base quality score below Q30, sequences with abnormal lengths, and low-quality data containing too many N bases, thus obtaining high-quality sequencing data after filtering. The filtered high-quality sequencing data is integrated and recombined to generate data quality statistics, thus obtaining quality-controlled data.
3. The paternity testing method based on SNP marker combinations according to claim 2, characterized in that, The quality-controlled data was compared with the human reference genome to obtain sequence alignment data. The sequence alignment data was then deduplicated to remove repetitive sequences, resulting in deduplicated data, including: The quality-controlled data is compared with the human reference genome, an index is built based on the human reference genome, and sequence alignment is performed to obtain sequence alignment data with alignment position information. By evaluating the alignment quality of the sequence alignment data and filtering out alignment results with alignment quality scores below a preset threshold, a high-quality sequence alignment dataset is obtained. Based on the identification of repetitive sequences in high-quality sequence alignment datasets, and a comprehensive judgment based on the sequence start position, end position, and alignment quality, repetitive sequencing fragments are marked and removed. The deduplicated sequence alignment data are then integrated to obtain the deduplicated data.
4. The paternity testing method based on SNP marker combinations according to claim 3, characterized in that, By filtering the optimized SNP genotyping data, SNP loci that meet the sequencing depth requirements and whose allele frequencies are in the middle range in the reference population database are selected, resulting in the final analysis data for alignment, including: Based on the optimized SNP genotyping dataset, the sequencing depth of each SNP site is evaluated, and SNP sites whose sequencing depth meets the preset depth standard are selected to obtain a subset of SNP sites that meet the depth requirements. A subset of SNP loci that meet the depth requirements are correlated with a reference population database to calculate the allele frequency of each SNP locus in the reference population. Based on allele frequencies, SNP sites with allele frequencies within a preset intermediate frequency range are selected. The selected SNP sites are integrated to generate final analysis data containing genotype information of the tested and suspected subjects.
5. The paternity testing method based on SNP marker combinations according to claim 4, characterized in that, Based on the final analysis data, the common SNP loci between the tested subjects and suspected subjects were counted, the proportion of loci conforming to Mendelian inheritance laws was calculated, and a comprehensive analysis was performed in conjunction with genetic similarity. A preset threshold was used to determine whether a parent-child relationship existed between the two individuals, including: By analyzing the dataset, SNP genotype data of the tested subjects and suspected subjects were extracted; the extracted SNP genotype data were matched by locus coordinates to identify SNP loci with the same genotype locus coordinates between the tested subjects and suspected subjects, and the total number of common SNP loci was counted. Based on shared SNP loci, a Mendelian inheritance conformity analysis was performed on the genotype of each locus, the number of SNP loci conforming to Mendelian inheritance was counted, and the proportion of loci conforming to Mendelian inheritance was calculated. Using shared SNP locus data, a genotype conversion value summary matrix and an allele frequency summary matrix were constructed between the two samples, and the genetic similarity index between the two samples was calculated. The proportion of loci conforming to Mendelian inheritance laws is compared with a first preset threshold, and the genetic similarity index is compared with a second preset threshold to obtain the comparison results. Based on the comparison results, when the proportion of loci conforming to Mendelian inheritance laws is greater than the first preset threshold and the genetic similarity index is greater than the second preset threshold, it is determined that the tested object and the suspected object have a parent-child relationship; otherwise, it is determined that there is no parent-child relationship.
6. A paternity testing system based on SNP marker combinations, wherein the system implements the method as described in any one of claims 1 to 5, characterized in that, include: The acquisition module is used to construct libraries and sequence samples from the tested object and the suspected object, respectively, to obtain the raw sequencing data of the samples; The filtering module is used to split the raw sequencing data to obtain individual sample data of the tested object and the suspected object; and to remove low-quality sequencing data by quality control filtering of the individual sample data to obtain quality-controlled data. The alignment module is used to align the quality-controlled data with the human reference genome to obtain sequence alignment data; the sequence alignment data is then deduplicated to eliminate duplicate sequences, resulting in deduplicated data. The analysis module is used to perform SNP typing analysis on the deduplicated data to obtain SNP typing data within a preset target area; The optimization module is used to map SNP genotyping data within a preset target area into feature vectors in a multidimensional feature space, construct a multidimensional reference structure and divide the analysis partitions; classify the feature vectors into the corresponding analysis partitions, calculate the dynamic calibration coefficients of the feature vector distribution based on the shape diameter function, and calibrate the SNP genotyping data through the dynamic calibration coefficients to obtain optimized SNP genotyping data; The screening module is used to filter the optimized SNP genotyping data to select SNP sites that meet the sequencing depth requirements and whose allele frequencies are in the middle range in the reference population database, so as to obtain the final analysis data for comparison. The processing module is used to count the common SNP loci between the tested object and the suspected object based on the final analysis data, calculate the proportion of loci that conform to Mendelian inheritance laws, and perform a comprehensive analysis in combination with genetic similarity to determine whether there is a parent-child relationship between the two through a preset threshold.
Citation Information
Patent Citations
Method of determining kinship using gene sequence variation
KR102391084B1
Method and apparatus for unfolding three-dimensional model, and device and storage medium
WO2023197779A1