Pedigree tracing method based on whole genome resequencing SNP big data and deep learning
Patent Information
- Application Number
- CN202410211362.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-27
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-02-27
AI Technical Summary
[0005]针对上述进化树方法每次问询都需要消耗大量算力以及简单机器学习算法或浅层神经网络算法无法充分挖掘和利用全基因组大数据的问题,本发明提供一种基于深度学习构建全基因组单核苷酸多态性(SNP)与不同地理种群或亚群体之间联系的方法,以实现对待测生物样品进行谱系溯源的目的
[0030]此外,本发明提出的遗传位点可解释流程在实际应用中具有多种用途。首先,这种方法可以用于指导下游设计探针或条形码方法。通过分析遗传位点的信息,我们可以确定哪些位点对于特定研究或应用最为关键。基于这些关键位点,我们可以设计特定的探针或条形码,用于针对性地检测或标记感兴趣的基因或序列,这样可以高效地筛选和分析相关的基因或序列,减少测序的成本和复杂性。其次,这种方法还可以帮助优化全基因组测序的成本和复杂性。通过解释遗传位点的信息,我们可以确定哪些位点是关键的,对于特定的研究或应用而言是最有意义的。相比于整个基因组的测序,只针对这些关键位点进行测序可以节省大量的成本和时间,同时降低数据分析的复杂性。本发明提出的遗传位点可解释流程在指导下游设计探针或条形码方法以及优化全基因组测序成本和复杂性方面具有巨大的潜力。它不仅有助于节省资源和精力,还可以提高研究和应用的效率。
Smart Images

Figure CN118248210B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of deep learning and genealogy tracing, and in particular to a method for genealogy tracing using whole-genome resequencing single nucleotide polymorphism (SNP) data and deep learning algorithms. Background Technology
[0002] Genealogical tracing is a technology applied across various fields and has been widely used and developed. Initially, it was primarily used to trace human genealogies and family lineages, helping people understand their family history and identity. However, with advancements in technology and changes in application techniques, the scope of genealogical tracing has expanded to other areas. Agricultural product traceability is a crucial application of genealogical tracing. By utilizing this technology, authentic agricultural products or important local wild medicinal herbs can be traced back to their source, ensuring the product's origin and authenticity. During the sale and distribution of agricultural products and medicinal herbs, counterfeit and substandard products are highly likely to be introduced, seriously impacting consumer rights and health. By tracing the distribution process, the true origin of products can be confirmed, reducing the circulation of counterfeit and substandard products. This allows consumers to purchase agricultural products with greater confidence, while also helping regulatory authorities track and prevent quality and food safety issues. The pedigree tracing of wild and domesticated animals is another important application area. By tracing the genes and genetic characteristics of animals, background information such as lineage, species, and geographical subgroup of origin can be understood. This is of great significance for wildlife conservation and species breeding, contributing to the protection of endangered species and genetic diversity. For example, distinguishing between artificially bred and wild Chinese sturgeon lineages not only protects wildlife genetic diversity but also allows artificially bred populations to enter the market under orderly government control, effectively reducing poaching of wild populations. It also allows local restaurants to offer local specialties, boosting the local economy. Using pollen characteristics in honey to trace nectar sources is crucial; local honey products, especially those collected in nature reserves, are often marketed as high-end honey, making accurate identification of their nectar source essential. Furthermore, differentiating wild local populations from different sources is vital for genetic research, maintaining biodiversity and genetic diversity, protecting local germplasm resources, and conserving endemic species. Taking the Eastern honeybee in the test dataset as an example, the Eastern honeybee is a unique native bee species to my country, highly adaptable, and far more resistant to cold and predators than the Western honeybee. The Eastern honeybee plays a vital role in maintaining my country's native ecological balance and ensuring food and oil security. Accurate pedigree tracing is a prerequisite for the preservation, propagation, and breeding of Eastern honeybees. Pest and disease tracing, by studying the transmission routes and spread of pests and diseases, can identify the source of diseases and diseases, allowing for appropriate control measures. Pedigree tracing helps identify and track the origin and transmission paths of diseases, providing effective response strategies and reducing damage to humans, livestock, and crops. In summary, the application of pedigree tracing has gradually expanded to multiple fields, including agricultural product tracing, pedigree tracing of wild and domesticated animals, and pest and disease tracing.
[0003] The molecular characteristics of the whole genome's nucleic acid material (DNA) provide us with a method for pedigree tracing: gene flow is restricted between different regions, while gene flow from the same region or geographically close regions is more frequent. Therefore, theoretically, populations in the same region should have similar DNA sequences, while their DNA molecular sequence composition should be significantly different compared to populations or groups that are geographically distant. Different populations can then be distinguished based on single nucleotide polymorphisms (SNPs). Therefore, by utilizing differential sequence sites on the whole genome of different geographical regions, comparisons between samples can theoretically determine their pedigree differences.
[0004] However, accurately tracing the lineage of a sample is a difficult task. Traditional phylogenetic methods, also known as phylogenetic methods, trace the lineage by constructing a phylogenetic tree that establishes the phylogenetic relationships between the sample and known samples. This method requires rebuilding the entire phylogenetic tree for each sample tested, consuming significant computational power, which contradicts current energy conservation and emission reduction goals. Furthermore, it requires specialized knowledge to infer the taxonomic group of the current sample from the phylogenetic tree, making full batch automation impossible. Besides phylogenetic analysis based on phylogenetic trees, some studies utilize machine learning and short gene fragments for honey origin tracing. These studies analyze the sequence characteristics of DNA in pollen to determine the composition of flowering plants in honey, and then trace the honey's origin based on the differences in flowering plants in different geographical regions. However, this method uses relatively simple models, cannot effectively extract large-scale whole-genome data, and only tests on short fragments without extending to whole-genome big data. Summary of the Invention
[0005] To address the issues of the high computational demands of phylogenetic tree methods for each query and the inability of simple machine learning algorithms or shallow neural network algorithms to fully mine and utilize whole-genome big data, this invention provides a method based on deep learning to construct connections between whole-genome single nucleotide polymorphisms (SNPs) and different geographical populations or subpopulations, thereby achieving the purpose of phylogenetic tracing of biological samples. This method can infer and identify phylogenetic relationships between individuals, revealing their genetic origins and evolutionary processes with high accuracy and reliability. It can provide important genetic information for researchers and scientists, and also contribute to phylogenetic research in fields such as agriculture, biomedicine, and ecology, promoting the development and application of phylogenetic tracing technology.
[0006] To achieve the above technical objectives, the present invention adopts the following technical solution:
[0007] A pedigree tracing method based on whole-genome resequencing SNP data and deep learning can be used for pedigree tracing of agricultural products, wild animals, domestic animals, and pests, including the following steps:
[0008] 1) Collect biological samples from known geographical populations or subpopulations, perform whole-genome resequencing on the samples to obtain single nucleotide polymorphism variation results, and use them to construct a reference database of samples from different geographical populations or subpopulations;
[0009] 2) Convert single nucleotide polymorphism data (VCF data of SNPs) from samples from different geographical populations or subpopulations into DNA Fasta format sequence files, and map them into feature value matrices based on one-hot encoding;
[0010] 3) Assign geographical location probability values to samples collected from different geographical populations or subpopulations. The samples obtain different geographical location probability values according to their different sources. The probability value of each sample from its own geographical population or subpopulation is 100%, and the probability value from other geographical populations or subpopulations is 0%, thus obtaining a probability value label matrix.
[0011] 4) The feature value matrix obtained in step 2) and the probability value label matrix obtained in step 3) are used for training the deep learning neural network model developed in this invention. The feature value matrix in step 2) is used as the input layer of the neural network for feature learning and extraction. The probability value label matrix in step 3) provides the learning target for the neural network, enabling the network to self-adjust and optimize according to the desired output. By learning these labels, the neural network can establish a mapping relationship between input and output.
[0012] 5) Input the feature value matrix of the sample to be tested, and output the genealogy of the geographical population or subpopulation to which the sample to be tested belongs after training the deep learning-based neural network model in step 4).
[0013] Specifically, step 1) above includes:
[0014] 1.1) Collect biological samples from known geographic populations or subpopulations to construct a reference database. Ideally, multiple samples should be included from each geographic population or subpopulation to adequately represent its genetic diversity and heterogeneity. Furthermore, in fresh samples, endogenous nucleases have relatively less impact on nucleic acid (DNA), resulting in better DNA quality. Therefore, priority should be given to collecting fresh, live samples from known geographic populations or subpopulations. Generally, DNA sequencing requires specific DNA fragment lengths; therefore, DNA breakage and degradation should be minimized during sample preservation to ensure DNA integrity. Collected samples should be immediately stored in alcohol-containing containers to ensure smooth subsequent experiments. To further preserve DNA integrity, aliquoting is recommended to avoid repeated freeze-thaw cycles. Then, whole-genome DNA should be isolated and extracted in the laboratory using kits or the phenol-chloroform extraction method.
[0015] 1.2) Sequencing the extracted whole-genome DNA. Currently, nucleic acid sequencing commonly uses second-generation sequencing technology, i.e., sequencing-as-synthesis. During synthesis, as the DNA strand lengthens, the efficiency of DNA polymerase gradually decreases, and its specificity also begins to deteriorate. This leads to a problem: as the strand lengthens, the error rate of base synthesis also gradually increases. In addition, sometimes the sequencer may experience fluctuations in quality values due to insufficient reaction stability in the early stages of synthesis. Therefore, it is necessary to perform standard quality control and filtering of the raw sequencing data, including using FastQC software to evaluate low-quality sequences and using Trimmomatic software to remove sequencing adapter sequences.
[0016] 1.3) Align all quality-controlled and filtered short reads back to the reference genome to obtain sequence variation information.
[0017] 1.4) Convert the SAM file obtained from step 1.3) into a BAM file and sort it. Then perform PCR repeat labeling and recalibrate the base mass fraction.
[0018] 1.5) Use GATK4 software to perform variant detection on the file after format conversion and tagging in step 1.4) to obtain the single nucleotide polymorphism variant result VCF file.
[0019] 1.6) Perform quality control and filtering on the whole genome resequencing single nucleotide polymorphism variant results obtained in step 1.5), including QualByDepth (QD) > 2, FisherStrand (FS) > 60, RMSMappingQuality (MQ) > 40, StrandOddsRatio (SOR) > 3, and minor allele count > 3.
[0020] In practice, step 1.1) above should be implemented in a way that ensures the collected samples come from different ethnic groups and sampling locations within the same geographic population. Furthermore, the accuracy and traceability of the source location should be ensured in the selection of samples to guarantee complete and accurate coverage of the local geographic population.
[0021] In step 1.3) above, the BWA mem algorithm is used to perform sequencing read alignment with the reference genome and a reference sequence index is built using BWA. The Samtools tool is then used to sort the alignment results. Simultaneously, to improve GATK indexing efficiency, Samtools is used to create an index for the reference sequence.
[0022] In step 1.4) above, the SAM file is converted to a BAM file using Samtools software, and then the SortSam tool in picard software is used to sort the BAM file. The MarkDuplicates algorithm in GATK software is used to mark duplicates in the library, the BaseRecalibrator method in GATK is used to build a calibration model, and then the quality values are calibrated using the ApplyBQSR algorithm.
[0023] In step 1.5 above, GATK is used to filter SNP and INDEL sites, the VariantRecalibrator method is used to construct a calibration model, and the ApplyVQSR algorithm is used to apply the model to correct the quality value.
[0024] Step 2) above constructs a reference database for samples from different geographic populations or subpopulations based on the high-quality SNP variant site VCF files obtained after quality control and filtering in step 1.6). First, the single nucleotide polymorphism data is converted into Fasta format DNA sequence files using the vcf2phylip software. Then, the DNA sequences are mapped to numerical matrices using one-hot encoding according to the IUPAC encoding format. For example, when constructing the feature numerical matrix for each sample in the reference database, the one-hot encoding method is as follows: adenine (A) is encoded as [[1,0,0,0,0,0,0,0,0,0,0,0,0,0,0]], cytosine (C) is encoded as [[0,1,0,0,0,0,0,0,0,0,0,0,0,0]], and the encoding forms of other bases in IUPAC follow the same pattern.
[0025] In the probability value label matrix described in step 3) above, the geographical locations of samples collected from different geographical populations or subpopulations are encoded: each sample is encoded as 100% in the geographical population or subpopulation from which it originates, and as 0% in other geographical populations or subpopulations.
[0026] In step 4), the numerical matrices obtained in steps 2) and 3) are input into a training model based on a deep learning neural network. The variables encoded by the whole genome SNP loci sequence of samples from different geographic populations are input. After machine learning based on the source information in step 3), the genealogy of the geographic population or subpopulation is output.
[0027] The neural network model is built based on the PyTorch deep learning neural network framework. In an embodiment of this invention, a ResNet-based deep residual neural network is constructed, including a convolutional attention module based on residual shortcut connections, a max-pooling layer, a fuzzy pooling layer, and an iSQRT-COV layer. The convolutional attention module based on residual shortcut connections is a fusion of convolutional attention modules, including channel attention and spatial attention modules, based on the residual convolutional modules of ResNet. The input nucleic acid sequence (in the form of a feature data matrix) is first fed into a convolutional layer adapted for nucleic acid sequence feature extraction. The resulting features are then fed into the convolutional attention module (in a cascaded order, sequentially fed into the channel attention module and spatial attention module for element-wise matrix multiplication). The result is added to the result obtained through the residual shortcut connections and fused. The fused result is then fed into subsequent convolutional layers, max-pooling layers, fuzzy pooling layers, and iSQRT-COV layers for further computation, ultimately yielding the phylogenetic tracing output. The convolutional layer adapted for nucleic acid sequence feature extraction is a specially modified two-dimensional convolutional layer. It performs convolution operations on features sequentially along the nucleic acid sequence. Since nucleic acid sequences are directional during subsequent transcription and translation (i.e., the information of mRNA corresponds to the sequence information in the 5'->3' direction of the nucleic acid coding strand), the model should ideally extract nucleic acid sequence features in this direction to extract features with corresponding biological significance. Residual shortcut connections are used to alleviate the gradient vanishing problem in complex deep neural networks. A convolutional attention module is fused to the residual convolutional module of ResNet to allow the model to adaptively focus on key discriminative sites that are more conducive to lineage tracing, thereby achieving dataset-specific adaptive learning. To address the issue of small input shifts or transformations caused by pooling downsampling leading to degraded output results, our model uses a low-pass filter to combat aliasing before downsampling the input signal by incorporating a fuzzy pooling layer, making the model more robust to variance in local feature regions. Because gene data is highly heterogeneous, with varying lengths of the same gene segment across different species, and further complicated by alternative splicing, our neural network uses iSQRT-COV layers to unify the dimensions of the extracted feature matrix. By replacing the first-order mean with the second-order covariance, the deep network can utilize the relevant information between the learned channels, improving model accuracy while accelerating training speed.
[0028] Preferably, after the training in step 4) is completed, the model is evaluated based on the ten-fold cross-validation method. The model with high overall accuracy and narrower confidence interval is the optimal model, which is used for genealogy tracing in step 5).
[0029] In a specific embodiment of this invention, three real datasets and two simulated datasets were used to evaluate the effectiveness of the above-mentioned technical solution, including individuals from different subspecies or geographical populations such as the Eastern honeybee, red imported fire ants, and domesticated chickens. Testing showed that this invention demonstrated high accuracy in tracing pedigrees and querying the geographical populations from which samples originated. This means that the method can provide researchers and agricultural and forestry personnel with effective and accurate pedigree information. This invention has wide applications and importance in practical applications. First, in the field of genetic research, this invention can provide important pedigree information for genomics research. Through accurate inferences from deep learning neural network models, researchers can identify family relationships, kinship, and evolutionary history between individuals, which is crucial for understanding the genetic diversity, population changes, and evolutionary processes of species. Second, in the agricultural and forestry field, the method of this invention is of great significance for the traceability of agricultural products. By accurately identifying and tracing pedigree relationships, the quality, safety, and traceability of agricultural products can be ensured. Furthermore, this method can also be used for livestock breeding and animal and plant protection, assisting agricultural and forestry departments in pedigree management and conservation efforts. In wildlife conservation and ecological research, by analyzing the phylogenetic relationships among individual wild animals, researchers can understand population structure, gene flow, and the evolutionary history of species, which helps in the development of conservation strategies and management measures.
[0030] Furthermore, the genetic locus interpretability workflow proposed in this invention has multiple applications in practice. First, this method can guide downstream probe or barcode design. By analyzing genetic locus information, we can identify which loci are most critical for specific research or applications. Based on these critical loci, we can design specific probes or barcodes to specifically detect or label genes or sequences of interest, thus efficiently screening and analyzing relevant genes or sequences and reducing sequencing costs and complexity. Second, this method can also help optimize the cost and complexity of whole-genome sequencing. By interpreting genetic locus information, we can identify which loci are critical and most meaningful for specific research or applications. Compared to sequencing the entire genome, sequencing only these critical loci can save significant costs and time while reducing the complexity of data analysis. The genetic locus interpretability workflow proposed in this invention has great potential in guiding downstream probe or barcode design and optimizing the cost and complexity of whole-genome sequencing. It not only helps save resources and effort but also improves the efficiency of research and applications.
[0031] In summary, the deep learning-based phylogenetic method based on whole-genome resequencing SNP data presented in this invention has significant practical application value in fields such as genetic research, agriculture, forestry, and ecological conservation. It provides a tool for accurately tracing phylogenetic relationships, thereby helping researchers and frontline staff in agricultural and forestry bureaus understand and protect species genetic diversity, and accelerating the development of scientific research and applications in related fields.
[0032] In summary, this invention provides a pedigree tracing method based on single nucleotide polymorphism (SNP) big data obtained from whole-genome resequencing and deep learning. By directly utilizing the variation information of the whole genome, it can obtain more discriminative information compared to existing short-sequence-based methods, thereby distinguishing geographically close subgroups and improving the resolution of pedigree tracing (further narrowing the geographical area). Furthermore, deep learning methods can better fit large datasets at the whole-genome level compared to shallow network models in machine learning. This invention can obtain pedigree origins end-to-end through training on a reference dataset. Therefore, compared to existing methods, this invention has more comprehensive feature coverage and higher feasibility in both the initial reference dataset information and subsequent feature extraction. This invention has significant implications for agricultural product pedigree tracing, livestock breeding, and animal and plant protection. Attached Figure Description
[0033] Figure 1 The genealogical tracing deep learning neural network model diagram (TraceNet) of this invention.
[0034] Figure 2 Factors and their impact on the accuracy of the final deep learning neural network model are described. The vertical axis represents accuracy, and the horizontal axes represent a) learning rate, b) batch size, and c) number of epochs, respectively.
[0035] Figure 3 The impact of highly variable SNP information sites on the accuracy of neural network models and the comparison of algorithm accuracy on real and simulated datasets are presented. The vertical axis represents accuracy. a) shows the impact of no filtering of SNP sites, filtering of SNP sites with a mutation value FST > 0.5, and filtering of SNP sites with a mutation value FST > 0.8 on model accuracy; b) shows the accuracy of the neural network model on two real datasets (red imported fire ants and domestic chickens) and two simulated datasets (Msprime population 3 and Msprime population 4). In the figure, TraceNet is the phylogenetic deep learning neural network model of this invention, and DeepFormer is a hybrid deep neural network model.
[0036] Figure 4 The TraceNet model of this invention is used to distinguish key information sites of various geographic populations, including: Aba, Central, Qinghai, Hainan, Northeast, Kashmir, Bomi, and Taiwan. The characters in the figure are IUPAC encoded. Detailed Implementation
[0037] The following is a more detailed description of the present invention. The parameters and specific implementation details are used to explain the feasibility and implementation effects of the present invention, and do not constitute a limitation of the present invention.
[0038] The method of this invention was tested on three real-world datasets and two simulated datasets to investigate its accuracy. The three real-world datasets are large-scale population genome data of the Eastern honeybee (Apis cerana), red imported fire ants, and domesticated chickens; the two simulated datasets were generated using msprime software. The Eastern honeybee is a major native pollinator in many parts of Asia, with a highly diverse geographical range of habitats from tropical to northern temperate zones. Due to varying levels of geographic isolation, incomplete phylogenetic classification, and hybridization introgression, different populations of the Eastern honeybee exhibit heterogeneous levels of genomic differentiation, thus representing a suitable population genome dataset to evaluate the performance of the method of this invention and to further explore which factors affect its accuracy. The other two widely used real-world datasets represent: red imported fire ants (Solenopsis invicta), one of the most concerning invasive species, which has spread to many countries worldwide; and domesticated chickens, representing one of the largest livestock meat consumption categories globally.
[0039] 1. Nucleic acid acquisition and sequencing after sample collection
[0040] Each collected sample was placed in a sterile centrifuge tube and preserved in alcohol for subsequent nucleic acid sequence extraction. Nucleic acid sequences were extracted and purified using the QIAamp DNA Micro Kit. The extracted DNA sequences were sent to a sequencing company for library construction, followed by bidirectional whole-genome sequencing.
[0041] 2. Detailed information and influencing factors of real and simulated datasets.
[0042] The Oriental Honeybee dataset, combining genomic, ecological, and morphological data, indicates that the continental lineage consists of eight evolutionarily independent native populations, including those in central and northeastern China, Qinghai, Aba, Bomi, Kashmir, Hainan, and Taiwan. The genomic dataset includes 362 worker honeybee individuals with an average genome coverage of 3.5–5 GB. The Red Imported Fire Ant dataset comprises five populations: three within their native range in northeastern Argentina and two invasive in the southeastern United States. Diploid female specimens were sampled from each population and from each distinct group, totaling 193 individuals. Samples collected from each population were used to construct seven restriction site (RAD) related libraries. The domestic chicken whole-genome sequencing dataset includes samples from 138 different breeds, ecotypes, and domestic populations, comprising five populations (Hainan subspecies of Red Junglefowl, Jiangxi Silkie Chicken, Yuanbao Chicken, Shigatse Domestic Chicken of Tibet, and Lhasa Domestic Chicken of Tibet). The parameters for the two simulated datasets are as follows: each population consists of 30 individuals, and the base mutation rate is 3 × 10⁻⁶. -9 The initial group size was 10 5 The time for group differentiation is 10. 10 generation.
[0043] 3. Acquisition of single nucleotide polymorphisms
[0044] Single nucleotide diversity (SNPs) are widely distributed throughout the genome, exhibiting rich diversity and genetic stability, and have been widely applied in genetics, functional genomics, and molecular breeding. Currently, SNP markers have become widely adopted DNA markers and are gradually becoming ideal molecular markers for species and variety identification. FastQC tools were used for data quality control and cleaning, and BWA software was used to construct an index of the reference genome for rapid searching and location when aligning sequences back to the reference genome. Sequences corresponding to the same chromosome in the SAM file were sorted by coordinate from smallest to largest. The MarkDuplicates algorithm in the picard tool was used to label PCR replication sequences in the BAM file using the FLAG information column. The HaplotypeCaller tool was used to detect variations in the BAM file, generating a VCF file for subsequent site quality control and annotation.
[0045] 4. Deep learning model training
[0046] This invention utilizes the PyTorch framework to construct a deep learning neural network algorithm to train numerical matrix data encoding nucleic acid sequences. The deep learning neural network in this invention is based on ResNet, a deep residual neural network (see...). Figure 1 Part 1: Convolutional Attention Module Based on Residual Shortcut Connections, including a large number of residual shortcut connections (see...). Figure 1Part 2: Residual Shortcut Connections) to obtain better representations of key features, thereby enhancing its learning ability. Theoretically, deeper neural networks are more powerful than ordinary neural networks containing a single hidden layer by representing richer features at different levels of abstraction. The complexity of the network model provides an important foundation for extracting key features from population genomic data. Focusing on key features through adaptive learning: The model learns to focus on which sites and where, which plays a crucial role in the final model performance. Especially when performing phylogenetic tracing on samples with geographically similar populations, the algorithm's ability to distinguish subtle key features is highly demanding. Convolutional Attention Module (CBAM, see Figure 1 Part 3: Convolutional Attention Module) is incorporated into the method of this invention, enabling the model to better adaptively select and extract features. CBAM includes two functional models: a channel attention model and a spatial attention model (…). Figure 1 Part 3 of the method not only tells the neural network where to focus its attention but also enhances its representation of regions of interest, enabling the network to focus more on important features and suppress unnecessary ones. While pooling layers (max pooling and average pooling) in convolutional neural networks are thought to collectively introduce some translation invariance, making the model more robust to variance in local regions of features, the reality is that model performance fluctuates significantly when the network's input changes slightly. During sequencing data processing, a few low-quality sites may be removed, but the entire sequence can still be used for analysis. The site-removed sequence has a "shift" compared to the complete standard sequence. From a biological perspective, both the complete and shifted sequences originate from the same source, and the model should produce consistent performance regardless of the input sequence, even with the shift. Therefore, the method of this invention addresses this issue by adding an anti-aliasing downsampling layer (also known as a blur pooling layer, BlurPool) (see [link to invention]). Figure 1 Part 4: Fuzzy Pooling Layer. This layer employs classic signal processing (low-pass filtering) to mitigate the jagged edges of deep neural networks. However, different gene lengths prevent the same model from utilizing multiple different genes from different populations for source tracing. To address this issue, this invention introduces iterative matrix square root normalization (iSQRT-COV) for the covariance pooling layer (see...). Figure 1 Part 5: iSQRT-COV. Furthermore, the iSQRT-COV layer can leverage information learned from different channels by deep neural networks to achieve better performance, generalization, finer classification, faster convergence, and more efficient computation.
[0047] 5. Hyperparameter settings for 10-fold cross-validation
[0048] Ten-fold cross-validation is used to evaluate the accuracy and performance of the trained model. Ten-fold cross-validation is a commonly used model testing method that divides the dataset into 10 parts, alternating between using 9 parts as training data and 1 part as test data. Each trial yields a corresponding accuracy rate, and the average of the 10 accuracy rates is used as an estimate of the algorithm's accuracy.
[0049] Figure 2 This invention demonstrates that the model is unaffected by batch size and number of rounds settings, and can be adjusted according to the user's machine configuration without affecting the final model's accuracy. According to... Figure 2 The effect of the learning rate is suggested to be 10. -3 As the initial setting value.
[0050] 6. Source of collected samples and interpretability of the model
[0051] The deep learning neural network model trained through the above steps can then be used to trace the lineage of each sample. Figure 3 Figure a) shows that when screening for whole-genome single nucleotide polymorphisms, the method of the present invention can not only improve the accuracy but also reduce the training time and memory consumption of the model. Figure 3 Figure b) shows that the method developed in this invention has high accuracy on both real and simulated datasets. Figure 4 The deep learning neural network model of this invention is used to distinguish the nucleic acid sites of different geographical populations of the Eastern honeybee. On the one hand, it can be used for model interpretability, and on the other hand, it can be used to guide downstream experiments, such as guiding the synthesis of upstream and downstream primers for specific sites, without having to perform pedigree tracing again through whole genome sequencing, which can further facilitate and reduce costs and time.
[0052] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made in accordance with the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A pedigree tracing method based on whole-genome resequencing SNP data and deep learning, comprising the following steps: 1) Collect biological samples from known geographical populations or subpopulations, perform whole-genome resequencing on the samples to obtain single nucleotide polymorphism variation results, and use them to construct a reference database of samples from different geographical populations or subpopulations; 2) Convert single nucleotide polymorphism data of samples from different geographical populations or subpopulations into DNA Fasta format sequence files, and map them into feature value matrices based on one-hot encoding; 3) Assign geographical location probability values to samples collected from different geographical populations or subpopulations. The samples obtain different geographical location probability values according to their different sources. The probability value of each sample from its own geographical population or subpopulation is 100%, and the probability value from other geographical populations or subpopulations is 0%, thus obtaining a probability value label matrix. 4) The feature value matrix obtained in step 2) and the probability value label matrix obtained in step 3) are used for training the deep learning neural network model. The feature value matrix obtained in step 2) serves as the input layer of the neural network for feature learning and extraction; the probability value label matrix obtained in step 3) provides the learning target for the neural network, enabling it to self-adjust and optimize based on the desired output, establishing a mapping relationship between input and output. The trained deep learning neural network model is a deep residual neural network based on ResNet, including a convolutional attention module based on residual shortcut connections, max pooling layers, fuzzy pooling layers, and iSQRT-COV layers. The convolutional attention module based on residual shortcut connections is a fusion of convolutional attention modules on top of ResNet's residual convolutional modules, including channel attention modules. Block and spatial attention modules; the input nucleic acid sequence is first fed into a convolutional layer adapted for nucleic acid sequence feature extraction. The convolutional layer adapted for nucleic acid sequence feature extraction is a specially modified two-dimensional convolutional layer. It performs convolutional operations on features along the nucleic acid sequence. The resulting features are then fed into the convolutional attention module, and then through the channel attention module and the spatial attention module to perform matrix element-wise multiplication operations. The result of the operation is added to the result of the residual shortcut connection and fused. The fused result is fed into the subsequent convolutional layer, max pooling layer, fuzzy pooling layer and iSQRT-COV layer for further operations to obtain the phylogenetic tracing output result; 5) Input the feature numerical matrix of the sample to be tested, and through the deep learning neural network model trained in step 4), output the phylogenetic of the geographical population or subpopulation to which the sample to be tested belongs.
2. The genealogical tracing method as described in claim 1, characterized in that, Step 1) includes: 1.1) Collect biological samples from known geographical populations or subpopulations, and isolate and extract the whole genome DNA from each sample; 1.2) Sequencing of the whole genome DNA of each sample, and quality control and filtering of the sequencing results; 1.3) Align the short reads obtained after quality control and filtering back to the reference genome to obtain sequence variation information; 1.4) Convert the SAM file obtained from step 1.3) into a BAM file and sort it. Then perform PCR repeat labeling and recalibrate the base mass fraction. 1.5) Use GATK4 software to perform variant detection on the file after format conversion and tagging in step 1.4) to obtain the single nucleotide polymorphism variant result VCF file; 1.6) Perform quality control and filtering on the whole genome resequencing single nucleotide polymorphism variation results obtained in step 1.5) to obtain the final SNP variation site VCF file and construct a reference database for samples from different geographical populations or subpopulations.
3. The genealogical tracing method as described in claim 2, characterized in that, In step 1.1), each geographic population or subpopulation source includes multiple samples; in step 1.2), quality control and filtering of sequencing results include the removal of low-quality sequences and sequencing adapter sequences.
4. The genealogical tracing method as described in claim 2, characterized in that, The quality control and filtering criteria for whole-genome resequencing single nucleotide polymorphism (SNP) variants in step 1.6) are: QualByDepth > 2, FisherStrand > 60, RMSMappingQuality > 40, StrandOddsRatio > 3, and minor allele count > 3.
5. The genealogical tracing method as described in claim 2, characterized in that, Step 2) Construct a reference database for different geographic populations or subpopulations based on the SNP variant site VCF files obtained after quality control and filtering in Step 1.6). First, use vcf2phylip software to convert the VCF files into Fasta format DNA sequence files, and then map the DNA sequences into a numerical matrix based on the one-hot encoding method according to the IUPAC encoding format.
6. The genealogical tracing method as described in claim 1, characterized in that, In the probability value label matrix of step 3), each sample is coded as 100% for its geographic population or subpopulation of origin, and 0% for other geographic populations or subpopulations.
7. The genealogical tracing method as described in claim 1, characterized in that, Step 4) After training, the model is evaluated using the ten-fold cross-validation method. The model with the highest overall accuracy and the narrowest confidence interval is the optimal model, which is then used for genealogy tracing in step 5).