Deep learning-based pig whole genome low-density SNP chip development method

By using deep learning methods to screen and construct low-density pig whole-genome SNP chips, the problems of cost and technical complexity of high-density chips in small-scale farms have been solved, achieving more economical, flexible and applicable pig breeding results, and improving production efficiency and health levels.

CN121709018APending Publication Date: 2026-03-20AGRICULTURAL GENOMICS INSTITUTE AT SHENZHEN CHINESE ACADEMY OF AGRICULTURAL SCIENCES (SHENZHEN BRANCH GUANGDONG LABORATORY FOR LINGNAN MODERN AGRICULTURE)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511607907.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-07
Filing Date
2025-11-05
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

High-density pig breeding chips are costly and technically complex in small-scale or decentralized farms, making them difficult to popularize. Existing low-density chips are insufficient in terms of flexibility and economy to meet the needs of personalized breeding.

Method used

Using a deep learning-based approach, significant associated SNP loci were screened through 50K microarray sequencing, GWAS analysis, GBLUP and Bayesian models, CNN models, and the deep explain algorithm. Low-density whole-genome SNP microarrays were constructed to optimize prediction accuracy and remove redundancy, ensuring the stability of genetic parameters.

Benefits of technology

It provides more economical, flexible and applicable pig breeding solutions, improves the production efficiency and animal health of small-scale farms, meets personalized breeding needs, and promotes the sustainable development of agricultural production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121709018A_ABST
    Figure CN121709018A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of bioinformatics, and discloses a pig whole genome low-density SNP chip development method based on deep learning, which comprises the following steps: (1) performing 50K chip sequencing to detect genotypes of different variety groups; (2) carrying out systematic arrangement and statistical analysis on genotype and phenotype data of the group samples; and (3) performing strict quality control on phenotypes and genotypes of group detection, wherein quality control standards comprise that SNP sites with deletion rate exceeding 5%, Hardy Weinberg equilibrium test P value less than 106 and minimum allele frequency less than 1% are eliminated. According to the pig whole genome low-density SNP chip development method based on deep learning, a high-density chip usually needs expensive infrastructures and a large amount of capital investment, which is possibly unpractical for a small-scale farm, and compared with the low-density chip, a more economical and practical choice is provided, the hardware cost and the operation cost are lower, and the development cost is lower. And the method is more suitable for small-scale or distributed farms.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics, specifically a pig whole genome low-density SNP chip development method based on deep learning. BACKGROUND

[0002] Pig breeding using whole genome high-density chips can accelerate the breeding process at the genetic level and promote genetic improvement of pig populations. Currently, commercially available high-density chips provide breeders with powerful tools due to their full genome coverage, high polymorphism, and high throughput. By analyzing large-scale SNP data, breeders can more accurately assess the genetic diversity of pig populations, select parents, screen superior individuals, and predict and improve the genetic characteristics of pigs.

[0003] Although whole genome high-density chips have great potential in pig breeding, their application also faces some challenges, such as high-density chips are usually used in large-scale pig farms, which are expensive and technically complex. In contrast, low-density chips provide a more economical and flexible solution for small-scale or decentralized breeding sites. The following are the importance and advantages of low-density chips in pig breeding. SUMMARY

[0004] To achieve the above-mentioned more economical and flexible effects, the present application provides the following technical solutions: a pig whole genome low-density SNP chip development method based on deep learning, comprising the following steps:

[0005] (1) Perform 50K chip sequencing to detect the genotypes of different breed populations;

[0006] (2) Systematically organize and statistically analyze the genotype and phenotype data of population samples;

[0007] (3) Strictly control the quality of the detected phenotypes and genotypes of the population, and the quality control standards include: excluding SNP sites with a missing rate of more than 5%, a Hardy-Weinberg equilibrium test P value less than 10^6, and a minimum allele frequency less than 1%, and excluding individual samples with a SNP site missing rate greater than 10%;

[0008] (4) Based on the data organized in step (2), perform whole genome association (GWAS) analysis for each phenotype to determine the SNP sites significantly associated with the target trait;

[0009] (5) For each trait, sort according to the association scores obtained from GWAS analysis, and construct SNP sets of different sizes for subsequent breeding value estimation;

[0010] (6) Using GBLUP (Genomic Best Linear Unbiased Prediction) and Bayesian model, breeding value estimation is performed on the SNP sets of different sizes screened in step (5), and comparative analysis is performed to identify the SNP sites with the greatest contribution to breeding value estimation;

[0011] (7) With the aid of a convolutional neural network (CNN) model, the genotypes of the minimum SNP set determined in step (6) are taken as input for phenotype prediction analysis, and the prediction accuracy is optimized by adjusting the model parameters;

[0012] (8) The contribution of each SNP site in step (7) is quantitatively evaluated using the deep explain algorithm, and the ranking is performed according to the score results;

[0013] (9) Based on the SNP site score ranking obtained in step (8), a different number of SNP site sets are selected for each trait, the accuracy of different size SNP sets in breeding value estimation is evaluated, and the minimum SNP set size required is determined;

[0014] (10) The minimum SNP sets obtained for each phenotypic trait in step (9) are taken and redundant items are removed to ultimately obtain a specific number of SNP sites;

[0015] (11) The SNP sites obtained in step (10) are further analyzed to eliminate abnormal sites, and a low-density whole-genome SNP chip is finally constructed.

[0016] In one specific embodiment, the SNP set constructed in step (5) includes 50k, 30k, 20k, 10k, 5k, 1k, etc. of different sizes.

[0017] In one specific embodiment, the SNP site with the greatest contribution to breeding value estimation is identified by comparing the breeding value estimation results of GBLUP and Bayesian model in step (6).

[0018] In one specific embodiment, the parameters of the convolutional neural network (CNN) model are adjusted in step (7) to minimize the mean square error (MSE) between the predicted value and the true value, thereby optimizing the prediction accuracy.

[0019] In one specific embodiment, the number of SNP sites ultimately obtained in step (10) is 13504.

[0020] In one specific embodiment, the further analysis in step (11) includes calculating the minimum allele frequency, adjacent SNP site interval, and linkage disequilibrium.

[0021] In one specific embodiment, the final constructed low-density whole genome SNP chip has a linkage disequilibrium of about 623601, a minimum allele frequency of about 248894, and a SNP interval of about 68367kb.

[0022] In one specific embodiment, the step (8) uses a deep explain algorithm to score the weight of each SNP site.

[0023] In one specific embodiment, the quality control standards in step (3) include removing SNP sites with a missing rate of more than 5%, SNP sites with a Hardy-Weinberg equilibrium test P value less than 10^6, SNP sites with a minimum allele frequency less than 1%, and individual samples with a SNP site missing rate greater than 10%.

[0024] In one specific embodiment, the phenotype data includes data of 4 important economic traits.

[0025] Compared with the prior art, the present application provides a pig whole genome low-density SNP chip development method based on deep learning, which has the following beneficial effects:

[0026] 1. The pig whole genome low-density SNP chip development method based on deep learning, high-density chips usually require expensive infrastructure and a large amount of capital investment, which may not be practical for small-scale farms. In contrast, low-density chips provide a more economical and affordable option, with lower hardware costs and operating expenses, making them more suitable for small-scale or decentralized farms.

[0027] 2. The pig whole genome low-density SNP chip development method based on deep learning, low-density chips have application potential in different sizes and types of pig farms. Low-density chips are usually designed to be more flexible and can adapt to different environments and scenarios. Whether in family farms or small-scale farms, low-density chips can help farmers achieve more precise breeding and management, improve production efficiency and animal health levels.

[0028] 3. The pig whole genome low-density SNP chip development method based on deep learning, due to the flexibility of low-density chips, it can better meet the needs of individualized breeding. According to the needs of specific pig populations, data can be collected and analyzed in a customized manner, providing more personalized and precise breeding solutions.

[0029] 4、The pig whole genome low-density SNP chip development method based on deep learning helps promote the sustainable development of agricultural production. Its low cost and applicability can promote more farmers to adopt modern technology, improve the efficiency and sustainability of agricultural production.

[0030] 5、The pig whole genome low-density SNP chip development method based on deep learning plays an important role in pig breeding. Its cost-effectiveness, wide applicability, flexibility and ease of use provide a practical breeding solution for small-scale farms, which helps improve the production performance and health level of pigs. BRIEF DESCRIPTION OF DRAWINGS

[0031] Figure 1 is a preparation flowchart of a pig whole genome low-density SNP chip;

[0032] Figure 2 is a distribution diagram of SNP sites used in the chip on the chromosome;

[0033] Figure 3 is a r2 frequency distribution diagram of adjacent SNP sites;

[0034] Figure 4 is a minimum allele frequency distribution diagram of SNP sites in the chip. DETAILED DESCRIPTION

[0035] The technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0036] Please refer to Figures 1-4, the present application provides a technical solution: a pig whole genome low-density SNP chip development method based on deep learning, comprising the following steps:

[0037] (1) Perform 50K chip sequencing to detect the genotypes of different breed populations;

[0038] (2) Systematically organize and statistically analyze the genotype and phenotype data of the population samples;

[0039] (3) Strictly control the quality of the detected phenotypes and genotypes of the population, and the quality control standards include: removing SNP sites with a missing rate of more than 5%, a Hardy-Weinberg equilibrium test P value less than 10^6 and a minimum allele frequency less than 1%, and removing individual samples with a SNP site missing rate of more than 10%;

[0040] (4) Based on the data sorted in step (2), perform genome-wide association (GWAS) analysis for each phenotype to determine the SNP sites significantly associated with the target traits;

[0041] (5) For each trait, sort according to the association scores obtained from the GWAS analysis, and construct SNP sets of different sizes for subsequent breeding value estimation;

[0042] (6) Use GBLUP (Genomic Best Linear Unbiased Prediction) and Bayesian models to estimate breeding values for different size SNP sets screened in step (5), and perform comparative analysis to identify the SNP sites that contribute most to breeding value estimation;

[0043] (7) Using a convolutional neural network (CNN) model, use the genotype of the smallest SNP set determined in step (6) as input for phenotype prediction analysis, and optimize prediction accuracy by adjusting model parameters;

[0044] (8) Use the deep explain algorithm to quantitatively evaluate the contribution of each SNP site in step (7), and sort according to the score results;

[0045] (9) Based on the SNP site score ranking obtained in step (8), select different numbers of SNP site sets for each trait, evaluate the accuracy of different size SNP sets in breeding value estimation, and determine the required minimum SNP set size;

[0046] (10) Take the union of the minimum SNP sets obtained for each phenotypic trait in step (9), remove redundant items, and finally obtain a specific number of SNP sites;

[0047] (11) Further analyze the SNP sites obtained in step (10) to eliminate abnormal sites, and finally construct a low-density whole-genome SNP chip.

[0048] Further, the SNP sets constructed in step (5) include 50k, 30k, 20k, 10k, 5k, 1k, etc. of different sizes.

[0049] Further, in step (6), the SNP sites that contribute most to breeding value estimation are identified by comparing the breeding value estimation results of GBLUP and Bayesian models.

[0050] Further, in step (7), the parameters of the convolutional neural network (CNN) model are adjusted to minimize the mean square error (MSE) between the predicted value and the true value, to optimize the prediction accuracy.

[0051] Further, the number of SNP sites obtained in step (10) is 13504.

[0052] Further, the further analysis in step (11) includes calculating the minimum allele frequency, the interval between adjacent SNP sites, and the linkage disequilibrium.

[0053] Further, the low-density whole-genome SNP chip finally constructed has a linkage disequilibrium of about 623601, a minimum allele frequency of about 248894, and an interval between SNPs of about 68367 kb.

[0054] Further, the weight of each SNP site is scored using the deep explain algorithm in step (8).

[0055] Further, the quality control standards in step (3) include removing SNP sites with a missing rate of more than 5%, SNP sites with a Hardy-Weinberg equilibrium test P value of less than 10^6, SNP sites with a minimum allele frequency of less than 1%, and individual samples with a SNP site missing rate of more than 10%.

[0056] Further, the phenotype data includes data of 4 important economic traits.

[0057] The pig whole-genome low-density SNP chip development method based on deep learning more effectively solves the problem in the following ways:

[0058] First, GWAS analysis is used to accurately identify SNP sites significantly associated with target traits, reducing redundancy. This step screens SNP sites highly related to key economic traits of pigs (such as growth performance) through whole-genome association analysis, ranks based on association scores (corrected P value), and ensures that the sites have biological significance. For example, in one embodiment, GWAS analysis is performed on 4 economic trait data of 3637 pigs, and a set of SNP sites significantly associated with the traits is screened, as described in steps (4) and (5), to construct different size sets (such as 50k to 1k), laying the foundation for subsequent optimization.

[0059] Second, combine GBLUP, Bayesian model, and CNN deep learning model to optimize SNP site screening and prediction accuracy. This step uses GBLUP and Bayesian model to evaluate the breeding value estimation ability of different size SNP sets, and uses CNN model for phenotype prediction to minimize mean square error (MSE) and improve prediction accuracy. For example, specifically, in steps (6) and (7), the minimum SNP set is input into the CNN model after training to optimize parameters such as adjusting the learning rate to reduce MSE, ensuring that the prediction accuracy of pig economic traits (such as meat yield) is close to the high-density chip level.

[0060] Third, quantify SNP site contribution through deep explain algorithm to ensure that the minimum set selected maintains high accuracy in breeding value estimation. This step is based on the weight score of the deep learning model to rank the importance of each SNP site and prioritize high-contribution sites. For example, in one embodiment, the deep explain algorithm ranks the site weight score output by the CNN model in step (8), and after sorting, selects different number sets (such as 1k-5k) for each trait (such as fertility), and verifies the accuracy through GBLUP to finally determine the minimum set size.

[0061] Finally, introduce formulaic conditional judgment (such as based on statistical threshold or genetic parameters) in key steps to dynamically adjust screening criteria, thereby maintaining genetic parameter stability while reducing density. This step sets threshold conditions such as removing SNP sites that do not meet genetic parameters to ensure linkage disequilibrium (LD) and minimum allele frequency (MAF) stability. For example, in the quality control of step (3), the formulaic conditions include: missing rate > 5% (range 0-100%, optimal value 5% to exclude low-quality data), Hardy-Weinberg equilibrium test P value < 10^{-6} (P value range 0-1, optimal threshold 10^{-6} to ensure population genetic balance), and MAF < 0.01 (range 0-0.5, optimal value 1% to avoid rare allele interference). These formulas are based on statistical genetics principles and dynamically filter abnormal sites (such as removing SNP with too low MAF), further analyze LD and MAF in step (11), and finally maintain LD at 0.623 (close to high-density 0.658) and MAF at 0.249 for low-density chips (such as 13504 sites), ensuring that genetic parameters do not degrade due to reduced density.

[0062] Step (1) involves genotype detection, meaning obtaining SNP site data through chip sequencing, for example, detecting pig population genotypes using a 50K chip, and the term SNP site uniformly refers to single nucleotide polymorphism positions. Step (2) organizes genotype and phenotype data, meaning systematically processing raw information, specifically in the context of pig breeding, phenotype data includes economic traits such as growth rate, ensuring consistency in phenotype and genotype terminology. Step (3) performs quality control, meaning filtering low-quality data, for example, removing SNP sites with a missing rate exceeding a threshold or a MAF below 1%, the parameter MAF (minimum allele frequency) is uniformly defined as the lower limit of allele frequency, ranging from 0 to 0.5, with an optimal value close to 0.05 to balance diversity and statistical efficiency. Step (4) performs GWAS analysis, meaning identifying trait-associated sites, association scores are based on P values, for example, P values are used to quantify the significance of SNPs to target traits, and the term significance threshold is uniformly represented by the letter alpha (e.g., alpha = 10^{-6}). Step (5) constructs SNP sets, meaning creating subsets of different sizes in order, for example, generating 1k-50k sets after sorting P values for pig economic traits, and the term association score is equivalent to the corrected P value. Step (6) estimates breeding values, meaning evaluating SNP contribution, using GBLUP and Bayesian models, and the accuracy of breeding value estimation A is uniformly represented by prediction accuracy, ranging from 0 to 1, with an optimal value approaching 1. Step (7) applies a CNN model, meaning phenotype prediction optimization, the formula mean squared error MSE = \frac{1} {n} \sum_ {i=1}^{n} (\hat{y}_i - y_i) ^2, parameter meaning: \hat{y}_i is the predicted phenotype value, y_i is the true phenotype value, n is the number of samples; n ranges from n>0, and the optimal value depends on the size of the data; minimizing MSE improves accuracy as it quantifies prediction bias. Step (8) quantifies SNP contribution, meaning sorting weight scores, for example, the output of the deep explain algorithm is uniformly referred to as association score. Step (9) evaluates the size of the SNP set, meaning determining the minimum number of effective sites, and the term LD (linkage disequilibrium) uniformly represents site correlation, ranging from 0 to 1, with an optimal value greater than 0.6 to maintain genome coverage. Step (10) integrates site sets, meaning merging the smallest set of traits, removing redundancies. Step (11) constructs a chip, meaning final filtering sites, for example, removing outliers after analyzing MAF and LD, ensuring parameter consistency.

[0063] While embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary of the principles and application of the present application. Numerous modifications and adaptions can be effected without departing from the spirit and scope of the present application, which is not limited to the exact construction and arrangement described. It is intended, therefore, to cover all modifications and adaptions that fall within the scope of the claims and their equivalents.

Claims

1. A method for developing a low-density SNP microarray for the entire porcine genome based on deep learning, characterized in that, Includes the following steps: (1) Perform 50K chip sequencing to detect the genotypes of different breed populations; (2) Systematically organize and statistically analyze the genotype and phenotypic data of the population samples; (3) Strict quality control is carried out on the phenotype and genotype of the population test. The quality control standards include: removing SNP sites with a deletion rate of more than 5%, a Hardy Weinberg equilibrium test P value of less than 10^6 and a minimum allele frequency of less than 1%, and removing individual samples with a SNP site deletion rate of more than 10%. (4) Based on the data collected in step (2), genome-wide association (GWAS) analysis is performed for each phenotype to identify SNP sites that are significantly associated with the target trait; (5) For each trait, the association scores obtained from the GWAS analysis are sorted and SNP sets of different sizes are constructed for subsequent breeding value estimation. (6) Using GBLUP (Genome Best Linear Unbiased Prediction) and Bayesian models, the breeding values ​​of the different sizes of SNP sets screened in step (5) are estimated and compared to identify the SNP sites that contribute the most to the breeding value estimation. (7) Using a convolutional neural network (CNN) model, the genotypes of the minimum SNP set determined in step (6) are used as input to perform phenotypic prediction analysis, and the prediction accuracy is optimized by adjusting the model parameters; (8) Use the deep explain algorithm to quantify the contribution of each SNP site in step (7) and sort them according to the scoring results; (9) Based on the SNP locus score ranking obtained in step (8), select different numbers of SNP locus sets for each trait, evaluate the accuracy of different sizes of SNP sets in estimating breeding values, and determine the minimum required SNP set size; (10) Take the union of the minimum SNP sets obtained for each phenotypic trait in step (9), remove redundant terms, and finally obtain a specific number of SNP sites; (11) Further analysis of the SNP sites obtained in step (10) was performed to remove abnormal sites and finally a low-density whole-genome SNP chip was constructed.

2. The method for developing a low-density SNP chip for the entire porcine genome based on deep learning according to claim 1, characterized in that, The SNP set constructed in step (5) includes different sizes such as 50k, 30k, 20k, 10k, 5k, and 1k.

3. The method for developing a low-density SNP chip for the entire pig genome based on deep learning according to claim 1, characterized in that, In step (6), the SNP sites that contribute the most to the breeding value estimation are identified by comparing the breeding value estimation results of GBLUP and Bayesian models.

4. The method for developing a low-density SNP chip for the entire pig genome based on deep learning according to claim 1, characterized in that, In step (7), the mean square error (MSE) between the predicted value and the true value is minimized by adjusting the parameters of the convolutional neural network (CNN) model in order to optimize the prediction accuracy.

5. The method for developing a low-density SNP chip for the entire pig genome based on deep learning according to claim 1, characterized in that, The final number of SNP sites obtained in step (10) is 13,504.

6. The method for developing a low-density SNP chip for the entire pig genome based on deep learning according to claim 1, characterized in that, The further analysis in step (11) includes calculating the minimum allele frequency, the interval between adjacent SNP sites, and linkage disequilibrium.

7. The method for developing a low-density SNP chip for the entire pig genome based on deep learning according to claim 1, characterized in that, The final low-density whole-genome SNP array has linkage disequilibrium of approximately 623,601, minimum allele frequency of approximately 248,894, and SNP interval of approximately 68,367 kb.

8. The method for developing a low-density SNP chip for the entire pig genome based on deep learning according to claim 1, characterized in that, In step (8), the deep explain algorithm is used to score the weight of each SNP site.

9. The method for developing a low-density SNP chip for the entire pig genome based on deep learning according to claim 1, characterized in that, The quality control criteria in step (3) include removing SNP sites with a deletion rate of more than 5%, SNP sites with a Hardy-Weinberg equilibrium test P value of less than 10^6, SNP sites with a minimum allele frequency of less than 1%, and individual samples with a SNP site deletion rate of more than 10%.

10. The method for developing a low-density SNP chip for the entire pig genome based on deep learning according to claim 1, characterized in that, The phenotypic data includes data on four important economic traits.