Cow and live pig high-quality breeding method based on AI genomics

By employing AI genomics methods and utilizing multidimensional data acquisition and neural network architecture, the problems of long cycles, high costs, and limited accuracy in traditional breeding have been solved, enabling efficient and accurate breeding of dairy cows and pigs, and significantly improving breeding efficiency and the accuracy of economic trait prediction.

CN121459916AInactive Publication Date: 2026-02-03SHENZHEN QINGGAN EDUCATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511370222.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-02-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional livestock and poultry breeding relies on pedigree records and phenotypic selection, which has a long breeding cycle, high cost and limited accuracy. Genomic selection technology is not accurate enough in predicting gene-environment interactions, non-additive effects and complex traits.

Method used

By employing an AI-based genomics approach, through multidimensional data acquisition, data preprocessing and feature enhancement, AI prediction model training, and optimized mating decisions, a neural network architecture is constructed to capture the complex non-additive effects between genotype and phenotype, as well as gene-environment interactions, thereby achieving efficient breeding.

Benefits of technology

It improves the prediction accuracy of important economic traits, enables earlier and more accurate selection and optimized mating, significantly accelerates genetic progress, reduces breeding costs, and improves breeding efficiency and benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121459916A_ABST
    Figure CN121459916A_ABST
Patent Text Reader

Abstract

The invention discloses a dairy cow and live pig high-quality breeding method based on AI genomics, and relates to the field of breeding. Comprising the following steps: multi-dimensional data acquisition: aiming at a target breeding group, acquiring whole genome variation data, various phenotype data and environment management data of each individual, establishing unique identification association for all the data, and storing the data in a central database; data preprocessing and feature enhancement: performing quality control, filling and standardization processing on the genome data, and constructing an effective feature set for model training from the original data based on statistics and machine learning methods; and training an AI prediction model. By introducing the artificial intelligence deep learning model, the complex non-additive effect and gene-environment interaction between the genotype and the phenotype can be efficiently captured, the prediction precision of important economic characters is greatly improved, and earlier and more accurate selection and optimized hybridization are realized, so that the genetic progress is greatly accelerated, the breeding cost is reduced, and the method is suitable for large-scale popularization and application. The breeding efficiency and benefits are comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of breeding, and more particularly to a method for high-quality breeding of dairy cows and pigs based on AI genomics. Background Technology

[0002] Traditional livestock and poultry breeding relies primarily on pedigree recording and phenotypic selection, which are time-consuming, costly, and have limited accuracy. Genomic selection (GS) technology has significantly improved breeding efficiency, but its core reliance on linear models (such as GBLUP and BayesA) still presents limitations in predicting gene-environment interactions, non-additive effects (such as dominance and epistatic effects), and complex traits. With decreasing sequencing costs, efficiently processing massive amounts of genomic, phenotypic, and environmental data to uncover deeper genetic patterns has become a pressing issue for modern breeding.

[0003] Therefore, this invention proposes a method for high-quality breeding of dairy cows and pigs based on AI genomics. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a high-quality breeding method for dairy cows and pigs based on AI genomics.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] AI-based genomics-based methods for high-quality breeding of dairy cows and pigs include the following steps:

[0007] S1: Multidimensional data collection, targeting the breeding population, collecting whole genome variation data, multiple phenotypic data and environmental management data for each individual, and establishing unique identifiers for all data, which are stored in a central database;

[0008] S2: Data Preprocessing and Feature Enhancement: Quality control, imputation, and standardization of genomic data, and construction of an effective feature set for model training from the raw data based on statistical and machine learning methods;

[0009] S3: AI Prediction Model Training: Build a neural network architecture, train the model using training set data, monitor the fitting status and adjust hyperparameters using the validation set;

[0010] S4: Model Application and Genome Selection: Genotyping of new candidate individuals and inputting their data into a trained AI prediction model to obtain the genomic estimated breeding value (GEBV) for multiple economic traits;

[0011] S5: Optimize mating decisions: Based on the GEBV and kinship of all candidate individuals, with the goal of maximizing the overall genetic gain of offspring and the constraint of controlling the level of inbreeding, an optimization algorithm is used to calculate the optimal mating combination scheme.

[0012] Preferably, in step S1, the whole genome variation data is one of SNPs, Indels, or CNVs.

[0013] Preferably, in step S1, the phenotypic data includes three-dimensional point cloud data of body structure obtained by a 3D body condition scanner, spectral data of meat components obtained by a near-infrared spectrometer, feed intake and weight gain data continuously recorded by an automated feeding station, and image data of intramuscular fat and backfat thickness obtained by an ultrasound device.

[0014] Preferably, in step S1, the environmental management data includes the temperature, humidity, ammonia concentration, stocking density, nutrient content of the diet, and individual health records within the enclosure.

[0015] Preferably, in step S2, the effective feature set includes significant SNP sites, genotype-derived features, and gene-environment interaction features.

[0016] Preferably, step S2 includes the following steps:

[0017] S21: Perform quality control on the collected raw genome data, remove genetic markers with low detection rate and low minimum allele frequency, and use a reference population for genotyping to improve data integrity.

[0018] S22: Use unsupervised learning algorithms to reduce the dimensionality of high-dimensional genome data and extract low-dimensional feature vectors containing population structure information and genetic variation;

[0019] S23: Quantify environmental data into numerical features and combine them with genotype data to construct new features that characterize the interaction between genes and the environment through statistical modeling or neural networks;

[0020] S24: Integrate the processed genomic features, environmental interaction features, and standardized phenotypic data into a structured feature-label dataset.

[0021] Preferably, the unsupervised learning algorithm in step S22 is principal component analysis or an autoencoder.

[0022] Preferably, in step S3, the neural network architecture includes a convolutional layer for capturing local genetic patterns, a recurrent layer or self-attention layer for capturing long-range dependencies, and a fully connected output layer.

[0023] Preferably, step S3, the model training method, includes the following steps:

[0024] S31: Divide the processed feature data into training set, validation set and test set;

[0025] S32: Use genotype and environmental characteristics as model inputs and phenotypic data as target output;

[0026] S33: Minimize the loss function between the predicted and the true values ​​using the backpropagation algorithm and the Adam optimizer;

[0027] S34: Use a validation set for validation, and employ early stopping and regularization to prevent overfitting.

[0028] Preferably, step S4 includes the following steps:

[0029] S41: Genotyping new candidate individuals (such as newborn piglets or young bulls to be selected);

[0030] S42: Input genotype data into a trained and validated AI prediction model;

[0031] S43: The model outputs the predicted breeding value (GEBV) or phenotypic value for multiple target traits of these individuals;

[0032] S44: Weighted integration of GEBV or phenotypic values ​​of multiple traits to form a comprehensive selection index.

[0033] The beneficial effects of this invention are as follows:

[0034] 1. By introducing an artificial intelligence deep learning model, this invention can efficiently capture the complex non-additive effects between genotype and phenotype, as well as gene-environment interactions, greatly improving the prediction accuracy of important economic traits. This enables earlier and more accurate selection and optimized mating, thereby significantly accelerating genetic progress, reducing breeding costs, and comprehensively improving breeding efficiency and benefits. Attached Figure Description

[0035] Figure 1 This is a flowchart of the AI ​​genomics-based high-quality breeding method for dairy cows and pigs proposed in this invention. Detailed Implementation

[0036] The technical solution of the present invention will be further described in detail below with reference to specific embodiments.

[0037] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "setting" should be interpreted broadly. For example, they can refer to a fixed connection or setting, a detachable connection or setting, or an integral connection or setting. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0038] Example 1:

[0039] The AI-based genomics-based high-quality breeding system for dairy cows and pigs includes the following steps:

[0040] S1: Multidimensional data collection, targeting the breeding population (dairy cows or pigs), collecting whole genome variation data, multiple phenotypic data and environmental management data for each individual, and establishing unique identifiers for all data, which are stored in a central database;

[0041] S2: Data Preprocessing and Feature Enhancement: Quality control, imputation, and standardization of genomic data, and construction of an effective feature set for model training from the raw data based on statistical and machine learning methods;

[0042] S3: AI Prediction Model Training: Build a neural network architecture, train the model using training set data, monitor the fitting status and adjust hyperparameters through the validation set to prevent overfitting, and finally obtain a stable high-performance prediction model.

[0043] S4: Model Application and Genome Selection: Genotyping of new candidate individuals and inputting their data into a trained AI prediction model to obtain the genomic estimated breeding value (GEBV) for multiple economic traits;

[0044] S5: Optimize mating decisions: Based on the GEBV and kinship of all candidate individuals, with the goal of maximizing the overall genetic gain of offspring and the constraint of controlling the level of inbreeding, an optimization algorithm is used to calculate the optimal mating combination scheme.

[0045] In step S1, the whole genome variation data are one of SNPs, Indels, or CNVs.

[0046] In step S1, the phenotypic data includes three-dimensional point cloud data of body structure obtained by a 3D body condition scanner, spectral data of meat components obtained by a near-infrared spectrometer, feed intake and weight gain data continuously recorded by an automated feeding station, and image data of intramuscular fat and backfat thickness obtained by an ultrasound device.

[0047] In step S1, the environmental management data includes the temperature, humidity, ammonia concentration, stocking density, nutrient content of the diet, and individual health records within the enclosure.

[0048] In step S2, the effective feature set includes significant SNP sites, genotype-derived features, and gene-environment interaction features.

[0049] In step S3, the neural network architecture includes convolutional layers for capturing local genetic patterns, recurrent layers or self-attention layers for capturing long-range dependencies, and fully connected output layers.

[0050] Example 2:

[0051] The AI-based genomics-based high-quality breeding system for dairy cows and pigs includes the following steps:

[0052] S1: Multidimensional data collection, targeting the breeding population (dairy cows or pigs), collecting whole genome variation data, multiple phenotypic data and environmental management data for each individual, and establishing unique identifiers for all data, which are stored in a central database;

[0053] S2: Data Preprocessing and Feature Enhancement: Quality control, imputation, and standardization of genomic data, and construction of an effective feature set for model training from the raw data based on statistical and machine learning methods;

[0054] S3: AI Prediction Model Training: Build a neural network architecture, train the model using training set data, monitor the fitting status and adjust hyperparameters through the validation set to prevent overfitting, and finally obtain a stable high-performance prediction model.

[0055] S4: Model Application and Genome Selection: Genotyping of new candidate individuals and inputting their data into a trained AI prediction model to obtain the genomic estimated breeding value (GEBV) for multiple economic traits;

[0056] S5: Optimize mating decisions: Based on the GEBV and kinship of all candidate individuals, with the goal of maximizing the overall genetic gain of offspring and the constraint of controlling the level of inbreeding, an optimization algorithm is used to calculate the optimal mating combination scheme.

[0057] In step S1, the whole genome variation data are one of SNPs, Indels, or CNVs.

[0058] In step S1, the phenotypic data includes three-dimensional point cloud data of body structure obtained by a 3D body condition scanner, spectral data of meat components obtained by a near-infrared spectrometer, feed intake and weight gain data continuously recorded by an automated feeding station, and image data of intramuscular fat and backfat thickness obtained by an ultrasound device.

[0059] In step S1, the environmental management data includes the temperature, humidity, ammonia concentration, stocking density, nutrient content of the diet, and individual health records within the enclosure.

[0060] In step S2, the effective feature set includes significant SNP sites, genotype-derived features, and gene-environment interaction features.

[0061] Step S2 includes the following steps:

[0062] S21: Perform quality control on the collected raw genome data, remove genetic markers with low detection rate and low minimum allele frequency, and use a reference population for genotyping to improve data integrity.

[0063] S22: Use unsupervised learning algorithms to reduce the dimensionality of high-dimensional genome data and extract low-dimensional feature vectors containing population structure information and genetic variation;

[0064] S23: Quantify environmental data into numerical features and combine them with genotype data to construct new features that characterize the interaction between genes and the environment through statistical modeling or neural networks;

[0065] S24: Integrate the processed genomic features, environmental interaction features, and standardized phenotypic data into a structured feature-label dataset.

[0066] In step S3, the neural network architecture includes convolutional layers for capturing local genetic patterns, recurrent layers or self-attention layers for capturing long-range dependencies, and fully connected output layers.

[0067] The model training method in step S3 includes the following steps:

[0068] S31: Divide the processed feature data into training set, validation set and test set;

[0069] S32: Use genotypic and environmental characteristics as model inputs and phenotypic data (such as feed conversion ratio) as target output;

[0070] S33: Minimize the loss function (such as mean squared error, MSE) between the predicted and true values ​​using the backpropagation algorithm and the Adam optimizer;

[0071] S34: Use a validation set for validation, and employ early stopping and regularization to prevent overfitting.

[0072] Step S4 includes the following steps:

[0073] S41: Genotyping new candidate individuals (such as newborn piglets or young bulls to be selected);

[0074] S42: Input genotype data into a trained and validated AI prediction model;

[0075] S43: The model outputs the predicted breeding value (GEBV) or phenotypic value for multiple target traits of these individuals;

[0076] S44: Weighted integration of GEBV or phenotypic values ​​of multiple traits to form a comprehensive selection index.

[0077] Example 3:

[0078] The AI-based genomics-based high-quality breeding system for dairy cows and pigs includes the following steps:

[0079] S1: Multidimensional data collection, targeting the breeding population (dairy cows or pigs), collecting whole genome variation data, multiple phenotypic data and environmental management data for each individual, and establishing unique identifiers for all data, which are stored in a central database;

[0080] S2: Data Preprocessing and Feature Enhancement: Quality control, imputation, and standardization of genomic data, and construction of an effective feature set for model training from the raw data based on statistical and machine learning methods;

[0081] S3: AI Prediction Model Training: Build a neural network architecture, train the model using training set data, monitor the fitting status and adjust hyperparameters through the validation set to prevent overfitting, and finally obtain a stable high-performance prediction model.

[0082] S4: Model Application and Genome Selection: Genotyping of new candidate individuals and inputting their data into a trained AI prediction model to obtain the genomic estimated breeding value (GEBV) for multiple economic traits;

[0083] S5: Optimize mating decisions: Based on the GEBV and kinship of all candidate individuals, with the goal of maximizing the overall genetic gain of offspring and the constraint of controlling the level of inbreeding, an optimization algorithm is used to calculate the optimal mating combination scheme.

[0084] In step S1, the whole genome variation data are one of SNPs, Indels, or CNVs.

[0085] In step S1, the phenotypic data includes three-dimensional point cloud data of body structure obtained by a 3D body condition scanner, spectral data of meat components obtained by a near-infrared spectrometer, feed intake and weight gain data continuously recorded by an automated feeding station, and image data of intramuscular fat and backfat thickness obtained by an ultrasound device.

[0086] In step S1, the environmental management data includes the temperature, humidity, ammonia concentration, stocking density, nutrient content of the diet, and individual health records within the enclosure.

[0087] In step S2, the effective feature set includes significant SNP sites, genotype-derived features, and gene-environment interaction features.

[0088] Step S2 includes the following steps:

[0089] S21: Perform quality control on the collected raw genome data, remove genetic markers with low detection rate and low minimum allele frequency, and use a reference population for genotyping to improve data integrity.

[0090] S22: Use unsupervised learning algorithms to reduce the dimensionality of high-dimensional genome data and extract low-dimensional feature vectors containing population structure information and genetic variation;

[0091] S23: Quantify environmental data into numerical features and combine them with genotype data to construct new features that characterize the interaction between genes and the environment through statistical modeling or neural networks;

[0092] S24: Integrate the processed genomic features, environmental interaction features, and standardized phenotypic data into a structured feature-label dataset.

[0093] In step S22, the unsupervised learning algorithm employs either principal component analysis or an autoencoder.

[0094] In step S3, the neural network architecture includes convolutional layers for capturing local genetic patterns, recurrent layers or self-attention layers for capturing long-range dependencies, and fully connected output layers.

[0095] The model training method in step S3 includes the following steps:

[0096] S31: Divide the processed feature data into training set, validation set and test set;

[0097] S32: Use genotypic and environmental characteristics as model inputs and phenotypic data (such as feed conversion ratio) as target output;

[0098] S33: Minimize the loss function (such as mean squared error, MSE) between the predicted and true values ​​using the backpropagation algorithm and the Adam optimizer;

[0099] S34: Use a validation set for validation, and employ early stopping and regularization to prevent overfitting.

[0100] Step S4 includes the following steps:

[0101] S41: Genotyping new candidate individuals (such as newborn piglets or young bulls to be selected);

[0102] S42: Input genotype data into a trained and validated AI prediction model;

[0103] S43: The model outputs the predicted breeding value (GEBV) or phenotypic value for multiple target traits of these individuals;

[0104] S44: Weighted integration of GEBV or phenotypic values ​​of multiple traits to form a comprehensive selection index.

[0105] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for high-quality breeding of dairy cows and pigs based on AI genomics, characterized in that, Includes the following steps: S1: Multidimensional data collection, targeting the breeding population, collecting whole genome variation data, multiple phenotypic data and environmental management data for each individual, and establishing unique identifiers for all data, which are stored in a central database; S2: Data Preprocessing and Feature Enhancement: Quality control, imputation, and standardization of genomic data, and construction of an effective feature set for model training from the raw data based on statistical and machine learning methods; S3: AI Prediction Model Training: Build a neural network architecture, train the model using training set data, monitor the fitting status and adjust hyperparameters using the validation set; S4: Model Application and Genome Selection: Genotyping of new candidate individuals and inputting their data into a trained AI prediction model to obtain the genomic estimated breeding value (GEBV) for multiple economic traits; S5: Optimize mating decisions: Based on the GEBV and kinship of all candidate individuals, with the goal of maximizing the overall genetic gain of offspring and the constraint of controlling the level of inbreeding, an optimization algorithm is used to calculate the optimal mating combination scheme.

2. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, In step S1, the whole genome variation data are one of SNPs, Indels, or CNVs.

3. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, In step S1, the phenotypic data includes three-dimensional point cloud data of body structure obtained by a 3D body condition scanner, spectral data of meat components obtained by a near-infrared spectrometer, feed intake and weight gain data continuously recorded by an automated feeding station, and image data of intramuscular fat and backfat thickness obtained by an ultrasound device.

4. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, In step S1, the environmental management data includes the temperature, humidity, ammonia concentration, stocking density, nutrient content of the diet, and individual health records within the enclosure.

5. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, In step S2, the effective feature set includes significant SNP sites, genotype-derived features, and gene-environment interaction features.

6. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, Step S2 includes the following steps: S21: Perform quality control on the collected raw genome data, remove genetic markers with low detection rate and low minimum allele frequency, and use a reference population for genotyping to improve data integrity. S22: Use unsupervised learning algorithms to reduce the dimensionality of high-dimensional genome data and extract low-dimensional feature vectors containing population structure information and genetic variation; S23: Quantify environmental data into numerical features and combine them with genotype data to construct new features that characterize the interaction between genes and the environment through statistical modeling or neural networks; S24: Integrate the processed genomic features, environmental interaction features, and standardized phenotypic data into a structured feature-label dataset.

7. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 6, characterized in that, In step S22, the unsupervised learning algorithm employs either principal component analysis or an autoencoder.

8. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, In step S3, the neural network architecture includes convolutional layers for capturing local genetic patterns, recurrent layers or self-attention layers for capturing long-range dependencies, and fully connected output layers.

9. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, The model training method in step S3 includes the following steps: S31: Divide the processed feature data into training set, validation set and test set; S32: Use genotype and environmental characteristics as model inputs and phenotypic data as target output; S33: Minimize the loss function between the predicted and the true values ​​using the backpropagation algorithm and the Adam optimizer; S34: Use a validation set for validation, and employ early stopping and regularization to prevent overfitting.

10. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, Step S4 includes the following steps: S41: Genotyping new candidate individuals (such as newborn piglets or young bulls to be selected); S42: Input genotype data into a trained and validated AI prediction model; S43: The model outputs the predicted breeding value (GEBV) or phenotypic value for multiple target traits of these individuals; S44: Weighted integration of GEBV or phenotypic values ​​of multiple traits to form a comprehensive selection index.