Cow and live pig high-quality breeding method based on AI genomics
By employing AI genomics methods and utilizing multidimensional data acquisition and neural network architecture, the problems of long cycles, high costs, and limited accuracy in traditional breeding have been solved, enabling efficient and accurate breeding of dairy cows and pigs, and significantly improving breeding efficiency and the accuracy of economic trait prediction.
Patent Information
- Application Number
- CN202511370222.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-02-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional livestock and poultry breeding relies on pedigree records and phenotypic selection, which has a long breeding cycle, high cost and limited accuracy. Genomic selection technology is not accurate enough in predicting gene-environment interactions, non-additive effects and complex traits.
By employing an AI-based genomics approach, through multidimensional data acquisition, data preprocessing and feature enhancement, AI prediction model training, and optimized mating decisions, a neural network architecture is constructed to capture the complex non-additive effects between genotype and phenotype, as well as gene-environment interactions, thereby achieving efficient breeding.
It improves the prediction accuracy of important economic traits, enables earlier and more accurate selection and optimized mating, significantly accelerates genetic progress, reduces breeding costs, and improves breeding efficiency and benefits.
Smart Images

Figure CN121459916A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of breeding, and more particularly to a method for high-quality breeding of dairy cows and pigs based on AI genomics. Background Technology
[0002] Traditional livestock and poultry breeding relies primarily on pedigree recording and phenotypic selection, which are time-consuming, costly, and have limited accuracy. Genomic selection (GS) technology has significantly improved breeding efficiency, but its core reliance on linear models (such as GBLUP and BayesA) still presents limitations in predicting gene-environment interactions, non-additive effects (such as dominance and epistatic effects), and complex traits. With decreasing sequencing costs, efficiently processing massive amounts of genomic, phenotypic, and environmental data to uncover deeper genetic patterns has become a pressing issue for modern breeding.
[0003] Therefore, this invention proposes a method for high-quality breeding of dairy cows and pigs based on AI genomics. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a high-quality breeding method for dairy cows and pigs based on AI genomics.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] AI-based genomics-based methods for high-quality breeding of dairy cows and pigs include the following steps:
[0007] S1: Multidimensional data collection, targeting the breeding population, collecting whole genome variation data, multiple phenotypic data and environmental management data for each individual, and establishing unique identifiers for all data, which are stored in a central database;
[0008] S2: Data Preprocessing and Feature Enhancement: Quality control, imputation, and standardization of genomic data, and construction of an effective feature set for model training from the raw data based on statistical and machine learning methods;
[0009] S3: AI Prediction Model Training: Build a neural network architecture, train the model using training set data, monitor the fitting status and adjust hyperparameters using the validation set;
[0010] S4: Model Application and Genome Selection: Genotyping of new candidate individuals and inputting their data into a trained AI prediction model to obtain the genomic estimated breeding value (GEBV) for multiple economic traits;
[0011] S5: Optimize mating decisions: Based on the GEBV and kinship of all candidate individuals, with the goal of maximizing the overall genetic gain of offspring and the constraint of controlling the level of inbreeding, an optimization algorithm is used to calculate the optimal mating combination scheme.
[0012] Preferably, in step S1, the whole genome variation data is one of SNPs, Indels, or CNVs.
[0013] Preferably, in step S1, the phenotypic data includes three-dimensional point cloud data of body structure obtained by a 3D body condition scanner, spectral data of meat components obtained by a near-infrared spectrometer, feed intake and weight gain data continuously recorded by an automated feeding station, and image data of intramuscular fat and backfat thickness obtained by an ultrasound device.
[0014] Preferably, in step S1, the environmental management data includes the temperature, humidity, ammonia concentration, stocking density, nutrient content of the diet, and individual health records within the enclosure.
[0015] Preferably, in step S2, the effective feature set includes significant SNP sites, genotype-derived features, and gene-environment interaction features.
[0016] Preferably, step S2 includes the following steps:
[0017] S21: Perform quality control on the collected raw genome data, remove genetic markers with low detection rate and low minimum allele frequency, and use a reference population for genotyping to improve data integrity.
[0018] S22: Use unsupervised learning algorithms to reduce the dimensionality of high-dimensional genome data and extract low-dimensional feature vectors containing population structure information and genetic variation;
[0019] S23: Quantify environmental data into numerical features and combine them with genotype data to construct new features that characterize the interaction between genes and the environment through statistical modeling or neural networks;
[0020] S24: Integrate the processed genomic features, environmental interaction features, and standardized phenotypic data into a structured feature-label dataset.
[0021] Preferably, the unsupervised learning algorithm in step S22 is principal component analysis or an autoencoder.
[0022] Preferably, in step S3, the neural network architecture includes a convolutional layer for capturing local genetic patterns, a recurrent layer or self-attention layer for capturing long-range dependencies, and a fully connected output layer.
[0023] Preferably, step S3, the model training method, includes the following steps:
[0024] S31: Divide the processed feature data into training set, validation set and test set;
[0025] S32: Use genotype and environmental characteristics as model inputs and phenotypic data as target output;
[0026] S33: Minimize the loss function between the predicted and the true values using the backpropagation algorithm and the Adam optimizer;
[0027] S34: Use a validation set for validation, and employ early stopping and regularization to prevent overfitting.
[0028] Preferably, step S4 includes the following steps:
[0029] S41: Genotyping new candidate individuals (such as newborn piglets or young bulls to be selected);
[0030] S42: Input genotype data into a trained and validated AI prediction model;
[0031] S43: The model outputs the predicted breeding value (GEBV) or phenotypic value for multiple target traits of these individuals;
[0032] S44: Weighted integration of GEBV or phenotypic values of multiple traits to form a comprehensive selection index.
[0033] The beneficial effects of this invention are as follows:
[0034] 1. By introducing an artificial intelligence deep learning model, this invention can efficiently capture the complex non-additive effects between genotype and phenotype, as well as gene-environment interactions, greatly improving the prediction accuracy of important economic traits. This enables earlier and more accurate selection and optimized mating, thereby significantly accelerating genetic progress, reducing breeding costs, and comprehensively improving breeding efficiency and benefits. Attached Figure Description
[0035] Figure 1 This is a flowchart of the AI genomics-based high-quality breeding method for dairy cows and pigs proposed in this invention. Detailed Implementation
[0036] The technical solution of the present invention will be further described in detail below with reference to specific embodiments.
[0037] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "setting" should be interpreted broadly. For example, they can refer to a fixed connection or setting, a detachable connection or setting, or an integral connection or setting. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0038] Example 1:
[0039] The AI-based genomics-based high-quality breeding system for dairy cows and pigs includes the following steps:
[0040] S1: Multidimensional data collection, targeting the breeding population (dairy cows or pigs), collecting whole genome variation data, multiple phenotypic data and environmental management data for each individual, and establishing unique identifiers for all data, which are stored in a central database;
[0041] S2: Data Preprocessing and Feature Enhancement: Quality control, imputation, and standardization of genomic data, and construction of an effective feature set for model training from the raw data based on statistical and machine learning methods;
[0042] S3: AI Prediction Model Training: Build a neural network architecture, train the model using training set data, monitor the fitting status and adjust hyperparameters through the validation set to prevent overfitting, and finally obtain a stable high-performance prediction model.
[0043] S4: Model Application and Genome Selection: Genotyping of new candidate individuals and inputting their data into a trained AI prediction model to obtain the genomic estimated breeding value (GEBV) for multiple economic traits;
[0044] S5: Optimize mating decisions: Based on the GEBV and kinship of all candidate individuals, with the goal of maximizing the overall genetic gain of offspring and the constraint of controlling the level of inbreeding, an optimization algorithm is used to calculate the optimal mating combination scheme.
[0045] In step S1, the whole genome variation data are one of SNPs, Indels, or CNVs.
[0046] In step S1, the phenotypic data includes three-dimensional point cloud data of body structure obtained by a 3D body condition scanner, spectral data of meat components obtained by a near-infrared spectrometer, feed intake and weight gain data continuously recorded by an automated feeding station, and image data of intramuscular fat and backfat thickness obtained by an ultrasound device.
[0047] In step S1, the environmental management data includes the temperature, humidity, ammonia concentration, stocking density, nutrient content of the diet, and individual health records within the enclosure.
[0048] In step S2, the effective feature set includes significant SNP sites, genotype-derived features, and gene-environment interaction features.
[0049] In step S3, the neural network architecture includes convolutional layers for capturing local genetic patterns, recurrent layers or self-attention layers for capturing long-range dependencies, and fully connected output layers.
[0050] Example 2:
[0051] The AI-based genomics-based high-quality breeding system for dairy cows and pigs includes the following steps:
[0052] S1: Multidimensional data collection, targeting the breeding population (dairy cows or pigs), collecting whole genome variation data, multiple phenotypic data and environmental management data for each individual, and establishing unique identifiers for all data, which are stored in a central database;
[0053] S2: Data Preprocessing and Feature Enhancement: Quality control, imputation, and standardization of genomic data, and construction of an effective feature set for model training from the raw data based on statistical and machine learning methods;
[0054] S3: AI Prediction Model Training: Build a neural network architecture, train the model using training set data, monitor the fitting status and adjust hyperparameters through the validation set to prevent overfitting, and finally obtain a stable high-performance prediction model.
[0055] S4: Model Application and Genome Selection: Genotyping of new candidate individuals and inputting their data into a trained AI prediction model to obtain the genomic estimated breeding value (GEBV) for multiple economic traits;
[0056] S5: Optimize mating decisions: Based on the GEBV and kinship of all candidate individuals, with the goal of maximizing the overall genetic gain of offspring and the constraint of controlling the level of inbreeding, an optimization algorithm is used to calculate the optimal mating combination scheme.
[0057] In step S1, the whole genome variation data are one of SNPs, Indels, or CNVs.
[0058] In step S1, the phenotypic data includes three-dimensional point cloud data of body structure obtained by a 3D body condition scanner, spectral data of meat components obtained by a near-infrared spectrometer, feed intake and weight gain data continuously recorded by an automated feeding station, and image data of intramuscular fat and backfat thickness obtained by an ultrasound device.
[0059] In step S1, the environmental management data includes the temperature, humidity, ammonia concentration, stocking density, nutrient content of the diet, and individual health records within the enclosure.
[0060] In step S2, the effective feature set includes significant SNP sites, genotype-derived features, and gene-environment interaction features.
[0061] Step S2 includes the following steps:
[0062] S21: Perform quality control on the collected raw genome data, remove genetic markers with low detection rate and low minimum allele frequency, and use a reference population for genotyping to improve data integrity.
[0063] S22: Use unsupervised learning algorithms to reduce the dimensionality of high-dimensional genome data and extract low-dimensional feature vectors containing population structure information and genetic variation;
[0064] S23: Quantify environmental data into numerical features and combine them with genotype data to construct new features that characterize the interaction between genes and the environment through statistical modeling or neural networks;
[0065] S24: Integrate the processed genomic features, environmental interaction features, and standardized phenotypic data into a structured feature-label dataset.
[0066] In step S3, the neural network architecture includes convolutional layers for capturing local genetic patterns, recurrent layers or self-attention layers for capturing long-range dependencies, and fully connected output layers.
[0067] The model training method in step S3 includes the following steps:
[0068] S31: Divide the processed feature data into training set, validation set and test set;
[0069] S32: Use genotypic and environmental characteristics as model inputs and phenotypic data (such as feed conversion ratio) as target output;
[0070] S33: Minimize the loss function (such as mean squared error, MSE) between the predicted and true values using the backpropagation algorithm and the Adam optimizer;
[0071] S34: Use a validation set for validation, and employ early stopping and regularization to prevent overfitting.
[0072] Step S4 includes the following steps:
[0073] S41: Genotyping new candidate individuals (such as newborn piglets or young bulls to be selected);
[0074] S42: Input genotype data into a trained and validated AI prediction model;
[0075] S43: The model outputs the predicted breeding value (GEBV) or phenotypic value for multiple target traits of these individuals;
[0076] S44: Weighted integration of GEBV or phenotypic values of multiple traits to form a comprehensive selection index.
[0077] Example 3:
[0078] The AI-based genomics-based high-quality breeding system for dairy cows and pigs includes the following steps:
[0079] S1: Multidimensional data collection, targeting the breeding population (dairy cows or pigs), collecting whole genome variation data, multiple phenotypic data and environmental management data for each individual, and establishing unique identifiers for all data, which are stored in a central database;
[0080] S2: Data Preprocessing and Feature Enhancement: Quality control, imputation, and standardization of genomic data, and construction of an effective feature set for model training from the raw data based on statistical and machine learning methods;
[0081] S3: AI Prediction Model Training: Build a neural network architecture, train the model using training set data, monitor the fitting status and adjust hyperparameters through the validation set to prevent overfitting, and finally obtain a stable high-performance prediction model.
[0082] S4: Model Application and Genome Selection: Genotyping of new candidate individuals and inputting their data into a trained AI prediction model to obtain the genomic estimated breeding value (GEBV) for multiple economic traits;
[0083] S5: Optimize mating decisions: Based on the GEBV and kinship of all candidate individuals, with the goal of maximizing the overall genetic gain of offspring and the constraint of controlling the level of inbreeding, an optimization algorithm is used to calculate the optimal mating combination scheme.
[0084] In step S1, the whole genome variation data are one of SNPs, Indels, or CNVs.
[0085] In step S1, the phenotypic data includes three-dimensional point cloud data of body structure obtained by a 3D body condition scanner, spectral data of meat components obtained by a near-infrared spectrometer, feed intake and weight gain data continuously recorded by an automated feeding station, and image data of intramuscular fat and backfat thickness obtained by an ultrasound device.
[0086] In step S1, the environmental management data includes the temperature, humidity, ammonia concentration, stocking density, nutrient content of the diet, and individual health records within the enclosure.
[0087] In step S2, the effective feature set includes significant SNP sites, genotype-derived features, and gene-environment interaction features.
[0088] Step S2 includes the following steps:
[0089] S21: Perform quality control on the collected raw genome data, remove genetic markers with low detection rate and low minimum allele frequency, and use a reference population for genotyping to improve data integrity.
[0090] S22: Use unsupervised learning algorithms to reduce the dimensionality of high-dimensional genome data and extract low-dimensional feature vectors containing population structure information and genetic variation;
[0091] S23: Quantify environmental data into numerical features and combine them with genotype data to construct new features that characterize the interaction between genes and the environment through statistical modeling or neural networks;
[0092] S24: Integrate the processed genomic features, environmental interaction features, and standardized phenotypic data into a structured feature-label dataset.
[0093] In step S22, the unsupervised learning algorithm employs either principal component analysis or an autoencoder.
[0094] In step S3, the neural network architecture includes convolutional layers for capturing local genetic patterns, recurrent layers or self-attention layers for capturing long-range dependencies, and fully connected output layers.
[0095] The model training method in step S3 includes the following steps:
[0096] S31: Divide the processed feature data into training set, validation set and test set;
[0097] S32: Use genotypic and environmental characteristics as model inputs and phenotypic data (such as feed conversion ratio) as target output;
[0098] S33: Minimize the loss function (such as mean squared error, MSE) between the predicted and true values using the backpropagation algorithm and the Adam optimizer;
[0099] S34: Use a validation set for validation, and employ early stopping and regularization to prevent overfitting.
[0100] Step S4 includes the following steps:
[0101] S41: Genotyping new candidate individuals (such as newborn piglets or young bulls to be selected);
[0102] S42: Input genotype data into a trained and validated AI prediction model;
[0103] S43: The model outputs the predicted breeding value (GEBV) or phenotypic value for multiple target traits of these individuals;
[0104] S44: Weighted integration of GEBV or phenotypic values of multiple traits to form a comprehensive selection index.
[0105] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for high-quality breeding of dairy cows and pigs based on AI genomics, characterized in that, Includes the following steps: S1: Multidimensional data collection, targeting the breeding population, collecting whole genome variation data, multiple phenotypic data and environmental management data for each individual, and establishing unique identifiers for all data, which are stored in a central database; S2: Data Preprocessing and Feature Enhancement: Quality control, imputation, and standardization of genomic data, and construction of an effective feature set for model training from the raw data based on statistical and machine learning methods; S3: AI Prediction Model Training: Build a neural network architecture, train the model using training set data, monitor the fitting status and adjust hyperparameters using the validation set; S4: Model Application and Genome Selection: Genotyping of new candidate individuals and inputting their data into a trained AI prediction model to obtain the genomic estimated breeding value (GEBV) for multiple economic traits; S5: Optimize mating decisions: Based on the GEBV and kinship of all candidate individuals, with the goal of maximizing the overall genetic gain of offspring and the constraint of controlling the level of inbreeding, an optimization algorithm is used to calculate the optimal mating combination scheme.
2. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, In step S1, the whole genome variation data are one of SNPs, Indels, or CNVs.
3. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, In step S1, the phenotypic data includes three-dimensional point cloud data of body structure obtained by a 3D body condition scanner, spectral data of meat components obtained by a near-infrared spectrometer, feed intake and weight gain data continuously recorded by an automated feeding station, and image data of intramuscular fat and backfat thickness obtained by an ultrasound device.
4. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, In step S1, the environmental management data includes the temperature, humidity, ammonia concentration, stocking density, nutrient content of the diet, and individual health records within the enclosure.
5. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, In step S2, the effective feature set includes significant SNP sites, genotype-derived features, and gene-environment interaction features.
6. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, Step S2 includes the following steps: S21: Perform quality control on the collected raw genome data, remove genetic markers with low detection rate and low minimum allele frequency, and use a reference population for genotyping to improve data integrity. S22: Use unsupervised learning algorithms to reduce the dimensionality of high-dimensional genome data and extract low-dimensional feature vectors containing population structure information and genetic variation; S23: Quantify environmental data into numerical features and combine them with genotype data to construct new features that characterize the interaction between genes and the environment through statistical modeling or neural networks; S24: Integrate the processed genomic features, environmental interaction features, and standardized phenotypic data into a structured feature-label dataset.
7. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 6, characterized in that, In step S22, the unsupervised learning algorithm employs either principal component analysis or an autoencoder.
8. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, In step S3, the neural network architecture includes convolutional layers for capturing local genetic patterns, recurrent layers or self-attention layers for capturing long-range dependencies, and fully connected output layers.
9. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, The model training method in step S3 includes the following steps: S31: Divide the processed feature data into training set, validation set and test set; S32: Use genotype and environmental characteristics as model inputs and phenotypic data as target output; S33: Minimize the loss function between the predicted and the true values using the backpropagation algorithm and the Adam optimizer; S34: Use a validation set for validation, and employ early stopping and regularization to prevent overfitting.
10. The method for high-quality breeding of dairy cows and pigs based on AI genomics according to claim 1, characterized in that, Step S4 includes the following steps: S41: Genotyping new candidate individuals (such as newborn piglets or young bulls to be selected); S42: Input genotype data into a trained and validated AI prediction model; S43: The model outputs the predicted breeding value (GEBV) or phenotypic value for multiple target traits of these individuals; S44: Weighted integration of GEBV or phenotypic values of multiple traits to form a comprehensive selection index.