Prediction method for measuring production health index of dairy cow

By constructing a multi-source data prediction model and employing nonlinear machine learning algorithms and heterogeneity correction factors, the problem of insufficient accuracy in predicting the health status of dairy cows caused by differences in management among pastures was solved, achieving precise quantification of the health status of dairy cows and stable improvement in production performance.

CN121687484APending Publication Date: 2026-03-17洛阳市种业发展中心(洛阳市农产品安全认证中心)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511849073.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of prediction models for dairy cow health status caused by differences in management between farms is insufficient, especially in the prediction of subclinical mastitis, which suffers from systematic bias.

Method used

By constructing a multi-source data prediction model and employing a nonlinear machine learning algorithm, a heterogeneity correction factor is generated to correct for management differences between ranches. This includes using IoT sensors, wearable sensors, management systems, and breeding record data to conduct health trait association analysis and health index calculation. By utilizing dynamic features, interaction term features, and domain knowledge features, a comprehensive health index is generated.

Benefits of technology

It improves the accuracy of predicting the health status of dairy cows and the generalization ability of the model, enabling early warning of potential disease risks, reducing economic losses, and improving the stability of dairy cow production performance and milk quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121687484A_ABST
    Figure CN121687484A_ABST
Patent Text Reader

Abstract

The invention discloses a prediction method for dairy cow production health index measurement. The prediction method comprises the following steps: data acquisition: acquiring multi-source data of a target dairy cow; data preprocessing and feature engineering: cleaning, standardization and feature generation processing are performed on the multi-source data, a feature set used for model input is generated based on the processed data, and feature generation comprises generation of heterogeneity correction factors used for correcting management differences between pastures; health character correlation analysis: inputting the standardized feature data set into a pre-trained health character correlation analysis model, and outputting a plurality of health problem prediction probabilities; and health index prediction: the plurality of health problem prediction probabilities and the heterogeneity correction factor are jointly input into a health index calculation model, a comprehensive health index is obtained through calculation, and the heterogeneity correction factor is explicitly introduced into the health index calculation model and is used for correcting prediction deviation caused by pasture management differences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent livestock farming technology, and in particular to a predictive method for measuring the health index of dairy cows. Background Technology

[0002] By constructing a predictive model based on multi-dimensional physiological indicators, behavioral characteristics, environmental parameters, and production performance data, a scientific and systematic measurement method is developed that can accurately quantify the health status of dairy cows, dynamically monitor fluctuations in production performance, and provide early warnings of potential disease risks. This method can effectively improve the health management level of individual dairy cows and herds, reduce economic losses such as decreased milk production and reduced reproductive efficiency caused by subclinical diseases or hidden health problems, achieve stable improvement in dairy cow production performance and continuous optimization of milk source quality, provide scientific decision support for dairy farming enterprises, promote the transformation and upgrading of the farming industry towards refinement and intelligence, and ultimately play an important technical support and practical guidance role in ensuring food safety, promoting sustainable agricultural development, and improving the economic and ecological benefits of farming. This forms a complete technical closed loop and value chain from individual health monitoring to herd production optimization, and from short-term benefit improvement to long-term sustainable development.

[0003] In existing technologies, the sample size is expanded simply by increasing the number of pastures. However, differences in management between pastures may introduce confounding variables, and the model does not quantify and correct for systematic differences between pastures. For example, the somatic cell count thresholds of different pastures may exhibit systematic biases due to different milking machine pressure settings, affecting the accuracy of subclinical mastitis prediction. Therefore, a predictive method for measuring the dairy cow production health index is proposed. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a predictive method for measuring the health index of dairy cows.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: A predictive method for measuring the production health index of dairy cows, comprising the following steps: Data Acquisition: Acquire multi-source data of the target dairy cows, including dairy cow production performance measurement data, environmental data collected by IoT sensors, behavioral data collected by wearable sensors, feeding management data obtained by the management system, genetic data obtained by breeding records, and rumen microbiome data collected by fecal or milk samples. Data preprocessing and feature engineering: The multi-source data is cleaned, standardized and feature generated, and a feature set for model input is generated based on the processed data. The feature generation includes generating a heterogeneity correction factor to correct for differences in management among pastures. Health trait association analysis: Dairy cows are grouped according to their lactation stage, and an independent health trait association analysis model is trained for each group. The standardized feature dataset is input into the pre-trained independent health trait association analysis model, which outputs the predicted probabilities of multiple health problems. The health trait association analysis model is trained using a nonlinear machine learning algorithm based on historical multi-source data and historical health records to establish a nonlinear association between multi-source data features and health traits. Health Index Prediction: The predicted probabilities of the multiple health problems and the heterogeneity correction factor are input into the health index calculation model to calculate the comprehensive health index. The health index calculation model explicitly introduces the heterogeneity correction factor to correct the prediction bias caused by differences in pasture management.

[0006] The above further includes: Furthermore, the method for generating the heterogeneity correction factor is to treat the pasture identifier as a random effect variable.

[0007] Furthermore, the method for generating the heterogeneity correction factor is to perform cluster analysis on the pastures to generate pasture category features, and use the pasture category features as the heterogeneity correction factor.

[0008] Furthermore, the feature set includes dynamic features, interaction item features, and domain knowledge features; The dynamic features are extracted from relevant time series data in the multi-source data, including moving average, rate of change, and seasonal trend; The interaction item features include the product or ratio between indicators from different data sources; The domain knowledge features include the heat stress accumulation index and the nutritional balance index.

[0009] Furthermore, dairy cows were grouped according to their lactation stage, including early, middle, and late lactation.

[0010] Furthermore, the nonlinear machine learning algorithms used in the health trait association analysis model include: generalized additive model, random forest, gradient boosting tree, XGBoost, LightGBM, support vector machine or neural network; The association analysis of the health traits specifically includes: Association analysis of reproductive disorders: Using a generalized additive model or random forest, with urea nitrogen, fat-to-protein ratio, environmental stress index and dairy cow activity as input features, we analyzed the nonlinear relationship between them and 21-day conception rate, and output the predicted probability of reproductive disorders. Nutritional metabolic disease association analysis: For early-stage dairy cows, a gradient boosting tree was used, with the change rates of β-hydroxybutyrate and urea nitrogen, fat-to-egg ratio, dietary formulation parameters, key microbiota, and rumination time as input features. Among these, the change rates of β-hydroxybutyrate and urea nitrogen were the main features. The nonlinear relationship between these factors and ketosis, hoof abnormalities, and rumen abnormalities was analyzed, and the predicted probability of nutritional metabolic diseases was output. The abundance of the key microbiota includes, but is not limited to, at least one of the following: the relative abundance of Prevotella, Vibrio butyricum, Ruminococcus, Filobacillus, Streptococcus, Lactobacillus, and Giant Ruminant Cocci. Association analysis of subclinical mastitis: For late-stage dairy cows, support vector machines or neural networks are used with somatic cell scores, milking equipment parameters, environmental hygiene scores and parity as input features, with somatic cell scores as the main feature. The nonlinear relationship between somatic cell scores and subclinical mastitis is analyzed, and the predicted probability of subclinical mastitis is output.

[0011] Furthermore, the performance of the health trait association analysis model and the health index calculation model is verified using an independent test set, including: Multiple data points from ≥5000 dairy cows randomly selected from independent ranches were used as the test set to calculate precision, recall, and F1 score. Time series validation was conducted to evaluate the model's ability to provide early warning of acute health events.

[0012] Furthermore, the health index calculation model calculates the comprehensive health index using the following formula: ,in, The comprehensive health index, As a heterogeneity correction factor, Let i be the predicted probability of the i-th health problem. The weights determined by the optimization algorithm for the i-th health problem are... The number of types of health problems.

[0013] Furthermore, the health index calculation model employs principal component analysis to reduce the dimensionality of the predicted probabilities of the multiple health problems and the heterogeneity correction factor, and uses the first principal component as the comprehensive health index.

[0014] Furthermore, this includes the model update step: New multi-source data and corresponding health records are periodically input into the health trait association analysis model and the health index calculation model, and online learning technology is used to update the parameters of the health trait association analysis model and the health index calculation model.

[0015] The present invention has the following beneficial effects: In this invention, ranch ID, management level, etc. are treated as random effect variables or generated features through cluster analysis, and represented as heterogeneity correction factors. These heterogeneity correction factors are explicitly incorporated into the prediction model, which enables the model to identify and remove systematic biases caused by management differences, thereby learning a purer and more universal relationship between health and indicators, and improving the model's generalization ability across different ranches. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the steps of a method for predicting the health index of dairy cows proposed in this invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Please see Figure 1 As shown, this invention provides a predictive method for determining the health index of dairy cows, comprising the following steps: Data Acquisition: Acquire multi-source data from the target dairy cows, including dairy cow production performance (DHI) data (milk yield, milk fat percentage, milk protein percentage, somatic cell count, etc., collected at a frequency of ≥ once a week), environmental data (temperature and humidity, light intensity, heat stress index, air quality) collected through IoT sensors, behavioral data (activity level, rumination time, feeding time, and lying down time) collected through wearable sensors, feeding management data (ration formulation, milking equipment parameters, and vaccination records) obtained from the management system, genetic data (breed, parity, pedigree, and genetic breeding values) obtained from breeding records, and rumen microbiome data collected from fecal or milk samples. Data preprocessing and feature engineering: The multi-source data is cleaned, standardized and feature generated, and a feature set for model input is generated based on the processed data. The feature generation includes generating a heterogeneity correction factor to correct for differences in management among pastures. Health trait association analysis: Dairy cows are grouped according to their lactation stage, and an independent health trait association analysis model is trained for each group. The standardized feature dataset is input into the pre-trained independent health trait association analysis model, which outputs the predicted probabilities of multiple health problems. The health trait association analysis model is trained using a nonlinear machine learning algorithm based on historical multi-source data and historical health records to establish a nonlinear association between multi-source data features and health traits. Health Index Prediction: The predicted probabilities of the multiple health problems and the heterogeneity correction factor are input into the health index calculation model to calculate the comprehensive health index. The health index calculation model explicitly introduces the heterogeneity correction factor to correct the prediction bias caused by differences in pasture management.

[0019] In this embodiment: Multi-source data collection: In the demonstration farm, 100 lactating dairy cows (including different parities and breeds) were selected as the target group for multi-source data collection. The collection period was 6 consecutive months, and the data collection frequency and content included: DHI measurement data: collected weekly, including milk yield (kg / d), milk fat percentage (%), milk protein percentage (%), somatic cell count (cells / mL), blood urea nitrogen (mg / dL), β-hydroxybutyrate (mmol / L), and fat-to-protein ratio. Data is automatically recorded using a DHI measuring instrument. Environmental data: collected in real time via an IoT sensor network; temperature and humidity sensors: record the temperature and relative humidity of the cattle shed, and calculate the heat stress index (THI, formula: Where T is temperature (°C) and RH is relative humidity (%); light sensor: records light intensity (lux); air quality sensor: monitors ammonia concentration (ppm). Behavioral data: collected via wearable neck tag sensors (e.g., accelerometer and gyroscope based), including: activity levels (steps / day), rumination time (minutes / day), feeding time (minutes / day), and lying time (minutes / day). Data is uploaded every 5 minutes. Feeding and management data: exported from the ranch management system, including: diet formulation: crude protein level (%), energy concentration (MJ / kg), mineral content; milking equipment parameters: milking pressure (kPa), hygiene score (1-5 points, based on visual assessment); vaccination records: vaccine type and vaccination date; Genetic data: obtained from breeding records, including variety (e.g., Holstein), parity, pedigree, and genetic breeding values ​​(e.g., TPI index).

[0020] Rumen microbiome data: Fecal samples were collected monthly, and microbial diversity (Shannon index) and abundance of key bacterial groups (such as Prevotella and Rumenococcus) were analyzed by 16S rRNA sequencing.

[0021] Sample Design: A stratified sampling strategy was adopted, stratifying the demonstration farms by size (large), region (temperate), and management level (high level), and randomly selecting dairy cows to track their entire lactation cycle (from calving to dry period). Additionally, to address heterogeneity, five similar farms were added as references, bringing the total sample size to 500 dairy cows.

[0022] Data preprocessing and feature engineering: Preprocessing and generating features from collected multi-source data. Data cleaning: For continuous variables (such as somatic cell count), use linear interpolation to fill the data; for categorical variables (such as variety), use mode filling, use the Z-score method to remove outliers with absolute values ​​greater than 3 (such as activity level suddenly dropping to 0), and perform Z-score standardization on all numerical features to make the mean 0 and the standard deviation 1. Feature generation: Dynamic features are extracted from time series data, including the 7-day moving average and rate of change (derivative) of somatic cell scores, seasonal trends in rumination time, periodic components are extracted using Fourier transform, and the product of urea nitrogen and fat-protein ratio, the ratio of β-hydroxybutyrate to activity level are created. Heat stress accumulation index (number of consecutive days with THI > 72) and nutritional balance index (dietary protein to energy ratio) are generated.

[0023] Heterogeneity correction: Pasture ID is introduced as a random effect variable and processed by a mixed-effects model during model training; alternatively, K-means clustering (k=3) is used to divide pastures into high, medium and low management level groups to generate pasture category features (such as "management level_high").

[0024] Association analysis of health traits: Dairy cows were grouped according to lactation stage (early stage: 0-100 days; mid stage: 101-200 days; late stage: >200 days), and an independent association analysis model was trained for each group. Association analysis of reproductive disorders (for the whole population and stages): A generalized additive model (GAM) was used, with the 21-day conception rate (a binary variable, yes / no) as the dependent variable and urea nitrogen, fat-to-protein ratio, environmental stress index (THI), and activity level as independent variables.

[0025] Nutritional metabolic disease association analysis (focusing on the early lactation group): Gradient boosting tree (GBDT, implemented with XGBoost) was used, with ketosis (yes / no) as the dependent variable and β-hydroxybutyrate, fat-to-protein ratio, diet formulation, rumination time, and microbial diversity (Shannon index) as independent variables.

[0026] Association analysis of subclinical mastitis (focusing on the late lactation group): Support vector machine (SVM, kernel function RBF) was used, with somatic cell score (continuous value) as the dependent variable and milking equipment parameters, environmental hygiene score, parity, and lying time as independent variables.

[0027] The model was trained using historical data (data from the previous four months as the training set), and 5-fold cross-validation was used to adjust hyperparameters (such as the smoothing parameter of GAM and the learning rate of XGBoost).

[0028] Health index prediction model construction and validation: The integrated model XGBoost was adopted, with all preprocessed features (including heterogeneity correction factors) as input, and the combined prediction probabilities (P1, P2, P3) of reproductive disorders, nutritional metabolic diseases and subclinical mastitis were output.

[0029] For time series data, an LSTM model is used to process continuous observations and predict the health index for the coming week. Health index integration: A weighted geometric mean is used to calculate the comprehensive health index (HI), and principal component analysis (PCA) is used to reduce the dimensionality of the joint probability. The first principal component is used as the health index.

[0030] Model validation: 5000 multi-source data records were randomly selected from an independent ranch (not involved in training) as the test set.

[0031] Evaluation metrics: accuracy (85%), recall (82%), F1 score (83%). Time series validation showed that the model could provide an early warning of mastitis outbreaks up to one week in advance (AUC = 0.89).

[0032] Model updates: The XGBoost model parameters are updated monthly with new data using online learning technology, and the weights are adjusted based on veterinary diagnostic results.

[0033] In one embodiment, the method for generating the heterogeneity correction factor is to treat the pasture identifier as a random effect variable.

[0034] In this embodiment: Model Building: Construct a logistic regression model that includes random effects. ,in, Let represent the probability that the j-th cow in the i-th pasture experiences a health event. It is a logical function, that is , This is the fixed effects intercept of the model, representing the average baseline level across all pastures. These are fixed-effects coefficients, corresponding to the effect size of predictor variables (such as urea nitrogen, fat-to-protein ratio, and activity level) shared by all pastures. It is the predictor value of the j-th cow in the i-th pasture. This is the random effects term, i.e., the heterogeneity correction factor. It specifically refers to the effect of the i-th pasture, assuming it follows a normal distribution. ,in, This represents the magnitude of variation between pastures. It is the error term, which also follows a normal distribution, i.e. ; Parameter estimation: Variance components were estimated using the restricted maximum likelihood (REML) method. (random effects variance of pasture) and (Residual variance), calculating ranch random effects using Best Linear Unbiased Prediction (BLUP). For example, if the breeding value of ranch A is +150 and that of ranch B is -30, it indicates that ranch A has a significant management advantage.

[0035] Model validation: Compare the model performance before and after calibration, or test the significance of the random effects of pastures using analysis of variance (ANOVA).

[0036] In one embodiment, the method for generating the heterogeneity correction factor is to perform cluster analysis on the pastures to generate pasture category features, and use the pasture category features as the heterogeneity correction factor.

[0037] In this embodiment: Data preprocessing: cleaning outliers (such as pasture area, extreme yield values), and standardizing features (such as Z-score standardized farmland and pasture area).

[0038] Determine the number of clusters k: Select the optimal value of k by using the silhouette coefficient (>0.5 indicates good clustering quality) or the elbow rule (SSE decreases as k increases).

[0039] K-means: Initialization: Randomly select k cluster centers; Sample allocation: Calculate the Euclidean distance from each pasture to each center. Assigned to the nearest center; Update centers: Recalculate the cluster mean as the new centers; Iterative convergence: until the center is stable or the maximum number of iterations (e.g., 100 times) is reached. Application of results: After clustering, pasture category features are generated and incorporated into the health index calculation model.

[0040] In one embodiment, the feature set includes dynamic features, interaction item features, and domain knowledge features; The dynamic features are extracted from relevant time series data in the multi-source data, including moving average, rate of change, and seasonal trend; The interaction features include the product or ratio between indicators from different data sources (product of urea nitrogen and fat-protein ratio, ratio of β-hydroxybutyrate to activity level). The domain knowledge features include the heat stress accumulation index and the nutritional balance index.

[0041] In one embodiment, dairy cows are grouped according to their lactation stage, including early, middle, and late lactation.

[0042] In one embodiment, the nonlinear machine learning algorithm used in the health trait association analysis model includes: Generalized Additive Model (GAM), Random Forest, Gradient Boosting Tree (GBDT), XGBoost, LightGBM, Support Vector Machine (SVM), or Neural Network. The association analysis of the health traits specifically includes: Association analysis of reproductive disorders: Using a generalized additive model or random forest, with urea nitrogen, fat-to-protein ratio, environmental stress index and dairy cow activity as input features, we analyzed the nonlinear relationship between them and 21-day conception rate, and output the predicted probability of reproductive disorders. Nutritional metabolic disease association analysis: For early-stage dairy cows, a gradient boosting tree was used, with the change rates of β-hydroxybutyrate and urea nitrogen, fat-to-egg ratio, diet formulation parameters, key microbiota, and rumination time as input features. Among these, the change rates of β-hydroxybutyrate and urea nitrogen were the main features. The nonlinear relationship between these features and ketosis, limb and hoof abnormalities, and rumen abnormalities was analyzed, and the predicted probability of nutritional metabolic diseases was output. The abundance of the key microbiota included, but was not limited to, at least one of the following: the relative abundance of Prevotella, Vibrio butyricum, Ruminococcus, Filobacillus, Streptococcus, Lactobacillus, and Giant Ruminant Cocci. Specifically, when predicting the risk of subacute rumen acidosis, the health trait association analysis model focuses on the increase in abundance of Streptococcus and Lactobacillus, as well as the decrease in ruminant giant cocci and microbial diversity; when predicting the risk of ketosis, the model focuses on the decrease in abundance of Prevotella, Vibrio butyricum and Ruminococcus. Association analysis of subclinical mastitis: For late-stage dairy cows, support vector machines or neural networks are used with somatic cell scores, milking equipment parameters, environmental hygiene scores and parity as input features, with somatic cell scores as the main feature. The nonlinear relationship between somatic cell scores and subclinical mastitis is analyzed, and the predicted probability of subclinical mastitis is output.

[0043] In one embodiment, the performance of the health trait association analysis model and the health index calculation model is verified using an independent test set, including: Multiple data points from ≥5000 dairy cows randomly selected from independent ranches were used as the test set to calculate precision, recall, and F1 score. Time series validation was conducted to assess the model's ability to provide early warning of acute health events, such as the week before a mastitis outbreak.

[0044] In one embodiment, the health index calculation model calculates the comprehensive health index using the following formula: ,in, The comprehensive health index, As a heterogeneity correction factor, Let i be the predicted probability of the i-th health problem. The weights determined by the optimization algorithm for the i-th health problem are... The number of types of health problems.

[0045] In this embodiment: The predicted probabilities of multiple health problems are obtained from association analysis of health traits, denoted as […]. .

[0046] Heterogeneity correction factors, derived from data preprocessing and feature engineering, are generated for the target dairy cow or its ranch and denoted as [missing information]. . The model was trained to derive a systematic effect on the health index of a specific ranch.

[0047] Calculation formula: .

[0048] Optimization objective: Minimize the model's predictions The difference between the actual veterinary diagnosis (such as labels like "healthy", "sub-healthy", or "ill").

[0049] This process is completed during the model training phase, and the optimized fixed weights are used directly during application.

[0050] In one embodiment, the health index calculation model employs principal component analysis (PCA) to perform dimensionality reduction and integration of the predicted probabilities of the multiple health problems and the heterogeneity correction factor, and uses the first principal component as the comprehensive health index.

[0051] In this embodiment: Data collection: Predicted probabilities of multiple health problems are obtained from association analysis of health traits, denoted as... The heterogeneity correction factor generated for the target dairy cow or its pasture, obtained from data preprocessing and feature engineering, is denoted as... . The model was trained to derive a systematic effect on the health index of a specific ranch.

[0052] Constructing input vectors: Construct an input vector for each cow. This vector contains all health probabilities and correction factors.

[0053] PCA projection: The PCA transformation fitted during the model training phase is applied to vector VV. PCA will find a new set of orthogonal coordinate axes (principal components), where the first principal component (PC1) is the direction with the largest data variance, that is, the comprehensive dimension that best distinguishes different health states.

[0054] Extraction index: the vector Projecting the value onto the first principal component yields the comprehensive health index. , represented as The mean vector and PC1 load vector are determined during PCA training.

[0055] In one embodiment, a model update step is included: New multi-source data and corresponding health records are periodically input into the health trait association analysis model and the health index calculation model, and online learning technology is used to update the parameters of the health trait association analysis model and the health index calculation model.

[0056] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method of predicting a cow production health index, characterized by, The method comprises the following steps: Data collection: obtaining multi-source data of target dairy cows, the multi-source data comprising dairy cow production performance measurement data, environmental data collected through Internet of Things sensors, behavioral data collected through wearable sensors, feeding management data obtained through a management system, genetic data obtained through breeding records, and rumen microbiome data collected through fecal or milk samples; Data preprocessing and feature engineering: cleaning, standardizing, and feature generating the multi-source data, and generating a feature set for model input based on the processed data, wherein the feature generation comprises generating a heterogeneity correction factor for correcting management differences between pastures; Health trait association analysis: grouping dairy cows according to lactation stages, and training an independent health trait association analysis model for each group, inputting the standardized feature data set into the pre-trained independent health trait association analysis model, and outputting a plurality of health problem prediction probabilities; wherein the health trait association analysis model is trained based on historical multi-source data and historical health records using a nonlinear machine learning algorithm, and is used to establish a nonlinear association between multi-source data features and health traits; Health index prediction: inputting the plurality of health problem prediction probabilities and the heterogeneity correction factor into a health index calculation model to calculate a comprehensive health index, wherein the health index calculation model explicitly introduces the heterogeneity correction factor to correct prediction bias caused by differences in pasture management.

2. A method of predicting a dairy cow production health index according to claim 1, characterized in that: The method for generating the heterogeneity correction factor is to process the pasture identification as a random effect variable.

3. A method of predicting a dairy cow production health index according to claim 1, characterized in that: The method for generating the heterogeneity correction factor is to perform cluster analysis on the pastures to generate pasture category features, and use the pasture category features as the heterogeneity correction factor.

4. A method of predicting a dairy cow production health index according to claim 1, characterized in that: The feature set comprises dynamic features, interaction term features, and domain knowledge features; The dynamic features are extracted from relevant time series data in the multi-source data, including moving average, rate of change, and seasonal trend; The interaction term features include the product or ratio of different data source indicators; The domain knowledge features include heat stress accumulation index and nutrient balance index.

5. A method of predicting a dairy cow production health index according to claim 1, characterized in that: The dairy cows are grouped according to lactation stages, including early, middle, and late stages.

6. A method of predicting a dairy cow production health index according to claim 6, characterized in that: The nonlinear machine learning algorithm used in the health trait association analysis model comprises a generalized additive model, a random forest, a gradient boosting tree, XGBoost, LightGBM, a support vector machine, or a neural network; The health trait association analysis specifically comprises: Reproduction disorder association analysis: using a generalized additive model or a random forest, using urea nitrogen, fat-to-protein ratio, environmental stress index, and dairy cow activity as input features, analyzing the nonlinear relationship between them and 21-day conception rate, and outputting reproduction disorder prediction probability; Nutrition metabolism disease association analysis: for early-stage cows, gradient boosting tree is used, with the change rate of β-hydroxybutyric acid and urea nitrogen, fat-protein ratio, ration formulation parameters, key flora and rumination time as input features, among which the change rate of β-hydroxybutyric acid and urea nitrogen is the main one, to analyze the nonlinear relationship between them and ketosis, limb and hoof abnormalities, and abnormal rumen, and output the prediction probability of nutrition metabolism disease; Subclinical mastitis association analysis: for late-stage cows, support vector machine or neural network is used, with somatic cell score, milking equipment parameters, environmental hygiene score and parity as input features, among which somatic cell score is the main one, to analyze the nonlinear relationship between them and subclinical mastitis, and output the prediction probability of subclinical mastitis.

7. A method of predicting a dairy cow production health index according to claim 1, characterized in that: The performance of the health trait association analysis model and the health index calculation model is verified by using an independent test set, including: Randomly selecting multi-source data of ≥5000 cows from an independent farm as a test set, calculating the accuracy, recall rate and F1 score; Time series verification is performed to evaluate the early warning ability of the model for acute health events.

8. A method of predicting a dairy cow production health index according to claim 1, characterized in that: The health index calculation model uses principal component analysis method to reduce and integrate the multiple health problem prediction probabilities and the heterogeneity correction factor, and takes the first principal component as the comprehensive health index.

9. A method of predicting a dairy cow production health index according to claim 1, characterized in that: The model updating step includes: Periodically input new multi-source data and corresponding health records into the health trait association analysis model and the health index calculation model, and update the parameters of the health trait association analysis model and the health index calculation model by using online learning technology.