Method and system for predicting nitrogen utilization efficiency of Hu sheep in fattening period based on machine learning algorithm

By using environmental parameter sensors to collect data in meat sheep breeding, a nitrogen utilization efficiency prediction model based on machine learning is constructed, which solves the problem of high cost and low accuracy of traditional methods, and realizes high-precision prediction and health monitoring of the nitrogen utilization efficiency of lake sheep.

CN120524801APending Publication Date: 2025-08-22NORTHEAST AGRICULTURAL UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510609415.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

In large-scale meat sheep breeding, traditional nitrogen utilization efficiency prediction methods are costly and have low accuracy, making it difficult to meet the management needs of meat sheep fattening stage. Existing machine learning algorithms are widely used in dairy cattle and beef cattle fields, and research on meat sheep is still blank.

Method used

Environmental parameter sensors are used to collect data such as temperature, humidity, ammonia concentration, body weight, nitrogen intake and fiber intake, and use machine learning algorithms to construct a prediction model of nitrogen utilization efficiency of lake sheep, including linear regression, elastic network, K-nearest neighbor and artificial neural network models, and improve prediction accuracy through feature selection and data set segmentation.

Benefits of technology

The prediction accuracy of the nitrogen utilization efficiency of meat sheep is improved, accurate monitoring of the health and nutrition of lake sheep is achieved, management costs are reduced, and the intelligence level of the breeding farm is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120524801A_ABST
    Figure CN120524801A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for predicting fattening Hu sheep nitrogen utilization efficiency based on a machine learning algorithm, and relates to the field of nitrogen utilization efficiency prediction. The problems that existing nitrogen utilization efficiency prediction is limited by technologies, equipment and cost, prediction accuracy is not high, and actual requirements are difficult to meet are solved, and the nitrogen utilization rate of individuals or sheep flock can be effectively monitored. The method comprises the following steps: acquiring T, RH, NH3, BW, NI, NDFI and ADFI data by using an environmental parameter sensor to form a winter data set, a summer data set and a mixed data set of the winter data set and the summer data set; establishing winter and summer data sets and a mixed data set based on an ML algorithm, and constructing a Hu sheep individual nitrogen utilization efficiency prediction model; and respectively comparing the growth performance of the Hu sheep in different ML algorithms and evaluating predictive variables by using the individual nitrogen utilization efficiency prediction model to complete the prediction of the nitrogen utilization efficiency of the Hu sheep in the fattening period based on the machine learning algorithm. The method is also suitable for the field of precise monitoring of mutton sheep health and nutrition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of nitrogen utilization efficiency prediction, and in particular to a method and system for predicting nitrogen utilization efficiency of Hu sheep during fattening period based on a machine learning algorithm. Background Art

[0002] In large-scale mutton sheep farming, feed costs remain high, and undigested nitrogen is excreted through excrement, causing environmental pollution. Relying solely on traditional techniques to adjust the composition of the diet is no longer an effective solution to the dilemma of low production efficiency. Nitrogen utilization efficiency, a key indicator reflecting this problem, decreases, leading to more undigested nitrogen excreted through feces and urine, increasing farming costs and causing environmental pollution. Traditional methods for predicting nitrogen utilization efficiency rely on manually collecting fecal and urine samples, calculating nitrogen excretion and utilization through laboratory analysis, and applying previously established linear regression equations or other linear models for prediction. While this method can be used for managing mutton sheep during the fattening phase, it is costly, and the model's prediction accuracy decreases significantly when the input features exhibit a nonlinear relationship with nitrogen utilization efficiency. Against this backdrop, the advantages of multidimensional information integration are becoming increasingly apparent. By increasing the types of input features and the amount of data, the predictive performance of machine learning models can be effectively improved, thereby addressing the problem of prediction accuracy caused by insufficient input variable information. With the promotion of smart farming policies, the maturity of pasture sensor technology provides basic data support for this study, while machine learning (ML) algorithms provide a new technical means for using production data to establish a nitrogen utilization efficiency monitoring model for Hu sheep during the fattening period.

[0003] Current applications of machine learning algorithms for predicting nitrogen use efficiency (NUE) are primarily focused on dairy and beef cattle. Studies in dairy cows have shown that when a modified linear model uses milk yield, crude protein content, and neutral detergent fiber content as input variables, NUE estimates are closer to actual measurements. Based on a dataset of 951 lactating dairy cows from 43 digestibility studies, researchers developed and evaluated a total nitrogen emission prediction model using multiple linear regression (MLR) and three machine learning algorithms (artificial neural network (ANN), random forest regression (RF), and support vector machine (SVM). Cross-validation results showed that the ANN model exhibited the best performance (RMSE = 34.7, CCC = 0.7), with significantly higher prediction accuracy than the MLR, RF, and SVR models (RMSEs of 44.7, 46.8, and 44.9, respectively). In addition, the classification model developed based on body weight and milk production divided nitrogen emissions into three groups: low (<300g / d / head), medium (300-450g / d / head), and high (>450g / d / head). The results showed that the prediction accuracy of the gradient boosting regression tree (BRT) was slightly better than that of the Bayesian network (BN). In beef cattle research, the researchers set three dietary protein gradients: low (84-143g / kgDM), medium (144-162g / kgDM), and high (163-217g / kgDM). The linear equation was used to predict the excretion and obtain the fecal R 2 =0.72, urine R 2 =0.83. A case study using Holstein cows showed that support vector regression (SVR) was superior to linear regression (LR) in predicting excretion, with the fecal nitrogen excretion prediction model developed based on the multivariate linear model achieving R 2 =0.94. Analysis of data from 286 beef cattle revealed that nitrogen excretion was significantly correlated with dry matter intake, body weight, nitrogen intake, and feed-to-weight ratio (P < 0.01). When nitrogen intake was used as a single predictor variable, joint modeling with body weight increased the coefficient of determination to 0.824 (MSE = 16.2). Integration of 751 observational data sets from 18 publications revealed that increasing dietary digestibility and reducing nitrogen and fiber content can improve nitrogen utilization efficiency. The best mean prediction errors (MPEs) for fecal nitrogen and urine nitrogen were 0.208 and 0.177, respectively.

[0004] Traditionally, nitrogen use efficiency (NUE) has been assessed primarily through digestibility measurements in animal studies. However, this approach faces significant challenges in practical application on large-scale sheep farms: individual-level feces and urine collection and feed intake records are difficult to implement, and the labor and material costs are prohibitive. While machine learning algorithms have made progress in predicting NUE in ruminants, existing research has primarily focused on dairy and beef cattle, with no research on NUE in sheep. Summary of the Invention

[0005] Nitrogen use efficiency (NUE) reflects the extent to which animals digest and utilize feed in large-scale livestock production, as well as the environmental impact of excreted nitrogen in urine and feces. However, due to limitations in technology, equipment, and costs, methods for effectively monitoring NUE at the individual or flock level are limited, and predictions are often inaccurate, making them difficult to meet practical needs.

[0006] To solve the above technical problems, the present invention is achieved through the following technical solutions:

[0007] Solution 1: The present invention proposes a method for predicting nitrogen utilization efficiency of Hu sheep during fattening period based on a machine learning algorithm, the method comprising the following steps:

[0008] Step 1: Use environmental parameter sensors to collect temperature T, humidity RH, ammonia NH3, body weight BW, nitrogen intake NI, neutral detergent fiber intake NDFI, and acid detergent fiber intake ADFI data to form a winter dataset, a summer dataset, and a mixed dataset of the two.

[0009] Step 2: Establish winter data set, summer data set and mixed data set of the two based on ML algorithm to construct individual nitrogen use efficiency prediction model of Hu sheep;

[0010] Step 3: Compare the individual nitrogen utilization efficiency prediction models of Hu sheep in different ML algorithms to evaluate the prediction variables and complete the prediction of nitrogen utilization efficiency of Hu sheep based on machine learning algorithm.

[0011] Furthermore, a preferred embodiment is provided, in which the environmental data described in step 1 is collected based on the equipment of the intelligent monitoring system for livestock and poultry breeding environment.

[0012] Furthermore, a preferred embodiment is provided, in which the construction of the individual nitrogen utilization efficiency prediction model of the lake sheep in step 2 includes the steps of constructing a linear regression model LR, an elastic network model EN, a K-nearest neighbor model KNN, and an artificial neural network model ANN.

[0013] Furthermore, a preferred embodiment is provided, in which the recursive feature elimination method is used to select features of the constructed test model LR, elastic network model EN, K-nearest neighbor model KNN, and artificial neural network model ANN, that is, by constructing the test model LR, the features to be screened are sequentially combined and introduced, and the RMSE, R 2 And the numerical value of MAE.

[0014] Furthermore, a preferred embodiment is provided, wherein the RMSE, R 2 And the method of calculating the numerical value of MAE is:

[0015]

[0016] Where N is the number of samples, Xi is the observed value, is the predicted value, is the mean of the observed values.

[0017] Furthermore, a preferred embodiment is provided, in which the winter dataset, summer dataset, and the mixed dataset of the two described in step 2 are respectively divided into a training set and a test set, and the ratio of the training set to the test set is 8:2.

[0018] Furthermore, a preferred embodiment is provided, wherein step 2 also includes the step of performing modeling analysis on the winter dataset, the summer dataset, and a mixed dataset of the winter dataset and the summer dataset.

[0019] Solution 2: A system for predicting nitrogen utilization efficiency of Hu sheep during fattening period based on a machine learning algorithm, the system comprising:

[0020] The data acquisition module is used to collect temperature T, humidity RH, ammonia NH3, body weight BW, nitrogen intake NI, neutral detergent fiber intake NDFI, and acid detergent fiber intake ADFI data using environmental parameter sensors to form a winter data set, a summer data set, and a mixed data set of the two;

[0021] A nitrogen utilization efficiency prediction model construction module is used to establish a winter dataset, a summer dataset, and a mixed dataset of the two based on the ML algorithm, and use the mixed dataset to construct a nitrogen utilization efficiency prediction model for individual Hu sheep;

[0022] The prediction module is used to compare the individual nitrogen utilization efficiency prediction models of Hu sheep in different ML algorithms to evaluate the prediction variables and complete the prediction of nitrogen utilization efficiency of Hu sheep based on machine learning algorithms.

[0023] The present invention is beneficial in that:

[0024] The method for predicting the nitrogen utilization efficiency of Hu sheep during the fattening period based on a machine learning algorithm described in the present invention collects NI, NDFI, ADFI, and BW, and uses sensors in the pasture to collect indoor environmental data information such as RH, T, and NH3 in the pasture. By integrating the multi-dimensional characteristics of Hu sheep and using the ML algorithm, a nitrogen utilization efficiency prediction regression model is established to improve the intelligent management and precise feeding level of the pasture, and realize accurate monitoring of the health and nutrition of meat sheep.

[0025] The method described in the present invention for predicting the nitrogen utilization efficiency of Hu sheep during the fattening period based on a machine learning algorithm, and the use of the ML algorithm to establish a NUE prediction model based on the multi-dimensional characteristic information of Hu sheep, has been verified to have certain feasibility through experiments. The characteristic weights of the N, NDF, ADF intake in the diet and the BW of Hu sheep are relatively high, which improves the prediction accuracy of the model. The RH, NH3, and T in the sheep house can improve the prediction accuracy of the model.

[0026] In the method for predicting nitrogen utilization efficiency of Hu sheep in the fattening period based on a machine learning algorithm described in the present invention, the comprehensive prediction performance of the KNN model is the best in the constructed winter data set, summer data set, and a mixed data set of the two.

[0027] The method for predicting nitrogen utilization efficiency of Hu sheep during fattening period based on machine learning algorithm described in the present invention can provide effective information for monitoring and regulating NUE of Hu sheep based on NI, NDFI, ADFI, BW and RH, T and NH3 in the house.

[0028] The present invention is also applicable to the field of precise monitoring of the health and nutrition of mutton sheep. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 This is a flowchart of the method for predicting nitrogen utilization efficiency of Hu sheep during fattening period based on machine learning algorithm as described in embodiment 1.

[0030] Figure 2 The test set described in the eleventh embodiment is a schematic diagram comparing the predicted values ​​and actual values ​​of nitrogen utilization efficiency using four models in a mixed data set.

[0031] Among them, (a) is a schematic diagram of the comparison between the predicted value of nitrogen utilization efficiency using the linear regression model LR and the true value in the test set as a mixed data set; (b) is a schematic diagram of the comparison between the predicted value of nitrogen utilization efficiency using the elastic network model EN and the true value in the test set as a mixed data set; (c) is a schematic diagram of the comparison between the predicted value of nitrogen utilization efficiency using the KNN model and the true value in the test set as a mixed data set; (d) is a schematic diagram of the comparison between the predicted value of nitrogen utilization efficiency using the artificial neural network model ANN and the true value in the test set as a mixed data set.

[0032] Figure 3 This is a schematic diagram comparing the importance of predicting variables using four models in a mixed data set for the test set as described in the eleventh embodiment.

[0033] Among them, (a) is a schematic diagram of the comparison of the importance of predicting variables using the linear regression model LR in the test set as a mixed data set; (b) is a schematic diagram of the comparison of the importance of predicting variables using the elastic network model EN in the test set as a mixed data set; (c) is a schematic diagram of the comparison of the importance of predicting variables using the KNN model in the test set as a mixed data set; (d) is a schematic diagram of the comparison of the importance of predicting variables using the artificial neural network model ANN in the test set as a mixed data set.

[0034] Figure 4 The test set described in the eleventh embodiment is a schematic diagram of the correlation between the features in the mixed data set.

[0035] Figure 5 The test set described in the eleventh embodiment is a schematic diagram of the predicted values ​​and true values ​​of nitrogen utilization efficiency of four models in the summer data set.

[0036] Among them, (a) is a schematic diagram of the comparison between the predicted value of nitrogen utilization efficiency using the linear regression model LR and the true value in the test set for the summer data set; (b) is a schematic diagram of the comparison between the predicted value of nitrogen utilization efficiency using the elastic network model EN and the true value in the test set for the summer data set; (c) is a schematic diagram of the comparison between the predicted value of nitrogen utilization efficiency using the KNN model and the true value in the test set for the summer data set; (d) is a schematic diagram of the comparison between the predicted value of nitrogen utilization efficiency using the artificial neural network model ANN and the true value in the test set for the summer data set.

[0037] Figure 6 This is a schematic diagram comparing the importance of predictive variables using four models in the summer dataset as the test set described in the eleventh embodiment.

[0038] Among them, (a) is a schematic diagram of the comparison of the importance of predicting variables using the linear regression model LR in the test set for the summer dataset; (b) is a schematic diagram of the comparison of the importance of predicting variables using the elastic network model EN in the test set for the mixed summer dataset; (c) is a schematic diagram of the comparison of the importance of predicting variables using the KNN model in the test set for the summer dataset; (d) is a schematic diagram of the comparison of the importance of predicting variables using the artificial neural network model ANN in the test set for the summer dataset.

[0039] Figure 7 The test set described in the eleventh embodiment is a schematic diagram of the correlation between the features in the summer data set.

[0040] Figure 8 The test set described in the eleventh embodiment is a schematic diagram comparing the predicted values ​​and actual values ​​of nitrogen utilization efficiency using four models in the winter data set.

[0041] Among them, (a) is a schematic diagram of the comparison between the predicted value of nitrogen utilization efficiency using the linear regression model LR and the true value in the test set for the winter data set; (b) is a schematic diagram of the comparison between the predicted value of nitrogen utilization efficiency using the elastic network model EN and the true value in the test set for the winter data set; (c) is a schematic diagram of the comparison between the predicted value of nitrogen utilization efficiency using the KNN model and the true value in the test set for the winter data set; (d) is a schematic diagram of the comparison between the predicted value of nitrogen utilization efficiency using the artificial neural network model ANN and the true value in the test set for the winter data set.

[0042] Figure 9 This is a schematic diagram comparing the importance of predictive variables using four models in the winter dataset as the test set described in the eleventh embodiment.

[0043] Among them, (a) is a schematic diagram of the comparison of the importance of predicting variables using the linear regression model LR in the test set for the winter dataset; (b) is a schematic diagram of the comparison of the importance of predicting variables using the elastic network model EN in the test set for the winter dataset; (c) is a schematic diagram of the comparison of the importance of predicting variables using the KNN model in the test set for the winter dataset; (d) is a schematic diagram of the comparison of the importance of predicting variables using the artificial neural network model ANN in the test set for the winter dataset.

[0044] Figure 10 The test set described in the eleventh embodiment is a schematic diagram of the correlation between the various features in the winter data set. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solutions and advantages of the implementation methods of this application clearer, the technical solutions in the implementation methods of this application will be clearly and completely described below in combination with the drawings in the implementation methods of this application. Obviously, the described implementation methods are only part of the implementation methods of this application, not all of the implementation methods.

[0046] Implementation method 1: This implementation method proposes a method for predicting nitrogen utilization efficiency of Hu sheep during fattening period based on a machine learning algorithm, the method comprising the following steps:

[0047] Step 1: Use environmental parameter sensors to collect temperature T, humidity RH, ammonia NH3, body weight BW, nitrogen intake NI, neutral detergent fiber intake NDFI, and acid detergent fiber intake ADFI data to form a winter dataset, a summer dataset, and a mixed dataset of the two.

[0048] Step 2: Establish a winter dataset, a summer dataset, and a mixed dataset of the two based on the ML algorithm, and use the mixed dataset to construct a prediction model for nitrogen use efficiency of individual mutton sheep;

[0049] Step 3: Compare the individual nitrogen utilization efficiency prediction models of Hu sheep in different ML algorithms to evaluate the prediction variables and complete the prediction of nitrogen utilization efficiency of Hu sheep based on machine learning algorithm.

[0050] Implementation method 2. This implementation method further limits the method for predicting nitrogen utilization efficiency of Hu sheep during fattening period based on machine learning algorithm described in implementation method 1. The environmental data described in step 1 is collected based on the intelligent monitoring system equipment of livestock and poultry breeding environment.

[0051] Implementation method three. This implementation method further limits the method for predicting the nitrogen utilization efficiency of Hu sheep during the fattening period based on the machine learning algorithm described in implementation method one. In step 2, constructing an individual nitrogen utilization efficiency prediction model for meat sheep includes the steps of constructing a linear regression model LR, an elastic network model EN, a K-nearest neighbor model KNN, and an artificial neural network model ANN.

[0052] Implementation method 4. This implementation method is a further limitation of the method for predicting nitrogen utilization efficiency of Hu sheep in fattening period based on machine learning algorithm described in implementation method 3. The recursive feature elimination method is used to select features of the constructed test model LR, elastic network model EN, K-nearest neighbor model KNN, and artificial neural network model ANN. That is, by constructing the test model LR, the features to be screened are combined and introduced in sequence, and the RMSE, R 2 And the numerical value of MAE.

[0053] Implementation mode 5. This implementation mode further limits the method for predicting nitrogen utilization efficiency of Hu sheep in fattening period based on machine learning algorithm described in implementation mode 1. The RMSE, R 2 And the method of calculating the numerical value of MAE is:

[0054]

[0055] Where N is the number of samples, Xi is the observed value, is the predicted value, is the mean of the observed values.

[0056] Implementation method six. This implementation method further limits the method for predicting the nitrogen utilization efficiency of Hu sheep during the fattening period based on the machine learning algorithm described in implementation method one. The winter data set, summer data set and the mixed data set of the two described in step 2 are divided into training sets and test sets respectively, and the ratio of training sets to test sets is 8:2.

[0057] Implementation method seven. This implementation method further limits the method for predicting nitrogen utilization efficiency of Hu sheep during fattening period based on machine learning algorithm described in implementation method one. Step 2 also includes the step of modeling and analyzing the winter data set, summer data set and the mixed data set of the two.

[0058] Implementation 8: This implementation provides a system for predicting nitrogen utilization efficiency of lake sheep based on a machine learning algorithm, the system comprising:

[0059] The data acquisition module is used to collect temperature T, humidity RH, ammonia NH3, body weight BW, nitrogen intake NI, neutral detergent fiber intake NDFI, and acid detergent fiber intake ADFI data using environmental parameter sensors to form a winter data set;

[0060] A nitrogen utilization efficiency prediction model construction module is used to establish a winter dataset, a summer dataset, and a mixed dataset of the two based on the ML algorithm, and use the mixed dataset to construct a nitrogen utilization efficiency prediction model for individual Hu sheep;

[0061] The prediction module is used to compare the individual nitrogen utilization efficiency prediction models of Hu sheep in different ML algorithms to evaluate the prediction variables and complete the prediction of nitrogen utilization efficiency of Hu sheep based on machine learning algorithms.

[0062] Implementation method 9. This implementation method proposes a computer device including a memory and a processor, wherein the memory stores a computer program. When the processor runs the computer program stored in the memory, the processor executes the method described in any one of implementation methods 1 to 7.

[0063] Implementation 10: This implementation proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the method described in any one of Implementations 1 to 7 are implemented.

[0064] Implementation 11: This implementation provides an example, which is used to explain the above implementations 1 to 10. Specifically, the example is as follows:

[0065] See also Figure 1 and Figure 10 This embodiment describes the present invention. This embodiment uses the effective T, RH, NH3 and BW, NI, NDFI, and ADFI collected by environmental parameter sensors to form a winter data set, a summer data set, and a mixed data set of the two. An individual nitrogen utilization efficiency prediction model for meat sheep based on the winter data set, the summer data set, and the mixed data set of the two is established. The performance of different ML algorithm models and the correlation between the prediction variables are explored. At the same time, compared with other data sets, the influence of different environmental parameters on the model performance is explored.

[0066] The test method is as follows:

[0067] This experiment was conducted in winter (n=75) and summer (n=75). A total of 150 two-month-old Huyang fattening male lambs were selected. Three dietary crude protein levels (17%, 15%, and 13%) were designed and fed to the sheep before, during, and after fattening in winter and summer, respectively. The experimental period was 168 days, with each stage lasting 28 days, including a 7-day transition period and a 21-day sampling period. Data collected from the two seasons were constructed into three datasets: winter data (n=225), summer data (n=225), and mixed data (n=450). Each dataset was randomly divided into a training set (80%) and a test set (20%) to construct four models and compare their predictive performance. The dietary nutrient levels are shown in Table 1.

[0068] Table 1 Nutritional levels of diets (dry matter basis)

[0069]

[0070] Among them, the nutritional levels are all measured values.

[0071] 1. Diet and feeding management

[0072] This experiment was conducted in winter (n=75) and summer (n=75) with a total of 150 male Hu sheep fattening lambs, approximately two months old. Before the experiment began, the sheep pens were thoroughly scrubbed and disinfected with quicklime. Each pen housed 15 sheep and had slatted floors. Environmental sensors and data collection equipment were installed in the middle of the flock, uploading real-time data to a computer. Each pen was equipped with a feed trough and automatic drinking water system to ensure free access to food and water for all Hu sheep during the experiment. All lambs received the same feeding and immunization schedule. Three dietary crude protein levels of 17%, 15%, and 13% were designed. During the pre-, mid-, and post-fattening stages of both seasons, pelleted feed was fed at a concentrate-to-roughage ratio of 7:3. Feed amounts were adjusted based on the previous day's feed intake to ensure at least 5% residual feed in each trough each day. Feeding times were 6:00 AM and 6:00 PM daily. During each phase, a transition period for feeding was conducted from days 1 to 7, and the main trial phase began from days 8 to 28. Digestion sampling was conducted from days 22 to 28. Body weight was measured before morning feeding at 6:00 AM daily, and feces, urine, and remaining feed intake were collected. Feed was then replaced with fresh feed, and feed intake was recorded. The trial was conducted in Zhaozhou County, Daqing City, Heilongjiang Province. Feed was provided by Harbin Aokai Huinong Feed Co., Ltd. Fattening sheep were selected from Heilongjiang Woyuan Animal Husbandry Co., Ltd.

[0073] 1.1 Index collection and detection methods

[0074] 1.1.1 Digestion test sampling and sample pretreatment process

[0075] The summer digestion experiment lasted seven days. On the first day, each sheep was weighed before morning feeding, and the initial weight was recorded. After weighing, the sheep were placed in pre-built and labeled digestion cages. A feed trough and water trough were also prepared, and small, frequent feedings were provided to allow the sheep to acclimate to the surrounding environment for a day. On the second day, 75 sheep were fitted with homemade feces and urine bags. A total feces and urine collection method was used to collect all feces and urine. Care was taken to avoid stress during the installation process, allowing the feces and urine bags to acclimate for a day. On the third day, before morning feeding, the inside and outside of each digestion cage were cleaned to prevent interference with the sampling process. The amount of pelleted feed was calculated and weighed based on 10% of each sheep's body weight, ensuring that at least 5% of the feed remained in each trough each day. The weighed feed was fed at 6:00 AM and 6:00 PM, with free access to water. The feces and urine bags were inspected for any loose or damaged contents. On the fourth day, before morning feeding, the remaining feed was collected and weighed for each sheep. Feed intake was calculated, and the total amount of feces and urine was recorded. The sheep were then fed. Weigh 500 grams of each sheep's feed and feces sample and place them in a No. 10 sealable bag. Add dilute sulfuric acid to the feces sample bag (20 mL of 10% dilute sulfuric acid per 100 g feces sample). Weigh 60 mL of each sheep's urine sample and place it in two 50 mL cryovials. Pour 3 mL of 10% dilute sulfuric acid into the cryovials (urine: dilute sulfuric acid = 10:1). After processing the above samples, place them in a -20°C refrigerator for subsequent testing. The fourth and fifth days are both adaptation stages. On the sixth day, sample according to the third day's operating procedures. On the seventh day, place the fattening sheep in the corresponding large feeding pen according to the cage number and weigh the final weight. Dismantle the digestion cage, clean the site, and disinfect.

[0076] The sampling process for the winter digestion test is the same as above.

[0077] 1.1.2 Environmental data sampling process

[0078] Environmental data is collected using an intelligent monitoring system for livestock and poultry house environments. This system comprises a sensing layer, a network layer, and an application layer. During the digestion experiment, temperature, humidity, and ammonia sensors were used daily to collect T, RH, and NH3 data. This data was transmitted to a central controller, which in turn transmitted it to a computer-based backend control system. A NOVA II intelligent gateway was also used to prevent data loss or inability to transmit to the backend system due to network lag. The sensors and central controller, sourced from Longteng Weiye Technology Co., Ltd., are both waterproofed to prevent condensation and water in the livestock and poultry house. The NOVA II intelligent gateway is manufactured by Fuhua Technology Co., Ltd. The daily collected environmental data was averaged.

[0079] 1.1.3 Sample testing method

[0080] The collected feed samples were oven-dried at 65°C for 48 hours, allowed to acclimate at room temperature for 24 hours, ground, passed through a 1 mm sieve, and placed in resealable bags. They were then stored at -20°C for subsequent determination of general nutritional composition. Fecal samples were oven-dried at 55°C for 48 hours, allowed to acclimate at room temperature for 24 hours, ground, passed through a 1 mm sieve, and placed in resealable bags. Together with urine samples, they were stored at -20°C for crude protein determination. The chemical composition of these processed samples was determined at the Animal Nutrition Laboratory of Northeast Agricultural University (Harbin, China). Dietary dry matter (DM), crude protein (CP), crude ash (ASH), crude fat (EE), calcium (Ca), and phosphorus (P) were determined according to AOAC (2006) methods. Neutral detergent fiber (NDF) and acid detergent fiber (ADF) contents were determined according to the Van Soest method. After the above determinations are completed, the indicators required for the test are calculated: neutral detergent fiber intake (NDFI), acid detergent fiber intake (ADFI), nitrogen intake (NI), fecal nitrogen (UN), urine nitrogen (FN), retained nitrogen (RN) and nitrogen utilization efficiency (NUE).

[0081] 1.2 Data Preprocessing

[0082] Three datasets were constructed using data from winter and summer: a winter dataset (n=225), a summer dataset (n=225), and a mixed dataset of both (n=450). Experimental models, including LR, EN, KNN, and ANN, were constructed for each dataset. Prior to model construction, the data was relatively scattered, so preprocessing and feature selection were performed to prevent interference with the model's predictive performance.

[0083] 1.2.1 Data Cleansing Methods

[0084] The data was cleaned to check for missing values ​​and outliers. A correlation coefficient heat map was used to determine the positive and negative correlation between each feature and NUE, and +0.5 was set as the correlation limit, that is, the absolute value of the correlation coefficient was greater than 0.5 for strong correlation, and less than 0.5 for correlation trend, which was convenient for subsequent result analysis. However, the units between the features are different, and they need to be scaled. Common methods are standardization and normalization. Normalization: Scale the data to a specific range (0,1) so that all features have the same scale and range, which helps to improve the performance of certain machine learning algorithms that are sensitive to distance metrics. Standardization: Convert the data to a distribution with a mean of 0 and a standard deviation of 1 to reduce the impact of dimensional inconsistency between features, making it easier for the algorithm to converge. During the experiment, it was found that although there was no significant difference in the scaling results of the two, the standardized model was more accurate.

[0085] 1.2.2 Feature Selection Methods

[0086] Use recursive feature elimination method to select features, that is, by building LR model, combine the features to be screened in sequence and bring them in, and compare the R in the model output results. 2 , RMSE, and feature weights are used to determine whether a feature or feature combination should be retained. Through an iterative optimization process, variable features are gradually selected. When model performance reaches optimality, meaning that further variable additions do not improve predictions, while removing variables reduces model accuracy, the final model architecture is determined. In this experiment, stepwise variable selection and direct input methods, as well as direct feature selection using the correlation coefficient method, were also tested. The stepwise variable selection method: During each iteration of model construction, the system rigorously screens candidate input variables. Variables not already included in the model are first evaluated. Variables with a significance level of P ≤ 0.1 are retained, while those with a significance level of P > 0.1 are excluded. If a variable significantly improves model predictive performance, it is immediately included in the model. Simultaneously, variables already included in the model are dynamically reviewed to ensure that removing any variable will not significantly degrade model performance, thereby ensuring a streamlined and efficient final model. The direct input method: This method utilizes a full-variable inclusion strategy. During model construction, input variables are not selectively selected, but all predictor variables are directly introduced into the regression equation. This method is the default. Correlation coefficient method: Construct a heat map of input variables and output variables, set the correlation coefficient to 0.7 as the limit, and retain a feature when the correlation coefficient with nitrogen utilization rate is greater than 0.7. The test results show that the stepwise method and the correlation coefficient method are more adaptable than the direct input method. However, when the above methods are met and the selected feature combination is brought into the model, the effect is not ideal. Some features may even increase the risk of overfitting and reduce the accuracy of model performance. The data of the three data sets were divided into a training set (80%) and a test set (20%). A 10-fold cross-validation was used, and the model was repeatedly trained 4 times: BW, NI, NDFI, ADFI and T, RH, NH3 were used as continuous input variables, and NUE was used as a continuous target variable. The three data sets were used to study the performance of the four models and the importance of predictive variables.

[0087] 1.3 Model adjustment and training

[0088] 1.3.1 Parameter Adjustment Process of Linear Regression Model

[0089] In the LR model, weights represent the contribution of each feature to the result and can also be understood as the feature coefficient or slope. Bias represents the base offset and can also be understood as the intercept. These parameters are automatically optimized by the model, finding their optimal values ​​through data training to minimize the error between the model's predicted and actual results. First, a parameter tuning method is selected, which typically involves adjusting the regularization coefficient (such as the parameters in L1 and L2 regression) and other parameters that may affect model performance. Furthermore, during the parameter tuning process, attention should be paid to overfitting and underfitting, and appropriate measures should be taken to address them. Parameter tuning methods can include manual tuning (based on experience or intuition), grid search, and random search. However, manual tuning is not authoritative and convincing, and the parameter adjustment process is time-consuming. Random search is a hyperparameter search method based on random sampling. It randomly samples a certain number of parameter combinations from the hyperparameter space. Then, by systematically training and evaluating all possible hyperparameter combinations, it ultimately determines the parameter configuration with the best predictive performance. While random search is computationally cheaper because it does not require an exhaustive search of all possible combinations, it also has some drawbacks. Since random search only randomly samples a subset of hyperparameter combinations for evaluation, it may miss some potentially better combinations and cannot guarantee finding the global optimal solution. Therefore, this experiment used grid search for parameter adjustment. This method is a method that finds the optimal solution by exhaustively enumerating all hyperparameter combinations. It systematically traverses multiple parameter combinations and determines the best performing parameters through cross-validation. However, as the number of hyperparameters and candidate values ​​increases, the search space expands, resulting in higher computational costs. However, due to its precise search, it achieves more accurate model predictions than other methods and was adopted in this experiment.

[0090] 1.3.2 Parameter Adjustment Process of the Elastic Network Model

[0091] The main parameters of the elastic net model are the regularization coefficient and the L1 regularization mixing ratio parameter. Regularization coefficient: used to control model complexity and prevent overfitting. The larger the value of the regularization coefficient, the greater the penalty for model complexity and the stronger the model's generalization ability may be. L1 regularization mixing ratio parameter: used to balance the contributions of L1 regularization and L2 regularization. The value of l1_ratio usually ranges from 0 to 1. When l1_ratio = 0, the elastic net only includes L2 regularization; when l1_ratio = 1, the elastic net only includes L1 regularization. By adjusting the value of l1_ratio, a trade-off can be made between L1 regularization and L2 regularization, thereby achieving better control over feature selection and model complexity. When using grid search to adjust parameters, it was found that this parameter approaches 1, indicating a greater bias towards L1 regularization.

[0092] 1.3.3 K-Nearest Neighbor Model Parameter Adjustment Process

[0093] An important parameter of the KNN model is the K value, which determines the number of training samples closest to the sample to be predicted in the feature space. A fixed value can be specified for K, but a smaller K value will make the model more sensitive to noise because the model will rely more on a few neighboring samples, which may lead to overfitting. A larger K value may lead to underfitting because the model will consider global information more and ignore local differences. Therefore, by setting the upper and lower limits of the parameter value range, the system can automatically optimize and select the optimal number of adjacent elements within this range. It should be pointed out that increasing the number of adjacent elements does not necessarily improve the accuracy of the model. In view of the uncertainty of the optimal number of adjacent elements, it is recommended to set the parameter search range to 2 <K<7。

[0094] 1.3.4 Parameter Adjustment Process of Artificial Neural Network Model

[0095] In the ANN model, a multilayer perceptron (MLP) is selected to construct the neural network. In the MLP, the neural network adopts a fully connected architecture, where neurons in adjacent layers transmit information through weighted connections. Data flow strictly follows a unidirectional path from the input layer through the hidden layer to the output layer, without any reverse feedback mechanism. Therefore, the number of network layers and the number of neurons in each layer of the MLP determine the complexity and capacity of the neural network. Deeper networks and more neurons can capture more complex features. Through the nonlinear activation functions in the hidden layers, complex nonlinear mappings between input and output can be learned. This gives it a significant advantage in handling nonlinear problems, but it can also lead to overfitting and training difficulties. The learning rate, a key hyperparameter, directly regulates the adjustment of network weights and biases during the gradient descent process, controlling the update step size of the model parameters. A small learning rate can slow training, while a large learning rate can lead to unstable training or even skipping the optimal solution. Since the maximum dataset in this experiment is small (n=450), we used lbfgs as the solver after careful consideration. The learning rate in the perceptron is only applicable when the solver is sgd. This avoids issues caused by adjusting the learning rate. Among common neural network loss function calculation methods, gradient descent calculates the gradient of the loss function to determine a reasonable parameter update direction, resulting in a gradual decrease in the loss function value. However, its descent rate is relatively slow, requiring a large number of iterations to converge. Furthermore, the gradient needs to be recalculated at each iteration, which can lead to unstable descent directions and oscillations. Furthermore, when the objective function is non-convex and contains multiple minima, gradient descent may become trapped in local minima and unable to continue descending, preventing the algorithm from finding the global optimal solution. In summary, this experiment set a random seed, used a grid parameter search to determine the hidden layer range, and used the activation function (identity, logistic, tanh, ReLU). All other parameters were left as default. The parameters of LR, EN, KNN, and ANN are shown in Table 2.

[0096] Table 2 Parameter sets for linear regression, elastic network, K-nearest neighbor, and artificial neural network

[0097]

[0098]

[0099] 1.4 Model Performance Analysis

[0100] The predicted values ​​of nitrogen use efficiency under different models were analyzed by RMSE and R 2Calculate and evaluate the MAE of the predictions of different models. Use Duncan analysis to perform a significance test to assess whether the differences between the models are significant. P < 0.05 indicates a significant difference. After normalizing the weights of the input variables of each model, sort them and evaluate their importance. Among them:

[0101]

[0102] N is the number of samples, Xi is the observed value, is the predicted value, is the mean of the observed values.

[0103] 2. Regarding the result analysis, the analysis of the mixed data set is as follows:

[0104] 2.1 Mixed Datasets

[0105] 2.1.1 Characteristics and growth performance data of each stage of sheep fattening period

[0106] The characteristic data of each fattening stage are shown in Table 3. BW and NUE showed significant differences among the three fattening stages, while other characteristics showed no significant differences.

[0107] Table 3 Characteristic data of each stage of fattening period

[0108]

[0109] The growth performance data of each stage of fattening are shown in Table 4. UN was higher than FN in each stage, and RN decreased from the early to the late fattening stage.

[0110] Table 4 Growth performance data at each stage of fattening period

[0111]

[0112]

[0113] 2.1.2 Model Performance

[0114] Table 5 shows the performance of the model after training on the mixed dataset and the performance of the model verified by the test set. From the performance of the trained model, we can see that the R 2 Same, both are 0.95, KNN R 2 The RMSE of the four models for NUE prediction ranged from 0.66 to 1.27 g2 / d, and the RMSE of LR was significantly (P<0.05) higher than that of KNN. The RSE of the four models for NUE prediction ranged from 0.66 to 1.27 g2 / d, and the RMSE of LR was significantly (P<0.05) higher than that of KNN. 2 All are greater than 0.9, and KNN's R 2The RMSE of the two methods was still the highest, at 0.98, and was significantly (P < 0.05) higher than that of the LR. The RMSE ranged from 0.85 to 1.35 g2 / d, and the RMSE of the LR method was significantly (P < 0.05) higher than that of the KNN method.

[0115] Table 5 Training performance and test set performance of four models

[0116]

[0117] Among them, different lowercase letters in the same data of a and b indicate significant differences (P<0.05), R 2 is the coefficient of determination, which reflects the degree of fit of the regression model to the data; RMSE, root mean square error, is the square root of the ratio of the square of the deviation between the predicted value and the true value to the number of observations n.

[0118] The predicted values ​​of nitrogen use efficiency of the four models are compared with the actual values. Figure 2 ,Compared with LR, EN and ANN, the KNN model has stronger ,prediction performance.

[0119] 2.1.3 Model evaluation and influencing factors

[0120] Depend on Figure 2 As can be seen, for LR, ADFI has the greatest influence, followed by NDFI. NI, T, RH, BW, and NH3 have relatively low impacts on NUE prediction. For EN, NDF has the greatest impact on feed intake prediction, followed by NI, while ADFI, NH3, BW, RH, and T have relatively small influences on model predictions. For the KNN model, BW is the most important, while RH, NI, NDFI, NH3, T, and ADFI have relatively low weights. For the ANN model, ADFI has the highest weight, while NDFI, BW, RH, T, and NH3 all have relatively low influences on model predictions. NI is particularly unimportant compared to the other six variables.

[0121] Depend on Figure 3 As can be seen, NUE is negatively correlated with BW and has a negative correlation trend with NDFI, ADFI, and NH3. NUE is positively correlated with RH and has a positive correlation trend with NI. BW is positively correlated with NDFI and ADFI, has a positive correlation trend with NI, and is negatively correlated with RH. NI is positively correlated with NDFI and ADFI, and has a positive correlation trend with NH3. RH has a negative correlation trend with NDFI, ADFI, and NH3. T has a positive correlation trend with all other characteristics. NH3 has a positive correlation trend with NDFI and ADFI.

[0122] Table 6 shows that the actual and predicted mean NUE values ​​were not significantly different. The MAE ranged from 0.61 to 1.12 g / d. KNN had the lowest MAE, 0.61. LR had a significantly higher MAE (P < 0.05) than KNN, at 1.12. EN and ANN showed no significant MAE differences from the other two models.

[0123] Table 6 Mean absolute error of four models

[0124]

[0125] 2.2 Summer Dataset

[0126] 2.2.1 Characteristics and growth performance data of each stage of the fattening period

[0127] The characteristic data of each fattening stage are shown in Table 7. BW, T and NUE showed significant differences among the three fattening stages. There were no significant differences in other characteristics.

[0128] Table 7 Characteristic data of each stage of fattening period

[0129]

[0130]

[0131] The growth performance data of each stage of fattening are shown in Table 8. UN was higher than FN in each stage, and RN decreased from the early to the late fattening stage.

[0132] Table 8 Growth performance data at each stage of fattening period

[0133]

[0134] 2.2.2 Model Performance

[0135] Table 9 shows the performance of the model after training on the summer dataset and the performance of the model verified by the test set. From the performance of the trained model, we can see that the R 2 The R of KNN was significantly (P<0.05) lower than that of KNN and ANN, which was 0.96. 2 The highest is 0.99, and all four models are above 0.9. The RMSE of the four models for predicting NUE ranges from 0.58 to 1.04 g2 / d, and the RMSE of LR is significantly higher (P<0.05) than that of KNN. The test set shows that the R 2 are all greater than 0.9, and the R 2 The highest was 0.98, which was significantly (P<0.05) higher than that of LR. The RMSE ranged from 0.76 to 1.10 g2 / d, and the RMSE of LR was significantly (P<0.05) higher than that of KNN and ANN.

[0136] Table 9 Training performance and test set performance of four models

[0137]

[0138] Among them, different lowercase letters in the same data of a and b indicate significant differences (P<0.05). 2 , determination coefficient, reflects the degree of fit of the regression model to the data; RMSE, root mean square error, is the square root of the ratio of the square of the deviation between the predicted value and the true value to the number of observations n. Figure 5 ,Compared with LR, EN, and KNN, the ANN model has stronger,prediction performance.

[0139] 2.2.3 Model evaluation and influencing factors

[0140] Depend on Figure 4 As can be seen, for LR, NDFI has the greatest impact, followed by ADFI. NI, NH3, BW, RH, and T have similar weights, all having relatively low impacts on NUE prediction. For EN, ADFI has the greatest impact on NUE prediction, followed by BW and NDFI. NH3, NI, RH, and T have relatively small influences on model predictions. For the KNN model, BW is the most important, while RH, NI, NDFI, NH3, T, and ADFI have lower weights. For the ANN model, NDFI has the highest weight, followed by ADFI. NI, BW, RH, T, and NH3 have similar weights and are particularly unimportant compared to the other two variables.

[0141] Depend on Figure 5 As can be seen, NUE is negatively correlated with BW, T, and NH3, with T having the lowest correlation coefficient, followed by NH3 and BW. NUE also shows a negative correlation trend with NDFI and ADFI. NUE is positively correlated with RH and has a positive correlation trend with NI. BW is positively correlated with NDFI, ADFI, and T, has a positive correlation trend with NI and NH3, and is negatively correlated with RH. NI is positively correlated with NDFI and ADFI, and has a positive correlation trend with NH3 and T. RH is negatively correlated with NDFI, ADFI, T, and NH3, and has a negative correlation trend with NI. T is positively correlated with NDFI, ADFI, and NH3.

[0142] As shown in Table 10, there was no significant difference between the actual and predicted mean NUE values. The MAE ranged from 0.69 to 0.96 g / d, with the ANN having the lowest MAE of 0.69. The LR model had a significantly higher MAE (P < 0.05) than the ANN model, at 0.96. The EN and KNN models showed no significant MAE differences from the other two models.

[0143] Table 10 Mean absolute error of four models

[0144]

[0145] 2.3 Winter Dataset

[0146] 2.3.1 Characteristics and growth performance data of each stage of the fattening period

[0147] The characteristic data of each fattening stage are shown in Table 11. BW, T and NUE showed significant differences among the three fattening stages. There were no significant differences in other characteristics.

[0148] Table 11 Characteristic data of each stage of fattening period

[0149]

[0150] The growth performance data of each stage of fattening are shown in Table 12. UN was higher than FN in each stage, and RN decreased from the early to the late fattening stage.

[0151] Table 12 Growth performance data at each stage of fattening period

[0152]

[0153] 2.3.2 Model Performance

[0154] Table 13 shows the model performance after training on the winter dataset and the model performance verified by the test set. From the performance of the trained model, we can see that the R 2 The lowest, R of 0.96, KNN and ANN 2 The highest, both are 0.99, and the R 2 The RMSE of the four models for NUE prediction ranged from 0.26 to 0.84 g2 / d, with ANN having the lowest RMSE of 0.26. The RMSE of EN was significantly (P<0.05) higher than that of KNN and ANN, at 0.84. The RSE of the four models for NUE prediction ranged from 0.26 to 0.84 g2 / d, with ANN having the lowest RMSE of 0.26. The RSE of EN was significantly (P<0.05) higher than that of KNN and ANN, at 0.84. 2 All are greater than 0.9, among which R 2 The lowest is 0.97, and the R 2 The highest was 0.99, which was significantly (P<0.05) higher than that of LR and EN. The RMSE ranged from 0.34 to 0.70 g2 / d, with the lowest RMSE of 0.34 for ANN. The RMSE of EN was significantly (P<0.05) higher than that of ANN.

[0155] Table 13 Training performance and test set performance of four models

[0156]

[0157] The predicted values ​​of nitrogen use efficiency of the four models are compared with the actual values. Figure 8 ,Compared with LR, EN, and KNN, the ANN model has stronger,prediction performance.

[0158] 2.3.3 Model evaluation and influencing factors

[0159] Depend on Figure 5 As can be seen, for LR, RH has the greatest influence, followed by the weight of NH3. NI, BW, NDFI, ADFI, and T all have a relatively low impact on NUE prediction. For EN, NI has the greatest impact on feed intake prediction, followed by T and NDFI, with equal weights, and then RH and ADFI. NH3 and BW are particularly unimportant compared to the other five variables. For the KNN model, BW is the most important, while RH, NI, NDFI, NH3, T, and ADFI have relatively low weights. For the ANN model, ADFI has the highest weight, followed by NI. RH, T, NH3, NDFI, and BW have relatively low influence on NUE prediction.

[0160] Depend on Figure 7 As can be seen, NUE is positively correlated with RH and T, with T having the highest correlation coefficient, followed by RH, and showing a positive correlation trend with NI and NH3. BW is positively correlated with NDFI and ADFI, and has a positive correlation trend with NI, but is negatively correlated with RH and T, and has a negative correlation trend with NH3. NI is positively correlated with NDFI and ADFI, and has a positive correlation trend with RH and NH3. RH is positively correlated with T and NH3, and has a negative correlation trend with NDFI and ADFI. T is negatively correlated with NDFI, has a negative correlation trend with ADFI, and has a positive correlation trend with NH3.

[0161] As shown in Table 14, there was no significant difference between the actual and predicted mean NUE values. The MAE ranged from 0.30 to 0.62 g / d, with the ANN having the lowest MAE of 0.30. The EN model had a significantly higher MAE (P < 0.05) than the ANN model, at 0.62. The LR and KNN models showed no significant MAE differences from the other two models.

[0162] Table 14 Mean absolute error of four models

[0163]

[0164] Based on this, this implementation method carried out a NUE prediction experiment for mutton sheep. In this experiment, we first used a mixed data set to build four models: LR, EN, KNN, and ANN, and calculated the weights of each input feature in the model. The results showed that NUE was related to environmental factors, and that excluding the indoor environmental parameters from the model during the feature combination selection process would reduce the accuracy and R of the four models. 2The results showed that BW, RH, NDFI, and ADFI were the most important predictors in the three datasets studied, while NI, T, and NH3 were effective in improving the performance and prediction accuracy of the model algorithm. The R values ​​of the training and validation datasets of the four models were 0. 2 The performance of KNN and ANN models in the three data sets is more outstanding. 2 All were above 0.95. In existing technology experiments, conventional information such as weight, activity level, and milk production performance was collected from various sensors on dairy cows on a pasture, and LR was used to predict lameness. The results showed that models with multiple feature variables (data from multiple sensors) consistently outperformed models with single feature variables (data from a single sensor).

[0165] In addition, compared with the above experimental results, the R of each model for each data set in this experiment is 2 They are all higher than the models in the above experiments. This may be because there are more arrays in the experimental data set, the abnormal data is properly smoothed, and the preprocessing steps of feature selection are better. In addition, in the actual livestock breeding production process, the NUE of meat sheep is usually affected by many factors, such as the cold and heat stress caused by seasonal climate, the nutritional composition of feed, the breed, and the nitrogen metabolism of meat sheep in different fattening periods, and many unexpected factors will also be encountered. The ML algorithm has a strong generalization ability and flexible and inclusive ability to handle various data. Compared with traditional algorithms, it has more advantages in prediction performance and is more suitable for actual production. However, there is still a lot of room for improvement in this prediction process. It is still necessary to further explore simpler and lighter data collection methods and improve the possibility of realizing accurate intelligent monitoring at the group level, so as to adapt to the precise regulation and feeding of large-scale sheep flocks.

[0166] Therefore, the above research results show that compared with traditional single linear models and multivariate linear models, the popular machine learning algorithms described in this embodiment have higher accuracy. At the same time, the results of this experiment show that the performance and true value prediction accuracy of popular machine learning algorithms such as KNN and ANN in the three data sets are far better than those in the above examples. This may be because the ML algorithm training model of this experiment uses a large amount of data, and the feature variable combination is appropriately selected to smooth the abnormal data exponentially and reduce the noise in the training process. Due to the large amount of data and good feature selection, it is closer to the true value, making the model more generalizable and performing well in both the training set and the test set. 2All are above 0.95. In addition, the ML method can provide better prediction results than the LR algorithm, and it is suggested that the selection of feature input combination and model algorithm will greatly affect the accuracy of the prediction results.

[0167] Those skilled in the art will understand that the above description is only a preferred embodiment of the present invention, and the features described in the various embodiments and / or claims of the present disclosure may be combined or coupled in various ways, even if such a combination or coupling is not explicitly described in the present disclosure. It is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art may still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.

[0168] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, the present invention is intended to include such changes and modifications as fall within the scope of the claims and their equivalents.

Claims

1. A method for predicting nitrogen utilization efficiency of Hu sheep during fattening period based on machine learning algorithm, characterized in that: The method comprises the following steps: Step 1: Use environmental parameter sensors to collect temperature T, humidity RH, ammonia NH3, body weight BW, nitrogen intake NI, neutral detergent fiber intake NDFI, and acid detergent fiber intake ADFI data to form winter, summer, and mixed data sets; Step 2: Establish winter data set, summer data set and mixed data set of the two based on ML algorithm to construct individual nitrogen use efficiency prediction model of Hu sheep; Step 3: Compare the growth performance of Hu sheep in different ML algorithms and use the individual nitrogen utilization efficiency prediction model of Hu sheep described in step 2 to evaluate the prediction variables, and complete the prediction of nitrogen utilization efficiency of Hu sheep in the fattening period based on machine learning algorithm.

2. The method for predicting nitrogen utilization efficiency of Hu sheep in fattening period based on machine learning algorithm according to claim 1, characterized in that: The environmental data described in step 1 is collected based on the livestock and poultry house breeding environment intelligent monitoring system equipment.

3. The method for predicting nitrogen utilization efficiency of Hu sheep in fattening period based on machine learning algorithm according to claim 1, characterized in that: The construction of the individual nitrogen utilization efficiency prediction model of Hu sheep in step 2 includes the steps of constructing a linear regression model LR, an elastic network model EN, a K-nearest neighbor model KNN, and an artificial neural network model ANN.

4. The method for predicting nitrogen utilization efficiency of Hu sheep in fattening period based on machine learning algorithm according to claim 3, characterized in that: The recursive feature elimination method is used to select features of the constructed test model LR, elastic network model EN, K-nearest neighbor model KNN, and artificial neural network model ANN. That is, by constructing the test model LR, the features to be screened are combined and introduced in sequence, and the RMSE and R values ​​in the output results of each model are compared. 2 And the numerical value of MAE.

5. The method for predicting nitrogen utilization efficiency of Hu sheep in fattening period based on machine learning algorithm according to claim 1, characterized in that: The RMSE and R 2 And the method of calculating the numerical value of MAE is: Where n is the number of samples, X i is the observed value, is the predicted value, is the mean of the observed values.

6. The method for predicting nitrogen utilization efficiency of Hu sheep in fattening period based on machine learning algorithm according to claim 1, characterized in that: The winter dataset, summer dataset, and the mixed dataset of the two described in step 2 are divided into training sets and test sets, respectively, with a training set and test set ratio of 8:

2.

7. The method for predicting nitrogen utilization efficiency of Hu sheep in fattening period based on machine learning algorithm according to claim 1, characterized in that: Step 2 also includes the steps of modeling and analyzing the winter dataset, the summer dataset, and a mixed dataset of the two.

8. A system for predicting nitrogen utilization efficiency of Hu sheep during fattening period based on machine learning algorithm, characterized in that: The system comprises: The data acquisition module is used to collect temperature T, humidity RH, ammonia NH3, body weight BW, nitrogen intake NI, neutral detergent fiber intake NDFI, and acid detergent fiber intake ADFI data using environmental parameter sensors to form a winter data set; A nitrogen utilization efficiency prediction model construction module is used to establish a winter dataset, a summer dataset, and a mixed dataset of the two based on the ML algorithm, and use the mixed dataset to construct a nitrogen utilization efficiency prediction model for individual Hu sheep; The prediction module is used to compare the individual nitrogen utilization efficiency prediction models of Hu sheep in different ML algorithms to evaluate the prediction variables, and complete the prediction of nitrogen utilization efficiency of Hu sheep during the fattening period based on machine learning algorithms.

9. A computer device comprising a memory and a processor, characterized in that A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Apparatus for use with memory means

    GB1560157A