Ruminant methane emission measurement method, system, electronic device and storage medium
By constructing a mechanism-driven and data-driven dual-model prediction system, the problems of the singleness and low data resolution of traditional ruminant methane emission monitoring methods were solved, and high-accuracy methane emission prediction was achieved.
Patent Information
- Application Number
- CN202510042269.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-01-10
AI Technical Summary
Traditional ruminant methane emission monitoring methods are single and have low data resolution, making it difficult to accurately reflect the dynamic changes and driving mechanisms of methane emissions, and the influencing factors are not fully considered.
A mechanism-driven and data-driven dual-model prediction system is constructed. By obtaining multi-source heterogeneous sample data, preprocessing and feature extraction are performed, and methods such as principal component analysis and factor analysis are used, combined with support vector regression, neural network and other algorithms, a methane emission process simulation model is constructed. Weighted average and discrete degree analysis are used to determine the final emission data.
It significantly improves the accuracy and reliability of methane emission predictions, combines the advantages of mechanistic and data-driven approaches, and overcomes the limitations of a single model.
Smart Images

Figure CN120105861B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of ruminant methane emission measurement, and in particular to a ruminant methane emission measurement method, system, electronic equipment and storage medium. Background Art
[0002] Ruminant methane emissions are a major source of agricultural greenhouse gas emissions. Accurately monitoring and quantifying ruminant methane emissions is crucial for assessing the carbon footprint of animal husbandry and developing emission reduction strategies. Traditional ruminant methane emission monitoring methods, including breathing masks, gas flow meters, and tracer methods, can obtain methane emission data. However, these methods often suffer from limitations such as limited monitoring methods, low data resolution, and incomplete consideration of influencing factors. These methods make it difficult to accurately reflect the dynamic changes and driving mechanisms of methane emissions, resulting in inaccurate monitoring data on ruminant methane emissions. Summary of the Invention
[0003] The present application provides a method, system, electronic device and storage medium for measuring methane emissions from ruminants, which are used to improve the accuracy of measuring methane emissions from ruminants.
[0004] In a first aspect, the present application provides a method for measuring methane emissions from ruminants, comprising:
[0005] Acquire multi-source heterogeneous sample monitoring data of ruminant methane emissions within a certain time span, sample data of factors affecting methane emissions, and sample data of the rumen environment, and construct a sample data set based on the multi-source heterogeneous sample monitoring data, the sample data of the factors affecting methane emissions, and the sample data of the rumen environment;
[0006] Preprocessing the sample data in the sample data set to obtain a standard sample data set, and preprocessing the standard sample data set to obtain a target sample data set;
[0007] Performing feature extraction on the sample data in the target sample data set to obtain sample feature data, and constructing a mechanism-driven first methane emission process simulation model and a data-driven second methane emission process simulation model based on the sample feature data;
[0008] Real-time collection of real-time multi-source heterogeneous emission monitoring data, real-time influencing factor data, and real-time key environmental parameter data of the ruminant to be tested, and construction of a real-time data set based on the real-time multi-source heterogeneous emission monitoring data, the real-time influencing factor data, and the real-time key environmental parameter data;
[0009] Inputting the real-time data set into the first methane emission process simulation model to obtain a first emission prediction result;
[0010] Inputting the real-time data set into the second methane emission process simulation model to obtain a second emission prediction result;
[0011] Target methane emission data is determined based on the first emission prediction result and the second emission prediction result.
[0012] In this technical solution, a sample dataset was constructed by acquiring multi-source, heterogeneous sample monitoring data on ruminant methane emissions over a specific time span, as well as sample data on factors influencing methane emissions and the rumen environment. This data laid a solid foundation for subsequent modeling. Systematic preprocessing of the sample dataset, including data cleaning and standardization, resulted in a standard sample dataset. Further, correlation analysis and other methods were used to identify target sample parameters highly correlated with methane emissions, resulting in a target sample dataset. This effectively improved the relevance and efficiency of modeling.
[0013] On this basis, this application uses methods such as principal component analysis and factor analysis to extract features from the target sample data set, and obtains sample characteristic data that comprehensively reflects the intrinsic structure and characteristics of the sample data. Using the sample characteristic data, a mechanism-driven first methane emission process simulation model and a data-driven second methane emission process simulation model are constructed. Among them, the first model starts from the conceptual model, introduces a kinetic model and is optimized through parameters, fully considering the biological mechanism of methane production and emission, and has the advantage of strong interpretability; the second model comprehensively uses data mining algorithms such as support vector regression and neural networks, which can accurately depict the complex nonlinear relationship between methane emissions and various influencing factors, reflecting the advantages of data-driven. The two types of models complement each other and comprehensively simulate the methane emission process from both the mechanism and data perspectives.
[0014] During the application phase, this application collects multi-source heterogeneous emission monitoring data, influencing factor data, and key environmental parameter data from the ruminants to be tested in real time, constructs a real-time data set, and inputs it into the first and second methane emission process simulation models, respectively, to obtain two sets of prediction results. Then, using weighted average and discreteness analysis methods, the two sets of prediction results are combined to determine the final target methane emission data. This strategy based on dual-model prediction and multi-result verification effectively combines the advantages of mechanistic and data-driven approaches, overcomes the limitations of a single model, and can significantly improve the accuracy and reliability of methane emission predictions.
[0015] In a second aspect of the present application, a ruminant methane emission measurement system is provided, the system comprising:
[0016] a sample data acquisition module for acquiring multi-source heterogeneous sample monitoring data of ruminant methane emissions, sample data of influencing factors affecting methane emissions, and sample data of the rumen environment within a certain time span, and constructing a sample data set based on the multi-source heterogeneous sample monitoring data, the sample data of influencing factors, and the sample data of the rumen environment;
[0017] a data processing module, configured to preprocess the sample data in the sample data set to obtain a standard sample data set, and preprocess the standard sample data set to obtain a target sample data set;
[0018] a modeling module, configured to extract features from the sample data in the target sample data set to obtain sample feature data, and construct a mechanism-driven first methane emission process simulation model and a data-driven second methane emission process simulation model based on the sample feature data;
[0019] A real-time data acquisition module is used to collect real-time multi-source heterogeneous emission monitoring data, real-time influencing factor data and real-time key environmental parameter data of the ruminant to be tested, and to construct a real-time data set based on the real-time multi-source heterogeneous emission monitoring data, the real-time influencing factor data and the real-time key environmental parameter data;
[0020] a first emission prediction module, configured to input the real-time data set into the first methane emission process simulation model to obtain a first emission amount prediction result;
[0021] a second emission prediction module, configured to input the real-time data set into the second methane emission process simulation model to obtain a second emission prediction result;
[0022] The third emission prediction module is used to determine target methane emission data based on the first emission prediction result and the second emission prediction result.
[0023] In a third aspect of the present application, a computer storage medium is provided. The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the above method steps.
[0024] In the fourth aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs the above method.
[0025] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0026] 1. This application constructs a sample dataset based on multi-source heterogeneous sample monitoring data of ruminant methane emissions over a certain time span, sample data on factors affecting methane emissions, and sample data on the rumen environment. This data provides a solid data foundation for subsequent modeling. Systematic preprocessing of the sample dataset, including data cleaning and standardization, yields a standard sample dataset. Further, through correlation analysis and other methods, target sample parameters highly correlated with methane emissions are screened to obtain a target sample dataset, effectively improving the pertinence and efficiency of modeling.
[0027] 2. This application uses methods such as principal component analysis and factor analysis to extract features from the target sample data set, and obtains sample characteristic data that comprehensively reflects the intrinsic structure and characteristics of the sample data. Using the sample characteristic data, a mechanism-driven first methane emission process simulation model and a data-driven second methane emission process simulation model are constructed. Among them, the first model starts from the conceptual model, introduces a kinetic model and is optimized through parameters, fully considering the biological mechanism of methane production and emission, and has the advantage of strong interpretability; the second model comprehensively uses data mining algorithms such as support vector regression and neural networks, which can accurately depict the complex nonlinear relationship between methane emissions and various influencing factors, reflecting the advantages of data-driven. The two types of models complement each other and comprehensively simulate the methane emission process from both the mechanism and data perspectives.
[0028] 3. This application constructs a real-time data set by collecting multi-source heterogeneous emission monitoring data, influencing factor data, and key environmental parameter data from the ruminants to be tested in real time. The data is then input into the first and second methane emission process simulation models to obtain two sets of prediction results. The weighted average and dispersion analysis methods are then used to combine the two sets of prediction results to determine the final target methane emission data. This strategy based on dual-model prediction and multi-result verification effectively combines the advantages of mechanistic and data-driven approaches, overcomes the limitations of a single model, and can significantly improve the accuracy and reliability of methane emission predictions. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 A schematic diagram of a process for measuring methane emissions from ruminants provided in an embodiment of the present application;
[0030] Figure 2 This is an architecture diagram of a ruminant methane emission measurement system provided in an embodiment of the present application;
[0031] Figure 3 This is a schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION
[0032] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0033] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.
[0034] In the description of the embodiments of the present application, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0035] In order to facilitate understanding of the method and system provided by the embodiments of the present application, before introducing the embodiments of the present application, the background of the embodiments of the present application is first introduced.
[0036] Ruminants, a collective term for herbivorous livestock such as cattle and sheep, emit methane, a significant component of agricultural greenhouse gas emissions. According to statistics, ruminant methane emissions account for nearly one-third of global methane emissions, making them the second-largest source of agricultural methane emissions after rice cultivation. Therefore, accurately monitoring and quantifying ruminant methane emissions is crucial for assessing the carbon footprint of the livestock industry and developing greenhouse gas reduction strategies.
[0037] For a long time, scientific researchers have conducted extensive research on monitoring methods for methane emissions from ruminants. Traditional monitoring methods mainly include the breathing mask method, the gas flow meter method, and the tracer method. Among them, the breathing mask method collects the exhaled gas over a certain period of time by covering the animal's head with a breathing mask, and measures the methane concentration and ventilation volume to calculate the total amount of methane emissions. The gas flow meter method uses a device such as a sac or air chamber to collect intestinal gas discharged by the animal over a period of time, and determines the methane emission rate by measuring the gas flow rate and methane concentration. The tracer method is to estimate the amount of methane produced by the animal by injecting a specific tracer substance (such as sulfur hexafluoride SF6) into the animal's body and using the ratio of the tracer substance to methane.
[0038] The above-mentioned traditional methods have played an important role in the field of methane emission monitoring in ruminants, but there are still some problems that need to be solved. First, most of these methods use a single monitoring method, which makes it difficult to comprehensively and systematically evaluate the level of methane emissions. For example, the breathing mask method can only obtain intermittent methane concentration data, and the gas flow meter method can only measure the amount of methane emitted from the intestine. Both cannot take into account the rumen respiration and intestinal emission processes. Secondly, the methane emission data obtained by the existing methods have a low temporal resolution, and usually only daily or weekly average emissions can be obtained, which makes it difficult to reflect the dynamic fluctuation characteristics of methane emissions. Third, most studies only focus on methane emissions themselves, and rarely consider the combined effects of multiple sources such as feed, genetics, and environment, and lack an explanation of the emission mechanism.
[0039] After the background introduction of the above content, those skilled in the art can understand the problems existing in the prior art. The technical solutions in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments.
[0040] On the basis of the above background technology, further, please refer to Figure 1 , Figure 1 This is a flow chart of a method for measuring methane emissions from ruminants provided in an embodiment of the present application. The system can be implemented by a computer program or can be run as an independent tool application. Specifically, in the embodiment of the present application, the method can be applied on a server, but can also be applied to electronic devices such as servers. A method for measuring methane emissions from ruminants includes the following steps:
[0041] S101, acquiring multi-source heterogeneous sample monitoring data of ruminant methane emissions within a certain time span, sample data of influencing factors affecting methane emissions, and sample data of the rumen environment, and constructing a sample data set based on the multi-source heterogeneous sample monitoring data, the sample data of influencing factors, and the sample data of the rumen environment;
[0042] Specifically, the first step of this application is to obtain multi-source, heterogeneous sample monitoring data on ruminant methane emissions over a certain time span, sample data on factors affecting methane emissions, and sample data on the rumen environment. Based on these three types of data, a sample dataset is constructed. This step aims to comprehensively collect historical data reflecting the patterns and influencing mechanisms of methane emissions, laying the data foundation for subsequent modeling and analysis.
[0043] In practice, this application utilizes a self-developed multi-parameter "electronic rumen pill" sensor array for data collection. This "electronic rumen pill" integrates a VFA sensor array, a methane sensor, and a rumen microbiome sensor, enabling simultaneous monitoring of dynamic changes in rumen VFA concentrations, CH4 concentrations, and key microbial populations. The VFA and CH4 concentration sensing units utilize highly sensitive, low-pH-tolerant electrochemical sensors. The rumen microbiome sensor utilizes specific recognition elements such as aptamers and antibodies to quantitatively analyze key bacterial groups, including methane-producing and propionic acid-utilizing bacteria. These sensor units are packaged into an orally ingestible capsule that can be administered to experimental animals for in vivo monitoring.
[0044] To accurately assess methane emissions from each animal, the sensor data generated by the "electronic rumen pill" must be precisely linked to the animal's individual information. To this end, each test animal is equipped with a unique electronic identification tag (such as an RFID ear tag). The animal's electronic tag is then matched with the data collected by its "electronic rumen pill" to ensure that the monitoring data can be traced back to the specific animal.
[0045] During the 12-month monitoring trial, an "electronic rumen pill" automatically collected daily data on each cow's rumen environment, providing high-resolution (hourly / daily) data on methane emissions. Simultaneously, daily feed records and regular body condition scoring were used to collect data on factors influencing methane emissions, such as diet composition, feed intake, body weight, and lactation stage. The trial covered the entire lactation and peripartum period, as well as the four seasons of spring, summer, autumn, and winter, striving to obtain representative sample monitoring data.
[0046] Through the above-mentioned experimental scheme, this application has obtained massive, multi-dimensional data on ruminant methane emissions. The data generated by the "electronic rumen pill" sensor array faithfully reproduces the dynamic process of methane synthesis in the rumen and can be correlated with changes in VFA metabolism and microbial composition, providing unprecedented data support for a deeper understanding of the mechanisms of methane production. Using data integration algorithms, raw data from diverse sources and structures are integrated into a unified database, forming a structured and easy-to-analyze sample dataset. This dataset covers all the parameters required for research and identifies multiple dimensions such as individual animals, time, and physiological status, making it ideal for big data mining and modeling analysis.
[0047] Based on the above embodiment, as an optional embodiment, preprocessing the sample data in the sample data set to obtain a standard sample data set includes:
[0048] Performing data cleaning on the sample data in the sample data set to obtain cleaned sample data;
[0049] The cleaned sample data is subjected to data standardization processing to obtain the standard sample data set.
[0050] Specifically, preprocessing the sample data in the sample dataset involves two sub-steps: first, cleaning the sample data to obtain cleaned sample data; then, standardizing the cleaned sample data to obtain a standardized sample dataset. These two sub-steps aim to improve the quality of the original sample data, eliminate potential noise and outliers, and unify data from different sources and dimensions to the same scale, laying a solid foundation for subsequent feature extraction and model training.
[0051] In the first sub-step, the system performs data cleaning on the sample dataset. Because raw sample data often contains missing values and outliers, direct use can affect the accuracy of subsequent analysis. Therefore, data cleansing is necessary. First, the system scans each data item in the sample dataset, identifies missing values, and selects an appropriate filling strategy based on the type of missing value and sample characteristics. For numeric missing values, interpolation is performed using the mean or median of similar samples. For categorical missing values, the mode is used based on the attribute's frequency distribution. Next, the system uses statistical methods such as boxplots and the 3σ principle to identify outliers in the sample data and further determine their causes. If the anomaly stems from a system error such as a sensor failure, the outlier is considered invalid and removed. If the anomaly reflects a rare extreme event, the outlier is retained and marked. Furthermore, the system checks the consistency and completeness of the sample data, identifying issues such as incorrect date formats and inconsistent numerical units, and performs appropriate format conversion and corrections. After this series of data cleaning processes, the system obtains cleaned sample data with higher quality and clearer characteristics.
[0052] In the second sub-step, the system normalizes the cleaned sample data. Because different attribute values in a sample dataset often have different dimensions and ranges, for example, weight in kg may have values in the hundreds, while methane concentration in ppm may have values between tens and hundreds. Using raw values directly for analysis can easily lead to an imbalance in the importance of different attributes, impacting subsequent model learning. To eliminate the impact of this dimensional inconsistency, attribute values need to be normalized. Common normalization methods include min-max normalization and Z-score normalization. Min-max normalization linearly maps raw values to the interval [0, 1], achieving uniform scaling of different attributes. Specifically, it first determines the maximum and minimum values of a particular attribute. Then, for each sample, the minimum value is subtracted from the value of that attribute, and the value is divided by the difference between the maximum and minimum values. This converts the raw values into relative values between 0 and 1. Z-score normalization uses the mean and standard deviation of the raw values to convert them into values that follow a standard normal distribution with a mean of 0 and a variance of 1. Specifically, the mean and standard deviation of a particular attribute are first calculated. The mean is then subtracted from the value of each sample on that attribute, and the resultant value is divided by the standard deviation. This converts the original value into a dimensionless indicator representing relative position. By standardizing all attributes sequentially, a standardized sample dataset with a uniform scale is obtained.
[0053] S102, preprocessing the sample data in the sample data set to obtain a standard sample data set, and preprocessing the standard sample data set to obtain a target sample data set;
[0054] Specifically, the data in the sample dataset was preprocessed to form a standard sample dataset. Because the raw data collected by the "electronic rumen pill" sensor array suffers from signal instability and baseline drift, it requires denoising and correction. Wavelet transforms were used to perform multi-scale decomposition of the sensor data, extracting different frequency components. An adaptive threshold filtering algorithm was designed to remove high-frequency noise. Furthermore, a Kalman filter algorithm was used to adaptively track and correct baseline drift in the signal, ensuring that the data accurately reflects changes in the rumen environment. Furthermore, for individual indicators with a high number of missing values, nearest neighbor interpolation was used to repair the data. For sensor data with significant outliers, distorted and invalid data was removed by cross-comparison with data from other animals collected during the same period. This processing significantly improved the integrity and reliability of the sample data.
[0055] Secondly, based on the standard sample data set, further data preprocessing was carried out to ultimately form a target sample data set for modeling and analysis. To achieve consistency and comparability of data from different sources, the Z-score standardization method was used to dimensionlessly process the data of each indicator, converting it to a standard normal distribution with a mean of 0 and a variance of 1. To reveal the inherent connections of the data, dimensionality reduction algorithms such as principal component analysis (PCA) were used to transform and compress high-dimensional sensor data, extract the main characteristic patterns of the data, and reduce redundant information. For classification indicators such as feed composition and lactation stage, one-hot encoding was used to convert qualitative descriptions into quantitative values, which were convenient for incorporating into mathematical models for quantitative analysis. At the same time, the sample data was divided into different categories such as high emissions and low emissions according to the methane emission intensity, and stratified sampling was used to select balanced training and test sets to prepare for the subsequent construction of machine learning classification models.
[0056] Through the two stages of data preprocessing described above, a standardized, normalized, and structurally balanced target sample dataset is ultimately obtained. Compared to the original sample dataset, the target sample dataset has the following advantages: ① The data is accurate and reliable, noise is effectively suppressed, and fidelity is significantly improved; ② The data is consistent and comparable, with data from different sources and scales unified and standardized; ③ The data information is concentrated, and the essential characteristics of the data are condensed through feature extraction and dimensionality reduction compression; ④ The data structure is optimized, and the sample categories are evenly distributed, facilitating the training of high-precision learning models. A high-quality target sample dataset is a prerequisite for data mining and knowledge discovery, and is crucial for fully realizing the value of big data and improving the efficiency and reliability of modeling and analysis.
[0057] Based on the above embodiment, as an optional embodiment, preprocessing the standard sample dataset to obtain the target sample dataset includes:
[0058] Analyzing the correlation between the various sample parameters in the standard sample data set using a correlation analysis method to obtain correlation values between the various sample parameters, and screening out target sample parameters that are highly correlated with methane emissions based on the correlation values and a preset correlation threshold;
[0059] The target sample data set is constructed based on the target sample parameters.
[0060] Specifically, preprocessing the standard sample dataset involves two sub-steps: first, using correlation analysis to analyze the correlations between the various sample parameters in the standard sample dataset, obtaining correlation values between the sample parameters. Based on these correlation values and a preset correlation threshold, target sample parameters highly correlated with methane emissions are screened; then, a target sample dataset is constructed based on the screened target sample parameters. The core purpose of these two sub-steps is to identify key factors that significantly influence methane emissions from the numerous sample parameters. Based on these factors, the feature space of the sample dataset is optimized, the data dimension is reduced, and a concise and informative dataset is provided for subsequent modeling and analysis.
[0061] In the first sub-step, the system uses correlation analysis to examine the correlation between each parameter in the standard sample data set and methane emissions. Correlation analysis is a statistical method for studying the degree of linear correlation between variables. Commonly used indicators include the Pearson correlation coefficient and the Spearman rank correlation coefficient. Among them, the Pearson correlation coefficient measures the strength of the linear correlation between two continuous variables, with values ranging from -1 to 1. The larger the absolute value, the stronger the linear relationship. Specifically, the system first calculates the Pearson correlation coefficient between each sample parameter and the target parameter of methane emissions, and then takes the absolute value as the correlation value. Next, the system compares the calculated correlation value with the preset correlation threshold, and screens out sample parameters with correlation values greater than the threshold. It is considered that they are highly correlated with methane emissions and are key factors reflecting the characteristics of methane emissions. The preset correlation threshold is usually 0.5 or 0.6, indicating a correlation of medium or above. For example, analysis revealed that correlations between methane emissions and parameters such as milk production, the proportion of roughage in the diet, and rumen pH were all greater than 0.6, indicating that these are important physiological indicators and feeding factors influencing methane emissions in dairy cows and warrant special attention. However, correlations between parameters such as lactation days and ambient temperature and emissions were less than 0.3, indicating that these factors have limited impact on methane emissions and can be considered for exclusion to reduce the complexity of subsequent analysis. Through correlation analysis and threshold screening, approximately ten target sample parameters with high correlations with methane emissions were identified from dozens of sample parameters.
[0062] In the second sub-step, the system constructs a target sample data set based on the screened target sample parameters. Specifically, it extracts the data columns corresponding to the target parameters in the standard sample data set to form a new sample matrix. Each row represents a sample, each column corresponds to a target parameter, and the element value of the matrix is the value of the corresponding sample on the parameter. By extracting a subset of target parameters, the system optimizes the sample feature space, effectively reduces the dimension of the sample data, and reduces the computational overhead of subsequent analysis. At the same time, since a large number of redundant and irrelevant parameters have been screened out, the target sample data set is more focused on the key factors that have a decisive effect on methane emissions, and the signal-to-noise ratio and characterization capabilities of the data are greatly improved. Taking the cow methane emission monitoring data as an example, the parameter dimensions of the original sample may be as high as dozens to hundreds, which not only puts pressure on real-time data transmission and storage, but also brings difficulties to model training and analysis. Through correlation analysis, it was found that 10 key parameters such as milk production, roughage ratio, rumen pH, VFA concentration, and methane-producing bacteria abundance had the highest correlation with methane emissions. Therefore, the target sample data set was constructed using these 10 parameters, which not only retained the main information of the original data but also achieved nearly 10 times data compression, saving a lot of time and computing power for subsequent modeling and analysis.
[0063] S103, performing feature extraction on the sample data in the target sample data set to obtain sample feature data, and constructing a mechanism-driven first methane emission process simulation model and a data-driven second methane emission process simulation model based on the sample feature data;
[0064] Feature engineering methods were used to extract sample feature data from the target sample dataset. Based on theories of animal nutrition and microbiology, combined with statistical analysis, characteristic parameters closely related to methane emissions were screened. For example, by analyzing the Pearson correlation coefficient between diet composition and methane emission intensity, fiber content was found to be a key feed factor influencing methane emissions, and this was subsequently incorporated into the model as a key characteristic parameter. Similarly, using Spearman rank correlation to analyze the relationship between VFA composition and methane concentration, the acetic acid / propionic acid ratio was found to be significantly positively correlated with methane biosynthesis capacity, and was therefore used as a characteristic parameter to characterize rumen fermentation type. Furthermore, the maximum information coefficient (MIC) algorithm was used to analyze the correlation between multi-parameter sensor data from the "electronic rumen pill" and methane emissions, screening key characteristic parameters such as the abundance of methane-producing bacteria and rumen pH. Through feature extraction, dozens of sample characteristic parameters closely related to methane emission processes were selected from thousands of monitored indicators while minimizing information loss. This significantly reduced the data dimensionality, allowing subsequent modeling to focus on key influencing factors and improving analysis efficiency.
[0065] Secondly, based on the extracted sample characteristic parameters, a mechanism-driven model capable of simulating the methane emission process (the first methane emission process simulation model) is constructed. Mechanistic models leverage existing biological knowledge, mathematically describing the key mechanisms of methane production and quantitatively explaining the functional relationship between characteristic parameters and methane emission rates. For example, by constructing a VFA metabolic flux balance equation to quantitatively describe the synthesis and decomposition of acetic and propionic acids, and coupling processes such as hydrogen partial pressure and microbial growth, a system of differential equations reflecting the methane production mechanism is ultimately established. By assigning sample parameters such as animal species, diet composition, pH, and VFA concentration, obtained through feature extraction, to the state variables in the equations, the methane synthesis rate under specific conditions can be simulated. Mechanistic models have clear biological significance, and their parameters have clear physical interpretations, making them useful for guiding the regulation of feeding processes. However, their applicability is often limited by the key processes under consideration, and some parameters are difficult to measure directly, resulting in certain limitations in their application.
[0066] Third, based on sample characteristic data, a data-driven modeling paradigm was used to construct an empirical model of the methane emission process (the second methane emission process simulation model). The data-driven model fully exploits the statistical laws and association patterns contained in the massive sample data and, through data mining techniques such as machine learning, constructs an empirical relationship between characteristic parameters and emission rates. For example, using the support vector machine (SVM) algorithm, with characteristic parameters such as diet composition, VFA composition, and microbial abundance as input, high-emission and low-emission sample data were trained for binary classification, resulting in a classifier model that discriminates sample emission intensity. In another example, using the random forest algorithm, with multidimensional characteristic parameters such as body weight, lactation stage, and rumen pH as input, multiple decision trees were trained through an ensemble learning strategy to form a second methane emission process simulation model to predict the hourly methane emission rate of animals. The data-driven model does not rely on mechanistic assumptions. By learning a small number of physical mechanism constraints, the model parameters are mainly trained from sample data, making it highly applicable. However, its generalization ability is often limited by the representativeness of the training samples. For "small sample" problems where the training set does not cover enough, the prediction effect may be poor.
[0067] Based on the above embodiment, as an optional embodiment, the feature extraction of the sample data in the target sample data set to obtain the sample feature data includes:
[0068] Using principal component analysis to perform feature dimensionality reduction on the target sample data set to extract the main features of the sample data;
[0069] Using factor analysis method to extract features of the target sample data set to obtain common factors reflecting the intrinsic structure of the sample data;
[0070] The sample feature data is constructed based on the main features and the common factors.
[0071] Specifically, feature extraction of sample data from a target sample dataset involves three sub-steps: first, principal component analysis is used to reduce the dimensionality of the target sample dataset and extract the primary features of the sample data; second, factor analysis is used to extract features from the target sample dataset to obtain common factors that reflect the inherent structure of the sample data; and finally, sample feature data is constructed based on the extracted primary features and common factors. The core purpose of these three sub-steps is to further explore the inherent characteristics and implicit structure of the sample data, while reducing data dimensionality and redundancy while maximally preserving the original information of the data and revealing the interrelationships between variables, thereby providing a more refined and high-level feature representation for subsequent emission process modeling and optimization of emission reduction measures.
[0072] In the first substep, the system uses principal component analysis (PCA) to reduce the dimensionality of the target sample dataset. PCA is a commonly used unsupervised linear dimensionality reduction method that transforms the original high-dimensional features into a set of linearly independent new features, known as principal components, through an orthogonal transformation. The first principal component is the direction with the highest variance in the original data and represents the most dominant characteristic of the data. Subsequent principal components are orthogonal to it and have successively decreasing variances. By selecting the first few principal components, the key information of the data can be summarized at a lower dimension. Specifically, the system first centers the target sample data, then calculates the sample covariance matrix and performs eigenvalue decomposition on it, resulting in an eigenvector matrix arranged in descending order of eigenvalues. Each eigenvector represents the direction of a principal component, and the size of the corresponding eigenvalue represents the variance contribution of that principal component. The system selects the first k eigenvectors whose summed variance contribution exceeds 80%, maps the original n-dimensional sample data into the space spanned by these k eigenvectors, and obtains a reduced k-dimensional principal component matrix, with each column representing a principal component. For example, principal component analysis (PCA) of ten methane emission-related parameters for dairy cows revealed that the sum of the variance contributions of the first three principal components exceeded 85%. Therefore, each sample in the original 10-dimensional space was projected onto these three principal component directions, resulting in a 3-row, n-column principal component matrix. Each sample is represented by a 3-dimensional vector, significantly reducing data storage and computational overhead. PCA not only reduces the dimensionality of the data but also reveals the correlation structure between the original variables, providing insights into the underlying mechanisms of methane emissions.
[0073] In the second sub-step, the system uses factor analysis to extract features from the target sample dataset. Like principal component analysis, factor analysis is a dimensionality reduction technique, but it focuses more on uncovering the hidden factors underlying variables. It assumes that the observed variables are linear combinations of a few potential common factors and independent factors. By exploring the covariance structure of the variables, it estimates the common factors, thereby revealing the conceptual structure underlying the variables. Specifically, the system first tests the factorability of the target sample data based on the sample correlation coefficient matrix, preliminarily determining the number of common factors based on the number of eigenvalues greater than 1. It then constructs a factor analysis model and estimates the factor loading matrix and unique variance using the maximum likelihood method or principal component method. Factor rotation simplifies the loading matrix structure to obtain interpretable common factors, each of which represents a combination of highly correlated variables. Finally, the sample scores on each common factor are used as new composite variables to characterize the sample's intrinsic characteristics. Using methane emission data from dairy cows as an example, factor analysis extracted three common factors, interpreted as "diet composition factor," "rumen fermentation factor," and "individual difference factor," indicating that methane emissions from dairy cows are primarily influenced by these three underlying factors. Using each sample's score on these three factors as new features not only reduces the data dimensionality but also provides a high-level abstraction of emission behavior from a mechanistic perspective. Factor analysis not only simplifies the data but also extracts essential attributes from a complex set of observed variables, helping to clarify the underlying mechanisms of methane emissions and providing theoretical guidance for emission reduction practices.
[0074] In the third sub-step, the system combines the principal features extracted by principal component analysis and the common factors obtained by factor analysis into sample feature data. The principal features characterize the projection and distribution of the sample in a specific mathematical space, while the common factors reflect the sample's conceptual structure and generative mechanism. These two represent the intrinsic characteristics of the sample from different perspectives and are highly complementary. The system concatenates each sample's k-dimensional principal component vector and m-dimensional factor score vector into a k+m-dimensional feature vector. The feature vectors for all samples form a sample feature matrix, with each column corresponding to a principal feature or common factor. This high-order feature representation comprehensively captures the mathematical properties and conceptual connotations of the sample, preserving the essence of the original information while revealing underlying patterns in the data. This provides a solid feature space foundation for the subsequent construction of a high-precision methane emission simulation model. For example, in dairy cow data, each sample is represented by three principal components and three factor scores, totaling a six-dimensional feature vector. These features are well-defined and complementary, significantly reducing the data dimension while significantly enhancing representational power.
[0075] Based on the above embodiment, as an optional embodiment, the constructing of a mechanism-driven first methane emission process simulation model based on the sample characteristic data includes:
[0076] S201, constructing a conceptual model of methane generation in the rumen based on the sample characteristic data, and introducing a kinetic model into the conceptual model of methane generation in the rumen to obtain a preliminary methane emission process simulation model;
[0077] Specifically, constructing a mechanism-driven first methane emission process simulation model based on sample characteristic data involves the following steps: first, constructing a conceptual model of rumen methane production based on the sample characteristic data; then, introducing a kinetic model into the conceptual model; and finally, obtaining a preliminary methane emission process simulation model. The core purpose of this step is to utilize the refined sample characteristics extracted earlier, combined with existing knowledge of the biological mechanisms of methane production, to construct a conceptual model that can depict the entire dynamic process of rumen methane production and emission. The model is then quantified by introducing kinetic equations, resulting in a preliminary process simulation model based on causal mechanisms to describe the dynamic characteristics of methane emissions from ruminants under specific physiological conditions and feeding conditions.
[0078] In implementation, the system first constructs a conceptual model of ruminal methane production based on sample characteristic data. Using methods such as factor analysis to extract potential influencing factors such as diet composition, ruminal fermentation, and individual differences from the sample data, and using methods such as principal component analysis to preserve the key distributed characteristics of the sample data, the system establishes a conceptual framework for ruminal methane production in ruminants. First, based on the extracted common factors, the system identifies the core component modules of the model. For example, factors such as diet composition and ruminal fermentation are abstracted into metabolic substrate input modules and microbial population modules. Second, using the variable correlations revealed by principal component analysis, the system characterizes the interactive relationships between modules and state variables, such as the substrate-substrate relationship between diet components and microbial populations. During the conceptual modeling process, the system draws heavily on existing knowledge of the biological mechanisms of methane production and employs symbolic representations (such as flow charts and relationship diagrams) to intuitively express the biological functions and causal relationships of each module. This results in a preliminary conceptual model of ruminal methane production with clear conceptual semantics, boundary conditions, and structural relationships.
[0079] Building on the conceptual model, the system further incorporates kinetic equations to quantitatively characterize the methane production and emission processes. By introducing differential equations representing material flow, energy conversion, and kinetic equilibrium between functional modules and state variables, supplemented by necessary algebraic equations, a comprehensive mathematical model for quantitatively describing the dynamics of methane production is established. For example, the system incorporates the Monod equation to describe the effect of substrate concentration on microbial growth, stoichiometric constraints to characterize energy conservation and metabolic balance, and logarithmic growth equations to describe microbial population dynamics. By embedding biological reaction kinetic equations between the various components of the conceptual model and supplementing them with appropriate initial and boundary conditions, a preliminary methane emission simulation model based on ODEs (ordinary differential equations) or DAEs (differential algebraic equations) is established, achieving a mathematical representation of the conceptual model. The differential equations in the model describe the instantaneous rates of change of each state variable, characterizing the dynamic evolution of the biological process; the algebraic equations characterize the constraints between different variables, reflecting the underlying mechanisms of the biological process. The kinetic model accurately describes the evolution of state variables on time and space scales, allowing the biological process of methane generation to be described in a strictly quantitative form.
[0080] The combination of mechanism-based conceptual modeling and quantitative modeling fully leverages the key influencing factors and inherent correlation structure embedded in the sample feature data, providing a robust data-driven and biological foundation for the constructed preliminary methane emission process simulation model. The conceptual model systematically represents the methane production process from an intuitive and semantic perspective, elucidating the mechanisms and interactions of key influencing factors. This model demonstrates its comprehensive understanding of the biological process of methane production and offers strong interpretability. Furthermore, the kinetic model introduced on this basis achieves a precise mapping of conceptual semantics to mathematical logic. Through a set of differential equations, it characterizes the factors influencing the methane production rate and its dynamics, endowing the conceptual model with quantitative predictive capabilities. The advantage of this mechanism-driven model lies in its ability to be extrapolated to new operating conditions, based on an understanding of the underlying biological processes. Its predictions are universal. Data-driven feature learning plays a key role in this process, automatically extracting condensed knowledge from massive sample data, providing data support and inspiration for the construction of the mechanistic model, embodying the organic integration of data intelligence and expert knowledge.
[0081] S202 , performing fitting optimization on key parameters of the kinetic differential equation group in the preliminary methane emission process simulation model using the sample characteristic data to obtain the first methane emission process simulation model.
[0082] Specifically, after obtaining a preliminary methane emission simulation model, the key parameters of the kinetic differential equations within the preliminary model need to be fitted and optimized using sample characteristic data, ultimately obtaining the first methane emission simulation model. The core purpose of this step is to fully leverage the inherent laws inherent in the sample data to calibrate and correct the key kinetic parameters within the preliminary model, thereby improving the model's accuracy in depicting the methane emission process and enabling it to more accurately predict and simulate the dynamics of methane production under different physiological conditions and feeding regimes.
[0083] During implementation, the system first determines key parameters based on the kinetic equations of the preliminary simulation model. Since the preliminary simulation model constructed in the previous step describes the kinetics of biological processes based on differential and algebraic equations, it necessarily includes a series of biokinetic parameters, such as the maximum reaction rate, half-saturation constant, and stoichiometric coefficient. These parameters quantitatively describe the intrinsic kinetic characteristics of the methane production process and determine the model's simulation accuracy. However, since these parameters are difficult to measure directly experimentally, they are usually solved by fitting sample observation data. Therefore, the system first sorts out the differential equations in the model, identifies the key kinetic parameters, and assigns them a reasonable initial value range. Generally speaking, parameters such as the maximum specific growth rate and half-saturation constant of the Monod equation and the specific growth rate of the logarithmic growth equation have a significant impact on the dynamics of methane production and should be the focus of optimization.
[0084] After determining the parameters to be optimized, the system performs parameter fitting based on the sample characteristic data. A large amount of high-quality sample characteristic data has been obtained through feature extraction, reflecting the actual methane emissions under different conditions and serving as a guide for parameter calibration. Optimization algorithms such as least-squares are typically used to construct a parameter solution model, with the goal of minimizing the sum of squared residuals between the model simulations and the sample observations. The model simulations are obtained by numerically solving the kinetic equations (e.g., the Runge-Kutta method), while the sample observations are derived from real measurements within the characteristic dataset. By searching within a reasonable parameter space, a set of parameter values is found that minimizes the objective function (i.e., the sum of squared residuals), thus obtaining the optimized kinetic parameters. To account for the potential variability in representativeness of different samples, sample weighting can be applied to improve the fitting. Furthermore, methods such as cross-validation can be used to prevent parameter overfitting and enhance model generalization. Multiple rounds of iterative optimization are typically used to continuously adjust and update key parameters until the simulations and observations agree well.
[0085] Through sample data-driven parameter fitting optimization, the descriptive power and predictive accuracy of the preliminary model can be significantly improved. The preliminary model, which originally only provided a rough depiction of the methane generation process, can now accurately capture subtle changes in emission dynamics under different conditions after parameter updates. The fit and reproducibility of sample characteristics are greatly improved, indicating that the model has a more accurate grasp of the inherent laws of methane emissions. At this point, although the differential equations and boundary conditions in the model have not changed in form, the embedded key kinetic parameters have accurately matched the inherent laws reflected in the sample data, giving it greater explanatory and predictive power.
[0086] Based on the above embodiment, as an optional embodiment, the step of constructing a data-driven second methane emission process simulation model based on the sample characteristic data includes:
[0087] S301, based on the sample characteristic data, using a support vector regression algorithm to construct a nonlinear regression model between methane emissions and various influencing factors;
[0088] Specifically, constructing a data-driven second methane emission simulation model based on sample characteristic data involves the following steps: Using a support vector regression algorithm, a nonlinear regression model is constructed between methane emissions and various influencing factors based on the sample characteristic data. The core purpose of this step is to fully explore the inherent correlations inherent in the sample data. Using machine learning algorithms, the complex nonlinear relationships between methane emissions and multiple factors, such as feeding level, diet composition, and individual differences, are directly learned and characterized. This allows for the establishment of a purely data-driven emissions prediction model that reflects the statistical correlation between the dependent and independent variables, serving as a useful complement to the kinetic-based emission simulation model.
[0089] During implementation, the system uses a large amount of high-quality sample feature data as a foundation and employs the Support Vector Regression (SVR) algorithm to construct a nonlinear regression model between methane emissions and various influencing factors. The selection of SVR as a regression learning tool is based on the following considerations: First, as a regression version of the Support Vector Machine (SVM), SVR inherits the excellent properties of SVM, such as maximizing classification intervals and using kernel techniques to handle nonlinearity. It possesses strong nonlinear modeling and generalization capabilities, and can effectively handle the complex nonlinear relationship between methane emissions and influencing factors. Second, SVR has a relatively low dependence on sample data, resulting in a concise and compact model that is less prone to overfitting, achieving good learning results even on medium-sized samples. Furthermore, SVR can effectively handle high-dimensional input spaces and automatically implement feature screening, which offers unique advantages for dealing with multi-source, heterogeneous aquaculture big data.
[0090] During the modeling process, the system randomly divides sample feature data into a training set and a test set. The methane emissions of each sample are set as the target variable, and various influencing factor characteristics (such as milk production, roughage ratio, diet composition, temperature and humidity, etc.) are set as input variables. The SVR model is trained on the training set. By introducing slack variables and penalty factors, SVR adds a regularization term to traditional least squares regression, controlling empirical risk while reducing structural risk, which helps improve the model's robustness and predictive accuracy. Furthermore, by solving the dual problem, the final SVR model is represented by a linear combination of a small number of support vector samples, which offers the advantages of concise expression and efficient prediction. During model training, key SVR parameters (such as kernel function type, penalty factor, kernel function parameters, etc.) are optimized through grid search using methods such as cross-validation to find the regression model with the optimal parameter combination. Model performance is evaluated on the test set. If the prediction accuracy meets the requirements, it is determined as the final model; otherwise, the model is returned to the training process with adjusted parameters. It is worth mentioning that the SVR modeling process is end-to-end. There is no need to screen data features in advance. Instead, it automatically discovers the intrinsic connections between features through techniques such as kernel functions, which greatly reduces the subjectivity of human participation.
[0091] S302, establishing a methane emission dynamic prediction model using a neural network algorithm based on the sample characteristic data;
[0092] Specifically, when constructing a data-driven second methane emission process simulation model, in addition to using the support vector regression algorithm to construct a nonlinear regression model between methane emissions and influencing factors, a neural network algorithm was also used to establish a dynamic prediction model for methane emissions based on sample characteristic data. The core purpose of this step is to fully utilize the powerful nonlinear fitting and dynamic modeling capabilities of neural networks to characterize and predict the methane emission process from a spatiotemporal dynamic perspective. By learning the dynamic evolution patterns of emissions contained in the sample data, a data-driven model guided by time series prediction is constructed to achieve early prediction and precise control of methane emission trends in the future, further expanding the application dimensions of the emission simulation model.
[0093] In specific implementation, the system uses a long short-term memory (LSTM) neural network model to model the dynamic evolution of methane emissions. LSTM is a special type of recurrent neural network (RNN) that incorporates a gating mechanism to overcome the long-term dependency issues faced by traditional RNNs. It is highly capable of processing and mining long-range dependencies in time series data and has achieved widespread success in the field of time series forecasting. The choice of LSTM as a time series modeling tool is based on the following considerations: First, methane emissions are a dynamic process. Emissions are not only related to the current state of various influencing factors, but also to historical states, exhibiting a certain "memory" effect. LSTM is well-suited to capturing and leveraging these temporal dependencies, automatically learning the dynamic evolution of methane emissions from time series data. Second, the factors influencing methane emissions are complex, with cross-influences and time lags between them, making them difficult to describe using simple physical models. The powerful nonlinear representation capabilities of LSTM are well-suited to addressing this complexity, adaptively building a deep nonlinear model that reflects the dynamic characteristics of methane emissions.
[0094] During the modeling process, the system divides continuously collected sample feature data into time slices. Each time slice contains hourly methane emissions data and monitoring data on various influencing factors at that time. Each time slice data is further divided into an input sequence (feature data from the previous period) and an output sequence (emissions data from the next period). Based on this, an LSTM network is constructed, consisting of an input layer, several LSTM hidden layers, a fully connected layer, and an output layer. The input layer receives the input sequence data for each time slice, and the output layer generates predicted methane emissions for several future time periods. During model training, the Truncated BPTT algorithm is used to optimize the network parameters at each layer by backpropagating the error over time, minimizing the error between the predicted output and the actual output. By training on a large number of time series samples, the model effectively learns and grasps the inherent dynamic characteristics of methane emissions, enabling prediction of future emission trends based on past environmental conditions and cattle physiological conditions. When applying the model, one only needs to input the online monitoring data of the most recent period to predict the dynamic changes in emissions in the subsequent period, providing a decision-making basis for precise regulation and proactive management of cow methane emissions.
[0095] S303: Combining the nonlinear regression model and the methane emission dynamic prediction model to form the second methane emission process simulation model.
[0096] Specifically, when constructing a data-driven methane emission simulation model, rather than simply adopting a single modeling paradigm, the team combined a nonlinear regression model based on support vector regression with a dynamic methane emission prediction model based on neural network learning to form a complete second methane emission simulation model. This approach aims to leverage the complementary strengths of the two models, creating a multi-perspective, multi-spatiotemporal characterization of the emission process and comprehensively enhancing the model's ability to characterize and predict complex emission behaviors. This model combination amplifies the strengths of the regression and time series models, while simultaneously compensating for their shortcomings. This results in more comprehensive and accurate simulation results, and a wider range of applications.
[0097] In implementation, the system uses a weighted averaging approach to combine the outputs of the support vector regression model and the LSTM prediction model to generate the final methane emissions simulation output. This combination strategy is based on the following considerations: First, the support vector regression model, based on the relationship between static characteristics at each sampling moment and emissions at that moment, excels at characterizing the overall emission level in equilibrium. The LSTM prediction model, based on the dynamic relationship between time series characteristics and future emissions, excels at predicting dynamic emissions processes in non-equilibrium states. The two models complement each other well in terms of modeling scale and focus. Second, the regression model requires only state characteristics at the current moment, while the prediction model also requires sequence characteristics within a historical time window. The two models differ in their data dependency and complexity in practical applications, making their combined use flexible to meet diverse needs. Furthermore, different models have varying sensitivities to sample quality and quantity. By using weighted averaging, when the predictive performance of one model degrades, the other can compensate and correct for any decline in performance. Overall, this model combination allows the final simulation output to balance both the static global and dynamic local characteristics of emissions behavior, expanding both predictive performance and applicable scenarios.
[0098] During the specific combination and fusion process, the weighted average weight parameters are not set subjectively but are adaptively adjusted based on the respective prediction accuracy of the two models. The system first uses an independent validation dataset to evaluate the predictive performance of the support vector regression model and the LSTM prediction model. The prediction deviation between the two models is measured using metrics such as mean squared error. The inverse of the deviation is normalized and used as the initial combination weight. The model with the smaller prediction deviation is assigned a higher weight coefficient. In actual use, the system regularly re-evaluates the two models using newly collected sample data, continuously tracking changes in the model's prediction performance and dynamically updating the combination weight coefficients. This allows the fusion output to automatically adapt to the optimal combination under different operating conditions. This adaptive combination and fusion strategy maximizes the potential of different models and demonstrates greater robustness in complex and changing real-world scenarios. Through this weighted fusion approach, the static global model and the dynamic local model achieve a "1+1>2" aggregation effect, significantly improving overall simulation and prediction performance.
[0099] S104, collecting real-time multi-source heterogeneous emission monitoring data, real-time influencing factor data, and real-time key environmental parameter data of the ruminant to be tested in real time, and constructing a real-time data set based on the real-time multi-source heterogeneous emission monitoring data, the real-time influencing factor data, and the real-time key environmental parameter data;
[0100] Specifically, this application uses an independently developed "electronic rumen pill" sensor array to conduct real-time monitoring of the animals to be tested in the target farm. Through the Internet of Things communication technology, the data collected by the "electronic rumen pill" is transmitted to the smart ranch management system in real time. The system automatically parses and classifies the received data stream to form structured real-time multi-source heterogeneous emission monitoring data, including real-time sequences of parameters such as methane concentration, rumen pH, VFA composition, and methane-producing bacteria abundance. Synchronously, through the integrated environmental sensor network deployed on the ranch, real-time key environmental parameters such as livestock house temperature and humidity, and air composition in the house are collected in real time. In addition, livestock management personnel use the smart ranch mobile terminal APP to timely record real-time influencing factor data such as feeding variety, feed intake, and milk production in routine management work.
[0101] Through the parallel collection of multiple channels, including an "electronic rumen pill" sensor array, a pasture environmental sensor network, and a mobile terminal app, this application obtains real-time data on three dimensions: individual animals, feeding management, and environmental conditions. The smart pasture management system utilizes data interface technology to aggregate and integrate heterogeneous data streams generated continuously 24 / 7, creating a real-time dataset reflecting the current livestock production process. Compared to static historical data, this real-time dataset is notable for its heterogeneity and dynamism, placing higher demands on data processing and analysis methods.
[0102] S105, inputting the real-time data set into the first methane emission process simulation model to obtain a first emission prediction result;
[0103] Specifically, key parameter values from the real-time data set are extracted and input into the first methane emission process simulation model in a suitable data structure. This mechanism-driven mathematical model, based on an abstract description of the methane production mechanism in the rumen of ruminants, quantitatively describes the nonlinear functional relationship between key influencing factors such as diet composition, rumen environment, and microbial population structure and the methane synthesis rate. It is primarily composed of a set of differential or algebraic equations.
[0104] Key parameters from the real-time data, such as the NDF and ADF content of the diet, rumen pH, VFA concentration, and abundance of methane-producing bacteria, are assigned to the corresponding state variables in the model. Parameters such as the animal's weight, lactation stage, and ambient temperature are then input into the model. After setting the initial and boundary conditions, the model can be used to calculate the methane synthesis rate for each animal under current breeding conditions using numerical integration. This instantaneous emission rate, as determined by the model, is the first emission prediction based on the real-time data.
[0105] The advantage of the first methane emission process simulation model is that its construction fully considers the biological mechanism of methane production, and each parameter has a clear physical meaning. By inputting real-time monitoring data, the model can respond to the dynamic changes in animal status and environmental conditions, and has a certain extrapolation ability for the prediction of emission rates. It shows good robustness in the case of data loss, system disturbances, etc. However, due to the limitations of the key processes and parameters considered, the model does not adequately depict the effects of some influencing factors (such as trace elements, feed additives, etc.), and is more sensitive to parameter uncertainty, which may overestimate or underestimate emissions in practical applications. Despite this, in the absence of other prior knowledge, the first emission prediction results can still serve as an important reference for judging the real-time emission status of animals.
[0106] S106, inputting the real-time data set into the second methane emission process simulation model to obtain a second emission prediction result;
[0107] Specifically, characteristic parameters matching the input structure of the second methane emission process simulation model are extracted from the real-time dataset. After preprocessing through feature encoding and data normalization, they are input into the empirical model as sample vectors. The model uses a machine learning algorithm trained on historical emission data to automatically learn the implicit correlation patterns between methane emission intensity and multiple influencing factors such as diet composition, rumen parameters, and microbial populations. It then constructs a multivariate nonlinear regression equation to predict methane emission rates for individual animals.
[0108] By inputting real-time data sets of dietary parameters such as NDF, ADF, and starch content; fermentation environment parameters such as rumen pH and VFA concentration; and microbial parameters such as the abundance of methane-producing and propionate-utilizing bacteria; along with metadata such as animal weight, lactation stage, and ambient temperature, into a trained machine learning model, a regression equation is used to calculate the predicted instantaneous methane emission rate for each individual animal based on the combination of these multiple input feature parameters. This emission rate prediction provided by the model is the second emission prediction result based on real-time data.
[0109] S107: Determine target methane emission data based on the first emission prediction result and the second emission prediction result.
[0110] Specifically, this application performs a confidence analysis on the first and second emissions forecasts, examining statistical characteristics such as the volatility of each model's output and the distribution of outliers, to preliminarily determine the credibility of the two forecasts. When the confidence levels are roughly consistent, the system fuses the two forecasts using a weighted average to form a preliminary estimate of emissions. The weighting coefficients are dynamically adjusted based on each model's historical performance and suitability for specific operating conditions, aiming to achieve a relatively robust fusion result.
[0111] In cases where confidence levels differ significantly, the system initiates a multi-source data cross-validation mechanism, integrating online monitoring readings, regional emission benchmarks, and historical data from similar animals to calibrate the predictions of the two models. Real-time monitoring data from online methane sensors is the most direct reference, but due to factors such as measurement error and missing data, it is not absolutely reliable and requires cross-correlation with other data sources. Regional emission benchmarks are statistically derived from historical emission levels of animals of the same breed and stage in the region, and can be used to determine whether model predictions fall outside the typical range. Historical data from similar animals demonstrates the distribution of individual emissions under specific dietary composition and feeding and management conditions, providing context for assessing the current animal's emission status. The system leverages this multi-source heterogeneous data, applying a series of mathematical statistics and machine learning algorithms to verify model predictions, estimate the probability of each prediction, and dynamically adjust the average weighting coefficient to ultimately determine the target methane emissions.
[0112] Based on the above embodiment, as an optional embodiment, determining target methane emission data based on the first emission prediction result and the second emission prediction result includes:
[0113] Calculating a weighted average of the first emission prediction result and the second emission prediction result to obtain a weighted average result of methane emissions;
[0114] Calculating the degree of dispersion of the first emission prediction result and the second emission prediction result, and when the degree of dispersion is greater than a preset threshold, determining the prediction result of the first emission prediction result and the second emission prediction result that deviates less from the historical emission data as the target methane emission data;
[0115] When the degree of dispersion is less than or equal to a preset threshold, the weighted average result of the methane emissions is determined as the target methane emission data.
[0116] Specifically, after obtaining the first mechanism-based emissions prediction result and the second data-based emissions prediction result, instead of simply combining the two, an adaptive weighted fusion strategy was designed to dynamically adjust the combination scheme based on the degree of discreteness of the two prediction results, thereby obtaining the final target methane emissions data. The core purpose of this is to maximize the integration of the prediction results of the two types of models, while adaptively weighting them according to the degree of consistency between the two, in order to obtain the optimal prediction value that is close to the actual emission level in different situations. By comprehensively considering the mean level and degree of discreteness of the prediction results and referring to historical real emission data, this method can flexibly adapt to scenarios with different model prediction performance and data quality, and independently select the optimal combination strategy, making the fusion result more robust and reliable.
[0117] During specific implementation, the system first calculates the weighted average of the first emission prediction result and the second emission prediction result as the basis for the combined prediction of the two models. The weight coefficient of the weighted average can be set according to the prediction accuracy of the two models on historical data. The model with a small prediction deviation will obtain a larger weight, so that its prediction result occupies a more important proportion in the average result. On this basis, the system further calculates the degree of dispersion of the two prediction results, that is, the size of the difference between the two. Usually, the degree of dispersion can be measured using indicators such as variance and standard deviation in a statistical sense to reflect the deviation of the prediction results of the two models. It is worth noting that since the first model and the second model are based on mechanisms and data respectively, they represent predictions of the same process from different perspectives. They have their own strengths and their applicability and reliability under different working conditions are not the same. Therefore, simply comparing the means of the two may cover up some valuable information, and the quality of the combination needs to be evaluated in combination with the degree of dispersion.
[0118] When the dispersion between the two prediction results is small—that is, their values are relatively close—it indicates that their predictions of methane emissions are relatively consistent. In this case, the weighted average can be used directly as the final target emissions. Small dispersion indirectly reflects the stable predictive performance of the two models and the high credibility of the combined result. However, when the dispersion is large, exceeding a preset threshold, it indicates that the emission predictions of the mechanism-based model and the data-based model diverge significantly, reducing the interpretability and reliability of the combined model. This can occur for a variety of reasons, such as insufficient ability of some models to characterize the current operating conditions or fluctuations in the representativeness or quality of the input data. To avoid significant bias introduced by the weighted average, the system evaluates the deviation of the two predictions from the actual emissions and selects the prediction with the smallest deviation from the historical observations as the final output. This utilizes prior knowledge provided by historical observations to assess and correct model prediction deviations, thereby reducing the uncertainty of the fused prediction. The thresholds described above should be determined through empirical analysis during the model validation phase and adjusted appropriately in practice based on changes in model performance.
[0119] Through this adaptive weighted fusion strategy, the present application has well balanced the combined effects of the mechanism-driven model and the data-driven model, which not only gives play to the complementarity of different models, but also can flexibly adjust the combination scheme according to actual conditions. When the prediction results of the two models are basically consistent, the weighted average strategy is directly adopted, which is equivalent to using the combined prediction results of the two models to offset their respective limitations, complement each other, and improve the overall prediction performance; and when the two models have obvious differences, they make full use of the reference information provided by the historical real data, and select the prediction result with the smallest deviation as the final output, avoiding the large deviation caused by blind averaging, and playing the role of real-time calibration and self-correction. This dynamic combination mechanism that takes into account the internal complementarity of the model and uses external reference data for quality assessment enables the fusion prediction results to adaptively correspond to different scenarios, and the overall prediction quality and robustness are significantly improved. The idea of external data correction embodied in this mechanism has important implications for improving the robustness of multi-model combinations of complex systems and reducing the uncertainty of combination results.
[0120] On the other hand, the present application also provides a ruminant methane emission measurement system, such as Figure 2 , the system comprises:
[0121] The sample data acquisition module 1 is used to acquire multi-source heterogeneous sample monitoring data of ruminant methane emissions within a certain time span, sample data of factors affecting methane emissions, and sample data of the rumen environment, and construct a sample data set based on the multi-source heterogeneous sample monitoring data, the sample data of the factors affecting methane emissions, and the sample data of the rumen environment;
[0122] Data processing module 2, used for preprocessing the sample data in the sample data set to obtain a standard sample data set, and preprocessing the standard sample data set to obtain a target sample data set;
[0123] Modeling module 3, used to extract features from the sample data in the target sample data set to obtain sample feature data, and to construct a mechanism-driven first methane emission process simulation model and a data-driven second methane emission process simulation model based on the sample feature data;
[0124] A real-time data acquisition module 4 is used to collect real-time multi-source heterogeneous emission monitoring data, real-time influencing factor data and real-time key environmental parameter data of the ruminant to be tested, and to construct a real-time data set based on the real-time multi-source heterogeneous emission monitoring data, the real-time influencing factor data and the real-time key environmental parameter data;
[0125] A first emission prediction module 5, configured to input the real-time data set into the first methane emission process simulation model to obtain a first emission amount prediction result;
[0126] A second emission prediction module 6, configured to input the real-time data set into the second methane emission process simulation model to obtain a second emission prediction result;
[0127] The third emission prediction module 7 is configured to determine target methane emission data based on the first emission prediction result and the second emission prediction result.
[0128] Please refer to Figure 3 The present application also discloses an electronic device. Figure 3 The electronic device 300 may include: at least one processor 301 , at least one network interface 304 , a user interface 303 , a memory 305 , and at least one communication bus 302 .
[0129] The communication bus 302 is used to implement the connection and communication between these components.
[0130] The user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.
[0131] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0132] The processor 301 may include one or more processing cores. Using various interfaces and circuits, the processor 301 connects to various components within the server. It executes instructions, programs, code sets, or instruction sets stored in the memory 305, as well as accesses data stored in the memory 305, to perform various server functions and process data. Optionally, the processor 301 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 301 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display screen; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 301 but implemented as a separate chip.
[0133] Among them, the memory 305 may include a random access memory (Random Access Memory, RAM) and may also include a read-only memory (Read~Only Memory). Optionally, the memory 305 includes a non-transitory computer-readable medium. The memory 305 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 305 may also optionally be at least one storage system located away from the aforementioned processor 301. Reference Figure 3 The memory 305 as a computer storage medium may include an operating system, a network communication module, a user interface module, and an application program for a flocculant addition analysis method.
[0134] exist Figure 3In the electronic device 300 shown, the user interface 303 is mainly used to provide an input interface for the user and obtain the data input by the user; and the processor 301 can be used to call the application program storing the road assessment method in the memory 305. When executed by one or more processors 301, the electronic device 300 executes one or more methods in the above-mentioned embodiments. It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that this application is not limited to the described order of actions, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for this application. In the above-mentioned embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0135] In the several embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, and the indirect coupling or communication connection of the system or unit can be electrical or other forms.
[0136] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0137] The present application also provides a computer storage medium that can store multiple instructions, which are suitable for being loaded and executed by a processor as described above. Figure 1 The road assessment method of the embodiment shown, the specific execution process can be found in Figure 1 The detailed description of the illustrated embodiment will not be repeated here.
[0138] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0139] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of this application. The aforementioned memory includes various media that can store program code, such as USB flash drives, mobile hard drives, magnetic disks, or optical disks.
[0140] The above are merely exemplary embodiments of the present disclosure and are not intended to limit the scope of the present disclosure. In other words, any equivalent variations and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the disclosure and the practical implications thereof.
[0141] This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not described herein. The description and examples are to be considered as exemplary only, and the scope and spirit of the present disclosure are to be defined by the claims.
Claims
1. A method for measuring methane emissions from ruminants, characterized in that: The method comprises: Acquire multi-source heterogeneous sample monitoring data of ruminant methane emissions within a certain time span, sample data of factors affecting methane emissions, and sample data of the rumen environment, and construct a sample data set based on the multi-source heterogeneous sample monitoring data, the sample data of the factors affecting methane emissions, and the sample data of the rumen environment; Preprocessing the sample data in the sample data set to obtain a standard sample data set, and preprocessing the standard sample data set to obtain a target sample data set; Performing feature extraction on the sample data in the target sample data set to obtain sample feature data, and constructing a mechanism-driven first methane emission process simulation model and a data-driven second methane emission process simulation model based on the sample feature data; Real-time collection of real-time multi-source heterogeneous emission monitoring data, real-time influencing factor data, and real-time key environmental parameter data of the ruminant to be tested, and construction of a real-time data set based on the real-time multi-source heterogeneous emission monitoring data, the real-time influencing factor data, and the real-time key environmental parameter data; Inputting the real-time data set into the first methane emission process simulation model to obtain a first emission prediction result; Inputting the real-time data set into the second methane emission process simulation model to obtain a second emission prediction result; determining target methane emission data based on the first emission prediction result and the second emission prediction result; Wherein, constructing a mechanism-driven first methane emission process simulation model based on the sample characteristic data includes: Constructing a conceptual model of methane generation in the rumen based on the sample characteristic data, and introducing a kinetic model into the conceptual model of methane generation in the rumen to obtain a preliminary methane emission process simulation model; Fitting and optimizing key parameters in the kinetic differential equation group in the preliminary methane emission process simulation model using the sample characteristic data to obtain the first methane emission process simulation model; Wherein, constructing a data-driven second methane emission process simulation model based on the sample characteristic data includes: Based on the sample characteristic data, a nonlinear regression model between methane emissions and various influencing factors is constructed using a support vector regression algorithm; Based on the sample characteristic data, a dynamic prediction model for methane emissions is established using a neural network algorithm; The nonlinear regression model and the methane emission dynamic prediction model are combined to form the second methane emission process simulation model.
2. The method according to claim 1, characterized in that The preprocessing of the sample data in the sample data set to obtain a standard sample data set includes: Performing data cleaning on the sample data in the sample data set to obtain cleaned sample data; The cleaned sample data is subjected to data standardization processing to obtain the standard sample data set.
3. The method according to claim 1, characterized in that The preprocessing of the standard sample data set to obtain the target sample data set includes: Analyzing the correlation between the various sample parameters in the standard sample data set using a correlation analysis method to obtain correlation values between the various sample parameters, and screening out target sample parameters that are highly correlated with methane emissions based on the correlation values and a preset correlation threshold; The target sample data set is constructed based on the target sample parameters.
4. The method according to claim 1, wherein The extracting features of the sample data in the target sample data set to obtain sample feature data includes: Using principal component analysis to perform feature dimensionality reduction on the target sample data set to extract the main features of the sample data; Using factor analysis method to extract features of the target sample data set to obtain common factors reflecting the intrinsic structure of the sample data; The sample feature data is constructed based on the main features and the common factors.
5. The method according to claim 1, wherein The determining target methane emission data based on the first emission prediction result and the second emission prediction result includes: Calculating a weighted average of the first emission prediction result and the second emission prediction result to obtain a weighted average result of methane emissions; Calculating the degree of dispersion of the first emission prediction result and the second emission prediction result, and when the degree of dispersion is greater than a preset threshold, determining the prediction result of the first emission prediction result and the second emission prediction result that deviates less from the historical emission data as the target methane emission data; When the degree of dispersion is less than or equal to a preset threshold, the weighted average result of the methane emissions is determined as the target methane emission data.
6. A ruminant methane emission measurement system, characterized in that: For implementing the method for measuring methane emissions from ruminants according to claim 1, the ruminant methane emission measuring system comprises: a sample data acquisition module for acquiring multi-source heterogeneous sample monitoring data of ruminant methane emissions, sample data of influencing factors affecting methane emissions, and sample data of the rumen environment within a certain time span, and constructing a sample data set based on the multi-source heterogeneous sample monitoring data, the sample data of influencing factors, and the sample data of the rumen environment; a data processing module, configured to preprocess the sample data in the sample data set to obtain a standard sample data set, and preprocess the standard sample data set to obtain a target sample data set; a modeling module, configured to extract features from the sample data in the target sample data set to obtain sample feature data, and construct a mechanism-driven first methane emission process simulation model and a data-driven second methane emission process simulation model based on the sample feature data; A real-time data acquisition module is used to collect real-time multi-source heterogeneous emission monitoring data, real-time influencing factor data and real-time key environmental parameter data of the ruminant to be tested, and to construct a real-time data set based on the real-time multi-source heterogeneous emission monitoring data, the real-time influencing factor data and the real-time key environmental parameter data; a first emission prediction module, configured to input the real-time data set into the first methane emission process simulation model to obtain a first emission amount prediction result; a second emission prediction module, configured to input the real-time data set into the second methane emission process simulation model to obtain a second emission prediction result; The third emission prediction module is used to determine target methane emission data based on the first emission prediction result and the second emission prediction result.
7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, which are suitable for being loaded by a processor and executed by a method according to any one of claims 1 to 5.
8. An electronic device, characterized in that: The electronic device comprises a processor, a memory and a transceiver, wherein the memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and system for monitoring and reducing ruminant methane production
US20110192213A1
Method, apparatus and system for detecting carbon emission-involved gas from ruminant
US20240040995A1