A method and system for detecting the environment of livestock breeding houses
By deploying infrared thermal imaging, voiceprint and gas sensors to generate space-time aligned multimodal data sets, a temperature-voiceprint-gas three-dimensional correlation matrix is constructed, combining temperature-humidity index and dynamic weight of gas concentration, the LSTM model is used to realize multimodal data fusion and risk prediction of the livestock farmhouse environment, solving the problems of insufficient perceptual dimensions and weak information fusion in the existing technology, and achieving efficient environmental regulation.
Patent Information
- Application Number
- CN202510474186.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing animal husbandry environment detection methods have not fully utilized animal sound data for state perception, multimodal perception information has not been effectively integrated under a unified spatial model, and the model lacks a dynamic weighting processing mechanism for multi-source characteristics, making it difficult to achieve joint prediction and dynamic regulation of the fusion stress risk and pollution risk of temperature, sound and gas.
Deploy infrared thermal imaging devices, voiceprint sensors and gas sensor arrays to generate space-time alignment multimodal data sets, build a temperature-voiceprint-gas three-dimensional correlation matrix through spatiotemporal grid processing, dynamically compensate the animal surface temperature with the temperature and humidity index, and input the gas concentration as a weight factor into the LSTM timing prediction model, output stress risk levels and pollution risk levels, and generate and implement environmental regulation strategies.
The joint identification and intelligent response to animal stress and environmental pollution risks in animal husbandry houses has been achieved, data integrity and prediction accuracy have been improved, and the timely regulatory response has been achieved, and the problems of insufficient perceptual dimensions and weak information fusion have been overcome.
Smart Images

Figure CN120008690B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent perception and regulation of animal husbandry environments, and specifically to a method and system for detecting the environment of animal husbandry houses. Background Art
[0002] With the development of smart agriculture, the livestock farming sector is gradually transitioning towards digitalization and intelligence, with environmental sensing and animal health monitoring technologies becoming research hotspots. Traditional livestock house environmental monitoring relies primarily on single physical sensors to collect parameters such as temperature, humidity, and ammonia for basic ventilation and exhaust control. However, the response of animals to environmental changes (such as stress responses) exhibits multidimensional and multimodal characteristics. In recent years, sensing methods such as infrared thermal imaging, biometric voiceprint recognition, and gas concentration sensing have been gradually introduced into livestock farming systems, enabling comprehensive assessments of animal behavior, health status, and environmental quality. In particular, the combination of multimodal sensing and deep learning models has become an emerging direction in intelligent livestock monitoring systems.
[0003] Although existing technologies have introduced image, gas and other sensing methods for environmental monitoring to a certain extent, there are still significant deficiencies in the fusion modeling and intelligent decision-making support of sound data. First, most livestock house monitoring systems ignore the key physiological and behavioral signals of animal sounds, resulting in a single monitoring dimension and an inability to perceive the true reaction state of animals. Even if individual studies involve sound recognition, they are usually limited to offline detection or simple classification, and lack the ability to link analysis with environmental data. Secondly, most existing multimodal systems have a "data side-by-side" structure, which does not effectively integrate temperature, sound and gas data in a unified spatial and temporal framework, making it difficult to construct a real and dynamic animal-environment interaction model. Third, existing time series prediction models generally ignore the weight of different modalities on the prediction effect, and especially lack the description of the relationship between sound characteristics and gas pollution. Summary of the Invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] Therefore, the technical problems solved by the present invention are: the existing methods for livestock breeding environment detection and animal health risk identification fail to fully utilize animal sound data for status perception, multimodal perception information fails to be effectively integrated under a unified spatial model, and the model lacks a dynamic weight processing mechanism for multi-source features, as well as how to achieve joint prediction and dynamic regulation of stress risk and pollution risk based on the fusion of temperature, sound and gas.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, an embodiment of the present invention provides a method for detecting the environment of a livestock breeding house, including deploying an infrared thermal imaging device, a voiceprint sensor, and a gas sensor array to synchronously collect raw thermal image data, raw audio signals, and raw gas data to generate a spatiotemporally aligned multimodal dataset;
[0008] Performing spatiotemporal gridding on the multimodal dataset, extracting temperature gradients, voiceprint MFCC features, and gas concentration spatial distribution, and constructing a temperature-voiceprint-gas three-dimensional correlation matrix;
[0009] Based on the temperature-voiceprint-gas three-dimensional correlation matrix, the animal's surface temperature is dynamically compensated in combination with the temperature and humidity index, and the gas concentration is input as a weight factor into the LSTM time series prediction model to output the stress risk level and the pollution risk level;
[0010] According to the stress risk level and the pollution risk level, a corresponding environmental control strategy is generated and executed.
[0011] As a preferred embodiment of the livestock breeding house environment detection method of the present invention, the deployment of infrared thermal imaging devices, voiceprint sensors, and gas sensor arrays includes setting up an array of infrared thermal imaging devices above the livestock house to periodically scan the animal group and collect raw thermal image data;
[0012] Install directional soundprint sensors in animal activity areas to collect the animals' original audio signals;
[0013] Deploy multiple gas sensor nodes in animal excretion areas and activity areas to synchronously collect raw gas data;
[0014] The raw gas data includes ammonia concentration, carbon dioxide concentration, and temperature and humidity parameters.
[0015] As a preferred embodiment of the livestock breeding house environment detection method of the present invention, the generating of a spatiotemporally aligned multimodal data set includes performing spatial coordinate calibration on the original thermal image data to obtain a temperature distribution data stream marked with spatial coordinates;
[0016] Extract the sound source orientation features from the original audio signal collected by the voiceprint sensor to obtain a voiceprint data stream with spatial orientation features;
[0017] Perform spatial mapping on the raw gas data collected by the gas sensor node to obtain a data stream of environmental parameters with coordinate markings;
[0018] Perform time stamp synchronization on the temperature distribution data stream, voiceprint data stream, and environmental parameter data stream to construct a multimodal perception data block with a unified time window and spatial reference system;
[0019] The multimodal perception data blocks are subjected to outlier removal, and the data segments exceeding the preset threshold range are screened out to obtain a cleaned spatiotemporally aligned multimodal dataset.
[0020] As a preferred embodiment of the livestock breeding house environment detection method of the present invention, the spatiotemporal gridding of the multimodal data set includes mapping the temperature distribution data stream into a preset three-dimensional spatial grid model, assigning each temperature sampling point to a corresponding spatial grid unit, and generating temperature grid data based on a three-dimensional grid structure;
[0021] The voiceprint data stream is divided into regions according to the azimuth and distance information of the sound source, and mapped to the corresponding spatial grid area to obtain the voiceprint grid data;
[0022] The environmental parameter data stream is distributed to each spatial grid node based on the physical location of the sensor node to obtain the environmental parameter spatial grid data;
[0023] The temperature grid data, voiceprint grid data and environmental parameter space grid data are aligned by grid numbering to generate a spatiotemporal gridded multimodal dataset with a unified spatial index structure.
[0024] As a preferred embodiment of the livestock breeding house environment detection method of the present invention, the extraction of temperature gradient, voiceprint MFCC features and gas concentration spatial distribution includes: from the temperature grid data in the spatiotemporal grid multimodal dataset, according to the temperature value changes of adjacent spatial grid cells, calculating the temperature difference between each grid, and extracting the temperature gradient information;
[0025] Extract the audio clips in each time window from the voiceprint grid data, perform pre-emphasis, frame division, windowing, fast Fourier transform and Mel filtering on the audio clips, calculate the Mel frequency cepstral coefficients, and obtain the voiceprint MFCC features;
[0026] Extract ammonia and carbon dioxide concentration data from the environmental parameter spatial grid data, and use a spatial interpolation algorithm to reconstruct a continuous gas concentration distribution map in a three-dimensional grid coordinate system to form gas concentration spatial distribution data;
[0027] The construction of the temperature-voiceprint-gas three-dimensional association matrix includes fusing the temperature gradient information, the voiceprint MFCC features and the gas concentration spatial distribution data based on the same spatial index number, and constructing a temperature-voiceprint-gas three-dimensional association matrix with the spatial grid unit as the index dimension, the temperature gradient as the first characteristic dimension, the voiceprint MFCC as the second characteristic dimension, and the gas concentration as the third characteristic dimension.
[0028] As a preferred embodiment of the livestock breeding house environment detection method described in the present invention, the method comprises: dynamically compensating the animal surface temperature based on the temperature-voiceprint-gas three-dimensional correlation matrix in combination with the temperature and humidity index, and inputting the gas concentration as a weight factor into the LSTM time series prediction model, utilizing the temperature gradient information in the temperature-voiceprint-gas three-dimensional correlation matrix and the temperature and humidity values in the environmental parameter spatial grid data, calculating the temperature and humidity index value of each spatial grid unit according to the temperature and humidity index calculation model, and dynamically correcting the temperature gradient based on the difference between the index and the set reference value in combination with a preset compensation coefficient to obtain a temperature data matrix after temperature and humidity compensation;
[0029] The ammonia and carbon dioxide concentration values of each grid cell are extracted from the spatial distribution data of gas concentrations. After normalization, they are used as weight factors to weight the Mel-frequency cepstral coefficients in the voiceprint features to generate a weighted voiceprint feature matrix reflecting the degree of gas pollution impact.
[0030] The temperature data matrix that has been compensated for temperature and humidity, the weighted voiceprint feature matrix reflecting the degree of gas pollution impact, and the spatial distribution data of gas concentration are spliced according to a unified spatial grid index to construct a time series input vector containing temperature, voiceprint and gas multimodal information. The vector is then arranged according to the set time window order to form a multimodal time series input sequence with temporal continuity and spatial consistency.
[0031] As a preferred embodiment of the livestock breeding house environment detection method of the present invention, the output of the stress risk level and the pollution risk level includes inputting a multimodal time series input sequence constructed according to a unified spatial index and a time window into a neural network model including a long short-term memory unit, wherein the neural network model structure includes an input layer for receiving the input sequence, a multi-layer bidirectional LSTM network layer for capturing time dependency, and a classification and regression joint output layer for outputting the results;
[0032] The neural network model is used to extract time series features and perform state modeling on the input sequence to obtain an output value of the animal stress risk level for each spatial grid unit in the current time window, wherein the stress risk level is divided into multiple level categories according to preset standards;
[0033] The regression output channel in the neural network model predicts and generates a pollution risk score for each spatial grid cell within the corresponding time window. The pollution risk score is a continuous risk index obtained by fitting the gas pollution characteristics and the voiceprint disturbance characteristics.
[0034] Finally, the corresponding animal stress risk level label and pollution risk level score in each time window are output.
[0035] As a preferred solution of the livestock breeding house environment detection method described in the present invention, the generation and execution of the corresponding environmental control strategy includes generating an environmental control strategy corresponding to the current livestock house environmental status based on the evaluation results of the stress risk level and the pollution risk level, and executing corresponding environmental adjustment operations according to the generated environmental control strategy.
[0036] In a second aspect, an embodiment of the present invention provides a livestock breeding house environment detection system, comprising:
[0037] Data acquisition and synchronization module: deploys infrared thermal imaging devices, voiceprint sensors, and gas sensor arrays to synchronously collect raw thermal image data, raw audio signals, and raw gas data to generate a spatiotemporally aligned multimodal dataset;
[0038] Multimodal grid modeling module: performs spatiotemporal gridding on the multimodal dataset, extracts temperature gradients, voiceprint MFCC features, and gas concentration spatial distribution, and constructs a temperature-voiceprint-gas three-dimensional correlation matrix;
[0039] Risk identification and prediction module: Based on the temperature-voiceprint-gas three-dimensional correlation matrix, the temperature and humidity index are combined to dynamically compensate for the animal's surface temperature, and the gas concentration is input as a weight factor into the LSTM time series prediction model to output the stress risk level and pollution risk level;
[0040] Strategy generation and execution module: generates and executes corresponding environmental control strategies according to the stress risk level and pollution risk level.
[0041] Beneficial effects of the present invention: The present invention realizes the joint identification and intelligent response of animal stress and environmental pollution risks in livestock breeding houses through multimodal data acquisition, three-dimensional space fusion, gas weight modeling and deep learning prediction, overcoming the problems of insufficient perception dimensions, weak information fusion and delayed risk identification in the existing technology, and has the beneficial effects of high data integrity, strong prediction accuracy and timely regulation response. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive work, among which:
[0043] Figure 1 This is an overall flow chart of a livestock breeding house environment detection method provided by the first embodiment of the present invention. DETAILED DESCRIPTION
[0044] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0045] Example 1, with reference to Figure 1 , as one embodiment of the present invention, provides a method for detecting the environment of a livestock breeding house, comprising:
[0046] S1: Deploy infrared thermal imaging devices, voiceprint sensors, and gas sensor arrays to synchronously collect raw thermal image data, raw audio signals, and raw gas data to generate a spatiotemporally aligned multimodal dataset.
[0047] S11: An array of infrared thermal imaging devices is installed above the livestock house to periodically scan the animal group and collect raw thermal image data;
[0048] Install directional soundprint sensors in animal activity areas to collect the animals' original audio signals;
[0049] Deploy multiple gas sensor nodes in animal excretion areas and activity areas to synchronously collect raw gas data;
[0050] The raw gas data includes ammonia concentration, carbon dioxide concentration, and temperature and humidity parameters.
[0051] In the embodiment of the present invention, in order to balance cost control and perception coverage quality, an economical optimization deployment strategy is adopted to deploy infrared thermal imaging, voiceprint collection and gas sensing equipment.
[0052] Thermal imaging devices (e.g., mid-range models like the FLIR T540 or Seek Shot Pro) are deployed every 5 meters along the main aisle of the barn, mounted on a 2.5-meter-high ceiling. Their horizontal field of view is set to cover an 8-meter area. To avoid blind spots, a 20–30% overlap is maintained between devices. This sparse deployment, combined with subsequent spatial interpolation (e.g., inverse distance weighted (IDW) or Gaussian regression), enables continuous estimation of spatially distributed temperatures at a manageable cost, making it particularly suitable for small and medium-sized barns. The image sampling frequency is set to 0.2 Hz (one frame every 5 seconds) to meet the needs of thermal image updates for dynamic animal populations.
[0053] For voiceprint collection, considering the complex acoustic reflections and high background interference in dynamic livestock houses, this embodiment uses a linear array microphone set (such as the ReSpeaker 4-Mic or MiniDSP UMA-8), deployed above the animal activity area. Compared to unidirectional microphones, this array design provides sound source direction of arrival estimation and localized sound source focus, enhancing the stability and robustness of sound source localization. Audio data is sampled in WAV format at a frequency of 16kHz. Delay estimation is performed using the GCC-PHAT algorithm, ensuring localization error within ±15°, sufficient to meet the requirements for spatial matching of sound source hotspots.
[0054] Gas sensor nodes are deployed in a regionalized strategy, with one node per 4 square meters in the excretion area and one node per 6 square meters in the activity area. They use multi-parameter electrochemical sensors (such as the Winsen ZE15-CO / NH3). The nodes have a sampling period of 20 seconds, and data is transmitted back via the low-power LoRa protocol, saving energy. Sensor locations are pre-calibrated and RFID tags are attached during deployment. Fixed UHF readers periodically collect tag location information, implementing a low-cost dynamic position offset correction mechanism, replacing the more expensive UWB system.
[0055] S12: Perform spatial coordinate calibration on the original thermal image data to obtain a temperature distribution data stream marked with spatial coordinates;
[0056] Extract the sound source orientation features from the original audio signal collected by the voiceprint sensor to obtain a voiceprint data stream with spatial orientation features;
[0057] Perform spatial mapping on the raw gas data collected by the gas sensor node to obtain a data stream of environmental parameters with coordinate markings;
[0058] Perform time stamp synchronization on the temperature distribution data stream, voiceprint data stream, and environmental parameter data stream to construct a multimodal perception data block with a unified time window and spatial reference system;
[0059] The multimodal perception data blocks are subjected to outlier removal, and the data segments exceeding the preset threshold range are screened out to obtain a cleaned spatiotemporally aligned multimodal dataset.
[0060] In this embodiment of the present invention, the system uses a unified time base for timing alignment. Considering the non-dedicated network environment within the barn, this solution uses the SNTP protocol for master-slave clock synchronization, combined with a sliding time window alignment mechanism, to achieve an alignment accuracy of ≤500ms, which is completely acceptable in scenarios where animal behavior monitoring does not involve high-speed changes.
[0061] To improve data quality and avoid the real-time operational pressures of the Isolation Forest algorithm, which requires high algorithmic complexity, this solution incorporates an adaptive threshold setting mechanism based on a historical window for outlier removal. For example, the body temperature threshold is dynamically updated based on the ±2σ mean of the 24-hour local temperature distribution. The ammonia concentration setting adjusts based on the daily mean and the time of excretion activity, accounting for environmental volatility. For sound data, anomalous event segments are identified based on sound pressure level (SPL) and the rate of change of the frequency spectrum shape, and anomalous segments are removed for subsequent modeling.
[0062] It should be noted that through the above deployment strategy and processing mechanism, this invention effectively reduces deployment density and system complexity without sacrificing multimodal spatial resolution and temporal consistency, making it highly feasible for implementation under existing farm conditions. In particular, the collaborative design of the voiceprint array and the dynamic gas concentration weighting mechanism provides a high-quality, low-error input foundation for the subsequent construction of a deep coupling model of animal behavior and environmental factors.
[0063] S2: Perform spatiotemporal gridding on the multimodal dataset, extract temperature gradient, voiceprint MFCC features, and gas concentration spatial distribution, and construct a temperature-voiceprint-gas three-dimensional correlation matrix.
[0064] S21: Mapping the temperature distribution data stream to a preset three-dimensional spatial grid model, assigning each temperature sampling point to a corresponding spatial grid unit, and generating temperature grid data based on the three-dimensional grid structure;
[0065] The voiceprint data stream is divided into regions according to the azimuth and distance information of the sound source, and mapped to the corresponding spatial grid area to obtain the voiceprint grid data;
[0066] The environmental parameter data stream is distributed to each spatial grid node based on the physical location of the sensor node to obtain the environmental parameter spatial grid data;
[0067] The temperature grid data, voiceprint grid data and environmental parameter space grid data are aligned by grid numbering to generate a spatiotemporal gridded multimodal dataset with a unified spatial index structure.
[0068] In this embodiment of the present invention, a three-dimensional spatial grid model is first constructed based on the spatial structure and dimensions of the barn. To ensure a balance between model accuracy and computational efficiency, a cube grid with a resolution of 0.5m × 0.5m × 0.5m is used for small barns (≤ 20m × 10m × 5m); a resolution of 1m × 1m × 1m is used for large barns. All grid cells use a unified spatial reference system, and a one-dimensional mapping of three-dimensional coordinates is performed using Morton encoding (Z-order curves) to optimize spatial locality in data access and storage.
[0069] Temperature data is derived from the spatially coordinate-labeled temperature data stream obtained after processing the infrared thermal imaging image in step S1. Temperature values are assigned to corresponding grid cells using the inverse distance weighted (IDW) interpolation method. The five closest temperature measurement points are used as input samples, with weights calculated based on the inverse square of the distance. The interpolation range is limited to the effective sensor sensing radius (e.g., 3 meters) to ensure that the temperature distribution reflects local animal density variations.
[0070] After obtaining the direction of the sound source through the microphone array, the voiceprint data is combined with a sound intensity attenuation model (based on the ISO 9613-2:1996 standard) to estimate the distance to the sound source, ultimately localizing the sound source within the three-dimensional coordinate system of the livestock house. After assigning this localization result to the corresponding grid cell, all audio frames within that cell are subjected to Mel-Frequency Cepstral Coefficient (MFCC) feature extraction and averaging to form the voiceprint feature vector for that cell.
[0071] Ammonia gas obtained by the gas sensor node ( ),carbon dioxide( ) concentration and temperature and humidity parameters are mapped to the grid according to the sensor location, and the adjacent grid values are diffused using the three-point IDW interpolation method. The concentration distribution was further fitted using the adaptive Kriging (AK) method. The spherical model was selected as the variogram, and the interpolation radius was dynamically adjusted to adapt to the diffusion differences between the discharge area and the ventilation area.
[0072] Finally, the constructed multimodal 3D grid data structure includes: spatial coordinates of each grid cell, temperature value, MFCC feature vector (13 dimensions) and 、 Concentration and temperature and humidity parameters are stored in a structured database with Morton index as the primary key.
[0073] S22: From the temperature grid data in the spatiotemporal gridded multimodal dataset, calculate the temperature difference between each grid according to the temperature value change of adjacent spatial grid cells, and extract the temperature gradient information;
[0074] Extract the audio clips in each time window from the voiceprint grid data, perform pre-emphasis, frame division, windowing, fast Fourier transform and Mel filtering on the audio clips, calculate the Mel frequency cepstral coefficients, and obtain the voiceprint MFCC features;
[0075] Extract ammonia and carbon dioxide concentration data from the environmental parameter spatial grid data, and use a spatial interpolation algorithm to reconstruct a continuous gas concentration distribution map in a three-dimensional grid coordinate system to form gas concentration spatial distribution data;
[0076] Constructing the temperature-voiceprint-gas three-dimensional correlation matrix includes fusing the temperature gradient information, voiceprint MFCC features and gas concentration spatial distribution data based on the same spatial index number, and constructing a temperature-voiceprint-gas three-dimensional correlation matrix with the spatial grid unit as the index dimension, the temperature gradient as the first characteristic dimension, the voiceprint MFCC as the second characteristic dimension, and the gas concentration as the third characteristic dimension.
[0077] In an embodiment of the present invention, feature calculation and correlation matrix construction are performed based on the above-mentioned grid data. The temperature gradient is calculated using the central difference method: in each grid unit, the temperature values of the adjacent grids in the front and back directions are taken respectively along the X, Y, and Z directions, and the temperature change rate is calculated to form a three-dimensional gradient vector. The edge grid is supplemented by the forward or backward difference method to avoid the gradient loss problem caused by data truncation. In order to improve the stability of the gradient calculation, the original temperature data is first subjected to a 3×3×3 grid Gaussian filter process to effectively suppress the high-frequency temperature noise generated by the infrared thermal imager.
[0078] Voiceprint features use 13-dimensional MFCC coefficients, extracted from audio data after pre-emphasis, framing (25ms), and Mel filtering. This dimension ensures sufficient information while avoiding the computational burden of high-dimensional differential features, effectively adapting to real-time analysis scenarios.
[0079] In terms of gas concentration characteristics, Because the sensor is significantly affected by humidity, a dynamic temperature and humidity compensation factor is introduced to correct concentration values in real time. This compensation model is based on experimental calibration results. For every 10% increase in humidity, the sensor output drifts by an average of approximately 1.5 ppm. A 3D median filter is applied to the corrected concentration distribution to remove abnormal spikes and improve the stability of subsequent modeling.
[0080] The above three types of features - temperature gradient vector (3D), MFCC features (13D) and gas concentration parameters (2D: and ) for unified grid index binding, ultimately generating a three-dimensional temperature-voiceprint-gas correlation matrix with dimensions N × (3 + 13 + 2) = N × 18, where N is the number of valid grids. This matrix is stored in the Compressed Sparse Row (CSR) format, which saves storage space and supports fast retrieval.
[0081] The constructed three-dimensional correlation matrix not only retains the spatial distribution characteristics of each perception dimension, but also has good data consistency, sparsity and fusion, providing a high-quality input basis for subsequent temperature and humidity compensation, feature weighting and LSTM model prediction.
[0082] S3: Based on the temperature-voiceprint-gas three-dimensional correlation matrix, the animal surface temperature is dynamically compensated in combination with the temperature and humidity index, and the gas concentration is input as a weight factor into the LSTM time series prediction model to output the stress risk level and the pollution risk level.
[0083] S31: Using the temperature gradient information in the temperature-voiceprint-gas three-dimensional correlation matrix and the temperature and humidity values in the environmental parameter spatial grid data, the temperature and humidity index value of each spatial grid unit is calculated according to the temperature and humidity index calculation model. Based on the difference between the index and the set reference value, the temperature gradient is dynamically corrected in combination with the preset compensation coefficient to obtain a temperature data matrix after temperature and humidity compensation;
[0084] The ammonia and carbon dioxide concentration values of each grid cell are extracted from the spatial distribution data of gas concentrations. After normalization, they are used as weight factors to weight the Mel-frequency cepstral coefficients in the voiceprint features to generate a weighted voiceprint feature matrix reflecting the degree of gas pollution impact.
[0085] The temperature data matrix that has been compensated for temperature and humidity, the weighted voiceprint feature matrix reflecting the degree of gas pollution impact, and the spatial distribution data of gas concentration are spliced according to a unified spatial grid index to construct a time series input vector containing temperature, voiceprint and gas multimodal information. The vector is then arranged according to the set time window order to form a multimodal time series input sequence with temporal continuity and spatial consistency.
[0086] In the embodiment of the present invention, a temperature and humidity index (THI) is constructed based on the temperature gradient data extracted by infrared thermal imaging and the ambient temperature and humidity information collected by the gas sensor to perform dynamic heat stress compensation. Specifically, a modified THI calculation model suitable for modeling the surface microenvironment of animals is adopted, and an interaction term between body surface temperature and relative humidity is introduced to improve the compensation accuracy. Based on the THI index, combined with the set optimal temperature and humidity threshold (such as = 68) and critical stress threshold (e.g. = 78), the original temperature gradient is dynamically adjusted by calculating a normalized compensation coefficient, thereby generating a compensated temperature data matrix that reflects the animal's actual heat stress state. To accommodate individual differences among livestock species, breed compensation factors (e.g., for cattle and pigs) can be further introduced to enhance the model's versatility.
[0087] In terms of voiceprint feature processing, Mel-frequency cepstral coefficients (MFCC) are used as representative behavioral voiceprint features. 13-dimensional static voiceprint features are extracted by pre-emphasis, frame windowing, Mel filtering and cepstral transformation of audio frames. On this basis, ammonia () and carbon dioxide () in the environment are introduced. ) concentration as a weighting factor to dynamically weight the MFCC feature matrix. Logarithmic transformation and min-max scaling are used in the weight normalization process to enhance sensitivity in low-concentration areas and avoid weight saturation. This processing method improves the sensitivity of voiceprint features to animal behavioral responses during periods of high pollution, enhancing the recognizability of stress behaviors. Experimental verification also maintains the structural stability of the voiceprint features themselves.
[0088] Finally, the temperature gradient (with THI compensation), weighted voiceprint features, and spatial distribution data of gas concentrations are concatenated according to a unified spatial grid index to form a multidimensional, multimodal feature vector. This vector is then sorted and ordered according to a set time window (e.g., 10 minutes with a 2-minute step), constructing a multimodal time series input sequence with temporal continuity and spatial consistency. The input tensor has the form N × T × F, where N is the number of grid cells, T is the time step, and F is the concatenated feature dimension (e.g., 18 dimensions: 3-dimensional temperature gradient, 13-dimensional MFCC features, and 2-dimensional gas concentration). This sequence provides high-quality input for subsequent LSTM time series modeling.
[0089] S32: Inputting the multimodal time series input sequence constructed according to the unified spatial index and time window into a neural network model including a long short-term memory unit, wherein the neural network model structure includes an input layer for receiving the input sequence, a multi-layer bidirectional LSTM network layer for capturing time dependencies, and a classification and regression joint output layer for outputting results;
[0090] The neural network model is used to extract time series features and perform state modeling on the input sequence to obtain an output value of the animal stress risk level for each spatial grid unit in the current time window, wherein the stress risk level is divided into multiple level categories according to preset standards;
[0091] The regression output channel in the neural network model predicts and generates a pollution risk score for each spatial grid cell within the corresponding time window. The pollution risk score is a continuous risk index obtained by fitting the gas pollution characteristics and the voiceprint disturbance characteristics.
[0092] Finally, the corresponding animal stress risk level label and pollution risk level score in each time window are output.
[0093] In an embodiment of the present invention, based on the constructed multimodal time series input, a long short-term memory neural network (LSTM) is used to establish a prediction model to output the animal stress risk level and the breeding environment pollution risk level.
[0094] In this example, the LSTM model uses a two-layer structure, with each layer containing 128 hidden units. It also uses a bidirectional time unfolding mechanism to capture both historical behavioral changes and potential signals of future trends. The model input is a batch tensor with dimensions N × T × F, where T is the time window length (e.g., 5 steps) and F is the feature dimension. The model uses a softmax function to output a classification result of the animal's stress risk level (three levels: Level I Normal, Level II Warning, and Level III Emergency). It also uses a sigmoid function to output a regression score for the contamination risk level, ranging from [0 to 1]. A high-risk warning is triggered when the contamination risk score exceeds a set threshold (e.g., 0.8).
[0095] To improve the predictive model's adaptability to class imbalance, the stress risk classification task uses Focal Loss training, while the contamination regression task uses Huber Loss, optimizing both model accuracy and robustness. During training, data augmentation strategies in both the temporal and spatial domains are introduced. For example, cropping the original time window, injecting noise, or randomly masking portions of spatial grid data enhances the model's tolerance to missing data or sensor anomalies.
[0096] Considering the limited edge computing resources in livestock farms, this embodiment uses knowledge distillation technology to compress the original LSTM model, converting the teacher model into a lightweight student model (such as a unidirectional LSTM with 64 units), and deploying it on NVIDIA Jetson-type devices. The end-to-end inference latency can be controlled within 500 milliseconds, meeting real-time monitoring needs.
[0097] S4: Generate and execute corresponding environmental control strategies according to the stress risk level and the pollution risk level.
[0098] Based on the assessment results of stress risk level and pollution risk level, an environmental control strategy corresponding to the current livestock house environmental status is generated, and corresponding environmental adjustment operations are performed according to the generated environmental control strategy.
[0099] In the embodiment of the present invention, in order to achieve efficient environmental control driven by multi-source perception, a combined strategy mapping mechanism based on animal stress risk level and air pollution risk score is proposed. Divided into Level I (normal), Level II (warning) and Level III (emergency), pollution risk score The risk level is a continuous value ranging from 0 to 1, categorized by risk level into three intervals: low (<0.6), medium (0.6-0.8), and high (≥0.8). The two are not bound together in the decision-making logic, meaning that the control strategy must consider all possible combinations of scenarios, not just those with simultaneous high values. This mechanism enables proactive decision-making and response based on both spatial and temporal perception of risk status.
[0100] In the specific strategy design, and The nine combinations of conditions are summarized into three control modes: energy-saving maintenance mode, active adjustment mode and emergency intervention mode. is level I and <0.6, the system maintains the most basic operating state, the ventilation speed is kept at the lowest gear (such as 0.5m / s), the light source color temperature is controlled to warm color temperature (3000K), and the brightness is controlled within 30%. If the pollution score rises to Level II or reaches medium, the active adjustment mechanism will be activated, including adjusting the directional ventilation angle (±30°) according to the thermal map gradient and increasing the wind speed to 1.5m / s, switching to a cool color temperature light source (4500K) and increasing the brightness to 70% to alleviate possible mild heat stress or pollution accumulation. If the stress risk reaches Level III or the pollution score exceeds 0.8, the emergency intervention mechanism is triggered, implementing full-power ventilation, increasing the lighting to full brightness of cool white (5000K), and increasing exhaust ventilation in local areas to ensure the rapid reduction of animal stress and the concentration of harmful gases in the air. In addition, considering the possibility of stress risk being Level III but the pollution score still being low (such as stress caused by light or noise), the system will still prioritize stress relief actions such as cooling and lighting adjustment, rather than activating the high-power exhaust system to avoid energy waste.
[0101] During the control execution process, all control behaviors are matched with the spatial grid as the basic unit. For example, when a certain area experiences high temperature and high humidity and is accompanied by intensive coughing soundprint signals, the wind speed in the area will be increased first, and the color temperature of the light source will be simultaneously shifted to a cooler color. The system updates the control strategy in a 2-minute cycle, and each strategy execution is adjusted based on the previous feedback indicators. If the environmental parameters do not improve effectively within the specified time window (such as 5 minutes), the system will automatically increase the control level and issue a local alarm to prompt manual inspection. All control behaviors are based on the joint judgment of sensor data to avoid false triggering; for example, the exhaust will only be activated when both the gas concentration and temperature and humidity sensors show abnormalities.
[0102] At the implementation guarantee level, the present invention adopts a modular actuator design to achieve distributed independent control of multiple areas in the livestock house, improving the accuracy of regulation and energy efficiency. At the same time, the system supports the coordination of local edge reasoning and cloud-based policy optimization, with real-time response and low-latency processing on the edge. The cloud regularly evaluates the effectiveness of policy execution and fine-tunes parameters based on multi-day historical data, thereby continuously optimizing the strategy selection for stress and pollution risk control. To ensure stable operation of the equipment, the system also has a fault identification and redundant switching mechanism. For example, when the main fan fails, it automatically switches to the backup unit. When the sensor loses connection, it can perform temporary estimation and compensation through the adjacent data grid to ensure uninterrupted environmental control.
[0103] In summary, based on the dynamic identification of environmental risks, the present invention effectively realizes hierarchical response control in response to different complex environmental scenarios by constructing a combined control logic of stress level and pollution score, and has the technical advantages of high control accuracy, fast response time and excellent resource utilization efficiency.
[0104] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
[0105] Example 2 is the second embodiment of the present invention, which is different from the previous embodiment in that:
[0106] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art or the portion of the current technical solution, can be embodied in the form of a software product. The current computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0107] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0108] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, and then editing, interpreting, or processing in another suitable manner as necessary, and then storing it in a computer memory.
[0109] Example 3 is an embodiment of the present invention, which provides a livestock breeding house environment detection system, including a data acquisition and synchronization module, a multimodal grid modeling module, a risk identification and prediction module, and a strategy generation and execution module.
[0110] Data acquisition and synchronization module: deploys infrared thermal imaging devices, voiceprint sensors, and gas sensor arrays to synchronously collect raw thermal image data, raw audio signals, and raw gas data to generate a spatiotemporally aligned multimodal dataset;
[0111] Multimodal grid modeling module: performs spatiotemporal gridding on the multimodal dataset, extracts temperature gradients, voiceprint MFCC features, and gas concentration spatial distribution, and constructs a temperature-voiceprint-gas three-dimensional correlation matrix;
[0112] Risk identification and prediction module: Based on the temperature-voiceprint-gas three-dimensional correlation matrix, the temperature and humidity index are combined to dynamically compensate for the animal's surface temperature, and the gas concentration is input as a weight factor into the LSTM time series prediction model to output the stress risk level and pollution risk level;
[0113] Strategy generation and execution module: generates and executes corresponding environmental control strategies according to the stress risk level and pollution risk level.
Claims
1. A method for detecting the environment of a livestock breeding house, characterized in that: include: Deploy infrared thermal imaging devices, voiceprint sensors, and gas sensor arrays to synchronously collect raw thermal image data, raw audio signals, and raw gas data to generate a spatiotemporally aligned multimodal dataset; Performing spatiotemporal gridding on the multimodal dataset, extracting temperature gradients, voiceprint MFCC features, and gas concentration spatial distribution, and constructing a temperature-voiceprint-gas three-dimensional correlation matrix; Based on the temperature-voiceprint-gas three-dimensional correlation matrix, the animal's surface temperature is dynamically compensated in combination with the temperature and humidity index, and the gas concentration is input as a weight factor into the LSTM time series prediction model to output the stress risk level and the pollution risk level; Generate and execute corresponding environmental control strategies according to the stress risk level and pollution risk level; The method comprises dynamically compensating the animal's surface temperature based on the temperature-voiceprint-gas three-dimensional correlation matrix in combination with the temperature and humidity index, and inputting the gas concentration as a weight factor into the LSTM time series prediction model, including utilizing the temperature gradient information in the temperature-voiceprint-gas three-dimensional correlation matrix and the temperature and humidity values in the environmental parameter spatial grid data, calculating the temperature and humidity index value of each spatial grid unit according to the temperature and humidity index calculation model, and dynamically correcting the temperature gradient based on the difference between the index and the set reference value in combination with a preset compensation coefficient to obtain a temperature data matrix after temperature and humidity compensation; The ammonia and carbon dioxide concentration values of each grid cell are extracted from the spatial distribution data of gas concentrations. After normalization, they are used as weight factors to weight the Mel-frequency cepstral coefficients in the voiceprint features to generate a weighted voiceprint feature matrix reflecting the degree of gas pollution impact. The temperature data matrix that has been compensated for temperature and humidity, the weighted voiceprint feature matrix reflecting the degree of gas pollution impact, and the spatial distribution data of gas concentration are spliced according to a unified spatial grid index to construct a time series input vector containing temperature, voiceprint, and gas multimodal information. These vectors are then arranged according to the set time window sequence to form a multimodal time series input sequence with temporal continuity and spatial consistency. Outputting the stress risk level and the contamination risk level includes inputting a multimodal time series input sequence constructed according to a unified spatial index and a time window into a neural network model including a long short-term memory unit, wherein the neural network model structure includes an input layer for receiving the input sequence, a multi-layer bidirectional LSTM network layer for capturing time dependencies, and a classification and regression joint output layer for outputting results; The neural network model is used to extract time series features and perform state modeling on the input sequence to obtain an output value of the animal stress risk level for each spatial grid unit in the current time window, wherein the stress risk level is divided into multiple level categories according to preset standards; The regression output channel in the neural network model predicts and generates a pollution risk score for each spatial grid cell within the corresponding time window. The pollution risk score is a continuous risk index obtained by fitting the gas pollution characteristics and the voiceprint disturbance characteristics. Finally, the corresponding animal stress risk level label and pollution risk level score in each time window are output.
2. The livestock breeding house environment detection method according to claim 1, characterized in that: The deployment of infrared thermal imaging devices, voiceprint sensors, and gas sensor arrays includes setting up an array of infrared thermal imaging devices above the livestock house to periodically scan the animal group and collect raw thermal image data; Install directional soundprint sensors in animal activity areas to collect the animals' original audio signals; Deploy multiple gas sensor nodes in animal excretion areas and activity areas to synchronously collect raw gas data; The raw gas data includes ammonia concentration, carbon dioxide concentration, and temperature and humidity parameters.
3. The livestock breeding house environment detection method according to claim 2, characterized in that: The generating of the spatiotemporally aligned multimodal data set includes performing spatial coordinate calibration on the original thermal image data to obtain a temperature distribution data stream marked with spatial coordinates; Extract the sound source orientation features from the original audio signal collected by the voiceprint sensor to obtain a voiceprint data stream with spatial orientation features; Perform spatial mapping on the raw gas data collected by the gas sensor node to obtain a data stream of environmental parameters with coordinate markings; Perform time stamp synchronization on the temperature distribution data stream, voiceprint data stream, and environmental parameter data stream to construct a multimodal perception data block with a unified time window and spatial reference system; The multimodal perception data blocks are subjected to outlier removal, and the data segments exceeding the preset threshold range are screened out to obtain a cleaned spatiotemporally aligned multimodal dataset.
4. The livestock breeding house environment detection method according to claim 3, characterized in that: The spatiotemporal gridding of the multimodal data set includes mapping the temperature distribution data stream into a preset three-dimensional spatial grid model, assigning each temperature sampling point to a corresponding spatial grid unit, and generating temperature grid data based on a three-dimensional grid structure; The voiceprint data stream is divided into regions according to the azimuth and distance information of the sound source, and mapped to the corresponding spatial grid area to obtain the voiceprint grid data; The environmental parameter data stream is distributed to each spatial grid node based on the physical location of the sensor node to obtain the environmental parameter spatial grid data; The temperature grid data, voiceprint grid data and environmental parameter space grid data are aligned by grid numbering to generate a spatiotemporal gridded multimodal dataset with a unified spatial index structure.
5. The livestock breeding house environment detection method according to claim 4, characterized in that: The extraction of temperature gradient, voiceprint MFCC features and gas concentration spatial distribution includes calculating the temperature difference between each grid according to the temperature value change of adjacent spatial grid cells from the temperature grid data in the spatiotemporal grid multimodal data set, and extracting temperature gradient information; Extract the audio clips in each time window from the voiceprint grid data, perform pre-emphasis, frame division, windowing, fast Fourier transform and Mel filtering on the audio clips, calculate the Mel frequency cepstral coefficients, and obtain the voiceprint MFCC features; Extract ammonia and carbon dioxide concentration data from the environmental parameter spatial grid data, and use a spatial interpolation algorithm to reconstruct a continuous gas concentration distribution map in a three-dimensional grid coordinate system to form gas concentration spatial distribution data; The construction of the temperature-voiceprint-gas three-dimensional association matrix includes fusing the temperature gradient information, the voiceprint MFCC features and the gas concentration spatial distribution data based on the same spatial index number, and constructing a temperature-voiceprint-gas three-dimensional association matrix with the spatial grid unit as the index dimension, the temperature gradient as the first characteristic dimension, the voiceprint MFCC as the second characteristic dimension, and the gas concentration as the third characteristic dimension.
6. The livestock breeding house environment detection method according to claim 1, characterized in that: The generating and executing the corresponding environmental control strategy includes generating an environmental control strategy corresponding to the current livestock house environmental state based on the evaluation results of the stress risk level and the pollution risk level, and executing corresponding environmental adjustment operations according to the generated environmental control strategy.
7. A livestock breeding house environment detection system for implementing the livestock breeding house environment detection method according to any one of claims 1 to 6, characterized in that: include: Data acquisition and synchronization module: deploys infrared thermal imaging devices, voiceprint sensors, and gas sensor arrays to synchronously collect raw thermal image data, raw audio signals, and raw gas data to generate a spatiotemporally aligned multimodal dataset; Multimodal grid modeling module: performs spatiotemporal gridding on the multimodal dataset, extracts temperature gradients, voiceprint MFCC features, and gas concentration spatial distribution, and constructs a temperature-voiceprint-gas three-dimensional correlation matrix; Risk identification and prediction module: Based on the temperature-voiceprint-gas three-dimensional correlation matrix, the temperature and humidity index are combined to dynamically compensate for the animal's surface temperature, and the gas concentration is input as a weight factor into the LSTM time series prediction model to output the stress risk level and pollution risk level; Strategy generation and execution module: generates and executes corresponding environmental control strategies according to the stress risk level and pollution risk level.
Citation Information
Patent Citations
A gas sensor array data fusion method and system
CN119760652A