A ruminant health status evaluation method based on multi-source sensor data fusion
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XINJIANG ACADEMY OF AGRI & RECLAMATION SCI
- Filing Date
- 2026-05-08
- Publication Date
- 2026-08-07
AI Technical Summary
[0005]有鉴于此,本发明实施提供了一种基于多源传感数据融合的反刍动物健康状态评估方法,解决相关技术无法准确和有效地对反刍动物的健康状态进行评估的问题
[0005] In view of this, the present invention provides a method for assessing the health status of ruminants based on multi-source sensor data fusion, which solves the problem that related technologies cannot accurately and effectively assess the health status of ruminants.
Smart Images

Figure CN122531768A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of precision animal husbandry technology, and relates to, but is not limited to, a method for assessing the health status of ruminants based on multi-source sensor data fusion. Background Technology
[0002] Ruminants, such as cattle, sheep, and deer, are a group of special mammals with complex, multi-chambered stomachs, especially the rumen. They can regurgitate partially swallowed food during rest for further chewing, a unique digestive method that allows them to efficiently utilize roughage. Their health not only directly affects their growth, reproduction, and production performance, such as milk or meat production, but is also closely linked to the balance of the vast and complex microbial ecosystem within the rumen, thus influencing feed conversion efficiency, greenhouse gas emissions, and even the economic benefits and public health safety of the entire farm. Therefore, continuous and precise health monitoring of ruminants is not only crucial for ensuring animal welfare and preventing disease outbreaks, but also an essential measure for optimizing husbandry management, ensuring the quality and safety of livestock products, maintaining ecological balance, and promoting the sustainable development of animal husbandry.
[0003] Among the related technologies, the main reliance on a single sensor can only roughly count the duration and frequency of rumination, and cannot distinguish between effective chewing and abnormal chewing patterns. Furthermore, due to the lack of cross-modal fusion analysis of sound and motion data, the accuracy drops significantly when faced with environmental interference or individual differences, making it difficult to meet the needs of precise health assessment.
[0004] Therefore, how to accurately and effectively assess the health status of ruminants has become an urgent problem to be solved. Summary of the Invention
[0005] In view of this, the present invention provides a method for assessing the health status of ruminants based on multi-source sensor data fusion, which solves the problem that related technologies cannot accurately and effectively assess the health status of ruminants.
[0006] According to a first aspect of the present invention, a method for assessing the health status of ruminants based on multi-source sensor data fusion is provided, comprising: The first audio feature sequence and the first motion feature sequence are constructed based on the collected audio data sequence and acceleration data sequence of ruminants, respectively; A feature extraction model incorporating a bidirectional cross-attention mechanism is used to extract features from the first audio feature sequence and the first motion feature sequence in parallel, yielding target feature sequences. Based on noise, stress behavior, and rumination stages, the target feature sequences are modulated. The target audio feature sequence and target motion feature sequence in the modulated target feature sequence are used alternately as query vectors, while another modality feature sequence is used as the key vector and value vector. Finally, a fused feature sequence is obtained through a query vector, key vector, value vector, and a gating fusion mechanism. The chewing features corresponding to each chewing event are extracted from the fused feature sequence; the physiological index values corresponding to the chewing events are determined based on the chewing features; and the physiological index values are predicted by a fully connected network to obtain health classification results, rumen pH value and blood β-hydroxybutyrate concentration; the physiological index values include effective chewing frequency, total number of abnormal chewing times and chewing frequency stability. Health assessment information for ruminants was determined using health classification results, rumen pH, and blood β-hydroxybutyrate concentration.
[0007] According to a second aspect of the present invention, an electronic device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; the memory is used to store at least one executable instruction, wherein the executable instruction causes the processor to perform an operation corresponding to the method of the first aspect.
[0008] According to the scheme provided by the embodiments of the present invention, a first audio feature sequence and a first motion feature sequence are constructed based on the collected audio data sequence and acceleration data sequence of ruminants, respectively; the first audio feature sequence and the first motion feature sequence are extracted in parallel by introducing a feature extraction model with a bidirectional cross-attention mechanism to obtain target feature sequences respectively; the target feature sequences are modulated based on noise, stress behavior and rumination stage, and the target audio feature sequence and the target motion feature sequence in the modulated target feature sequence are used alternately as query vectors, and another modality feature sequence is used as key vector and value vector; and a fused feature sequence is obtained through a query vector, key vector, value vector and gating fusion mechanism. The process extracts chewing features corresponding to each chewing event from the fused feature sequence; determines physiological index values corresponding to the chewing events based on these features; and predicts these physiological index values using a fully connected network to obtain health classification results, rumen pH, and blood β-hydroxybutyrate concentration. The physiological index values include effective chewing frequency, total abnormal chewing count, and chewing frequency stability. Health assessment information for ruminants is determined using the health classification results, rumen pH, and blood β-hydroxybutyrate concentration. In this process, after constructing the first audio feature sequence and the first motion feature sequence, feature extraction is performed using a feature extraction model incorporating a bidirectional cross-attention mechanism. Then, the bidirectional cross-attention mechanism and the gating fusion mechanism are used to weight the audio features, weight the motion features, and finally fuse the two types of weighted features, effectively improving the accuracy of feature representation. By calculating five physical quantities for each chewing event—duration, sound energy, motion energy, chewing frequency, and dominant sound frequency—three core indicators—effective chewing frequency, total abnormal chewing count, and chewing frequency stability—are further obtained. Effective chewing frequency reflects rumination efficiency, chewing frequency stability reflects the regularity of chewing rhythm, and total abnormal chewing count reflects the degree of pain or stress. By inputting these three indicators into a multi-task learning network, the classification branch outputs the health status, and the regression branch outputs the rumen pH value and blood β-hydroxybutyrate concentration. By outputting these two parameters and combining them with the output health classification results, the core state of metabolic health of ruminants can be directly reflected, thereby achieving an accurate assessment of the health status of ruminants. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein: Figure 1 A flowchart illustrating a method for assessing the health status of ruminants based on multi-source sensor data fusion, provided for the implementation of this invention; Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention. Based on the examples in the present invention, all other examples obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0011] In the following description, references are made to “some examples”, which describe a subset of all possible embodiments. However, it is understood that “some examples” may be the same subset or different subset of all possible embodiments and may be combined with each other without conflict.
[0012] It should be noted that the terms "first, second, third" used in the examples of this invention are only used to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" can be interchanged in a specific order or sequence where permitted, so that the examples of this invention described herein can be implemented in an order other than that illustrated or described herein.
[0013] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which these embodiments of the invention pertain. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0014] Figure 1 This is a flowchart illustrating a method for assessing the health status of ruminants based on multi-source sensor data fusion, as provided in an embodiment of the present invention. This method can be executed by an electronic device, which can be a server.
[0015] like Figure 1 As shown, the method for assessing the health status of ruminants based on multi-source sensor data fusion includes: S101. Construct a first audio feature sequence and a first motion feature sequence based on the collected audio data sequence and acceleration data sequence of ruminants, respectively.
[0016] In embodiments of the present invention, a triaxial MEMS accelerometer and a high-fidelity electret microphone can be installed on a ruminant. The fixing strap of the triaxial MEMS accelerometer is designed to be adjustable, allowing for preset wearing position references for different ruminants. For example, when the ruminant is a cow, it can be positioned 1-1.5 cm below the mandible of a dairy cow and 1.5-2 cm below the mandible of a beef cow. The accelerometer can collect acceleration data in real time, accurately capturing the chewing motion trajectory. The high-fidelity electret microphone can collect audio data in real time, providing clear chewing sound waves. After collecting the acceleration and audio data, they are time-aligned to ensure a time synchronization error of less than 1 millisecond. The accelerometer's sampling rate can be 10 Hz, and the microphone's sampling rate can be 16 Hz.
[0017] Furthermore, after the accelerometer and microphone collect data in real time, they are stored in a buffer. A sliding window is set, with a window length of 2 seconds and a window step size of 1 second. Whenever the buffer is full of 2 seconds of data, the audio data sequence and acceleration data sequence in the window are preprocessed by data cleaning, outlier handling, and standardization before feature extraction. When extracting features from the audio data sequence, 13-dimensional Mel-frequency cepstral coefficients (MFCC), 1-dimensional short-time energy, 1-dimensional spectral flatness, and 1-dimensional harmonic distortion are extracted from the audio data collected at each time point, for a total of 16 acoustic features. The first audio feature sequence is constructed based on the 16-dimensional acoustic features collected at each time point. At the same time, time-domain features and frequency-domain features are extracted from the acceleration collected at each time point. The time-domain features include 3-dimensional mean, 2-dimensional variance, 3-dimensional kurtosis, 1-dimensional peak acceleration, and 1-dimensional acceleration rate of change. The frequency-domain features include 1-dimensional motion energy, 1-dimensional rate of change of direction, and 1-dimensional dominant frequency, for a total of 11 features. The first motion feature sequence is constructed based on the 11-dimensional features at each time point.
[0018] Among them, the one-dimensional peak acceleration and the one-dimensional acceleration change rate are obtained by calculation. Specifically, the peak acceleration is the maximum absolute value of the mean of the three-dimensional acceleration at each time point; the acceleration change rate is obtained by calculating the absolute value of the difference between the peak accelerations at adjacent time points, that is, the absolute value of the peak acceleration at the current time point minus the peak acceleration at the previous time point is used as the acceleration change rate.
[0019] Among them, the Mel frequency cepstral coefficient is used to distinguish between normal and abnormal chewing; short-time energy is used to determine the timing and intensity of chewing action; spectral flatness describes the flatness of the spectrum, and high flatness indicates that the sound is close to noise, which is used to distinguish normal chewing sound from environmental noise or empty chewing sound; harmonic distortion is used to determine whether the chewing sound is clear and normal, and to help identify abnormal chewing state.
[0020] The mean reflects the average triaxial acceleration during jaw chewing movements and related head movements in ruminants, indicating the reference position of the jaw and the direction of head posture deviation. Variance reflects the fluctuation range of acceleration during jaw chewing movements and related head movements in ruminants, indicating the stability of chewing movements and the intensity of head movements. Kurtosis reflects the sharpness of the acceleration distribution during jaw chewing movements and related head movements in ruminants, used to detect sudden impact actions such as head shaking and neck friction. Peak acceleration reflects the maximum intensity of acceleration during jaw chewing movements and related head movements in ruminants, used to quantify the peak force of chewing and is a core indicator for early detection of rumen acidosis. Acceleration variability reflects the rate of change of acceleration during jaw chewing movements and related head movements in ruminants, used to capture the impact characteristics of stress behavior and distinguish between normal chewing and stressful head shaking. Kinetic energy reflects the overall intensity of the acceleration signal during jaw chewing movements and related head movements in ruminants, indicating the overall activity level of ruminant activity. The rate of directional change reflects the speed of directional changes in the jaw chewing movements and related head actions of ruminants, and is used to capture the jaw turning frequency during chewing. The dominant frequency reflects the main frequency components of the jaw chewing movements and related head actions of ruminants, and is used to determine the regularity of the chewing cycle.
[0021] S102. By introducing a feature extraction model with a bidirectional cross-attention mechanism, features are extracted in parallel from the first audio feature sequence and the first motion feature sequence to obtain target feature sequences. Based on noise, stress behavior and rumination stage, feature modulation is performed on the target feature sequences. The target audio feature sequence and the target motion feature sequence in the modulated target feature sequence are used alternately as query vectors, and the other modality feature sequence is used as key vector and value vector. The fused feature sequence is obtained through query vector, key vector, value vector and gating fusion mechanism.
[0022] In some embodiments of the present invention, the feature extraction model includes a bidirectional cross-attention mechanism, meaning the feature extraction model has two cross-attention branches. First, features are extracted from the first audio feature sequence to obtain the target feature sequence, i.e., the second audio feature sequence. Then, features are extracted from the first motion feature sequence in parallel to obtain the target feature sequence, i.e., the second motion feature sequence. When collecting data from both modalities, the audio data sequence may contain environmental noise, such as mechanical operations and thunderstorms. The acceleration data sequence may contain ruminant stress behaviors, i.e., abnormal reactions of ruminants when stimulated or uncomfortable by external stimuli, including sudden head and body movements such as head shaking, neck rubbing, and agitation. Each ruminant also has a rumination phase, including an initiation phase, a stable phase, and an end phase. During the initiation phase, the chewing frequency and intensity gradually increase; during the stable phase, a high-frequency, stable chewing state is maintained; and during the end phase, the chewing frequency and intensity gradually decrease until they stop. The second audio feature sequence and the second motion feature sequence are modulated using noise, stress behaviors, and interference from the rumination phase to obtain the modulated target audio feature sequence and the modulated target motion feature sequence.
[0023] Furthermore, in the dual cross-attention mechanism branch, one branch uses the target audio feature sequence as the query vector and the target motion feature sequence as the key vector and value vector; in parallel, the other branch uses the target motion feature sequence as the query vector and the target audio feature sequence as the key vector and value vector. Finally, based on the query vectors, key vectors, and value vectors of the two branches, and combined with the gating fusion mechanism, the outputs of the two branches are fused to obtain the target fused feature.
[0024] S103. Extract the chewing features corresponding to each chewing event from the fused feature sequence; determine the physiological index values corresponding to the chewing events based on the chewing features; and predict the physiological index values through a fully connected network to obtain the health classification results, rumen pH value and blood β-hydroxybutyrate concentration; the physiological index values include effective chewing frequency, total number of abnormal chewing times and chewing frequency stability.
[0025] In an embodiment of the present invention, the fused feature sequence is segmented to obtain the chewing features corresponding to each independent chewing event. Then, the parameter value corresponding to each chewing event is predicted based on the chewing features, and a set of physiological index values corresponding to all chewing events is calculated based on the parameter values. These physiological index values include the effective chewing frequency, the total number of abnormal chewing events, and the stability of the chewing frequency.
[0026] Furthermore, the stability of chewing frequency, the total number of abnormal chewing events, and the effective chewing frequency are input into a fully connected network. This fully connected network first extracts common features through two cascaded shared fully connected layers. The first shared fully connected layer contains 32 neurons, and the second shared fully connected layer contains 16 neurons, both using the ReLU activation function. Subsequently, classification and regression branches are connected. The classification branch consists of two fully connected layers: the first fully connected layer contains 8 neurons and uses the ReLU activation function, and the second fully connected layer contains 3 neurons and uses the Softmax activation function. The classification branch outputs a health classification result, including healthy, subclinical abnormal, and clinical abnormal. Subclinical abnormality is characterized by the absence of obvious symptoms but with metabolic risk, while clinical abnormality is characterized by the presence of obvious clinical symptoms.
[0027] Meanwhile, the regression branch consists of two fully connected layers. The first fully connected layer contains 8 neurons and uses the ReLU activation function, while the second fully connected layer contains 2 neurons and directly outputs the predicted values of rumen pH and blood β-hydroxybutyrate concentration.
[0028] S104. Determine the health assessment information of ruminants based on health classification results, rumen pH, and blood β-hydroxybutyrate concentration.
[0029] In embodiments of the present invention, an early warning mechanism can be constructed based on health classification results, predicted values of rumen pH, and blood β-hydroxybutyrate (BHT) concentration. When a ruminant is in a subclinical abnormal state for a certain period of time, such as 2 hours, a subclinical risk warning is triggered, and a feed adjustment plan is determined from a preset adjustment plan library, such as suggesting increasing the roughage ratio by 5% to 8% and adding a buffer. When a ruminant is in a clinical abnormal state, a clinical intervention warning is triggered, and a feed adjustment plan is determined from a preset adjustment plan library based on the clinical abnormality results, rumen pH, and blood BHT concentration. For example, when the rumen pH is 5.4 and the blood BHT concentration is 2.1 mmol / L, the animal should be fasted for 4 hours, given 200 grams of sodium bicarbonate orally, and 500 ml of glucose intravenously, and a veterinarian should be notified to arrive within 1 hour. When an early warning is triggered, the aforementioned health classification results, predicted values of rumen pH and blood BHT concentration, warning information, and corresponding adjustment plan are sent to relevant personnel.
[0030] It is understood that, in the embodiments of the present invention, after constructing the first audio feature sequence and the first motion feature sequence, feature extraction is performed using a feature extraction model that incorporates a bidirectional cross-attention mechanism. Then, the audio features are weighted, the motion features are weighted, and the two types of features are finally fused using the bidirectional cross-attention mechanism and the gating fusion mechanism, respectively, effectively improving the accuracy of feature representation. By calculating five physical quantities—duration, sound energy, motion energy, chewing frequency, and dominant sound frequency—for each chewing event, three core indicators are further obtained: effective chewing frequency, total number of abnormal chewing events, and chewing frequency stability. Effective chewing frequency reflects rumination efficiency, chewing frequency stability reflects the regularity of chewing rhythm, and total number of abnormal chewing events reflects the degree of pain or stress. These three indicators are input into a multi-task learning network. The classification branch outputs the health status, and the regression branch outputs the rumen pH value and blood β-hydroxybutyrate concentration. By outputting these two parameters and combining them with the output health classification results, the core state of metabolic health in ruminants can be directly reflected, thereby achieving an accurate assessment of the health status of ruminants.
[0031] In some embodiments of the present invention, the feature extraction model that introduces a bidirectional cross-attention mechanism in S102 performs feature extraction on the first audio feature sequence and the first motion feature sequence in parallel to obtain the target feature sequence. This can be achieved through the following steps: inputting the first audio feature sequence into a bidirectional long short-term memory network to obtain the target audio feature sequence; and inputting the first motion feature sequence into a convolutional neural network to obtain the target motion feature sequence.
[0032] Specifically, the first audio feature sequence is input into a standard bidirectional long short-term memory network, where the hidden layer can be 64-dimensional, ultimately yielding a 128-dimensional target audio feature sequence. The convolutional neural network contains three convolutional layers with kernel sizes of 3, 5, and 7, a stride of 1, and channel numbers of 32, 64, and 128, respectively, ultimately resulting in a 128-dimensional target motion feature sequence.
[0033] In some embodiments of the present invention, before performing feature modulation on the target feature sequence based on noise, stress behavior, and rumination stage in S102, and alternately using the target audio feature sequence and target motion feature sequence in the modulated target feature sequence as query vectors, and another modal feature sequence as key vector and value vector, the method further includes the following steps: extracting a short-time energy sequence from the first audio feature sequence; calculating the difference between the short-time energy at each time point and the short-time energy at the previous time point in the short-time energy sequence; if there exists a difference corresponding to a target time point that is not less than an absolute noise threshold and the spectral flatness is not less than a region threshold, then the audio feature at the target time point is characterized by noise interference; extracting the peak acceleration at each time point from the first motion feature sequence; if the peak acceleration at the target time point is not less than a stress behavior threshold, then the motion feature at the target time point is characterized by stress behavior interference; fusing the first audio feature sequence and the first motion feature sequence to obtain a fused feature sequence, and determining the rumination stage at each time point through the fused feature sequence.
[0034] Specifically, short-time energy sequences composed of different time points can be extracted from the first audio feature sequence. Then, the difference between the short-time energy at each time point and the short-time energy at the previous time point is calculated to obtain the difference in short-time energy at each time point. An absolute noise threshold, such as 60 dB, is set, and the difference in short-time energy at each time point is compared with 60 dB. If the difference at a time point is not less than the absolute noise threshold, the spectral flatness corresponding to that time point is obtained, and a spectral flatness threshold is set, which can be 0.6. If the spectral flatness is not less than the spectral flatness threshold, it proves that the short-time energy at that time point, i.e., the target time point, is subject to noise interference from the surrounding environment, further characterizing the audio features at that time point as being subject to noise interference. When setting the spectral flatness threshold, the average value and standard deviation of the spectral flatness of ruminants in a healthy state over 72 hours are calculated. After calculating 1.5 times the standard deviation, the product and the average value are summed to obtain the spectral flatness threshold. This 1.5 times is the optimal parameter set by artificial experiments.
[0035] Furthermore, the peak acceleration at each time point is extracted from the first motion feature sequence and compared with a preset stress behavior threshold. If the peak acceleration is not less than the stress behavior threshold, it indicates that the motion feature at that time point is subject to stress behavior interference. The audio and motion features from the first audio feature sequence and the first motion feature sequence are concatenated according to time points to obtain a fused feature sequence. This fused feature sequence is then input into a trained rumination stage prediction model, such as a multilayer perceptron, to obtain the predicted rumination stage at each time point. The stress behavior threshold can be 0.8, which can be the optimal experimental parameter determined through experiments.
[0036] In some embodiments of the present invention, the feature modulation of the target feature sequence based on noise, stress behavior, and rumination stage in S102, and the alternating use of the target audio feature sequence and target motion feature sequence in the modulated target feature sequence as query vectors, and the use of another modal feature sequence as key vector and value vector, can be achieved through the following steps: coupling the first modulation coefficient corresponding to the stress behavior and the second modulation coefficient corresponding to the rumination stage to obtain the first target modulation coefficient, and modulating the second motion feature sequence with the first target modulation coefficient to obtain the target motion feature sequence; coupling the third modulation coefficient corresponding to the noise and the second modulation coefficient to obtain the second target modulation coefficient, and modulating the second audio feature sequence with the second target modulation coefficient to obtain the target audio feature sequence; inputting the target audio feature sequence and the target motion feature sequence into the first cross-attention mechanism, using the target audio feature sequence as the query vector and the target motion feature sequence as the key vector and value vector; and inputting the target audio feature sequence and the target motion feature sequence into the second cross-attention mechanism in parallel, using the target motion feature sequence as the query vector and the target audio feature sequence as the key vector and value vector.
[0037] Specifically, when stress behavior is present, the first modulation coefficient corresponding to the stress behavior can be 0.2; otherwise, it is 1.0. When noise interference is present, the corresponding third modulation coefficient can be 0.2; otherwise, it is 1.0. When the rumination stage is a stable period, the corresponding second modulation coefficient can be 1.5, and the second modulation coefficient for the start and end stages can be 0.7. When modulating the second motion feature sequence, the corresponding modulation coefficient can be selected according to whether stress behavior exists at each time point and the corresponding rumination stage. Then, the second modulation coefficient and the first modulation coefficient corresponding to each time point are multiplied to obtain the first target modulation coefficient. Then, the first target modulation coefficient at each time point is multiplied with the motion feature at each time point to obtain the target motion feature sequence.
[0038] Meanwhile, when modulating the second audio feature sequence, the corresponding modulation coefficient can be determined based on whether there is noise at each time point and the corresponding rumination stage. Then, the third modulation and the second modulation coefficient are multiplied to obtain the second target modulation coefficient. Then, the second target modulation coefficient at each time point is multiplied to obtain the audio feature sequence at each time point in the second audio feature sequence.
[0039] Furthermore, the bidirectional cross-attention mechanism consists of two cross-attention branches. In the first cross-attention mechanism, the target audio feature sequence is used as the query vector, and the target motion feature sequence is used as the key vector and value vector. In the second cross-attention mechanism, the target motion feature sequence is used as the query vector, and the target audio feature sequence is used as the key vector and value vector.
[0040] In some embodiments of the present invention, obtaining the target fusion feature in S102 through query vector, key vector, value vector and gating fusion mechanism can be implemented through the following steps: In the first cross-attention mechanism, a weighted audio feature sequence is determined according to the corresponding query vector, key vector and value vector, and the weighted audio feature sequence and the corresponding query vector are residually connected to obtain an enhanced audio feature sequence; In the second cross-attention mechanism, a weighted motion feature sequence is determined according to the corresponding query vector, key vector and value vector, and the weighted motion feature sequence and the corresponding query vector are residually connected to obtain an enhanced motion feature sequence; The enhanced audio feature sequence and the enhanced motion feature sequence are fused through gating fusion mechanism to obtain a fused feature sequence.
[0041] Specifically, in the first cross-attention mechanism, after calculating the attention weights from the query vector and key vector, the value vectors are weighted and summed using the attention weights to obtain a weighted audio feature sequence. Then, a residual connection is performed between the weighted audio feature sequence and the query vector (which has undergone a linear transformation of the audio feature sequence) to obtain an enhanced audio feature sequence. In the second cross-attention mechanism, after calculating the attention weights from the query vector and key vector, the value vectors are weighted and summed using the attention weights to obtain a weighted motion feature sequence. Then, a residual connection is performed between the weighted motion feature sequence and the query vector (which has undergone a linear transformation of the motion feature sequence) to obtain an enhanced motion feature sequence.
[0042] Furthermore, a gating fusion mechanism is proposed, with the following formula: In the above formula, To fuse feature sequences, For gating fusion coefficient, For the enhanced audio feature sequence, This is the enhanced motion feature sequence.
[0043] In the above formula, For learnable bias vectors, it conforms to This represents a vector concatenation operation. It is the Sigmoid activation function. This is a learnable weight matrix.
[0044] Specifically, in a quiet environment, i.e. when the enhanced audio feature sequence is reliable, the learned... The learning frequency tends towards 1, primarily preserving the enhanced audio feature sequence, supplemented by a small amount of enhanced motion feature sequence information to ensure no loss of sound details; in noisy environments, i.e., when the enhanced audio feature sequence is unreliable: the learned frequency... When the value approaches 0, the enhanced audio feature sequence is automatically suppressed, and the decision is made entirely based on the enhanced motion feature sequence, thereby achieving robust noise-resistant fusion.
[0045] In some embodiments of the present invention, determining the physiological index values of chewing events based on the fused feature sequence in S103 can be achieved through the following steps: determining the energy sequence based on the mean sequence in the fused feature sequence; comparing the energy in the energy sequence with the energy threshold; and taking the feature segments that are continuous and not less than the energy threshold as chewing features; performing global pooling processing on the chewing features to obtain the pooled chewing features; performing prediction processing on the chewing features through a fully connected network to obtain the parameter values corresponding to each chewing event; and determining the physiological index values through the parameter values; the parameter values include duration, sound energy, motion energy, chewing frequency, and sound dominant frequency; and the physiological index values include effective chewing frequency, chewing frequency stability, and total number of abnormal chewing events.
[0046] Specifically, after the gated fusion mechanism outputs the fused feature sequence, a mean sequence is determined in the motion feature sequence. Then, the modulus of the 3D mean at each time point is calculated to obtain the energy at each time point, which reflects the intensity of mandibular movement. The calculated energy is compared with an energy threshold; regions in the energy sequence that are continuously higher than the energy threshold represent the time period of a chewing event. Feature segments corresponding to the time period are extracted from the fused feature sequence. These feature segments are processed by a global average pooling layer to obtain a pooled feature vector. This pooled feature vector is then input into a fully connected network, which includes an input layer, two hidden layers, and an output layer. The number of neurons in the two hidden layers can be 64 and 32, respectively. The pooled feature vector is then input into the fully connected network to obtain chewing events. Each chewing event includes duration, sound energy, motion energy, chewing frequency, and sound dominant frequency.
[0047] In some embodiments of the present invention, determining physiological index values through parameter values can be achieved through the following steps: determining the total duration of all chewing events based on their duration; for each chewing event, when the duration meets a preset time condition, the Pearson correlation coefficient between sound energy and motion energy meets a preset coefficient condition, and the difference between the chewing frequency and the dominant sound frequency meets a preset frequency difference, the corresponding chewing event is considered a valid chewing event; otherwise, it is considered an abnormal chewing event; calculating the effective chewing frequency based on the valid chewing events and the total duration; summing the number of all abnormal chewing events to obtain the total number of abnormal chewing events; and determining the stability of the chewing frequency based on the duration of all chewing events.
[0048] Specifically, based on the duration of each chewing event, the total duration is calculated from the start time of the first chewing event to the end time of the last chewing event. For each chewing event, a time threshold is set, such as 0.7 and 1.5. The duration of each chewing event is compared with the time threshold. If the duration is between 0.7 and 1.5, energy synchronization is then judged. Energy synchronization conditions are set, and the Pearson correlation coefficient between sound energy and motion energy is calculated. The calculated Pearson correlation coefficient is compared with the set preset condition coefficient, such as 0.6. If the Pearson correlation coefficient is not less than 0.6, the absolute value of the difference between the chewing frequency and the main sound frequency is further compared with the preset frequency difference, which can be 0.3. When the absolute value of the difference is not greater than 0.3, the corresponding chewing event is judged as a valid chewing event, that is, a valid chewing event is one that meets all the above judgment conditions; otherwise, it is an abnormal chewing event.
[0049] Furthermore, after the assessment is completed, the ratio of the number of effective chewing events to the total duration is calculated to obtain the effective chewing frequency. The effective chewing frequency reflects rumination activity; a decrease indicates an imbalance in energy metabolism. Then, the number of all abnormal chewing events is summed to obtain the total number of abnormal chewing events. When calculating the stability of the chewing frequency, the average duration of all chewing events is calculated, and then the standard deviation is calculated based on the average. The standard deviation is used as the stability of the chewing frequency.
[0050] Example 1: Data collection was conducted at a large-scale, commercially operated dairy farm with a complete management system and veterinary team. Experimental cattle were selected from the farm's own herd, Holstein breed, aged 2-5 years, weighing 500-700 kg. The experimental samples included three states: healthy, subclinically abnormal, and clinically abnormal, with 30% healthy, 40% subclinically abnormal, and 30% clinically abnormal. Healthy cattle were directly selected from the farm's normal herd; subclinically and clinically abnormal cattle were screened through routine health monitoring, selecting individuals with naturally occurring diseases such as rumen acidosis and hoof diseases for inclusion in the experiment. All cattle remained in their original farm environment to ensure the naturalness and authenticity of the data collection. To address the insufficient coverage of complex clinical samples such as latent rumen acidosis and mixed diseases, naturally occurring individuals meeting the criteria were selected, or subclinical acidosis was induced through feed adjustments to obtain latent samples. Ultimately, the proportion of these two extreme types of samples was no less than 15% of the total sample, ensuring the model's generalization ability in complex clinical scenarios. The collected data were divided into an 8:1:1 ratio. Table 1 shows the test set used to evaluate the final performance of each method. Each method includes a traditional rumination monitor, which outputs rumination duration and rumination frequency, requiring manual judgment based on the data. Single-modal model 1 and single-modal model 2 have the same model structure, which are hybrid models consisting of an LSTM model and a fully connected network, respectively. The only difference is that the input data consists of only a single audio feature sequence and a single motion feature sequence. The experimental results are as follows: Table 1 Prediction results of different methods In Table 1 above, because this invention uses a bidirectional cross-attention mechanism to dynamically interact and enhance audio and motion features, and then uses gated fusion to adaptively balance the contributions of original and fused features, it more accurately captures the subtle differences between effective and abnormal chewing. This provides higher-quality feature input for subsequent fully connected networks to predict the physical quantities of chewing events, thereby improving the accuracy of health status classification and physiological parameter regression. Therefore, its accuracy and false alarm rate are far superior to other methods. By extracting 16-dimensional acoustic features and 11-dimensional motion features from audio and motion data, covering multi-dimensional information in the spectrum, energy, time domain, and frequency domain, and accurately distinguishing between effective chewing and abnormal patterns such as empty chewing through three conditions of time, energy synchronization, and spectrum consistency, this invention achieves a disease risk warning accuracy of 82.7%, far higher than other methods. Because the model of this invention is the most complex, the sample processing speed is slower, but the time difference is small, all in the millisecond range, which fully meets the needs of real-time monitoring. Considering all indicators, this invention can more effectively and accurately assess the state of ruminants.
[0051] Reference Figure 2The diagram shows a structural schematic of an electronic device according to an embodiment of the present invention. The specific examples of the present invention do not limit the specific implementation of the electronic device.
[0052] like Figure 2 As shown, the electronic device may include: a processor 502, a communications interface 504, a memory 506, and a communications bus 508.
[0053] in: The processor 502, communication interface 504, and memory 506 communicate with each other via communication bus 508.
[0054] Communication interface 504 is used to communicate with other electronic devices or servers.
[0055] The processor 502 is used to execute program 510, which can specifically execute the relevant steps in the above-described server-side or user-side method embodiments.
[0056] Specifically, program 510 may include program code that includes computer operation instructions.
[0057] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The smart device may include one or more processors of the same type, such as one or more CPUs; or it may include processors of different types, such as one or more CPUs and one or more ASICs.
[0058] Memory 506 is used to store program 510. Memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0059] Specifically, program 510 can be used to cause processor 502 to perform the operations corresponding to the methods described in the above method embodiments.
[0060] The specific implementation of each step in program 510 can be found in the corresponding descriptions of the steps and units in the above method embodiments, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.
[0061] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of the present invention can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present invention.
[0062] The methods described above according to embodiments of the present invention can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored on a remote recording medium or a non-transitory machine-readable medium and subsequently stored on a local recording medium, downloaded via a network. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.
[0063] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the embodiments of the present invention.
[0064] The above embodiments are only used to illustrate the embodiments of the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the patent protection scope of the embodiments of the present invention should be defined by the claims.
Claims
1. A method for assessing the health status of ruminants based on multi-source sensor data fusion, characterized in that, include: The first audio feature sequence and the first motion feature sequence are constructed based on the collected audio data sequence and acceleration data sequence of ruminants, respectively; By introducing a feature extraction model with a bidirectional cross-attention mechanism, features are extracted in parallel from the first audio feature sequence and the first motion feature sequence to obtain the target feature sequence. Based on noise, stress behavior, and rumination stages, the target feature sequence is modulated. The target audio feature sequence and target motion feature sequence in the modulated target feature sequence are used alternately as query vectors, and another modality feature sequence is used as key vector and value vector. The fused feature sequence is obtained through query vector, key vector, value vector and gating fusion mechanism. The chewing features corresponding to each chewing event are extracted from the fused feature sequence; the physiological index values corresponding to the chewing events are determined based on the chewing features; and the physiological index values are predicted by a fully connected network to obtain health classification results, rumen pH value and blood β-hydroxybutyrate concentration; the physiological index values include effective chewing frequency, total number of abnormal chewing times and chewing frequency stability. Health assessment information for ruminants was determined using health classification results, rumen pH, and blood β-hydroxybutyrate concentration.
2. The method according to claim 1, characterized in that, The feature extraction model includes a bidirectional long short-term memory network and a convolutional neural network; The feature extraction model, which introduces a bidirectional cross-attention mechanism, extracts features from the first audio feature sequence and the first motion feature sequence in parallel, respectively, to obtain the target feature sequence, including: The first audio feature sequence is input into a bidirectional long short-term memory network to obtain the target audio feature sequence; The first motion feature sequence is input into a convolutional neural network to obtain the target motion feature sequence.
3. The method according to claim 1, characterized in that, Before performing feature modulation on the target feature sequence based on noise, stress behavior, and rumination stage, and alternately using the target audio feature sequence and target motion feature sequence in the modulated target feature sequence as query vectors, and another modal feature sequence as key vector and value vector, the method further includes: Extract the short-time energy sequence from the first audio feature sequence; and calculate the difference between the short-time energy at each time point and the short-time energy at the previous time point in the short-time energy sequence; If there exists a difference at the target time point that is not less than the absolute noise threshold and the spectral flatness is not less than the regional threshold, then the audio features representing the target time point are subject to noise interference. The peak acceleration at each time point is extracted from the first motion feature sequence. If the peak acceleration at the target time point is not less than the stress behavior threshold, it indicates that the motion feature at the target time point is subject to stress behavior interference. The first audio feature sequence and the first motion feature sequence are fused to obtain a fused feature sequence, and the rumination stage at each time point is determined by the fused feature sequence.
4. The method according to claim 1, characterized in that, The target feature sequence is a second motion feature sequence and a second audio feature sequence; The feature modulation of the target feature sequence based on noise, stress behavior, and rumination stage, wherein the target audio feature sequence and target motion feature sequence in the modulated target feature sequence are alternately used as query vectors, and another modal feature sequence is used as key vector and value vector, including: The first modulation coefficient corresponding to the stress behavior and the second modulation coefficient corresponding to the rumination stage are coupled to obtain the first target modulation coefficient, and the second motion feature sequence is modulated by the first target modulation coefficient to obtain the target motion feature sequence. The third modulation coefficient and the second modulation coefficient corresponding to the noise are coupled to obtain the second target modulation coefficient, and the second audio feature sequence is modulated by the second target modulation coefficient to obtain the target audio feature sequence. The target audio feature sequence and the target motion feature sequence are input into the first cross-attention mechanism, with the target audio feature sequence as the query vector and the target motion feature sequence as the key vector and value vector. In parallel, the target audio feature sequence and the target motion feature sequence are input into the second cross-attention mechanism, with the target motion feature sequence as the query vector and the target audio feature sequence as the key vector and value vector.
5. The method according to claim 4, characterized in that, The process of obtaining fused features through query vectors, key vectors, value vectors, and a gating fusion mechanism includes: In the first cross-attention mechanism, the weighted audio feature sequence is determined based on the corresponding query vector, key vector and value vector, and the weighted audio feature sequence and the corresponding query vector are residually connected to obtain the enhanced audio feature sequence. In the second cross-attention mechanism, the weighted motion feature sequence is determined based on the corresponding query vector, key vector and value vector, and the weighted motion feature sequence and the corresponding query vector are residually connected to obtain the enhanced motion feature sequence. The enhanced audio feature sequence and the enhanced motion feature sequence are fused using a gating fusion mechanism to obtain a fused feature sequence.
6. The method according to claim 1, characterized in that, The determination of physiological index values corresponding to chewing events based on chewing characteristics includes: The energy sequence is determined based on the mean sequence in the fusion feature sequence. The energy in the energy sequence is compared with the energy threshold, and the feature segments that are continuous and not less than the energy threshold are taken as chewing features. The chewing features are globally pooled to obtain pooled chewing features; the chewing features are then predicted using a fully connected network to obtain the parameter values corresponding to each chewing event. Physiological index values are determined by parameter values, including duration, sound energy, movement energy, chewing frequency, and dominant sound frequency; physiological index values include effective chewing frequency, chewing frequency stability, and total number of abnormal chewing events.
7. The method according to claim 6, characterized in that, The process of determining physiological index values through parameter values includes: The total duration of a chewing event is determined by the duration of all chewing events. For each chewing event, if the duration meets the preset time condition, the Pearson correlation coefficient between sound energy and motion energy meets the preset coefficient condition, and the difference between the chewing frequency and the dominant sound frequency meets the preset frequency difference, the corresponding chewing event is considered a valid chewing event; otherwise, it is considered an abnormal chewing event. The effective chewing frequency is calculated by combining the effective chewing events and the total duration; the total number of abnormal chewing events is obtained by summing the counts of all abnormal chewing events; and the stability of the chewing frequency is determined by the duration of all chewing events.