Sound recognition system for sheep ingestion behavior

By designing a sound recognition system that integrates sound acquisition, voice enhancement, voiceprint feature extraction and multimodal classification, the problem of monitoring of sheep feeding behavior in complex environments is solved, accurate identification and privacy protection are achieved, and the needs of modern intelligent ranches are met.

CN120220699AActive Publication Date: 2025-06-27ANHUI AGRICULTURAL UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510331294.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-27
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively monitor sheep feeding behavior in complex environments, and lacks effective privacy protection and efficient energy supply solutions, which cannot meet the needs of modern intelligent ranches for precise monitoring and scientific management.

Method used

A sound recognition system for sheep feeding behavior is designed, including a sound acquisition module, a voice enhancement module, a voiceprint feature extraction module, a multimodal classification module and a data fusion unit. The deep neural network and convolutional neural network are used to combine acceleration sensor data and environmental perception to achieve accurate identification and analysis of sheep feeding behavior. At the same time, technologies such as soundprint desensitization schemes based on homomorphic encryption and flexible piezoelectric fiber arrays are adopted to ensure data privacy protection and self-energy of the equipment.

Benefits of technology

It realizes accurate identification and analysis of sheep feeding behavior in complex environments, ensures the privacy protection of sheep voiceprint data, improves the energy utilization efficiency of equipment, and meets the needs of modern intelligent ranches for precise monitoring and scientific management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220699A_ABST
    Figure CN120220699A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice recognition, in particular to a voice recognition system for sheep feeding behaviors. According to the technical scheme, the system comprises a sound collection module, a voice enhancement module, a voiceprint feature extraction module, a multi-mode classification module and a data fusion unit, the sound collection module is a wearable device deployed at the neck of a sheep, and the wearable device is provided with a wind noise resistant directional microphone array and is used for collecting environment sound signals in real time; and the voice enhancement module is connected with the sound acquisition module. According to the invention, accurate identification and analysis of the sheep feeding behavior are realized, the problems of sound signal processing, individual and group monitoring, privacy protection and energy supply in a complex environment are effectively solved, the intelligent level and efficiency of pasture management are improved, the labor cost is reduced, the health problem of sheep can be found in time, the health of sheep flock is guaranteed, and the economic benefit is improved. Meanwhile, data privacy of the sheep is protected, and the energy utilization efficiency of equipment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sound recognition, and particularly to a sound recognition system for sheep feeding behavior. Background Art

[0002] In the modern large-scale livestock farming mode, the number of sheep is numerous. The traditional method of relying on manual observation of sheep feeding behavior faces many challenges. On the one hand, manual monitoring not only consumes a large amount of manpower, material resources and time, with low efficiency, but is also easily affected by subjective factors, making it difficult to achieve accurate and real-time monitoring. For example, in large-scale pastures, it is difficult for breeders to observe the feeding status of each sheep simultaneously, and it is easy to miss the abnormal feeding behavior of sheep, and it is impossible to detect the health problems of sheep in time.

[0003] On the other hand, the pasture environment where sheep live is complex and changeable, with various noise interferences. The sounds of wind and rain in the natural environment, as well as the mechanical equipment sounds in the pasture, etc., will interfere with the feeding sounds of sheep, making it difficult to obtain clear and accurate feeding sound signals. This poses a huge obstacle to monitoring sheep feeding behavior based on sound recognition technology, and traditional sound collection and processing technologies are difficult to effectively extract the feeding sound characteristics of sheep in such a complex environment.

[0004] In addition, with the continuous improvement of data security and privacy protection awareness, data such as the voiceprint of individual sheep contains important information. How to ensure data security and privacy in the process of processing and analyzing these data has become a key issue. At present, the privacy protection technology for sheep voiceprint data is not yet mature in the application of animal husbandry, and there is a lack of effective solutions.

[0005] At the same time, the energy supply of sheep wearable devices is also an urgent problem to be solved. If the device needs to be charged or the battery needs to be replaced frequently, it will bring great inconvenience to the actual use, increasing the breeding cost and management difficulty. Therefore, developing efficient self-powered technologies to enable wearable devices to operate continuously and stably is crucial for realizing long-term monitoring of sheep feeding behavior.

[0006] In the prior art, although there are some studies on animal behavior monitoring, there are few systems specifically for sound recognition of sheep feeding behavior that can effectively address the above complex problems. Most technologies have deficiencies in aspects such as complex environment adaptability, privacy protection, and energy supply, and cannot meet the requirements of modern intelligent pastures for accurate monitoring and scientific management of sheep. There is an urgent need for an innovative sound recognition system for sheep feeding behavior to solve the actual problems in the current livestock breeding process and promote the development of intelligent breeding technologies.

[0007] In summary, the present application proposes a sound recognition system for sheep feeding behavior. Summary of the Invention

[0008] The object of the present invention is to address the problem in the background art that the needs of modern intelligent pastures for precise monitoring and scientific management of sheep cannot be met, and to propose a sound recognition system for sheep feeding behavior.

[0009] The technical solution of the present invention: A sound recognition system for sheep feeding behavior, comprising:

[0010] A sound acquisition module, deployed on a wearable device around the neck of a sheep, configured with an anti-wind-noise directional microphone array for real-time acquisition of ambient sound signals;

[0011] A voice enhancement module, connected to the sound acquisition module, using a voice noise reduction algorithm based on a deep neural network to perform environmental noise suppression and voice feature enhancement processing on the original audio signal;

[0012] A voiceprint feature extraction module, integrating an improved Mel-frequency cepstral coefficient (MFCC) and a convolutional neural network (CNN) to extract voiceprint feature vectors with individual differences from the enhanced voice signal;

[0013] A multi-modal classification module, including a pre-trained voice classification model and a voice retrieval unit. The voice classification model constructs a feeding voiceprint feature library through a contrastive learning mechanism, and uses an attention mechanism to dynamically allocate voice feature weights. The voice retrieval unit performs similarity matching between the real-time audio and the feature library through a dynamic time warping (DTW) algorithm;

[0014] A data fusion unit, which fuses the voiceprint recognition result with the acceleration sensor data and generates a feeding behavior probability output through a Bayesian inference algorithm.

[0015] Optionally, the voice enhancement module includes:

[0016] A dual-modal noise reduction unit: constructs a dual-path recurrent convolutional network (DPRNN), where the first path processes the time-frequency masking of the original audio signal, using a complex-domain mask estimation technique to retain the voice phase information; the second path synchronously receives the 3-8 Hz chewing vibration signal collected by the acceleration sensor, and eliminates the resonance noise caused by mechanical vibration through an adaptive notch filter, and its transfer function is:

[0017]

[0018] where, f c is the notch center frequency dynamically adjusted according to the chewing vibration frequency detected by the acceleration sensor in real time, F s is the audio sampling rate, r is the damping coefficient, r = 0.9 controls the bandwidth, z -1 is the unit delay operator;

[0019] Environmental perception sub-module: Deploy a lightweight random forest classifier to identify the pasture environment type (such as open air / shed) based on temperature and humidity sensor data, and dynamically switch the noise reduction mode: When the humidity > 70%, enable the high-frequency compensation filter to increase the signal-to-noise ratio in the 4 - 6 kHz frequency band by more than 3 dB;

[0020] Speech codec: Adopt the Vector Quantized Variational Autoencoder (VQ-VAE) to compress the 24 kHz sampled audio into an 8 kbps bitstream. When training its codebook, introduce an adversarial spectral loss function:

[0021]

[0022] Among them, Mel k (x) is the energy of the k-th frequency band of the Mel spectrum of the original audio, is the energy value of the reconstructed audio in the k-th frequency band on the Mel scale, and K is the total number of Mel filter banks.

[0023] Ensure that the reconstruction error in the key frequency band (200 - 4000 Hz) < 2 dB, and the speech intelligibility is maintained at MOS ≥ 4.0 when the compression rate is increased by 50%.

[0024] Optionally, the voiceprint feature extraction module specifically includes:

[0025] Hybrid ResNet: The first branch uses a dynamic Mel filter bank, and its center frequency is adaptively adjusted according to the vocal cord length of individual sheep. The adjustment formula is:

[0026]

[0027] Among them, f m is the m-th center frequency of the standard Mel filter bank, L vocal is the vocal cord feature length calculated by inverting the glottal pulse, L avg and L std are the population average and standard deviation, and 0.05 is the frequency adjustment amplitude scale factor;

[0028] Deformable convolutional layer: Deploy a deformable convolutional kernel (Deformable CNN) in the second branch. Its offset learning module is driven by an attention mechanism to capture the microscopic time-varying features of non-steady feeding sounds, and the deformation range of the convolutional kernel is controlled within ±15% of the sampling points;

[0029] Multi-scale feature distillation unit: Adopt a teacher-student network architecture. The teacher network uses EfficientNet-B3 to extract 256-dimensional high-precision features, and the student network uses a spectral perception distillation loss function: Among them, F T ,FS are the feature vectors output by the teacher network and the student network, Mel(x) is the Mel spectrogram of the original audio, and MSE(F T , F S ) is the mean square error of the feature vectors of the teacher network and the student network, ‖·‖1 is the first-order absolute error in the logarithmic domain of the Mel spectrogram, and α and β are the weight coefficients of the loss terms.

[0030] When the model compression rate reaches 60%, the equal error rate (EER) only increases by 0.8%.

[0031] Optionally, the multi-modal classification module includes:

[0032] Meta-learning classifier: Construct a contrastive learning framework based on the prototype network, and its dynamic feature library update strategy is:

[0033]

[0034] where γ is the retention ratio of the historical prototype weight, N is the number of new samples, γ = 0.95 is the decay factor for controlling the historical prototype weight, and f(x i ) is the feature vector of the new sample. When a new individual is detected, a subclass prototype is automatically created to support incremental learning;

[0035] Gated attention mechanism: Design a spatio-temporal dual attention unit. The time-domain attention weight is calculated through the LSTM hidden state, and the frequency-domain attention uses a learnable wavelet basis function. The fusion formula of the two is:

[0036] w final = σ(W t ·h t + W f ·ψ(f))

[0037] where W t is the time-domain attention transformation matrix, h t is the temporal context feature extracted by the bidirectional LSTM network, W f is the frequency-domain attention transformation matrix, ψ(f) is the output of the Morlet wavelet transform, σ is the Sigmoid function, and the output is compressed to the [0,1] interval to achieve a 30%-50% increase in the weight of key features;

[0038] Adaptive DTW retrieval unit: Propose a segmented dynamic time warping algorithm (Segmented DTW), which divides the audio stream into 50ms segments and introduces a penalty factor when calculating the local similarity matrix:

[0039] D(i,j) = d(i,j) + λ·|i / N - j / M|

[0040] Among them, D(i,j) is the Euclidean distance between frame i and frame j, λ is the temporal misalignment penalty intensity, λ = 0.3 is the temporal alignment penalty coefficient to control the temporal drift tolerance, and N, M are the number of frames of the audio to be matched and the template.

[0041] Optionally, the data fusion unit further includes:

[0042] Multi-source evidence fusion engine: It uses the D-S evidence theory to fuse the voiceprint confidence C v , the periodic characteristic P of the acceleration signal a and the ambient light intensity L e , and its basic probability assignment function is:

[0043]

[0044] Among them, C v is the voiceprint confidence, P a is the periodic characteristic of the acceleration signal, L e is the ambient light intensity, σ v , σ a , σ L is the standard deviation of each sensor data, Z is the normalization factor to ensure that the sum of probabilities is 1, and when the conflict factor K > 0.3, it triggers the voice navigation unit to re-collect data;

[0045] Behavior probability calibration module: It establishes a Bayesian network dynamic inference model, and the nodes include the feeding duration, chewing frequency, and head movement trajectory, and updates the posterior probability in real time through variational inference:

[0046]

[0047] Among them, P prior (Eat) is the initial feeding probability based on historical data, and P(D|Eat) is the probability of observing data D when the feeding behavior occurs;

[0048] ∑ s is the sum over all possible behavior states (such as resting, walking, ruminating);

[0049] The observed data D includes the voiceprint matching degree and the acceleration peak value, and the model update period ≤ 1 second.

[0050] Optionally, it further includes:

[0051] Voice navigation and positioning module: It integrates a UWB / Bluetooth multi-mode positioning engine. When the GPS signal is lost, it uses the RFID tag array deployed in the pasture for position fingerprint matching, and the positioning error < 1.5 meters;

[0052] Spatio-temporal correlation analyzer: Construct a graph neural network (GNN) model, where nodes represent the voiceprint features of individual sheep, and the edge weights are jointly determined by the physical distance and the synchronization of feeding behaviors. When an abnormality is detected in a certain node, it automatically analyzes the changes in the behavior patterns of its 3-hop neighbor nodes to identify early signs of the spread of group diseases;

[0053] Energy optimization unit: Adopt dynamic voltage and frequency scaling (DVFS) technology. When the microphone array detects that the silent period is > 5 seconds, it automatically switches the processor to the low-power mode, and still maintains a real-time response delay < 200 ms when the power consumption is reduced by 70%.

[0054] Optionally, the system further includes:

[0055] Adversarial training module: Inject time-frequency adversarial samples during the model training phase, including band-pass noise pulses (200 - 800 Hz), time-domain stretching (±15%), and frequency perturbation (±50 Hz offset), to improve the robustness of the model in windy and rainy weather, and maintain the recognition accuracy above 89% in extreme environments;

[0056] Privacy protection unit: Adopt a voiceprint desensitization scheme based on homomorphic encryption, and perform random orthogonal projection f enc = f · R on the feature vector f, where R is a random orthogonal matrix locally generated by the edge device;

[0057] Self-powered device: Integrate a flexible piezoelectric fiber array in the wearable device, whose resonance frequency matches the neck movement spectrum of sheep (2 - 5 Hz), the energy conversion efficiency ≥ 18%, and cooperate with a micro-supercapacitor to achieve continuous operation for 30 days without charging.

[0058] Optionally, the adversarial training module includes:

[0059] Adaptive adversarial sample generator: Construct a generative adversarial network (GAN) framework, where the generator adopts a conditional WaveGAN architecture, and dynamically synthesizes time-frequency domain adversarial samples according to the characteristics of the real-time collected environmental noise spectrum (when the signal-to-noise ratio ≤ 10 dB), and its generation strategy is:

[0060]

[0061] where ∈ is the perturbation intensity coefficient, which is adaptively adjusted according to the current environmental noise power (0.1 - 0.3), f env is the dominant frequency of the environmental noise, is the gradient direction of the model with respect to the input x, is the Gaussian noise centered on the main frequency f env of the environmental noise;

[0062] Meta - learning Defense Mechanism: Introduce the MAML meta - learning algorithm in the model fine - tuning stage. Through second - order gradient update, the classifier can quickly adapt to new types of adversarial attacks. Its objective function is:

[0063]

[0064] Among them, is the meta - task for different adversarial attack scenarios, α is the inner - loop parameter update step size, and λ is the constraint parameter change amplitude;

[0065] Multi - scale Robustness Verification Unit: Deploy a cascaded detection network. The first - stage uses STFT time - frequency analysis to detect abnormal energy pulses, and the second - stage extracts high - level semantic features through a pre - trained VGGish network. When the confidence - level difference between the two stages > 0.3, trigger the model - parameter rollback mechanism.

[0066] Optionally, it also includes the collaborative optimization of the privacy - protection unit and the self - power supply device, specifically including:

[0067] Dynamic Key Management: The generation of the random orthogonal matrix R uses a physical unclonable function (PUF) based on the piezoelectric energy waveform. Use the microsecond - level voltage fluctuation sequence {v t} of the piezoelectric fiber array as the entropy source, through the hash - chain operation:

[0068]

[0069] Among them, v t is the microsecond - level voltage fluctuation of the piezoelectric fiber, is rounding to an integer, ‖ is bit - string concatenation, and the projection matrix is updated every 30 minutes, increasing the cracking difficulty to the order of 2 128 magnitude;

[0070] Energy - aware Encryption Scheduling: Establish an energy - state machine model. When the supercapacitor power < 15%:

[0071] Activate the lightweight encryption mode, and only project the first 128 dimensions of the voiceprint feature vector;

[0072] Turn off the online learning function of the adversarial training module;

[0073] Adjust the piezoelectric resonance frequency to the optimal energy - harvesting point (4.2Hz ± 0.5Hz), so that the privacy - protection maintenance rate under low power is ≥ 95% and the battery life is extended by 40%.

[0074] Optionally, it also includes a resonance - frequency self - optimization unit: Construct a closed - loop control system. Monitor the neck - motion acceleration a(t) through a Hall sensor, and dynamically adjust the pre - stress F of the piezoelectric fiber using the particle - swarm algorithm p :

[0075]

[0076] Among them, f res is the desired resonance frequency, T is the period duration for observing neck movement, k(F p ) is the stiffness coefficient related to the prestress, m is the equivalent mass of the piezoelectric oscillator, and the stiffness coefficient k0 is the initial stiffness coefficient, and 0.23 and 1.2 are the stiffness-prestress relationship parameters determined by experiments.

[0077] Compared with the prior art, the present application includes at least one of the following beneficial technical effects:

[0078] The environmental perception sub-module of the voice enhancement module can dynamically switch the noise reduction mode according to environmental types such as the temperature and humidity in the pasture; the adversarial training module injects time-frequency adversarial samples during the model training stage to improve the robustness of the model in extreme weather such as wind and rain, ensuring that the recognition accuracy rate remains above 89% in extreme environments.

[0079] The voiceprint feature extraction module can realize the individual identification of sheep, facilitating the precise management of each sheep. The spatio-temporal correlation analyzer constructs a graph neural network model, which can identify the early signs of the spread of group diseases by analyzing the voiceprint features and the synchronization of feeding behaviors of individual sheep, ensuring the overall health of the flock.

[0080] Through the privacy protection unit, a voiceprint desensitization scheme based on homomorphic encryption is adopted to perform random orthogonal projection on the feature vectors, protecting the privacy of sheep voiceprint data. The dynamic key management further enhances the randomness and security of key generation, improving the privacy protection intensity.

[0081] By integrating a flexible piezoelectric fiber array into the self-powered device, the movement energy of the sheep's neck is converted into electrical energy, and continuous operation for 30 days without charging can be achieved in cooperation with a micro-supercapacitor. The energy optimization unit adopts dynamic voltage and frequency adjustment technology to reduce power consumption while ensuring real-time response. The energy-aware encryption scheduling intelligently adjusts the encryption strategy and functions according to the battery level, extending the battery life of the device.

[0082] The present invention realizes the precise identification and analysis of sheep feeding behaviors, effectively solving problems such as sound signal processing, individual and group monitoring, privacy protection, and energy supply in complex environments. It not only improves the intelligent level and efficiency of pasture management, reduces labor costs, but also can timely detect the health problems of sheep, ensure the health of the flock, protect the privacy of sheep data, and improve the energy utilization efficiency of the device. Description of the Drawings

[0083] Figure 1 It is a schematic structural diagram of a sound recognition system for sheep feeding behaviors. Detailed Embodiments

[0084] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0085] Embodiment 1

[0086] As Figure 1 shown, a voice recognition system for sheep feeding behavior proposed by the present invention includes a voice acquisition module, a voice enhancement module, a voiceprint feature extraction module, a multi-modal classification module, and a data fusion unit. Each module will be described in detail below.

[0087] The voice acquisition module is deployed on a wearable device around the neck of the sheep and is equipped with an anti-wind-noise directional microphone array for real-time acquisition of environmental sound signals; it collects the sounds around the sheep, and the anti-wind-noise design can reduce environmental noise interference, ensuring clear and accurate sound signals are collected, providing reliable data for subsequent analysis.

[0088] The voice enhancement module is connected to the voice acquisition module and uses a voice noise reduction algorithm based on a deep neural network to perform environmental noise suppression and voice feature enhancement processing on the original audio signal; the voice enhancement module includes:

[0089] The dual-modal noise reduction unit: constructs a dual-path recurrent convolutional network (DPRNN), where the first path processes the time-frequency masking of the original audio signal, using complex domain mask estimation technology to retain the voice phase information; the second path synchronously receives the 3-8 Hz chewing vibration signal collected by the acceleration sensor, and eliminates the resonance noise caused by mechanical vibration through an adaptive notch filter, and its transfer function is:

[0090]

[0091] where, f c is the notch center frequency dynamically adjusted according to the chewing vibration frequency detected by the acceleration sensor in real time, F s is the audio sampling rate, r is the damping coefficient, r = 0.9 controls the bandwidth, and z -1 is the unit delay operator; the dual-path design can process noise from different angles, more comprehensively suppress noise, retaining the voice phase information helps to restore a more real voice signal, and eliminating the resonance noise can further improve the purity of the audio.

[0092] Environmental perception sub-module: Deploy a lightweight random forest classifier to identify the pasture environment type (such as open-air / shed) based on temperature and humidity sensor data, and dynamically switch the noise reduction mode: When the humidity > 70%, enable a high-frequency compensation filter to increase the signal-to-noise ratio in the 4 - 6 kHz frequency band by more than 3 dB; Automatically adjust the noise reduction strategy according to different environments to improve the pertinence of the noise reduction effect, increase the signal-to-noise ratio in specific frequency bands in a humid environment, ensure the quality of the voice signal, and improve the system's adaptability to environmental changes.

[0093] Voice codec: Adopt a vector quantization variational autoencoder (VQ-VAE) to compress the 24 kHz sampled audio into an 8 kbps bitstream. When training its codebook, introduce an adversarial spectral loss function:

[0094]

[0095] Among them, Mel k (x) is the energy of the k-th frequency band of the Mel spectrum of the original audio, is the energy value of the reconstructed audio in the k-th frequency band on the Mel scale, and K is the total number of Mel filter banks. Achieve efficient audio compression while ensuring speech intelligibility, reduce data transmission and storage pressure, and the adversarial spectral loss function helps improve the quality of the reconstructed audio and ensure that key information is not lost.

[0096] Ensure that the reconstruction error in the key frequency band (200 - 4000 Hz) < 2 dB, and the speech intelligibility maintains MOS ≥ 4.0 when the compression ratio is increased by 50%. Ensure that at a high compression ratio, the key frequency band information of the speech can still be accurately restored, maintain a high speech intelligibility, and do not affect subsequent speech-based analysis and judgment.

[0097] Voiceprint feature extraction module, integrating an improved Mel-frequency cepstral coefficient (MFCC) and a convolutional neural network (CNN), extracts voiceprint feature vectors with individual differences from the enhanced voice signal; The voiceprint feature extraction module specifically includes:

[0098] Hybrid ResNet: The first branch adopts a dynamic Mel filter bank, whose center frequency is adaptively adjusted according to the vocal cord length of individual sheep. The adjustment formula is:

[0099]

[0100] Among them, f m is the m-th center frequency of the standard Mel filter bank, L vocal is the vocal cord feature length calculated by inverting the glottal pulse, L avg and L std are the population average and standard deviation, and 0.05 is the frequency adjustment amplitude proportionality factor;

[0101] Deformable Convolution Layer: Deploy a deformable convolution kernel (Deformable CNN) in the second branch. Its offset learning module is driven by an attention mechanism to capture the microscopic time-varying features of non-steady feeding sounds. The deformation range of the convolution kernel is controlled within ±15% of the sampling points;

[0102] Multi-scale Feature Distillation Unit: Adopt a teacher-student network architecture. The teacher network uses EfficientNet-B3 to extract 256-dimensional high-precision features. The student network uses a spectral perception distillation loss function: where F T , F S are the feature vectors output by the teacher network and the student network, Mel(x) is the Mel spectrogram of the original audio, MSE(F T , F S ) is the mean square error of the feature vectors of the teacher network and the student network, ‖·‖1 is the first-order absolute error in the logarithmic domain of the Mel spectrogram, and α, β are the weight coefficients of the loss terms. When the model compression rate reaches 60%, the equal error rate (EER) only increases by 0.8%. While achieving model compression, it maintains a high accuracy, reduces computational resource consumption, improves the system operation efficiency, and enables the system to still maintain good performance under resource constraints.

[0103] Multi-modal Classification Module, including a pre-trained speech classification model and a speech retrieval unit. The speech classification model constructs a feeding soundprint feature library through a contrastive learning mechanism and uses an attention mechanism to dynamically allocate speech feature weights. The speech retrieval unit matches the similarity between the real-time audio and the feature library through the dynamic time warping (DTW) algorithm; Comprehensive use of multiple technologies realizes accurate classification and rapid retrieval of sheep feeding behaviors, improving the accuracy and efficiency of system recognition. The multi-modal classification module includes:

[0104] Meta-learning Classifier: Construct a contrastive learning framework based on a prototype network. Its dynamic feature library update strategy is:

[0105]

[0106] where γ is the retention ratio of the historical prototype weights, N is the number of new samples, γ = 0.95 is the decay factor for controlling the historical prototype weights, f(x i ) is the feature vector of the new sample. When a new individual is detected, a subclass prototype is automatically created to support incremental learning;

[0107] Gated Attention Mechanism: Design a spatio-temporal dual attention unit. The time-domain attention weight is calculated through the LSTM hidden state, and the frequency-domain attention uses a learnable wavelet basis function. The fusion formula of the two is:

[0108] w final= σ(W t ·h t + W f ·ψ(f))

[0109] where W t is the time-domain attention transformation matrix, h t is the temporal context feature extracted by the bidirectional LSTM network, W f is the frequency-domain attention transformation matrix, ψ(f) is the output of the Morlet wavelet transform, σ is the Sigmoid function, compressing the output to the interval [0, 1], achieving a 30%-50% increase in the weights of key features; highlighting key features, increasing the attention to the features of the feeding sound, enhancing the expression ability of features, further improving the accuracy of classification, and making the system's judgment of the feeding behavior more reliable.

[0110] Adaptive DTW retrieval unit: The segmented dynamic time warping algorithm (Segmented DTW) is proposed. The audio stream is segmented into 50 ms segments, and a penalty factor is introduced when calculating the local similarity matrix:

[0111] D(i, j) = d(i, j) + λ·|i / N - j / M|

[0112] where D(i, j) is the Euclidean distance between frame i and frame j, λ is the temporal misalignment penalty intensity, λ = 0.3 is the temporal alignment penalty coefficient, controlling the temporal drift tolerance, and N, M are the number of frames of the audio to be matched and the template. It improves the accuracy of audio matching, effectively controls temporal drift, more accurately measures the similarity between real-time audio and the audio in the feature library, reduces misjudgment, and improves the recognition performance of the system.

[0113] In this embodiment, the data fusion unit fuses the speaker recognition result with the acceleration sensor data and generates the output of the feeding behavior probability through the Bayesian inference algorithm. The data fusion unit further includes:

[0114] Multi-source evidence fusion engine: The D-S evidence theory is used to fuse the speaker confidence C v , the periodic characteristics P a of the acceleration signal, and the environmental light intensity L e . Its basic probability assignment function is:

[0115]

[0116] where C v is the speaker confidence, P a is the periodic characteristics of the acceleration signal, L e is the environmental light intensity, σ v , σ a , σ Lis the standard deviation of the data of each sensor, and Z is the normalization factor to ensure that the sum of probabilities is 1. When the conflict factor K > 0.3, the voice navigation unit is triggered to re-collect data;

[0117] Behavior probability calibration module: Establish a Bayesian network dynamic inference model. The nodes include feeding duration, chewing frequency, and head movement trajectory. The posterior probability is updated in real time through variational inference:

[0118]

[0119] Among them, P prior (Eat) is the initial feeding probability based on historical data, and P(D|Eat) is the probability of observing data D when the feeding behavior occurs;

[0120] ∑ s is the sum over all possible behavioral states (such as resting, walking, ruminating);

[0121] The observed data D includes voiceprint matching degree and acceleration peak value. The model update period ≤ 1 second. By integrating multiple data information, the reliability and accuracy of the judgment of the feeding behavior are improved, and the error and uncertainty brought by single data are reduced. The feeding behavior probability is dynamically updated according to real-time data, which can more accurately reflect the current feeding state of the sheep and provide more real-time and accurate information for pasture management.

[0122] In this embodiment, it further includes:

[0123] Voice navigation and positioning module: Integrate a UWB / Bluetooth multi-mode positioning engine. When the GPS signal is lost, use the RFID tag array deployed in the pasture for location fingerprint matching, and the positioning error < 1.5 meters;

[0124] Spatio-temporal correlation analyzer: Construct a graph neural network (GNN) model. The nodes represent the voiceprint features of individual sheep, and the edge weights are jointly determined by the physical distance and the feeding behavior synchronization. When an abnormality is detected in a certain node, automatically analyze the changes in the behavior patterns of its 3-hop neighbor nodes to identify early signs of the spread of group diseases; it helps to timely detect the possible disease spread risks in the pasture, take preventive and control measures in advance, ensure the health of the sheep flock, and reduce economic losses.

[0125] Energy optimization unit: Adopt dynamic voltage and frequency scaling (DVFS) technology. When the microphone array detects that the silent period > 5 seconds, automatically switch the processor to the low-power mode. When the power consumption is reduced by 70%, the real-time response delay is still maintained < 200 ms. While ensuring the system performance, the power consumption is reduced, the battery life of the wearable device is extended, the frequency of device charging or battery replacement is reduced, and the stability and sustainability of the system are improved.

[0126] Embodiment 2

[0127] Based on Embodiment 1, this embodiment further includes an adversarial training module, a privacy protection unit, and a self-powered device.

[0128] Among them, the adversarial training module: injects time-frequency adversarial samples during the model training stage, including bandpass noise pulses (200 - 800 Hz), time-domain stretching (±15%), and frequency perturbation (±50 Hz offset), to improve the robustness of the model in rainy and windy weather, and maintain the recognition accuracy above 89% in extreme environments; the adversarial training module includes:

[0129] Adaptive adversarial sample generator: constructs a generative adversarial network (GAN) framework, where the generator adopts a conditional WaveGAN architecture, and dynamically synthesizes time-frequency domain adversarial samples according to the spectral characteristics of the ambient noise collected in real time (when the signal-to-noise ratio ≤ 10 dB), and its generation strategy is:

[0130]

[0131] where ∈ is the perturbation intensity coefficient, adaptively adjusted according to the current ambient noise power (0.1 - 0.3), f env is the dominant frequency of the ambient noise, is the gradient direction of the model with respect to the input x, is the Gaussian noise centered on the main frequency f of the ambient noise env ; dynamically generates adversarial samples according to the actual ambient noise, specifically improves the robustness of the model in different noise environments, and enhances the generalization ability of the model.

[0132] Meta-learning defense mechanism: introduces the MAML meta-learning algorithm during the model fine-tuning stage, and enables the classifier to quickly adapt to new types of adversarial attacks through second-order gradient updates. Its objective function is:

[0133]

[0134] where, is the meta-task for different adversarial attack scenarios, α is the inner-loop parameter update step size, and λ is the constraint parameter change amplitude; enables the system to quickly respond to new types of adversarial attacks, improves the security and stability of the model, and ensures that the system can still work properly in the face of changing attack methods.

[0135] Multi-scale robustness verification unit: deploys a cascaded detection network. The first stage uses STFT time-frequency analysis to detect abnormal energy pulses, and the second stage extracts high-level semantic features through a pre-trained VGGish network. When the confidence difference between the two stages > 0.3, a model parameter rollback mechanism is triggered. Effectively detects abnormal situations during the training and operation of the model, timely discovers potential problems and repairs them, and ensures the reliability and accuracy of the model.

[0136] In this embodiment, the privacy protection unit: adopts a voiceprint desensitization scheme based on homomorphic encryption to perform random orthogonal projection f on the feature vector f enc = f·R, where R is a random orthogonal matrix locally generated by the edge device;

[0137] Furthermore, the self-powered device: integrates a flexible piezoelectric fiber array in the wearable device, whose resonance frequency matches the sheep neck movement spectrum (2 - 5 Hz), the energy conversion efficiency ≥ 18%, and cooperates with a micro-supercapacitor to achieve continuous operation for 30 days without charging.

[0138] Embodiment 3

[0139] This embodiment further includes the collaborative optimization of the privacy protection unit and the self-powered device based on Embodiment 2, specifically including:

[0140] Dynamic key management: The generation of the random orthogonal matrix R adopts a physical unclonable function (PUF) based on the piezoelectric energy waveform, and uses the microsecond-level voltage fluctuation sequence {v t} of the piezoelectric fiber array as the entropy source, through hash chain operation:

[0141]

[0142] where v t is the microsecond-level voltage fluctuation of the piezoelectric fiber, is rounding to an integer, ‖ is bit string concatenation, and the projection matrix is updated every 30 minutes, increasing the cracking difficulty to 2 128 order of magnitude; enhancing the randomness and security of key generation, improving the intensity of privacy protection, making it difficult for attackers to crack encrypted information, and better protecting data privacy.

[0143] Energy-aware encryption scheduling: Establish an energy state machine model. When the power of the supercapacitor < 15%:

[0144] Activate the lightweight encryption mode and only project the first 128 dimensions of the voiceprint feature vector;

[0145] Turn off the online learning function of the adversarial training module;

[0146] Adjust the piezoelectric resonance frequency to the optimal energy harvesting point (4.2 Hz ± 0.5 Hz), so that the privacy protection maintenance rate at low power ≥ 95% and the battery life is extended by 40%. Improve the energy harvesting efficiency at low power, extend the device battery life, and at the same time maintain a high level of privacy protection, ensuring the normal operation and data security of the system in the low-power state.

[0147] Resonant frequency self-optimization unit: Construct a closed-loop control system. Monitor the neck movement acceleration a(t) through a Hall sensor, and use the particle swarm algorithm to dynamically adjust the prestress F of the piezoelectric fiber p :

[0148]

[0149] Among them, f res is the resonant frequency to be achieved, T is the period of observing neck movement, k(F p ) is the stiffness coefficient related to prestress, m is the equivalent mass of the piezoelectric oscillator, and the stiffness coefficient k0 is the initial stiffness coefficient, and 0.23 and 1.2 are the stiffness-prestress relationship parameters determined by experiments. Optimize the resonant frequency of the piezoelectric fiber in real time, improve the energy conversion efficiency, further enhance the performance of the self-powered device, and provide more sufficient energy support for the stable operation of the system.

[0150] The above specific embodiments are only several alternative embodiments of the present invention. Based on the technical solution of the present invention and the relevant revelations of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.

Claims

1. A sound recognition system for sheep feeding behavior, characterized in that: include: The sound collection module is deployed on the wearable device on the sheep's neck and is equipped with a wind-noise-resistant directional microphone array to collect environmental sound signals in real time; A speech enhancement module is connected to the sound collection module and uses a speech noise reduction algorithm based on a deep neural network to suppress environmental noise and enhance speech features on the original audio signal; The voiceprint feature extraction module integrates the improved Mel-type cepstral coefficient and convolutional neural network to extract voiceprint feature vectors with individual differences from the enhanced speech signal; A multimodal classification module, including a pre-trained speech classification model and a speech retrieval unit. The speech classification model constructs a feeding voiceprint feature library through a contrastive learning mechanism and dynamically allocates speech feature weights using an attention mechanism. The speech retrieval unit matches the similarity between real-time audio and the feature library through a dynamic time warping algorithm. The data fusion unit fuses the voiceprint recognition results with the acceleration sensor data and generates the probability output of feeding behavior through the Bayesian reasoning algorithm.

2. The sound recognition system for sheep feeding behavior according to claim 1, characterized in that: The speech enhancement module comprises: Dual-mode noise reduction unit: A dual-path recurrent convolutional network is constructed, in which the first path processes the time-frequency masking of the original audio signal and uses complex domain mask estimation technology to retain the speech phase information; the second path synchronously receives the 3-8Hz chewing vibration signal collected by the acceleration sensor and eliminates the resonance noise caused by mechanical vibration through an adaptive notch filter. Its transfer function is: Among them, f c F is the notch center frequency dynamically adjusted according to the chewing vibration frequency detected by the acceleration sensor in real time. s is the audio sampling rate, r is the damping coefficient, r = 0.9 controls the bandwidth, z -1 is the unit delay operator; Environmental perception submodule: deploys a lightweight random forest classifier to identify the pasture environment type based on the temperature and humidity sensor data, and dynamically switches the noise reduction mode: when the humidity is greater than 70%, the high-frequency compensation filter is enabled to improve the signal-to-noise ratio in the 4-6kHz frequency band by more than 3dB; Speech codec: A vector quantized variational autoencoder is used to compress 24kHz sampled audio to an 8kbps bitstream. The adversarial spectrum loss function is introduced during codebook training: Among them, Mel k (x) is the kth frequency band energy of the Mel spectrum of the original audio, To reconstruct the energy value of the kth frequency band of the audio in the Mel scale, K is the total number of Mel filter banks.

3. The sound recognition system for sheep feeding behavior according to claim 1, characterized in that: The voiceprint feature extraction module specifically includes: Hybrid deep residual network: The first branch uses a dynamic Mel filter bank, whose center frequency is adaptively adjusted according to the length of the vocal cords of individual sheep. The adjustment formula is: Among them, f' m is the adjusted center frequency of the mth Mel filter, f m is the mth center frequency of the standard Mel filter bank, L vocal is the characteristic length of the vocal cords calculated by glottal pulse inversion, L avg and L std is the group mean and standard deviation, and 0.05 is the frequency adjustment amplitude scaling factor; Deformable convolution layer: Deformable convolution kernels are deployed in the second branch, and its offset learning module is driven by the attention mechanism to capture the microscopic time-varying characteristics of non-steady-state feeding sounds; Multi-scale feature distillation unit: adopts the teacher-student network architecture. The teacher network uses EfficientNet-B3 to extract 256-dimensional high-precision features, and the student network uses spectrum-aware distillation loss function: Among them, F T ,F S is the feature vector output by the teacher network and the student network, Mel(x) is the Mel spectrum of the original audio, MSE(F T ,F S ) is the mean square error between the teacher network and the student network feature vector, ‖·‖1 is the first-order absolute error in the Mel spectrum logarithm domain, and α, β are the weight coefficients of the loss term.

4. The sound recognition system for sheep feeding behavior according to claim 1, characterized in that: The multimodal classification module comprises: Meta-learning classifier: Construct a contrastive learning framework based on the prototype network, and its dynamic feature library update strategy is: Among them, p new is the updated feature library prototype vector, p old is the prototype vector of the historical feature library, γ is the retention ratio of the historical prototype weight, N is the number of new samples, γ = 0.95 is the attenuation factor for controlling the historical prototype weight, f(x i ) is the feature vector of the newly added sample, and a subclass prototype is automatically created when a new individual is detected; Gated attention mechanism: Design a dual-time and space attention unit. The time domain attention weight is calculated through the LSTM hidden state, and the frequency domain attention uses a learnable wavelet basis function. The fusion formula of the two is: w final =σ(W t ·h t +W f ·ψ(f)) Among them, w final is the final attention weight, W t is the time domain attention transformation matrix, h t is the temporal context feature extracted by the bidirectional LSTM network, W f is the frequency domain attention transformation matrix, ψ(f) is the Morlet wavelet transform output, σ is the Sigmoid function, and the output is compressed to the [0,1] interval; Adaptive DTW retrieval unit: proposes a segmented dynamic regularization algorithm, divides the audio stream into 50ms segments, and introduces a penalty factor when calculating the local similarity matrix: D(i,j)=d(i,j)+λ·|i / Nj / M| Where D(i,j) is the penalized Euclidean distance between frame i and frame j, d(i,j) is the original Euclidean distance between frame i and frame j, λ is the penalty intensity for timing misalignment, λ=0.3 is the timing alignment penalty coefficient, which controls the tolerance for timing drift, and N, M are the number of frames of the audio to be matched and the template.

5. The sound recognition system for sheep feeding behavior according to claim 1, characterized in that: The data fusion unit further comprises: Multi-source evidence fusion engine: Using DS evidence theory to fuse voiceprint confidence C v , acceleration signal periodic characteristics P a and ambient light intensity L e , its basic probability distribution function is: Among them, m(A) is the basic probability distribution value of evidence A, C v is the voiceprint confidence, P a Acceleration signal periodicity, L e is the ambient light intensity, μ a is the mean value of the periodic characteristic of the acceleration signal, σ v is the standard deviation of the voiceprint confidence data, σ a is the standard deviation of the acceleration signal periodicity data, σ L is the standard deviation of the ambient light intensity data, L ideal is the ideal light intensity reference value, Z is the normalization factor, and the probability sum is guaranteed to be 1. When the conflict factor K>0.3, the voice navigation unit is triggered to re-collect data; Behavior probability calibration module: Establish a Bayesian network dynamic inference model, with nodes including feeding time, chewing frequency and head movement trajectory, and update the posterior probability in real time through variational inference: Where P(Eat|D) is the posterior probability of eating behavior under the observed data D, P prior (Eat) is the initial feeding probability based on historical data, P(D|Eat) is the probability of observing data D when feeding behavior occurs; ∑ s is the sum of all possible behavioral states, s is other behavioral states, including rest, walking, and rumination; The observation data D includes voiceprint matching degree and acceleration peak value, and the model update cycle is ≤1 second.

6. The sound recognition system for sheep feeding behavior according to claim 1, characterized in that: Also includes: Voice navigation and positioning module: integrated with UWB / Bluetooth multi-mode positioning engine. When GPS signal is lost, it uses RFID tag array deployed on the ranch for location fingerprint matching, with positioning error less than 1.5 meters. Spatiotemporal correlation analyzer: Build a graph neural network model, where nodes represent the voiceprint characteristics of individual sheep, and edge weights are determined by physical distance and synchronization of feeding behavior. When an abnormal node is detected, the behavior pattern changes of its 3-hop neighbor nodes are automatically analyzed to identify early signs of the spread of group diseases. Energy Optimization Unit: Using dynamic voltage and frequency adjustment technology, when the microphone array detects a silent period of more than 5 seconds, it automatically switches the processor to low power mode.

7. The sound recognition system for sheep feeding behavior according to claim 1, characterized in that: The system further comprises: Adversarial training module: Inject time-frequency adversarial samples during the model training phase, including bandpass noise pulses, time domain stretching, and frequency perturbations; Privacy protection unit: adopts a voiceprint desensitization scheme based on homomorphic encryption to perform random orthogonal projection f on the feature vector f enc =f·R, where R is a random orthogonal matrix generated locally on the edge device; Self-powered device: Integrating a flexible piezoelectric fiber array into a wearable device, whose resonant frequency matches the spectrum of sheep neck movements.

8. The sound recognition system for sheep feeding behavior according to claim 7, characterized in that: The adversarial training module includes: Adaptive adversarial sample generator: Construct a generative adversarial network framework, in which the generator adopts the conditional WaveGAN architecture to dynamically synthesize time-frequency domain adversarial samples based on the spectrum characteristics of environmental noise collected in real time. Its generation strategy is: Among them, x adv is the generated adversarial sample, x clean is the original clean audio signal, ∈ is the disturbance intensity coefficient, which is adaptively adjusted according to the current environmental noise power, sign(·) is the sign function, and f env is the dominant frequency of environmental noise, is the gradient direction of the model for the input x, θ is the model parameter, y is the true label, The main frequency of the ambient noise is f env is a Gaussian noise with a center and a standard deviation of 0.2; Meta-learning defense mechanism: The MAML meta-learning algorithm is introduced in the model fine-tuning stage to enable the classifier to quickly adapt to new adversarial attacks through second-order gradient updates. Its objective function is: Among them, θ is the initial parameter of the model, is the loss function of the ith meta-task, θ' i is the updated parameter of the inner loop, that is is the meta-task for different adversarial attack scenarios, α is the inner loop parameter update step size, λ is the constraint parameter change amplitude, is the square of the L2 norm; Multi-scale robustness verification unit: deploy a cascade detection network. The first level uses STFT time-frequency analysis to detect abnormal energy pulses. The second level extracts high-level semantic features through a pre-trained VGGish network. When the confidence difference between the two levels is greater than 0.3, the model parameter rollback mechanism is triggered.

9. The sound recognition system for sheep feeding behavior according to claim 7, characterized in that: It also includes the coordinated optimization of the privacy protection unit and the self-powered device, including: Dynamic key management: The random orthogonal matrix R is generated by using a physical unclonable function based on the piezoelectric energy waveform, using the microsecond voltage fluctuation sequence {v t } as entropy source, through hash chain operation: Among them, R t+1 is the updated random orthogonal matrix, SHA256 is the secure hash algorithm, R t is the current random orthogonal matrix, v t is the microsecond voltage fluctuation of the piezoelectric fiber, is rounded to an integer, ‖ is the bit string concatenation, the projection matrix is ​​updated every 30 minutes, and the cracking difficulty is increased to 2 128 Magnitude; Energy-aware encryption scheduling: Establish an energy state machine model. When the supercapacitor power is less than 15%, Activate lightweight encryption mode, projecting only the first 128 dimensions of the voiceprint feature vector; Disable the online learning function of the adversarial training module; Adjust the piezoelectric resonance frequency to the optimal energy collection point, so that the privacy protection maintenance rate at low power is ≥ 95% and the battery life is extended by 40%.

10. The sound recognition system for sheep feeding behavior according to claim 9, characterized in that: It also includes a resonant frequency self-optimization unit: a closed-loop control system is constructed to monitor the neck movement acceleration a(t) through a Hall sensor, and a particle swarm algorithm is used to dynamically adjust the prestress F of the piezoelectric fiber. p : Among them, f res is the desired resonant frequency, T is the period of the observed neck movement, k(F p ) is the stiffness coefficient related to prestress, m is the equivalent mass of the piezoelectric vibrator, and the stiffness coefficient k0 is the initial stiffness coefficient, and 0.23 and 1.2 are the stiffness-prestress relationship parameters determined experimentally.

Citation Information

Patent Citations

  • Pig cough identification method, device and equipment and readable storage medium

    CN113488071A

  • Gunshot detection and identification method and system based on comparative learning pre-training

    CN118262727A

  • Animal voiceprint monitoring method and device, medium and product

    CN119091890A

  • Method and apparatus for automatically identifying animal species from their vocalizations

    US20050049877A1