A sound recognition system for sheep feeding behavior

By deploying wind-noise-resistant microphone arrays and deep neural network processing technology on the necks of sheep, combined with self-powered devices and privacy protection, the problems of environmental noise interference, privacy protection, and energy supply in sheep grazing behavior monitoring have been solved. This has enabled accurate identification and management of sheep grazing behavior, and improved ranch management efficiency and equipment stability.

CN120220699BActive Publication Date: 2025-11-28ANHUI AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510331294.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-11-28
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

Existing technologies are insufficient for accurately monitoring sheep grazing behavior in complex environments, and suffer from problems such as high labor costs, significant noise interference, insufficient privacy protection, and unstable energy supply, thus failing to meet the needs of modern intelligent ranches.

Method used

Sound is collected using a wind-noise directional microphone array. The system combines a deep neural network speech denoising algorithm with an improved Mel-frequency cepstral coefficient and a convolutional neural network to extract voiceprint features. Through multimodal classification and data fusion, the system recognizes sheep feeding behavior and ensures stable operation through a self-powered device and privacy protection technology.

Benefits of technology

It enables accurate identification of sheep grazing behavior in complex environments, improves the level of intelligence in ranch management, reduces labor costs, detects health problems in a timely manner, protects sheep data privacy, and extends the device's battery life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220699B_ABST
    Figure CN120220699B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of sound recognition, and particularly relates to a sound recognition system for sheep feeding behavior. The technical scheme comprises a sound collection module, a voice enhancement module, a voiceprint feature extraction module, a multi-modal classification module, a data fusion unit, the sound collection module is arranged on a wearable device on the neck of a sheep, is provided with an anti-wind-noise directional microphone array, and is used for collecting environmental sound signals in real time; the voice enhancement module is connected with the sound collection module. The present application realizes accurate recognition and analysis of sheep feeding behavior, effectively solves the problems of sound signal processing in a complex environment, individual and group monitoring, privacy protection and energy supply, not only improves the intelligent level and efficiency of pasture management, reduces labor costs, but also can timely find health problems of sheep, protect the health of the sheep, protect the privacy of sheep data, and improve the energy utilization efficiency of equipment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of sound recognition, in particular to a sound recognition system for sheep feeding behavior. BACKGROUND

[0002] Under the modern large-scale breeding mode of livestock farming, the number of sheep is large. The traditional method of observing sheep feeding behavior by artificial observation faces many challenges. On the one hand, artificial monitoring not only consumes a lot of manpower, material resources and time, but also is low in efficiency and is easily affected by subjective factors, making it difficult to achieve accurate and real-time monitoring. For example, in a large-scale pasture, it is difficult for a shepherd to observe the feeding state of each sheep at the same time, and it is easy to miss the abnormal feeding behavior of the sheep and fail to discover the health problems of the sheep in time.

[0003] On the other hand, the pasture environment where the sheep live is complex and changeable, and there are various noise interferences. The sound of wind and rain in the natural environment, the sound of wind, and the sound of mechanical equipment in the pasture will interfere with the feeding sound of the sheep, making it difficult to obtain clear and accurate feeding sound signals. This brings great obstacles to the monitoring of sheep feeding behavior based on sound recognition technology, and the traditional sound collection and processing technology cannot effectively extract the feeding sound features of the sheep in such a complex environment.

[0004] In addition, with the continuous improvement of data security and privacy protection awareness, the voiceprint data of individual sheep contains important information, and how to ensure data security and privacy in the process of processing and analyzing these data has become a key problem. At present, the privacy protection technology for sheep voiceprint data is not mature enough in the application of livestock farming, and there is a lack of effective solutions.

[0005] At the same time, the energy supply of the wearable device of the sheep is also a problem to be solved. If the device needs to be frequently charged or the battery needs to be replaced, it will bring great inconvenience to the actual use and increase the breeding cost and management difficulty. Therefore, it is crucial to develop efficient self-powered technology to enable the wearable device to operate continuously and stably for the long-term monitoring of sheep feeding behavior.

[0006] In the prior art, although there are some researches on animal behavior monitoring, there are few systems that are specifically used for sheep feeding behavior sound recognition and can effectively deal with the above-mentioned complex problems. Most of the technologies have deficiencies in adaptability to complex environments, privacy protection, energy supply, etc., and cannot meet the needs of modern intelligent pasture for accurate monitoring and scientific management of sheep. Therefore, there is an urgent need for an innovative sheep feeding behavior sound recognition system to solve the practical problems in the current livestock breeding process and promote the development of intelligent breeding technology.

[0007] In view of the above, the present application provides a sound recognition system for sheep feeding behavior. SUMMARY

[0008] The purpose of the present application is to address the problem that the existing technology cannot meet the demand of modern intelligentized farms for precise monitoring and scientific management of sheep, and to provide a sound recognition system for sheep feeding behavior.

[0009] The technical scheme of the present application is a sound recognition system for sheep feeding behavior, comprising:

[0010] A sound collection module is arranged on a wearable device on the neck of a sheep, and is configured with an anti-wind-noise directional microphone array for real-time collection of environmental sound signals;

[0011] A voice enhancement module is connected to the sound collection module, and uses a deep neural network-based voice noise reduction algorithm to perform environmental noise suppression and voice feature enhancement processing on the original audio signal;

[0012] A voiceprint feature extraction module integrates an improved Mel frequency cepstrum coefficient (MFCC) and a convolutional neural network (CNN) to extract a voiceprint feature vector with individual differences from the enhanced voice signal;

[0013] A multi-modal classification module includes a pre-trained voice classification model and a voice search unit, the voice classification model constructs a feeding voiceprint feature library through a contrast learning mechanism, and uses an attention mechanism to dynamically allocate voice feature weights, and the voice search unit matches the similarity of real-time audio and the feature library through a dynamic time warping (DTW) algorithm;

[0014] A data fusion unit fuses the voiceprint recognition result with the acceleration sensor data to generate a feeding behavior probability output through a Bayesian inference algorithm.

[0015] Optionally, the voice enhancement module comprises:

[0016] A dual-mode noise reduction unit: a dual-path recurrent convolutional network (DPRNN) is constructed, wherein the first path processes the time-frequency masking of the original audio signal, and uses complex domain mask estimation technology to preserve the phase information of the voice; the second path synchronously receives the 3-8Hz chewing vibration signal collected by the acceleration sensor, and eliminates the resonance noise caused by mechanical vibration through an adaptive notch filter, and the transfer function is:

[0017]

[0018] wherein, is the notch center frequency dynamically adjusted according to the real-time detection of the chewing vibration frequency of the acceleration sensor, is the audio sampling rate, is the damping coefficient, is a unit delay operator;

[0019] Environmental perception submodule: Deploys a lightweight random forest classifier to identify pasture environment type (such as open field / shed) based on temperature and humidity sensor data, and dynamically switches noise reduction modes: when humidity > 70%, a high-frequency compensation filter is enabled to improve the signal-to-noise ratio of the 4-6kHz frequency band by more than 3dB;

[0020] The speech codec employs a vector quantization variational autoencoder (VQ-VAE) to compress 24kHz sampled audio into an 8kbps bitstream. An adversarial spectral loss function is introduced during codebook training.

[0021]

[0022] in, The Mel spectrum of the original audio Energy of each frequency band To reconstruct audio at the Mel scale Energy value of each frequency band This represents the total number of Mel filter banks.

[0023] Ensure reconstruction error of <2dB in key frequency bands (200-4000Hz), and maintain speech intelligibility MOS ≥4.0 when compression rate is increased by 50%.

[0024] Optionally, the voiceprint feature extraction module specifically includes:

[0025] Hybrid ResNet: The first branch uses a dynamic Mel filter bank, whose center frequency is adaptively adjusted according to the vocal cord length of an individual sheep. The adjustment formula is as follows:

[0026]

[0027] in, The m-th center frequency of a standard Mel filter bank. The characteristic length of the vocal cords is calculated by glottic pulse inversion. and The values ​​represent the population mean and standard deviation, and 0.05 is the frequency adjustment factor.

[0028] Deformable Convolutional Layer: Deformable CNN kernels are deployed in the second branch. Their offset learning module is driven by an attention mechanism to capture the microscopic time-varying features of non-steady feeding sounds. The deformation range of the convolutional kernel is controlled within ±15% of the sampling points.

[0029] Multi-scale feature distillation unit: Employing a teacher-student network architecture, the teacher network uses EfficientNet-B3 to extract 256-dimensional high-precision features, while the student network uses a spectrum-aware distillation loss function. in, feature vectors output by the teacher network and the student network, Mel-spectrogram of the original audio, mean square error of the feature vectors of the teacher network and the student network, first-order absolute error of the Mel-spectrogram in the logarithmic domain, is a loss term weight coefficient. When the model compression rate is 60%, the equal error rate (EER) only increases by 0.8%.

[0030] Optionally, the multi-modal classification module comprises:

[0031] Meta-learning classifier: a contrastive learning framework based on a prototype network is constructed, and the dynamic feature library update strategy is: wherein, is a reservation ratio of historical prototype weights, is the number of new samples, is a feature vector of the new sample, a sub-class prototype is automatically created when a new individual is detected, and incremental learning is supported;

[0032] Gated attention mechanism: a spatio-temporal dual attention unit is designed, the time domain attention weight is calculated through the LSTM hidden state, the frequency domain attention adopts a wavelet basis function, and the fusion formula of the two is: wherein, is a time domain attention transformation matrix, is a time series context feature extracted through a bidirectional LSTM network, is a frequency domain attention transformation matrix, is a Morlet wavelet transform output, is a Sigmoid function, which compresses the output to the [0, 1] interval, and realizes a key feature weight improvement of 30%-50%;

[0033] Adaptive DTW retrieval unit: a segmented dynamic time warping algorithm (Segmented DTW) is proposed, which divides the audio stream into 50ms segments, and introduces a penalty factor when calculating the local similarity matrix: wherein, is the Euclidean distance between frame i and frame j, is the time series misalignment penalty strength, , is the number of frames of the audio to be matched and the template.

[0034] Optionally, the data fusion unit further comprises:

[0035] Multi-source evidence fusion engine: D-S evidence theory is used to fuse the voiceprint confidence , the periodic characteristics of the acceleration signal and the ambient light intensity​ with the basic probability assignment function: where, is the voiceprint confidence, is the acceleration signal periodicity, is the ambient light intensity, , , is the standard deviation of each sensor data, is the normalization factor, ensuring the probability sum is 1, when the conflict factor >0.3 triggers the voice navigation unit to re-collect data;

[0036] Behavior probability calibration module: Establish a Bayesian network dynamic inference model, nodes include feeding duration, chewing frequency and head trajectory, update the posterior probability in real time through variational inference: where, is the initial feeding probability based on historical data, is the probability of observing data D when the feeding behavior occurs; is the sum of all possible behavior states (such as rest, walk, rumination);

[0037] Observation data includes voiceprint matching degree and acceleration peak value, model update period ≤1 second.

[0038] Optionally, it also includes:

[0039] Voice navigation positioning module: integrate UWB / Bluetooth multi-mode positioning engine, when GPS signal is lost, use the RFID tag array deployed in the pasture for location fingerprint matching, positioning error <1.5 meters;

[0040] Space-time correlation analyzer: build a graph neural network (GNN) model, nodes represent the voiceprint features of individual sheep, edge weights are determined by physical distance and feeding behavior synchronicity, when an abnormal node is detected, automatically analyze the behavior pattern changes of its 3-hop neighbor nodes, identify early signs of group disease transmission;

[0041] Energy optimization unit: use dynamic voltage frequency adjustment (DVFS) technology, when the microphone array detects a silent period >5 seconds, automatically switch the processor to low-power mode, power consumption is reduced by 70% while still maintaining real-time response delay <200ms.

[0042] Optionally, the system further includes:

[0043] Adversarial training module: inject time-frequency adversarial samples in the model training stage, including band-pass noise pulse (200-800Hz), time domain stretching (±15%) and frequency disturbance (±50Hz offset), improve the robustness of the model in rainy weather, and maintain the recognition accuracy in extreme environment above 89%;

[0044] Privacy protection unit: adopt a voiceprint desensitization scheme based on homomorphic encryption to randomly orthogonally project the feature vector , wherein is a random orthogonal matrix generated locally by the edge device;

[0045] Self-powered device: integrate flexible piezoelectric fiber array in wearable device, whose resonance frequency matches the neck movement frequency spectrum of sheep (2-5Hz), energy conversion efficiency ≥18%, cooperate with micro super capacitor to realize continuous work for 30 days without charging.

[0046] Optionally, the adversarial training module comprises:

[0047] Adaptive adversarial sample generator: construct a generative adversarial network (GAN) framework, wherein the generator adopts a conditional WaveGAN architecture, dynamically synthesizes time-frequency domain adversarial samples according to the real-time collected environmental noise spectrum characteristics (when signal-to-noise ratio ≤10dB), and its generation strategy is:

[0048] , wherein is the disturbance intensity coefficient, which is adaptively adjusted according to the current environmental noise power (0.1-0.3), is the dominant frequency of environmental noise, is the gradient direction of the model to the input x, is the Gaussian noise centered on the dominant frequency of environmental noise;

[0049] Meta-learning defense mechanism: introduce MAML meta-learning algorithm in the model fine-tuning stage, and make the classifier quickly adapt to new adversarial attacks through second-order gradient update, and the objective function is: , wherein , is the meta-task of different adversarial attack scenarios, is the inner loop parameter update step, is the constraint parameter change amplitude;

[0050] Multi-scale robustness verification unit: deploy a cascade detection network, the first stage adopts STFT time-frequency analysis to detect abnormal energy pulses, and the second stage extracts high-level semantic features through a pre-trained VGGish network, and when the confidence difference of the two stages is >0.3, the model parameter rollback mechanism is triggered.

[0051] ​​Optionally, further comprising a privacy protection unit cooperated with the self-powered device for optimization, specifically comprising:

[0052] Dynamic key management: random orthogonal matrix The generation adopts a physical unclonable function (PUF) based on a piezoelectric energy waveform, and utilizes a microsecond voltage fluctuation sequence output by a piezoelectric fiber array As an entropy source, through a hash chain operation: Wherein, is a microsecond voltage fluctuation of the piezoelectric fiber, is rounded to an integer, is a bit string splicing, the projection matrix is updated every 30 minutes, and the cracking difficulty is increased to orders of magnitude;

[0053] Energy-aware encryption scheduling: an energy state machine model is established, and when the super capacitor power is <15%:

[0054] Activate the lightweight encryption mode, and only project the first 128 dimensions of the voiceprint feature vector;

[0055] Turn off the online learning function of the adversarial training module;

[0056] Adjust the piezoelectric resonance frequency to the optimal energy collection point (4.2Hz±0.5Hz), so that the privacy protection maintenance rate under low power is ≥95% and the endurance is prolonged by 40%.

[0057] Optionally, further comprising a resonant frequency self-optimization unit: a closed-loop control system is constructed, and the neck movement acceleration is monitored through a Hall sensor The particle swarm algorithm is used to dynamically adjust the prestress of the piezoelectric fiber :

[0058]

[0059] Wherein, is the expected resonance frequency, is the cycle length of the observed neck movement, is the stiffness coefficient related to the prestress, is the equivalent mass of the piezoelectric vibrator, and the stiffness coefficient , is the initial stiffness coefficient, and 0.23 and 1.2 are the stiffness-prestress relationship parameters determined by experiments.

[0060] Compared with the prior art, the present application includes at least one of the following beneficial technical effects:

[0061] The environment perception sub-module of the voice enhancement module can dynamically switch the noise reduction mode according to the environment type such as the temperature and humidity of the pasture; the adversarial training module injects time-frequency adversarial samples in the model training stage, improves the robustness of the model under extreme weather such as wind and rain, and ensures that the recognition accuracy under extreme environment is maintained above 89%.

[0062] The voiceprint feature extraction module can realize individual sheep identification, facilitating accurate management of each sheep. The spatio-temporal correlation analyzer constructs a graph neural network model, which can identify early signs of group disease transmission by analyzing individual sheep voiceprint features and feeding behavior synchronicity, and protect the overall health of the sheep herd.

[0063] The privacy protection unit adopts a voiceprint desensitization scheme based on homomorphic encryption to protect the privacy of sheep voiceprint data by randomly orthogonally projecting the feature vector. Dynamic key management further enhances the randomness and security of key generation, improving the strength of privacy protection.

[0064] Through the integration of a flexible piezoelectric fiber array in the self-powered device, the motion energy of the sheep's neck is converted into electrical energy, and a micro super capacitor is used to realize continuous operation for 30 days without charging. The energy optimization unit uses dynamic voltage frequency adjustment technology to reduce power consumption while ensuring real-time response. Energy-aware encryption scheduling intelligently adjusts encryption strategies and functions based on power, extending the device's battery life

[0065] The present application realizes accurate identification and analysis of sheep feeding behavior, effectively solving the problems of sound signal processing, individual and group monitoring, privacy protection and energy supply in complex environments. Not only improves the intelligent level and efficiency of pasture management, reduces labor costs, but also can timely find out the health problems of sheep, protect the health of sheep herd, protect the privacy of sheep data, and improve the energy utilization efficiency of the device. BRIEF DESCRIPTION OF DRAWINGS

[0066] Figure 1 It is a structural schematic diagram of a voice recognition system for sheep feeding behavior. DETAILED DESCRIPTION

[0067] The technical solutions of the present application will be further described below in combination with the drawings and specific embodiments.

[0068] Embodiment 1

[0069] As shown in Figure 1 A voice recognition system for sheep feeding behavior is proposed, which includes a sound collection module, a voice enhancement module, a voiceprint feature extraction module, a multi-modal classification module, and a data fusion unit. Each module will be described in detail below.

[0070] The sound collection module is arranged on the wearable device on the neck of the sheep and is configured with an anti-wind-noise directional microphone array for collecting environmental sound signals in real time; the sound collection module collects the sound around the sheep, and the anti-wind-noise design can reduce environmental noise interference and ensure that clear and accurate sound signals are collected to provide reliable data for subsequent analysis.

[0071] The voice enhancement module is connected to the sound collection module and adopts a voice noise reduction algorithm based on a deep neural network to perform environmental noise suppression and voice feature enhancement processing on the original audio signal; the voice enhancement module comprises:

[0072] The dual-mode noise reduction unit comprises a dual-path recurrent convolutional network (DPRNN), wherein a first path processes time-frequency masking of the original audio signal and adopts a complex domain mask estimation technology to retain voice phase information; and a second path synchronously receives a 3-8Hz chewing vibration signal collected by the acceleration sensor and eliminates resonance noise caused by mechanical vibration through an adaptive notch filter, and a transfer function of the adaptive notch filter is:

[0073]

[0074] wherein, is a notch center frequency dynamically adjusted according to the chewing vibration frequency detected by the acceleration sensor in real time, is an audio sampling rate, is a damping coefficient, is a unit delay operator; the dual-path design can process noise from different angles, more comprehensively suppress noise, retain voice phase information to help restore more authentic voice signals, and eliminate resonance noise to further improve the purity of the audio.

[0075] The environmental perception sub-module comprises a lightweight random forest classifier arranged to identify the type of pasture environment (such as open air or shed) according to the temperature and humidity sensor data and dynamically switch the noise reduction mode: when the humidity is greater than 70%, a high-frequency compensation filter is enabled to improve the signal-to-noise ratio of the 4-6kHz frequency band by more than 3dB; the noise reduction strategy is automatically adjusted according to different environments to improve the pertinence of the noise reduction effect, improve the signal-to-noise ratio of a specific frequency band in a humid environment, ensure the quality of the voice signal, and improve the adaptive ability of the system to environmental changes.

[0076] The voice codec adopts a vector quantization variational autoencoder (VQ-VAE) to compress 24kHz sampling audio to an 8kbps code stream, and an adversarial spectral loss function is introduced during codebook training:

[0077]

[0078] wherein, is the energy of the mth frequency band of the mel spectrum of the original audio, ​to reconstruct the energy value of the audio in the mel scale frequency band, is the total number of mel filter banks. To achieve efficient audio compression while ensuring speech intelligibility, reduce data transmission and storage pressure, and improve the quality of reconstructed audio, the adversarial spectrum loss function helps to ensure that key information is not lost.

[0079] Ensure that the reconstruction error of the key frequency band (200-4000 Hz) is less than 2 dB, and the speech intelligibility is maintained at MOS≥4.0 when the compression ratio is increased by 50%. It ensures that the key frequency band information of the speech can still be accurately restored at a high compression ratio, maintaining a high speech intelligibility, and not affecting subsequent speech-based analysis and judgment.

[0080] The voiceprint feature extraction module integrates the improved mel frequency cepstral coefficient (MFCC) and the convolutional neural network (CNN) to extract the voiceprint feature vector with individual differences from the enhanced speech signal. The voiceprint feature extraction module specifically includes:

[0081] Hybrid ResNet: The first branch uses a dynamic mel filter bank, and the center frequency is adaptively adjusted according to the individual sheep vocal cord length, and the adjustment formula is:

[0082]

[0083] wherein, is the mth center frequency of the standard mel filter bank, is the vocal cord feature length calculated by glottal pulse inversion, and is the group average and standard deviation, and 0.05 is the frequency adjustment amplitude proportion factor;

[0084] Deformable convolution layer: In the second branch, a deformable convolution kernel (Deformable CNN) is deployed, and its offset learning module is driven by an attention mechanism to capture the microscopic time-varying features of non-stationary feeding sound, and the convolution kernel deformation range is controlled within ±15% sampling points;

[0085] Multi-scale feature distillation unit: A teacher-student network architecture is used, and the teacher network uses EfficientNet-B3 to extract 256-dimensional high-precision features, and the student network uses the spectrum perception distillation loss function: wherein, is the feature vector output by the teacher network and the student network, is the mel spectrum of the original audio, is the mean square error of the feature vectors of the teacher network and the student network, is the first-order absolute error of the log domain of the mel spectrum, , is a loss term weight coefficient. When the model compression rate is 60%, the equal error rate (EER) only increases by 0.8%. The implementation of the model compression while maintaining a high accuracy, reducing the consumption of computing resources, improving the efficiency of system operation, so that the system can still maintain good performance in the case of resource constraints.

[0086] The multi-modal classification module includes a pre-trained speech classification model and a speech retrieval unit. The speech classification model constructs a foraging voiceprint feature library through a contrast learning mechanism and dynamically allocates speech feature weights using an attention mechanism. The speech retrieval unit matches the similarity of real-time audio and the feature library through a dynamic time warping (DTW) algorithm. A variety of technologies are integrated to accurately classify and quickly retrieve the foraging behavior of sheep, improving the accuracy and efficiency of system identification. The multi-modal classification module includes:

[0087] Meta-learning classifier: a contrast learning framework based on prototype network is constructed, and the dynamic feature library update strategy is: wherein, is the retention ratio of historical prototype weight, is the number of new samples, is the feature vector of the new sample, and a sub-class prototype is automatically created when a new individual is detected, supporting incremental learning;

[0088] Gated attention mechanism: a spatio-temporal dual attention unit is designed, the time domain attention weight is calculated through the LSTM hidden state, and the frequency domain attention uses the wavelet basis function, and the fusion formula of the two is: wherein, is the time domain attention transformation matrix, is the time series context feature extracted through the bidirectional LSTM network, is the frequency domain attention transformation matrix, is the Morlet wavelet transform output, is the Sigmoid function, which compresses the output to the [0, 1] interval, and the key feature weight is improved by 30%-50%; highlight the key features, improve the attention to foraging sound features, enhance the expression ability of the features, further improve the accuracy of classification, and make the system more reliable in judging the foraging behavior.

[0089] Adaptive DTW retrieval unit: a segmented dynamic warping algorithm (Segmented DTW) is proposed, which divides the audio stream into 50ms segments, and introduces a penalty factor when calculating the local similarity matrix: wherein, is the Euclidean distance between frame i and frame j, is the time series misalignment penalty strength, ​The frame number of the audio to be matched and the template. The accuracy of audio matching is improved, the timing drift is effectively controlled, the similarity between real-time audio and feature library audio is more accurately measured, false positives are reduced, and the system recognition performance is improved.

[0090] In the embodiment, the data fusion unit fuses the voiceprint recognition result and the acceleration sensor data, and generates a foraging behavior probability output through a Bayesian inference algorithm. The data fusion unit further includes:

[0091] Multi-source evidence fusion engine: D-S evidence theory is used to fuse voiceprint confidence , acceleration signal periodicity and environmental light intensity , and the basic probability assignment function is: wherein, is the voiceprint confidence, is the acceleration signal periodicity, is the environmental light intensity, , , is the standard deviation of each sensor data, is a normalization factor, ensuring that the probability sum is 1, and when the conflict factor >0.3, the voice navigation unit reacquires data;

[0092] Behavior probability calibration module: a Bayesian network dynamic inference model is established, the nodes include foraging time, chewing frequency and head movement trajectory, and the posterior probability is updated in real time through variational inference: wherein, is the initial foraging probability based on historical data, is the probability of observing data D when the foraging behavior occurs, is the summation of all possible behavior states (such as rest, walking, and rumination);

[0093] Observation data includes voiceprint matching degree and acceleration peak value, and the model update period is ≤1 second. By integrating various data information, the reliability and accuracy of foraging behavior judgment are improved, and the error and uncertainty caused by single data are reduced. According to real-time data, the foraging behavior probability is dynamically updated, which more accurately reflects the current foraging state of the sheep, and provides more real-time and accurate information for pasture management.

[0094] In the embodiment, it also includes:

[0095] Voice navigation positioning module: integrated UWB / Bluetooth multi-mode positioning engine, when the GPS signal is lost, the RFID tag array deployed in the pasture is used for location fingerprint matching, and the positioning error is <1.5 meters;

[0096] Spatiotemporal correlation analyzer: Constructs a graph neural network (GNN) model where nodes represent the voiceprint features of individual sheep, and edge weights are determined by physical distance and synchronicity of grazing behavior. When an anomaly is detected in a node, it automatically analyzes the behavioral pattern changes of its three-hop neighbor nodes to identify early signs of disease transmission in the flock. This helps to promptly detect potential disease transmission risks in pastures, take preventive measures in advance, protect the health of sheep, and reduce economic losses.

[0097] Energy Optimization Unit: Employing Dynamic Voltage Frequency Scaling (DVFS) technology, when the microphone array detects a silence period >5 seconds, it automatically switches the processor to a low-power mode, maintaining a real-time response latency of <200ms while reducing power consumption by 70%. This reduces power consumption while ensuring system performance, extending the battery life of wearable devices, reducing the frequency of device charging or battery replacement, and improving system stability and sustainability.

[0098] Example 2

[0099] This embodiment, based on embodiment 1, also includes an adversarial training module, a privacy protection unit, and a self-powered device.

[0100] The adversarial training module injects time-frequency adversarial examples during model training, including bandpass noise impulses (200-800Hz), time-domain stretching (±15%), and frequency perturbations (±50Hz shift), to improve the model's robustness in windy and rainy weather, maintaining the recognition accuracy above 89% in extreme environments. The adversarial training module includes:

[0101] Adaptive Adversarial Example Generator: A Generative Adversarial Network (GAN) framework is constructed, in which the generator adopts a conditional WaveGAN architecture. It dynamically synthesizes time-frequency domain adversarial examples based on the real-time acquired environmental noise spectrum characteristics (when the signal-to-noise ratio is ≤10dB). Its generation strategy is as follows:

[0102] in, This is the disturbance intensity coefficient, which is adaptively adjusted (0.1~0.3) based on the current ambient noise power. The dominant frequency of environmental noise. This represents the gradient direction of the model with respect to the input x. Based on the dominant frequency of environmental noise The model is centered on Gaussian noise; adversarial examples are dynamically generated based on actual environmental noise to specifically improve the robustness of the model under different noise environments and enhance the model's generalization ability.

[0103] Meta-learning defense mechanism: The MAML meta-learning algorithm is introduced during the model fine-tuning stage. Second-order gradient updates enable the classifier to quickly adapt to new adversarial attacks. Its objective function is: in, , Meta task for different adversarial attack scenarios, Inner loop parameter update step size, Constrained parameter change amplitude; enable the system to quickly respond to new adversarial attacks, improve the security and stability of the model, and ensure that the system can still work normally when facing constantly changing attack methods.

[0104] Multi-scale robustness verification unit: deploy a cascade detection network, the first level uses STFT time-frequency analysis to detect abnormal energy pulses, the second level extracts high-level semantic features through a pre-trained VGGish network, and triggers the model parameter rollback mechanism when the confidence difference between the two levels is >0.3. Effectively detect abnormal situations during training and running of the model, discover potential problems in time and repair them, and ensure the reliability and accuracy of the model.

[0105] In this embodiment, the privacy protection unit: adopts a voiceprint desensitization scheme based on homomorphic encryption, and randomly orthogonally projects the feature vector on the random orthogonal matrix generated by the edge device locally;

[0106] Further, the self-powered device: integrates a flexible piezoelectric fiber array in the wearable device, whose resonance frequency matches the neck movement frequency spectrum of the sheep (2-5Hz), and the energy conversion efficiency is ≥18%, which cooperates with a micro super capacitor to realize continuous work for 30 days without charging.

[0107] Embodiment 3

[0108] This embodiment is based on the basis of embodiment 2 and further includes a privacy protection unit and a self-powered device for collaborative optimization, specifically including:

[0109] Dynamic key management: the generation of the random orthogonal matrix uses a physical unclonable function (PUF) based on the piezoelectric energy waveform, and uses the microsecond voltage fluctuation sequence output by the piezoelectric fiber array as an entropy source, and performs hash chain operation: wherein, is the microsecond voltage fluctuation of the piezoelectric fiber, is rounded to an integer, is a bit string splicing, and the projection matrix is updated every 30 minutes, and the cracking difficulty is increased to order of magnitude; enhance the randomness and security of key generation, improve the strength of privacy protection, make it difficult for attackers to crack encrypted information, and better protect data privacy.

[0110] Energy-aware encryption scheduling: an energy state machine model is established, and when the super capacitor power is <15%:​

[0111] Activate lightweight encryption mode, only project the first 128 dimensions of the voiceprint feature vector;

[0112] Turn off the online learning function of the adversarial training module;

[0113] Adjust the piezoelectric resonance frequency to the optimal energy harvesting point (4.2 Hz ± 0.5 Hz), so that the privacy protection maintenance rate is ≥95% and the endurance is extended by 40% under low power. Improve the energy harvesting efficiency under low power, prolong the device endurance time, while maintaining a high level of privacy protection, and ensure the normal operation and data security of the system under low power.

[0114] Resonance frequency self-optimization unit: build a closed-loop control system, monitor the acceleration of neck movement through a Hall sensor , dynamically adjust the prestress of piezoelectric fibers using a particle swarm algorithm :

[0115]

[0116] where, is the desired resonance frequency, is the period of the observed neck movement, is the stiffness coefficient related to the prestress, is the equivalent mass of the piezoelectric vibrator, and the stiffness coefficient , is the initial stiffness coefficient, and 0.23 and 1.2 are the stiffness-prestress relationship parameters determined by experiments. Real-time optimization of the resonance frequency of piezoelectric fibers improves energy conversion efficiency and further enhances the performance of self-powered devices, providing more abundant energy support for stable operation of the system.

[0117] The above specific embodiments are only a few optional embodiments of the present application. Based on the technical solutions of the present application and the related inspiration of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.

Claims

1. A sound recognition system for sheep feeding behavior, characterized in that, Comprise: A sound acquisition module is deployed on the neck of the sheep and is configured with an anti-wind noise directional microphone array for real-time acquisition of environmental sound signals; A voice enhancement module is connected to the sound acquisition module and uses a deep neural network-based voice noise reduction algorithm to suppress environmental noise and enhance voice features of the original audio signal; The voice enhancement module comprises: A dual-mode noise reduction unit: a double-path recurrent convolutional network is constructed, wherein a first path processes the time-frequency masking of the original audio signal, and a complex domain mask estimation technique is used to retain the phase information of the voice; the second path synchronously receives the 3-8Hz chewing vibration signal collected by the acceleration sensor, and eliminates the resonance noise caused by mechanical vibration through an adaptive notch filter, and the transfer function is: ; wherein, is a notch center frequency dynamically adjusted according to the real-time detected chewing vibration frequency of the acceleration sensor, is an audio sampling rate, is a damping coefficient, is a unit delay operator; A voice coding decoder: a vector quantization variational autoencoder is used to compress the 24kHz sampled audio to an 8kbps code stream, and an adversarial spectral loss function is introduced during codebook training: ; wherein, is the energy of the original audio in the i-th frequency band in the Mel scale, is the energy of the reconstructed audio in the i-th frequency band in the Mel scale, is the total number of the Mel filter banks.​​ A voiceprint feature extraction module integrates an improved mel-frequency cepstral coefficient and a convolutional neural network to extract a voiceprint feature vector with individual differences from the enhanced voice signal; A multi-modal classification module includes a pre-trained voice classification model and a voice search unit, the voice classification model constructs a feeding voiceprint feature library through a contrast learning mechanism, and uses an attention mechanism to dynamically allocate voice feature weights, and the voice search unit matches the similarity of real-time audio and the feature library through a dynamic time warping algorithm; A data fusion unit fuses the voiceprint recognition result and the acceleration sensor data to generate a feeding behavior probability output through a Bayesian inference algorithm; The multi-modal classification module comprises: Meta-learning classifier: build a contrastive learning framework based on prototype network, its dynamic feature library update strategy is: Wherein, is the updated feature library prototype vector, is the historical feature library prototype vector, is the reservation ratio of historical prototype weight, is the number of new samples, is the feature vector of the new sample, and a subclass prototype is automatically created when a new individual is detected; Gated attention mechanism: a spatiotemporal dual attention unit is designed, the time domain attention weight is calculated through the LSTM hidden state, the frequency domain attention adopts the wavelet basis function, and the fusion formula is: wherein, is the final attention weight, is the time domain attention transformation matrix, is the time sequence context feature extracted through the bidirectional LSTM network, is the frequency domain attention transformation matrix, is the Morlet wavelet transform output, is the Sigmoid function, which compresses the output to the [0, 1] interval; Adaptive DTW retrieval unit: A piecewise dynamic time warping algorithm is proposed, which divides the audio stream into 50ms segments and introduces a penalty factor when calculating the local similarity matrix: wherein, is the band-penalty Euclidean distance between frame i and frame j, is the original Euclidean distance between frame i and frame j, is the time misalignment penalty strength , is the frame number of the audio to be matched and the template.

2. The sound recognition system for sheep feeding behavior according to claim 1, characterized in that, The voice enhancement module comprises: An environmental perception sub-module: a lightweight random forest classifier is deployed to identify the type of pasture environment based on temperature and humidity sensor data, and dynamically switch the noise reduction mode: when the humidity is greater than 70%, a high-frequency compensation filter is enabled to improve the signal-to-noise ratio in the 4-6kHz frequency band by more than 3dB.

3. The sheep feeding behavior voice recognition system according to claim 1, wherein The voiceprint feature extraction module specifically comprises: A hybrid deep residual network: the first branch uses a dynamic mel filter bank, and the center frequency is adaptively adjusted according to the vocal cord length of the individual sheep, and the adjustment formula is: ; wherein, is the adjusted mth center frequency of the mel filter, is the mth center frequency of the standard mel filter bank, is the vocal cord characteristic length calculated by glottal pulse inversion, and is the group average and standard deviation, and 0.05 is the frequency adjustment amplitude proportion factor. A deformable convolution layer: a deformable convolution kernel is deployed in the second branch, and the offset learning module is driven by an attention mechanism to capture the micro time-varying features of non-stationary feeding sound; Multi-scale feature distillation unit: adopt teacher-student network architecture, the teacher network uses EfficientNet-B3 to extract 256-dimensional high-precision features, and the student network uses the spectral perception distillation loss function: Wherein, is the feature vector output by the teacher network and the student network, is the mel spectrum of the original audio, is the mean square error of the feature vectors of the teacher network and the student network, is the first-order absolute error of the mel spectrum logarithmic domain, , is the loss term weight coefficient.

4. The sound recognition system for sheep feeding behavior according to claim 1, characterized in that, The data fusion unit further comprises: Multi-source evidence fusion engine: Employs DS evidence theory to fuse voiceprint confidence. Periodic characteristics of acceleration signals and ambient light intensity Its basic probability allocation function is: in, As evidence The basic probability allocation value, For voiceprint confidence, periodic characteristics of acceleration signals For ambient light intensity, The mean value of the periodic characteristics of the acceleration signal. The standard deviation of the voiceprint confidence data. The standard deviation of the periodic data of the acceleration signal. The standard deviation of ambient light intensity data. This is a reference value for ideal light intensity. As a normalization factor, ensuring the sum of probabilities is 1, when the conflict factor When the value is >0.3, the voice navigation unit is triggered to re-collect data; Behavior probability calibration module: a Bayesian network dynamic inference model is established, nodes include feeding duration, chewing frequency and head movement trajectory, and the posterior probability is updated in real time through variational inference: wherein is the posterior probability of feeding behavior under observation data D, is the initial feeding probability based on historical data, is the probability of observing data D when feeding behavior occurs; is the summation of all possible behavior states, is the other behavior state, and the other behavior state includes rest, walking and rumination; Observation data Including voiceprint matching degree and acceleration peak value, model updating period ≤ 1 second.

5. The sound recognition system for sheep feeding behavior according to claim 1, characterized in that, Also comprising: A voice navigation positioning module: an UWB / Bluetooth multi-mode positioning engine is integrated, and when the GPS signal is lost, an RFID tag array deployed in the pasture is used for location fingerprint matching, and the positioning error is less than 1.5 meters; A spatio-temporal correlation analyzer: a graph neural network model is constructed, a node represents the voiceprint feature of an individual sheep, and the edge weight is determined by the physical distance and the synchronicity of the feeding behavior, when an abnormal node is detected, the behavior pattern change of the 3-hop neighbor nodes is automatically analyzed to identify early signs of group disease transmission; Energy optimization unit: adopt dynamic voltage frequency adjustment technology, when the microphone array detects that the silent period is greater than 5 seconds, the processor is automatically switched to the low power consumption mode.

6. The sound recognition system for sheep feeding behavior according to claim 1, characterized in that, The system further comprises: Adversarial training module: inject time-frequency adversarial samples in the model training stage, including band-pass noise pulse, time domain stretching and frequency disturbance; The privacy protection unit adopts a voiceprint desensitization scheme based on homomorphic encryption to perform random orthogonal projection on the feature vector wherein, is a random orthogonal matrix generated locally by the edge device.​ Self-powered device: integrate flexible piezoelectric fiber array in wearable device, whose resonance frequency matches the movement frequency spectrum of sheep neck.

7. The sound recognition system for sheep feeding behavior according to claim 6, characterized in that, The adversarial training module comprises: Adaptive adversarial sample generator: build a generative adversarial network framework, in which the generator adopts a conditional WaveGAN architecture to dynamically synthesize time-frequency domain adversarial samples according to the real-time collected environmental noise spectrum characteristics, and the generation strategy is: ; wherein, is the generated adversarial sample, is the original clean audio signal, is the perturbation strength coefficient, which is adaptively adjusted according to the current ambient noise power, is a sign function, is the dominant frequency of the ambient noise, is the gradient direction of the model to the input x, is the model parameter, is the true label, is the dominant frequency of the ambient noise is a Gaussian noise with a center and a standard deviation of 0.

2. Meta-learning defense mechanism: Introduce MAML meta-learning algorithm in the model fine-tuning stage, and make the classifier quickly adapt to new adversarial attacks through second-order gradient update. The objective function is: where, is the initial parameter of the model, is the loss function of the th meta-task, is the updated parameter after the inner loop, that is, , is the meta-task of different adversarial attack scenarios, is the inner loop parameter update step, is the constraint parameter change amplitude, is the square of the L2 norm. Multi-scale robustness verification unit: deploy a cascade detection network, the first stage adopts STFT time-frequency analysis to detect abnormal energy pulses, and the second stage extracts high-level semantic features through a pre-trained VGGish network, and when the confidence difference between the two stages is greater than 0.3, a model parameter rollback mechanism is triggered.

8. The sound recognition system for sheep feeding behavior according to claim 6, characterized in that, Further comprising a privacy protection unit and a self-powered device optimization, specifically comprising: Dynamic key management: random orthogonal matrix The generation of the random orthogonal matrix employs a physically unclonable function based on piezoelectric energy waveform, utilizing the microsecond voltage fluctuation sequence output by the piezoelectric fiber array As an entropy source, through hash chain operation: Wherein, is the updated random orthogonal matrix, is the secure hash algorithm, is the current random orthogonal matrix, is the microsecond voltage fluctuation of the piezoelectric fiber, is rounded to an integer, is the bit string splicing, the projection matrix is updated every 30 minutes, and the cracking difficulty is increased to orders of magnitude; Energy-aware encryption scheduling: establish an energy state machine model, when the super capacitor power is less than 15%: Activate lightweight encryption mode, only project the first 128 dimensions of the voiceprint feature vector; Turn off the online learning function of the adversarial training module; Adjust the piezoelectric resonance frequency to the optimal energy collection point, so that the privacy protection maintenance rate under low power is greater than or equal to 95% and the endurance is prolonged by 40%.

9. The sound recognition system for sheep feeding behavior according to claim 8, characterized in that, Also include the resonant frequency self-optimizing unit: build a closed-loop control system, through the Hall sensor monitoring neck movement acceleration , using particle swarm algorithm dynamic adjustment of piezoelectric fiber prestress : wherein, is the resonance frequency to be achieved, is the period of the observed neck motion, is the stiffness coefficient related to the pre-stress, is the equivalent mass, stiffness coefficient of the piezoelectric vibrator , is the initial stiffness coefficient, 0.23 and 1.2 are the stiffness-pre-stress relationship parameters determined experimentally.

Citation Information

Patent Citations

  • Gunshot detection and identification method and system based on comparative learning pre-training

    CN118262727A

  • Animal voiceprint monitoring method and device, medium and product

    CN119091890A