Cow sound denoising and classifying method and system

Through multi-channel audio data processing technology, combined with beamforming and coherent spectral subtraction, the problem of poor performance of traditional denoising methods in complex ranch environments is solved, and more efficient denoising and classification accuracy of cattle sound signal is achieved.

CN119964580APending Publication Date: 2025-05-09INNER MONGOLIA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510118825.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In complex pasture environments, traditional denoising methods are difficult to effectively remove multiple noise interferences, resulting in a decrease in the accuracy of extracting and classification of cattle sound signals.

Method used

Multi-channel audio data processing technology is adopted to remove noise signals through a combination of beamforming and coherent spectral subtraction, and use deep learning classification models to classify audio data.

Benefits of technology

It significantly improves the denoising effect and classification accuracy of the bull sound signal, and can effectively retain the detailed information of the audio signal in complex noise environments and reduce signal distortion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964580A_ABST
    Figure CN119964580A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a cattle sound denoising and classifying method, and the method comprises the steps: receiving multi-channel cattle sound audio data, and carrying out the channel separation of the multi-channel cattle sound audio data, and obtaining a plurality of single-channel audio data; combining a plurality of single-channel audio data into one channel by using a beam forming technology, performing short-time Fourier transform on the audio data, and calculating a coherence coefficient of the audio data and a noise signal; according to the coherence coefficient, adaptively adjusting an alpha parameter so as to control the proportion of removing the noise signal in the audio data, and obtaining de-noised audio data; and classifying the de-noised audio data by using a classification model to obtain category information of the audio data. According to the method, processing of multi-channel audio data is realized, audio signals and noise are effectively distinguished by calculating coherence of the signals, and detailed information of cattle sound is better reserved, so that the accuracy of cattle sound classification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio denoising and classification, and in particular to a method and system for denoising and classifying cow sounds. Background Art

[0002] With the development of smart ranch management technology, sound data has gradually become an important means of monitoring cattle health and behavior. However, complex non-stationary noise in the ranch environment (such as wind noise, mechanical noise, and the sounds of other animals) can seriously affect the extraction and classification of sound signals. Although traditional denoising methods such as spectral subtraction and Wiener filtering can reduce noise to a certain extent, they will introduce speech distortion or noise residue in strong noise interference and non-stationary environments, thereby reducing the accuracy of sound classification.

[0003] In the existing technology, the denoising method based on beamforming can enhance the target signal, but the effect is limited in dynamic noise scenes. In addition, most sound classification systems rely only on single-channel signals and cannot effectively combine multi-channel information to improve positioning and classification performance. For this reason, there is an urgent need for a system that combines multi-channel data processing, coherent spectral subtraction optimization and deep learning classification technology to achieve more efficient noise suppression, sound source localization and sound classification. Summary of the invention

[0004] The purpose of the embodiment of the present invention is to provide a method for denoising and classifying cattle voices, which is used to denoise multi-channel audio data, thereby improving the accuracy of audio classification.

[0005] In order to achieve the above-mentioned purpose, an embodiment of the present invention provides a method for denoising and classifying cow sounds, which includes: receiving multi-channel cow sound audio data, and performing channel separation to obtain multiple single-channel audio data; using beamforming technology to merge the multiple single-channel audio data into one channel, and obtaining the noise signal in the merged audio data through noise audio estimation; performing STFT short-time Fourier transform and smoothing processing on the merged audio data to obtain the spectrum amplitude of the merged audio data; according to the spectrum amplitude of the audio data, calculating the coherence coefficient of the merged audio data and the noise signal; according to the coherence coefficient, adaptively adjusting the α parameter to control the proportion of the noise signal removed from the audio data to obtain the denoised audio data; and classifying the denoised audio data using a classification model to obtain the category information of the audio data.

[0006] Optionally, before merging the multiple single-channel audio data into one channel using beamforming technology, the method further includes: performing signal processing on the multiple single-channel audio data respectively to obtain position information of the multiple single-channel audio data.

[0007] Optionally, the signal processing includes: obtaining the position information of the multiple single-channel audio data by delay estimation based on the time difference of receiving the multiple single-channel audio data and the position information of the audio data acquisition device that collects the multi-channel cow sound audio data; performing delay compensation on the multiple single-channel audio data after delay estimation to adjust the time difference between the multiple single-channel audio data.

[0008] Optionally, the method also includes: receiving cattle video data, performing frame-by-frame depth estimation on the video data using a depth estimation model and generating a depth map; and mapping the position information to a video frame image of the video data in combination with the depth map, and based on the position information, marking category information corresponding to the position information to a corresponding position in the video frame image.

[0009] Optionally, calculating the coherence coefficient of the merged audio data and the noise signal according to the frequency spectrum amplitude of the audio data includes: calculating the coherence degree of the merged audio data and the noise signal at a frequency f, and the calculation formula is:

[0010]

[0011] Wherein, γ(f) is the coherence coefficient, X(f) represents the complex spectrum of the audio signal in the audio data at the frequency f, N(f) represents the complex spectrum of the noise signal at the frequency f, and E[·] represents the expected value between X(f) and N(f). represents the average approximation of the complex spectrum of the noise signal at frequency f.

[0012] Optionally, according to the coherence coefficient, the α parameter is adaptively adjusted, and its calculation formula is:

[0013]

[0014] Among them, α0 is the spectral reduction coefficient, which is used to adjust the intensity of noise reduction, γ(f) is the coherence coefficient calculated at frequency f, ∈ is a set value, which is used to avoid the situation where the denominator is zero, and α max It is the maximum value of α(f), which is used to prevent α(f) from being too large and causing excessive noise reduction.

[0015] In a second aspect, an embodiment of the present invention provides a cattle voice denoising and classification system, the system comprising:

[0016] A multi-channel audio acquisition module is used to collect multi-channel cattle sound audio data and perform channel separation to obtain multiple single-channel audio data; a denoising module is used to merge the multiple single-channel audio data into one channel by using beamforming technology, obtain the noise signal in the merged audio data by noise audio estimation, perform STFT short-time Fourier transform and smoothing processing on the merged audio data to obtain the spectrum amplitude of the merged audio data, calculate the coherence coefficient of the merged audio data and the noise signal according to the spectrum amplitude of the audio data, and adaptively adjust the α parameter according to the coherence coefficient to control the proportion of the noise signal removed from the audio data to obtain the denoised audio data; and a classification module is used to classify the denoised audio data using a classification model to obtain category information of the audio data.

[0017] Optionally, the system also includes: a video processing module, which collects cattle video data, uses a depth estimation model to perform frame-by-frame depth estimation on the video data and generate a depth map; a visualization module, which is used to map the position information to the video frame image of the video data in combination with the depth map, and based on the position information, annotate the category information corresponding to the position information to the corresponding position in the video frame image.

[0018] In a third aspect, an embodiment of the present invention provides a machine-readable storage medium, on which instructions are stored, and the instructions are used to enable a machine to execute any of the above-mentioned methods for denoising and classifying cow sounds in the present application.

[0019] In a fourth aspect, an embodiment of the present invention provides a processor, characterized in that it is used to run a program, wherein the program, when run, is used to execute: the cattle voice denoising and classification method as described in any one of claims 1-6.

[0020] Through the above technical solution, the embodiment of the present invention can effectively combine multi-channel audio information to improve the performance of audio classification by receiving multi-channel audio data, and processing multiple single-channel audio data after channel separation of the multi-channel audio data. The beamforming technology is used to merge multiple single-channel audios into one channel and then perform denoising processing to achieve the purpose of enhancing the audio data. De-noising based on coherent spectral subtraction is used, combined with a dynamically adjusted adaptive parameter α, and the denoising intensity is flexibly adjusted according to the coherence characteristics of the signal and noise, which significantly improves the denoising effect in a complex noise environment, effectively retains the detailed information of the audio signal, and reduces signal distortion. In addition, the denoised audio data is used for subsequent audio classification to improve the accuracy of classification.

[0021] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific implementations, they are used to explain the embodiments of the present invention, but do not constitute a limitation on the embodiments of the present invention. In the accompanying drawings:

[0023] Figure 1 It is a flowchart of a method for denoising and classifying cattle voices provided by an embodiment of the present disclosure;

[0024] Figure 2 is a schematic diagram of a circular 6-microphone array provided in an embodiment of the present disclosure;

[0025] Figure 3 is a schematic diagram comparing the sound pickup distances of microphone arrays provided in an embodiment of the present disclosure;

[0026] Figure 4 is a schematic diagram of the structure of a fixing device of the annular microphone array provided in an embodiment of the present disclosure;

[0027] Figure 5 is a schematic diagram of the structure of the Jetson AGX Orin computer provided in an embodiment of the present disclosure;

[0028] Figure 6 is a process schematic diagram of a denoising method provided by an embodiment of the present disclosure;

[0029] Figure 7 It is a filtering effect diagram after the audio signal is processed using different denoising algorithms provided by the embodiments of the present disclosure;

[0030] Figure 8 is a schematic diagram of the structure of the ECAPA-TDNN classification model provided by the embodiment of the present disclosure;

[0031] Fig. 9 is a schematic diagram of locating a sound source using time delay estimation provided by an embodiment of the present disclosure;

[0032] Fig.10 is a depth map obtained by performing depth estimation on a scene provided by an embodiment of the present disclosure;

[0033] Fig.11 is a schematic diagram of mapping the position information of audio data into a three-dimensional scene provided by an embodiment of the present disclosure;

[0034] Fig.12 is a schematic diagram of mapping the position information of audio data into a two-dimensional plane provided by an embodiment of the present disclosure;

[0035] Fig.13It is a schematic diagram of the visualization effect of audio data classification and positioning provided by the embodiment of the present disclosure;

[0036] Fig.14 is an operation flow chart of the cattle sound denoising and classification method in a cattle farm environment provided by an embodiment of the present disclosure;

[0037] Fig.15 is a workflow diagram of the ranch management robot provided by the embodiment of the present disclosure;

[0038] Fig.16 Schematic diagram of the structure of a ranch management robot carrying a microphone array used in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0039] The specific implementation of the embodiment of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the embodiment of the present invention, and is not used to limit the embodiment of the present invention.

[0040] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application are in compliance with the relevant provisions of national laws and regulations. In the embodiments of this application, some existing solutions in the industry such as certain software, components, and models may be mentioned, which should be considered as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of this application, but it does not mean that the applicant has or will necessarily use the solution.

[0041] The existing technologies have been studied and applied to the cattle call classification system. For example, the cattle call classification technology based on signal processing and machine learning algorithms has been able to identify the different states of cattle, such as courtship, hunger and pain. However, the performance of these systems is significantly affected by noise interference in practical applications. The existing denoising technologies mainly include the following categories:

[0042] Spectral subtraction: This is one of the most common denoising techniques. It uses the spectral characteristics of the noise signal to reduce the noise by subtracting the signal's spectrum. However, spectral subtraction is prone to introduce speech distortion and residual noise when facing non-stationary noise, especially when processing cattle call signals with complex spectral characteristics, the effect is not ideal.

[0043] Wiener filtering: This method is based on the minimum mean square error criterion and uses the statistical characteristics of the signal and noise for filtering. However, Wiener filtering also performs poorly in non-stationary noise environments and is highly dependent on the statistical characteristics of the signal, making it difficult to adapt to the changing cattle farm environment.

[0044] Machine learning denoising: In recent years, deep learning and other machine learning methods have been gradually applied to denoising tasks, especially in the field of speech signal processing. However, these methods usually require a large amount of labeled data for training, and the trained models may be difficult to generalize to different noise scenarios in practical applications. In addition, the training process is complex and requires high computing resources.

[0045] Although these existing technologies have improved the performance of the cattle call classification system to varying degrees, they all have their own limitations. In particular, when faced with complex noise environments, the effect of traditional denoising methods is significantly reduced. Based on this, an embodiment of the present invention proposes a cattle call denoising and classification method, which acquires multi-channel cattle call audio data and combines a beamforming algorithm to enhance the sound source signal in a strong noise and non-stationary environment, greatly improve the strength of the audio signal, and provide high-quality input data for the denoising algorithm. The denoising technology based on coherent spectral subtraction, combined with the dynamically adjusted adaptive parameter α, flexibly adjusts the denoising intensity according to the coherence characteristics of the audio signal and the noise signal, significantly improves the denoising effect in a complex noise environment, effectively retains the detail information of the target signal, and reduces speech distortion.

[0046] Figure 1 FIG. 1 is a flow chart of a method for denoising and classifying cow sounds provided by an embodiment of the present disclosure. Figure 1 As shown, the method includes steps S1 to S6.

[0047] Step S1: receiving multi-channel cattle sound audio data, and performing channel separation on the data to obtain a plurality of single-channel audio data;

[0048] Step S2: after merging the plurality of single-channel audio data into one channel by using beamforming technology, obtaining a noise signal in the merged audio data by noise audio estimation;

[0049] Step S3: performing STFT short-time Fourier transform and smoothing processing on the merged audio data to obtain the frequency spectrum amplitude of the merged audio data;

[0050] Step S4: calculating the coherence coefficient between the combined audio data and the noise signal according to the frequency spectrum amplitude of the audio data;

[0051] Step S5: adaptively adjusting the α parameter according to the coherence coefficient to control the proportion of the noise signal removed from the audio data to obtain the denoised audio data;

[0052] Step S6: using a classification model to classify the denoised audio data to obtain category information of the audio data.

[0053] Specifically, in step S1, a microphone array is used to obtain multi-channel audio data of cow sounds, and the WAV files in the obtained multi-channel audio data are separated into channels according to the number of microphone arrays. Assuming that a 6-microphone array is used, the audio file will be divided into 6 channels, and each channel corresponds to the position of a microphone. Figure 2 Compared with the first generation microphone array, the second generation microphone array used in the embodiment of the present invention has a sound pickup distance increased from 3.5 meters to 10 meters. The comparison diagram of the sound pickup distance of the microphone array is shown in Figure 3 As shown. It can output unprocessed 8-channel raw audio files, which is convenient for verification experiments. Using the universal UAC protocol, it can be directly connected to the computer without the need for special driver adaptation. It is compatible with the latest version of Nano motherboard, Raspberry Pi and other main controllers. The structural diagram of the fixing device of the ring microphone array is shown in Figure 4 As shown, the annular microphone array can be fixed to the movable device by a fixing device. In the embodiment of the present disclosure, the cow sound denoising and classification method provided by the present invention is run by using a Jetson AGX Orin computer. The structural diagram of the Jetson AGX Orin computer is shown in FIG. Figure 5 As shown. Jetson AGX Orin is designed to meet the needs of smart ranches and provide strong support for cattle call classification and health monitoring. The PWM speed-adjustable cooling fan on the top of the device can automatically adjust according to environmental changes in the ranch to ensure long-term stable operation. The power button, forced recovery button, and reset button simplify device operation and facilitate daily maintenance and emergency handling of the ranch. The Micro SD card slot expands storage space for long-term audio data collection and analysis. The network cable interface provides high-speed network connection to ensure data transmission efficiency. The USB interface is combined with the microphone array to achieve accurate collection and processing of multi-channel cattle calls and improve classification accuracy. The USB micro-B interface facilitates debugging and system monitoring, ensures stable operation of the device in complex environments, and is used to provide efficient and reliable hardware platform support.

[0054] Those skilled in the art can set different numbers of microphone arrays and the number of divided channels according to actual application scenarios and needs. The annular 6-microphone array provided in the embodiment of the present invention is only used for exemplary purposes. The channel separation method can be: channel separation based on pulse code modulation PCM, channel separation based on spectrum processing, or channel separation based on deep learning, etc. Since channel separation technology is relatively mature and there are many ways to achieve this goal, those skilled in the art can choose different channel separation methods according to actual needs, and the embodiment of the present invention will not repeat the relevant content.

[0055] In step S2, multiple single-channel audio data are merged into one channel through beamforming technology, which can suppress interference signals from non-target directions and enhance speech signals in the target direction. This feature can significantly improve the signal-to-noise ratio of the received signal in complex environments, such as when there are multiple sound sources, noise interference or reverberation effects, and merge the sound signals collected by multiple microphones into a higher-quality, clearer target sound source signal.

[0056] By performing noise audio estimation on the merged audio data, it is used to identify and estimate the state of noise in advance before denoising the audio data, thereby providing key parameters for subsequent denoising processing and achieving a more accurate denoising effect.

[0057] Steps S3 to S5 are the specific process of the denoising method proposed in the embodiment of the present invention. Figure 6 As shown, the spectral subtraction method is improved to the cyclic coherent spectral subtraction method, and the adaptive parameter α is introduced to form the cyclic adaptive coefficient coherent spectral subtraction method proposed in the embodiment of the present disclosure. The audio data is converted to the frequency domain by performing STFT short-time Fourier transform and smoothing processing on the merged audio data to obtain the spectrum amplitude of the merged audio data. According to the spectrum amplitude of the audio data, the coherence coefficient of the merged audio data and the noise signal is calculated.

[0058] In some embodiments, calculating the coherence coefficient of the merged audio data and the noise signal according to the frequency spectrum amplitude of the audio data includes: calculating the coherence degree of the merged audio data and the noise signal at frequency f, and the calculation formula is:

[0059]

[0060] Wherein, γ(f) is the coherence coefficient, X(f) represents the complex spectrum of the audio signal in the audio data at the frequency f, N(f) represents the complex spectrum of the noise signal at the frequency f, and E[·] represents the expected value between X(f) and N(f). represents the average approximate value of the complex spectrum of the noise signal at frequency f. By introducing the average value of multiple frame signals, the robustness of coherence calculation is improved.

[0061] In some embodiments, the α parameter is adaptively adjusted according to the coherence coefficient, and its calculation formula is:

[0062]

[0063] Among them, α0 is the spectral reduction coefficient, which is used to adjust the intensity of noise reduction, γ(f) is the coherence coefficient calculated at frequency f, ∈ is a set value, which is used to avoid the situation where the denominator is zero, and αmax It is the maximum value of α(f), which is used to prevent excessive noise reduction caused by α(f) being too large. Through this adaptive adjustment, α(f) can be reduced when the coherence is high (the audio signal and the noise spectrum are more correlated), thereby protecting the signal. At frequencies with lower coherence (the audio signal and the noise spectrum are uncorrelated), α(f) is increased to more effectively remove noise. According to the coherence of the signal, the α parameter is dynamically adjusted to make the denoising process more flexible and achieve better denoising effects in different noise environments.

[0064] In some embodiments, according to the adaptive adjustment of the α parameter, the calculated noise spectrum is subjected to a spectrum subtraction operation to remove the noise signal, and the audio data after the noise signal is removed is subjected to an inverse short-time Fourier transform to convert the audio data from the frequency domain to the time domain for subsequent classification of the audio data. In addition, steps S3 to S5 are repeated for a preset number of cycles, and when the number of cycles is reached, denoising is completed. The number of cycles is obtained by experiments.

[0065] The disclosed embodiments experimentally verify the denoising method provided, aiming to evaluate the performance of various audio denoising algorithms. Various classic and improved denoising algorithms including spectral subtraction, cyclic spectral subtraction, coherent spectral subtraction, Wiener filtering, LMS filtering, and adaptive coefficient coherent spectral subtraction are tested in the experiment. The experiment comprehensively analyzes the denoising effect and signal fidelity of each algorithm by comparing key indicators such as signal-to-noise ratio (SNR), root mean square error (RMSE), and signal smoothness before and after denoising. The experimental results show that the cyclic spectral subtraction method performs well in improving the SNR through multiple iterations, while the adaptive coefficient coherent spectral subtraction method provides excellent signal smoothness while maintaining a high SNR, which is suitable for application in smart ranch management systems. Table 1 is a comparison of the denoising effects of different post-filtering algorithms.

[0066] Table 1 Comparison of denoising effects of different post-filtering algorithms

[0067]

[0068] In the experiment, cyclic spectral subtraction*8 and cyclic adaptive coefficient coherent spectral subtraction*2 showed significant denoising effects. Cyclic spectral subtraction*8 achieved the highest signal-to-noise ratio (SNR 16.78dB) while maintaining the lowest root mean square error (RMSE 0.033106) and good smoothness (0.017145), indicating that this method can effectively remove noise without damaging the signal in multiple iterations. Cyclic adaptive coefficient coherent spectral subtraction*2, with excellent performance of SNR 14.77dB, RMSE 0.041745, and smoothness 0.015025 in a small number of iterations, showed higher signal smoothness and adaptability, and is suitable for application scenarios with high requirements for sound quality. Both methods have demonstrated excellent noise reduction capabilities in different application scenarios and are ideal choices for efficient denoising.

[0069] The present invention also compares the filtering effects of using different denoising algorithms to process the audio signal. Figure 7 As shown. In this experiment, the effects of the original audio signal after being processed by different noise reduction algorithms, including cyclic coherence spectral subtraction, multiple cyclic spectral subtraction, and Wiener filtering, were compared. The results show that after multiple iterations (1 and 2 times), cyclic coherence spectral subtraction can significantly reduce noise and maintain signal integrity, especially in dual-channel speech processing. It shows excellent noise reduction capabilities. In contrast, although Wiener filtering also improves signal quality, it is slightly inferior in noise suppression and signal fidelity. Overall, cyclic coherence spectral subtraction shows excellent denoising performance and high signal fidelity under multiple iterations, which is suitable for smart ranch scenarios with high requirements for audio processing.

[0070] The embodiments of the present invention also compare the denoising methods and application scenarios in different patents, and summarize the differences therein, as shown in Table 2.

[0071] Table 2 Comparison of qualitative effects of various patents

[0072]

[0073]

[0074] Compared with other mentioned patents, this invention shows unique advantages in application scenarios and technical methods. First, it is designed for dynamic and unpredictable pasture environments and can adapt to changing external conditions through scene recognition, which is rare in other patents, most of which are aimed at relatively stable or specific environments (such as classrooms, music production rooms, etc.). This scene adaptability makes it more advantageous when facing sudden environmental noise.

[0075] In addition, the flexibility of the strategy is to dynamically adjust the denoising strategy through specific scene analysis instead of adopting a general denoising method, which makes efficient and accurate noise management possible. In contrast, although the music and vocal separation method of invention patent CN116153327A has a clear application scenario in music production, its method may not have the same effect in non-music scenarios, and the denoising method of deep learning cannot guarantee the real-time requirements in the ranch environment.

[0076] The present invention does not rely too much on complex hardware configuration, but achieves the purpose of noise reduction through algorithm optimization, which is different from the noise reduction headphones in utility model CN215499532U that require specific physical equipment. Although the latter provides physical noise isolation, it may be more limited in cost and implementation. Such an approach not only saves costs, but also increases the practicality and flexibility of the system, and is suitable for a wide range of ranch environments and possible other similar application scenarios.

[0077] In summary, the present invention demonstrates significant advantages in addressing audio processing requirements in complex and changing environments, and provides a highly adaptive and cost-effective solution.

[0078] The audio data after denoising is classified by using a classification model as mentioned in step S6. In some embodiments, the classification function is implemented by a deep learning model ECAPA-TDNN model. The ECAPA-TDNN model includes a part consisting of multiple sequentially linked SE-Res2Block modules. The core components of this module include key components such as TDNN+ReLU+BN, SE-Res2Block, Attentive Statistics Pooling (ASP) and Multilayer Feature Aggregation (MFA). And the AAM-Softmax loss function is used for optimization. The structural schematic diagram of the ECAPA-TDNN classification model provided in an embodiment of the present invention is shown in the figure. Figure 8 As shown. It should be noted that the classification model provided by the present invention is only for illustrative purposes, and those skilled in the art may use any other classification model that can achieve the audio classification effect. Since the technical effects of the components used in the classification model of the present invention are the same as those of the prior art, they will not be described in detail. It can be understood that the category information obtained by the classification model includes the category and confidence of the cattle sound audio data, such as chewing sounds, low calls, roars, etc.

[0079] Randomly and uniformly distributed unbiased noise is added to all channels of the collected multi-channel original audio data of cattle voices, and the noisy audio is uniformly denoised using the denoising method provided in the embodiment of the present invention. The ECAPA-TDNN classification model is used on the processed audio files to obtain the classification accuracy and loss value. The comparison is performed with different denoising methods, and the comparison results are shown in Table 3.

[0080] Table 3 Comparison of classification effects after denoising by different post-filtering algorithms

[0081]

[0082] In some embodiments, before merging the multiple single-channel audio data into one channel using beamforming technology, the method further includes: performing signal processing on the multiple single-channel audio data respectively to obtain position information of the multiple single-channel audio data. The signal processing includes: obtaining the position information of the multiple single-channel audio data using delay estimation according to the time difference of receiving the multiple single-channel audio data and the position information of the audio data acquisition device that collects the multi-channel cattle sound audio data; performing delay compensation on the multiple single-channel audio data after delay estimation to adjust the time difference between the multiple single-channel audio data.

[0083] The time delay estimation sound source localization method is currently the mainstream positioning method. Time delay estimation can be based on the generalized cross-correlation function time delay estimation method. Due to the presence of noise and reverberation, the peak of the traditional cross-correlation function is not obvious, resulting in inaccurate estimation. Consider using the generalized cross-correlation method. That is, first convert the time domain to the frequency domain, perform normalization in the frequency domain to achieve the effect of noise reduction, and then perform inverse Fourier transform to the time domain. The generalized cross-correlation function can reduce the impact of noise and reverberation in the actual environment.

[0084] Suppose there is a sound source in space, denoted as s(t), and its position in space is S. The two microphones are denoted as m1 and m2, and their positions in space are denoted as M1 and M2 respectively. The received audio signals are denoted as x1(t) and x2(t).

[0085] The audio signals received by microphones m1 and m2 are:

[0086] x1(t)=s(t-τ1)+n1(t)

[0087] x2(t)=s(t-τ2)+n2(t)

[0088] Where τ1 and τ2 are the delay times of the audio signal emitted by the sound source reaching the two microphones, and n1(t) and n2(t) are additive noises. Then the arrival time difference TDOA of the audio signal emitted by the sound source reaching the two microphones is:

[0089] τ=τ1-τ2

[0090] τ1 and τ2 can be calculated by the following formula:

[0091]

[0092] Where c is the speed of sound. Select the position of a microphone as the reference position, for example, M2 as the reference position, then τ2 = 0. When the geometry of the microphone array is known, the sound source localization problem becomes a delay estimation problem.

[0093] In order to detect the azimuth and distance of the sound source, the microphone array is arranged in an equilateral triangle, with the center of the triangle as the coordinate origin. The distance from the sound source to each microphone MIC is R1, R2, and R3, as shown in Fig. 9 When the length of the triangle side is known, the azimuth and distance R can be calculated for the position of any sound source by using the time delay difference from the sound source to each MIC.

[0094] Create a coordinate system such as Fig. 9 The dotted line on the right, the center of gravity is the origin O(0,0), the side length of the equilateral triangle is L, the distance difference from the sound source p to each vertex is: D1=R2-R1; D2=R3-R1; D3=R3-R2, the polar coordinates of the sound source point S to be determined are (R, Φ).

[0095] The polar coordinates of the vertices of the three MICs are Mic a: (r,π / 2), Mic b: (r.4π / 3), and Mic c: (r,1π / 6).

[0096] The relationship between the sides and angles of a triangle is:

[0097]

[0098] Substituting the known variables into the following formula, any two expressions can be connected to immediately solve for the values ​​of Φ and R.

[0099]

[0100] Since different audio signals are received by the microphone at different times, before the multiple single-channel audio data are subsequently beamformed, the time differences between the different signals are aligned through delay compensation.

[0101] In some embodiments, the method also includes: receiving cattle video data, performing frame-by-frame depth estimation on the video data using a depth estimation model and generating a depth map; and mapping the position information to a video frame image of the video data in combination with the depth map, and based on the position information, marking the category information corresponding to the position information to a corresponding position in the video frame image.

[0102] Specifically, by using a video acquisition device to collect videos of cattle, the depth estimation technology is used to accurately predict and visualize the distance information in the video scene. Based on the LapDepth depth estimation model (LDRN), the depth of the video frame is estimated frame by frame and a pseudo-color depth map is generated. By processing the input video frame image, the model successfully predicts the depth distribution of the scene, such as Fig.10 As shown in the figure, the depth information of the target at different distances is clearly displayed. Experimental results show that the model can effectively capture the level of details in the scene, such as the clear separation of background and foreground. The generated depth map maintains good accuracy in high-contrast areas, especially for the geometric contours of close objects, providing strong support for subsequent 3D scene reconstruction and target detection.

[0103] According to the acquired position information of the audio data, combined with the above-mentioned depth map, the position information of the audio data is mapped to the depth map accordingly, and the position information of the audio data in the three-dimensional scene is obtained as follows: Fig.11 As shown in the figure, the effect of mapping the position information of the audio data to the two-dimensional plane is as follows Fig.12 As shown. The position information in the three-dimensional scene and the position information in the two-dimensional plane are matched with the video frame image, so as to realize the mapping of the position information of the acquired audio data to the video frame image. According to the position information of the audio data and the category information corresponding to the position information obtained by the classification model, the corresponding position in the video frame image is marked, and the sound source category and position are visualized by color marking. The visualization result is shown as follows: Fig.13 shown.

[0104] In some embodiments, the cattle sound denoising and classification method provided by the present invention is applied in an actual cattle farm environment. The cattle sound signal is collected by a circular microphone array. After the collected signal is subjected to delay compensation and beamforming, it is input into the coherent spectrum subtraction module for denoising. The processed signal is then subjected to a subsequent classification algorithm for audio classification, thereby accurately identifying different types of cattle sounds, such as chewing sounds, low moans, roars, etc. The operation flow chart of the cattle sound denoising and classification method in a cattle farm environment is shown in FIG. Fig.14 shown.

[0105] In some embodiments, the cattle sound denoising and classification method provided by the present invention is also applied to the automatic monitoring system of the ranch. By using a ranch management robot equipped with a microphone array to collect and process cattle sounds in real time, abnormal conditions, such as changes in the health of cattle, can be automatically detected. By using coherent spectrum subtraction, the system can still maintain a high recognition accuracy rate in a strong noise environment, providing reliable support for ranch management. The workflow diagram of the ranch management robot is shown in the figure below: Fig.15 The schematic diagram of the structure of the ranch management robot with microphone array used in the embodiment of the present disclosure is shown in FIG. Fig.16 As shown. Among them, it includes: Jetson AGX Orin protective case, power storage, chassis, Mecanum wheels, retractable camera bracket, microphone array fixing place and screen bracket. The Jetson AGX Orin protective case is designed to protect the Jetson AGX Orin processor and other key control components, and also serves as a chassis for carrying microphone arrays and other equipment brackets. The power storage is used to place power units, such as "electric second" or other power supply equipment. It is convenient for users to replace the power supply and ensure good electrical insulation and safety protection. As the structural foundation of the equipment, the chassis not only needs to provide physical support, but also has sufficient structural strength and stability to withstand the load and external impact of the equipment during operation. Mecanum wheels allow the robot to move in any direction without turning, which makes the robot highly maneuverable in narrow or complex environments. The design of Mecanum wheels needs to ensure sufficient durability and the ability to adapt to different ground conditions. The retractable camera bracket design allows the camera to be extended and adjusted when necessary, providing a wider viewing angle and flexible monitoring options, suitable for different monitoring needs. The microphone array fixing part is a fixing device for the microphone array to ensure that the microphone array is stably installed on the robot. There are four drill holes for installing the microphone array. The screen bracket is used to install display devices, such as operation panels or monitoring displays, so that the operator can directly view the device status or control the robot.

[0106] The second aspect of the embodiment of the present invention also provides a cow sound denoising and classification system, the system comprising: a multi-channel audio acquisition module, used to collect multi-channel cow sound audio data, and perform channel separation to obtain multiple single-channel audio data; a denoising module, used to merge the multiple single-channel audio data into one channel by using beamforming technology, obtain the noise signal in the merged audio data by noise audio estimation, perform STFT short-time Fourier transform and smoothing processing on the merged audio data to obtain the spectrum amplitude of the merged audio data, calculate the coherence coefficient of the merged audio data and the noise signal according to the spectrum amplitude of the audio data, and adaptively adjust the α parameter according to the coherence coefficient to control the proportion of the noise signal in the audio data to be removed to obtain the denoised audio data; and a classification module, used to classify the denoised audio data using a classification model to obtain category information of the audio data.

[0107] In some embodiments, the system also includes: a video processing module, which collects cattle video data, uses a depth estimation model to perform frame-by-frame depth estimation on the video data and generate a depth map; a visualization module, which is used to map the position information to the video frame image of the video data in combination with the depth map, and based on the position information, annotates the category information corresponding to the position information to the corresponding position in the video frame image.

[0108] A third aspect of an embodiment of the present invention further provides a machine-readable storage medium, on which instructions are stored, and the instructions are used to enable a machine to execute any of the above-mentioned methods for denoising and classifying cow sounds in the present application.

[0109] A fourth aspect of an embodiment of the present invention further provides a processor, characterized in that it is used to run a program, wherein the program is used to execute the cattle voice denoising and classification method when it is run.

[0110] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0111] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0112] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0113] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0114] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0115] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0116] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0117] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0118] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A method for denoising and classifying cattle voices, characterized in that: The method comprises: Receive multi-channel audio data and perform channel separation to obtain multiple single-channel audio data; After merging the plurality of single-channel audio data into one channel by using beamforming technology, obtaining a noise signal in the merged audio data by noise audio estimation; Performing STFT short-time Fourier transform and smoothing processing on the merged audio data to obtain the frequency spectrum amplitude of the merged audio data; Calculating a coherence coefficient between the combined audio data and the noise signal according to a frequency spectrum amplitude of the audio data; Adaptively adjusting an α parameter according to the coherence coefficient to control the proportion of the noise signal removed from the audio data to obtain the denoised audio data; and The denoised audio data is classified using a classification model to obtain category information of the audio data.

2. The method for denoising and classifying cattle sounds according to claim 1, characterized in that: Before combining the plurality of single-channel audio data into one channel using beamforming technology, the method further includes: Signal processing is performed on the multiple single-channel audio data respectively to obtain position information of the multiple single-channel audio data.

3. The method for denoising and classifying cattle sounds according to claim 2, characterized in that: The signal processing includes: Obtaining the position information of the plurality of single-channel audio data by using time delay estimation according to the time difference of receiving the plurality of single-channel audio data and the position information of the audio data acquisition device that acquires the multi-channel cattle sound audio data; Delay compensation is performed on the multiple single-channel audio data after delay estimation to adjust the time difference between the multiple single-channel audio data.

4. The method for denoising and classifying cattle sounds according to claim 2, characterized in that: The method further comprises: Receiving cattle video data, performing frame-by-frame depth estimation on the video data using a depth estimation model and generating a depth map; and The position information is mapped to the video frame image of the video data in combination with the depth map, and category information corresponding to the position information is marked at a corresponding position in the video frame image according to the position information.

5. The method for denoising and classifying cattle sounds according to claim 1, characterized in that: Calculating the coherence coefficient of the combined audio data and the noise signal according to the frequency spectrum amplitude of the audio data, including: The coherence degree between the combined audio data and the noise signal at frequency f is calculated using the following calculation formula: Wherein, γ(f) is the coherence coefficient, X(f) represents the complex spectrum of the audio signal in the audio data at the frequency f, N(f) represents the complex spectrum of the noise signal at the frequency f, and E[·] represents the expected value between X(f) and N(f). represents the average approximation of the complex spectrum of the noise signal at frequency f.

6. The method for denoising and classifying cattle sounds according to claim 1, characterized in that: According to the coherence coefficient, the α parameter is adaptively adjusted, and its calculation formula is: Among them, α0 is the spectral reduction coefficient, which is used to adjust the intensity of noise reduction, γ(f) is the coherence coefficient calculated at frequency f, ∈ is a set value, which is used to avoid the situation where the denominator is zero, and α max It is the maximum value of α(f), which is used to prevent α(f) from being too large and causing excessive noise reduction.

7. A cattle voice denoising and classification system, characterized in that: The system comprises: A multi-channel audio acquisition module is used to collect multi-channel cattle sound audio data and separate its channels to obtain multiple single-channel audio data; A denoising module is used to combine the multiple single-channel audio data into one channel by using beamforming technology, and obtain a noise signal in the combined audio data by noise audio estimation, Performing STFT short-time Fourier transform and smoothing processing on the merged audio data to obtain the frequency spectrum amplitude of the merged audio data, Calculating the coherence coefficient between the combined audio data and the noise signal according to the frequency spectrum amplitude of the audio data, Adaptively adjusting an α parameter according to the coherence coefficient to control the proportion of the noise signal removed from the audio data to obtain the denoised audio data; and The classification module is used to classify the denoised audio data using a classification model to obtain category information of the audio data.

8. The cattle sound denoising and classification system according to claim 7, characterized in that: The system further comprises: A video processing module collects cattle video data, uses a depth estimation model to perform frame-by-frame depth estimation on the video data and generates a depth map; A visualization module is used to map the position information to the video frame image of the video data in combination with the depth map, and to annotate the category information corresponding to the position information to the corresponding position in the video frame image according to the position information.

9. A machine-readable storage medium having instructions stored thereon, the instructions being used to enable a machine to execute any of the above-mentioned methods for denoising and classifying cattle voices of the present application.

10. A processor, characterized in that: Used to run a program, wherein the program, when run, is used to execute: the cattle sound denoising and classification method according to any one of claims 1-6.