A method for recognizing the vocalizations of golden pomfret schools in cages based on dual-channel time-varying Wiener filtering
Through the dual-channel time-varying Wiener filtering method, noise interference is eliminated and the sound signal of golden pomfret is enhanced, achieving efficient identification and positioning in complex noise environments, and solving the problem of low monitoring efficiency in existing technologies.
Patent Information
- Application Number
- CN202510974052.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-07-15
AI Technical Summary
Existing acoustic recognition methods are difficult to effectively identify the sounds of caged golden pomfret schools, especially in non-stationary noise environments. They are unable to effectively remove noise interference and require a large amount of training data, resulting in low monitoring efficiency.
A method based on dual-channel time-varying Wiener filtering is adopted. Through dual-channel frequency domain signal processing, equalizer and compensation factor are used to eliminate noise interference, and a posteriori and a priori signal-to-noise ratio calculations are combined to achieve target signal enhancement and noise suppression without the need for a large amount of training data.
It can quickly and accurately identify the sounds of golden pomfret in complex noise environments, has the ability to locate targets, and achieve near-real-time monitoring, solving the problems of poor performance and dependence on training data of traditional methods under non-stationary noise conditions.
Smart Images

Figure CN120472913B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of sound recognition of golden pomfret schools, and in particular to a method for sound recognition of golden pomfret schools in cages based on dual-channel time-varying Wiener filtering. Background Art
[0002] In cage aquaculture, it is crucial to keep abreast of the status of fish schools. By deeply analyzing the sounds of fish schools through passive acoustic technology and utilizing their sound characteristics, the relationship between fish school sounds and their status can be established. This can better grasp the status of fish schools and contribute to more scientific and efficient cage aquaculture management.
[0003] Current speech recognition methods are mainly divided into six categories:
[0004] (1) Spectral subtraction and its extensions. Based on the estimation of the noise spectrum, noise control is achieved by subtracting the noise spectrum from the spectrum of the noisy audio signal. This method is easy to implement, but has obvious disadvantages: it can only cope with scenes with relatively stable noise and is not very effective for non-stationary noise.
[0005] (2) Minimum Statistics Method. This method achieves optimal noise suppression by tracking the minimum spectral amplitude of each signal frequency band and minimizing the mean square error between the signal and noise. It can effectively recover signals with a high signal-to-noise ratio, but it requires a good estimate of the statistical characteristics of the signal and noise, has high computational complexity, and is less efficient when processing data on a large time scale.
[0006] (3) Probabilistic model estimation methods based on pure target signals and noise-contaminated signals. Relatively successful examples include minimum mean square error spectrum estimation and maximum likelihood spectrum amplitude estimation. When available adaptive data is limited, such methods can suffer from overfitting, resulting in poor classification or regression performance.
[0007] (4) Modeling based on the maximum a posteriori probability criterion. This method considers a certain prior distribution of the input signal during modeling, effectively avoiding overfitting. However, this method has a significant disadvantage: it can cause severe distortion of the target signal when the input signal-to-noise ratio is high.
[0008] (5) Deep learning-based methods. Using machine learning methods such as deep neural networks (DNNs) and acoustic signal enhancement generative adversarial networks (SEGANs), the mapping relationship between a large number of noise-contaminated signals and their corresponding pure signals is learned to achieve target signal recognition. This method achieves excellent denoising results under complex noise conditions, but its drawbacks include its high reliance on data, the requirement for large amounts of high-quality labeled data for training, and a lack of transparency in its principles. In actual fish monitoring, preparing large amounts of training data is not only labor-intensive, but also impractical to obtain pure golden pomfret vocalizations in blind conditions. Furthermore, the extensive training time significantly impacts overall efficiency, and its lack of interpretability adds uncertainty to subsequent analysis.
[0009] (6) Non-negative matrix factorization (NMF). This method separates noise from the signal source by modeling signal-related features and encoding matrices. This method can effectively handle non-stationary noise backgrounds, but it has similar shortcomings to deep learning methods. It requires a large amount of supervised training data to achieve optimal performance, and the model training process takes a long time. This, in actual monitoring work, will also seriously affect the timeliness of monitoring results.
[0010] Under natural environmental conditions, the collected acoustic signals are often mixed with environmental noise and man-made noise. When monitoring underwater environmental noise on deep-sea aquaculture platforms, the sound of waves on the sea surface, the noise of platform equipment operating, ship noise, and even distant aircraft passing by can all contribute to noise interference. Due to the combined effects of these factors, the cage monitoring signal appears as a non-stationary signal with a high signal-to-noise ratio. Removing this type of non-stationary noise has always been a difficult problem in acoustic signal processing. Currently, many common denoising methods, such as spectral subtraction, Wiener filtering, minimum mean square error spectrum estimation, and maximum likelihood spectrum amplitude estimation, can effectively suppress the influence of noise in certain specific environments. However, they are unable to effectively identify target signals when faced with non-stationary noise conditions.
[0011] In recent years, various deep learning-based extended methods and non-negative matrix factorization (NMF) have demonstrated strong performance in noise separation. However, they all suffer from a significant drawback: they require a large amount of training data to achieve optimal results. Specifically, when using these two methods or their extended approaches for denoising, it is necessary to pre-collect the original data of the object under study and separate the pure target signal to serve as training data for the model. In natural environments, noise interference is ubiquitous, and collecting the original data and then separating the pure signal source is costly, time-consuming, and difficult, making this premise difficult to achieve under real-world conditions.
[0012] In conclusion, existing acoustic recognition methods are difficult to use in the vocalization recognition of golden pomfret schools in cages. Summary of the Invention
[0013] In view of the above-mentioned deficiencies in the prior art, the present invention provides a method for identifying the vocalizations of golden pomfret schools in cages based on dual-channel time-varying Wiener filtering, which solves the problem that the existing acoustic recognition methods are difficult to use in identifying the vocalizations of golden pomfret schools in cages.
[0014] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:
[0015] A method for recognizing the vocalizations of a school of golden pomfret in a cage based on a dual-channel time-varying Wiener filter comprises the following steps:
[0016] The dual-channel acoustic signals from the golden pomfret breeding area are collected and frame segmented. A Hamming window is added to each frame and converted into a frequency domain signal through fast Fourier transform to obtain a dual-channel frequency domain signal; the dual channels are the left channel and the right channel;
[0017] Based on the dual-channel frequency domain signal, the left and right equalizers are trained separately by the normalized least mean square method to make the target signal parts in the left and right channels consistent.
[0018] Based on the left and right equalizers, the left and right channel frequency domain signals are spectrally aligned and the target signal is eliminated as much as possible to obtain an ideal estimation result of the background noise interference signal of the left and right channels;
[0019] By introducing a compensation factor into the ideal estimation results of the background noise interference signals of the left and right channels, the distortion caused by the equalizer on the background noise interference signal is corrected, and the estimated results of the compensated background noise interference signal are obtained;
[0020] The posterior signal-to-noise ratio is calculated based on the estimation results of the dual-channel acoustic signal of the target area and the compensated background noise interference signal; the prior signal-to-noise ratio is calculated based on the posterior signal-to-noise ratio through a decision-guided method;
[0021] A gain function is constructed based on the prior signal-to-noise ratio, and the gain function is multiplied by the frequency domain signals of the left and right channels respectively to obtain enhanced frequency domain signals of the left and right channels;
[0022] The frequency domain signals of the left and right channels are respectively subjected to inverse fast Fourier transform and smoothed together using a Hanning window to obtain the reconstructed time domain signals, that is, the sound signals of the left and right channels of the golden pomfret.
[0023] Furthermore, when performing frame segmentation, each frame is 20 ms and the frame shift is 10 ms.
[0024] Furthermore, the expression converted into frequency domain signal by fast Fourier transform is:
[0025]
[0026]
[0027] in for Frame Frequency: left channel frequency domain signal; for Frame Frequency: right channel frequency domain signal; represents the fast Fourier transform conversion; is the left channel acoustic signal; is the right channel acoustic signal.
[0028] Furthermore, the expressions for training the left equalizer and the right equalizer respectively by the normalized least mean square method are:
[0029]
[0030]
[0031] in express Left equalizer under the frame; express Right equalizer under the frame; express Left equalizer under the frame; express Right equalizer under the frame; is a constant; represents the L2 norm; for Left channel frequency domain signal under the frame; for The right channel frequency domain signal under the frame; the superscript T represents the transpose of the matrix; during the training process, a white noise signal is used as a proxy for the sound signal of the golden pomfret, and the direction of the target signal source is simulated by the transfer function to ensure training accuracy.
[0032] Furthermore, based on the left and right equalizers, the left and right channel frequency domain signals are spectrally aligned and the target signal is eliminated as much as possible. The expression for the ideal estimation result of the background noise interference signal of the left and right channels is obtained as follows:
[0033]
[0034]
[0035] in and They are Frame The ideal estimation result of the background noise interference signal of the left channel and Frame Frequency ideal estimation result of right channel background noise interference signal; for Frame Frequency: left channel frequency domain signal; for Frame Frequency: right channel frequency domain signal; and They are Frame Frequency left and right equalizers.
[0036] Furthermore, the expression of the compensation factor is:
[0037]
[0038] in for Frame Frequency compensation factor; Expressing hope; express Frame Frequency dual-channel frequency domain signal; Represents the ideal estimation result of the background noise interference signal of the left and right channels The complex conjugate of It means taking the square of the magnitude; Indicates the left channel; Indicates the right channel.
[0039] Furthermore, the expression for introducing the compensation factor into the ideal estimation result of the background noise interference signal of the left and right channels is:
[0040]
[0041]
[0042] in and After compensation Frame Frequency estimation results of left channel background noise interference signal and compensated Frame Frequency estimation result of right channel background noise interference signal; for Frame Frequency ideal estimation result of left channel background noise interference signal; for Frame Frequency is the ideal estimation result of the background noise interference signal in the right channel.
[0043] Furthermore, the calculation expressions of the posterior signal-to-noise ratio and the prior signal-to-noise ratio are:
[0044]
[0045]
[0046] in for Frame the posterior signal-to-noise ratio of the frequency; for Frame the a priori signal-to-noise ratio of the frequency; It is a dual-channel acoustic signal for the golden pomfret breeding area; for Frame The estimation result of the background noise interference signal of frequency is given by and composition; is the smoothing factor; For the current Frame Frequency enhancement results; for Frame The estimation result of the background noise interference signal frequency; Indicates taking the maximum value; the calculation expression of the priori signal-to-noise ratio of the first frame is: .
[0047] Furthermore, the expression of the gain function based on the prior signal-to-noise ratio is:
[0048]
[0049] in for Frame Gain function of frequency.
[0050] Furthermore, the expression for multiplying the gain function by the left and right channel frequency domain signals is:
[0051]
[0052]
[0053] in for Frame The frequency domain signal of the left channel after frequency enhancement; for Frame The frequency domain signal of the right channel after frequency enhancement; for Frame Frequency: left channel frequency domain signal; for Frame Frequency: right channel frequency domain signal.
[0054] The beneficial effects of the present invention are:
[0055] 1. Modeling the left and right channels of the input signal separately, offsetting the target signal to analyze the signal-to-noise ratio characteristics, and combining the Wiener filtering method to achieve the enhancement of the target biological sound and the suppression of the noise signal. It can identify the vocalizations of golden pomfret in complex noise environments and has the ability to locate the target.
[0056] 2. The present invention introduces a dual-channel signal equalization and interference compensation mechanism, integrating a posteriori and a priori signal-to-noise ratio estimation. It can effectively identify the vocalizations of golden pomfret schools under complex noise conditions without the need for training data. This solves the problem that traditional noise separation methods have poor performance in separating non-stationary noise and require a large amount of training data. It provides a technical prerequisite for cage acoustic monitoring and creates the possibility for subsequent rapid, efficient, and accurate analysis of cage fish vocalizations.
[0057] 3. The present invention has its own preprocessing step. The required input data is a time domain signal. The raw data obtained by the mainstream monitoring means in the industry can be directly input into the algorithm (if there are other requirements, the user can filter the data according to their actual needs before inputting the algorithm). There is no need for tedious preprocessing, and on-site sampling and on-site analysis can be realized. Various problems can be handled in a timely manner, achieving the effect of nearly real-time monitoring of the vocal status of fish schools in cages. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 Schematic diagram of the process of this method;
[0059] Figure 2 It is a schematic diagram of the overall working structure of the present invention. DETAILED DESCRIPTION
[0060] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
[0061] like Figure 1 and Figure 2 As shown, the method for recognizing the sound of a school of golden pomfret in a cage based on dual-channel time-varying Wiener filtering includes the following steps:
[0062] S1, collecting dual-channel acoustic signals from the golden pomfret breeding area and performing frame segmentation, adding a Hamming window to each frame and converting it into a frequency domain signal through fast Fourier transform to obtain a dual-channel frequency domain signal; wherein the dual channels are a left channel and a right channel;
[0063] S2. Based on the dual-channel frequency domain signal, the left equalizer and the right equalizer are trained separately by the normalized least mean square method to make the target signal parts in the left and right channels tend to be consistent;
[0064] S3, based on the left equalizer and the right equalizer, aligning the spectrum of the left and right channel frequency domain signals and eliminating the target signal as much as possible to obtain an ideal estimation result of the background noise interference signal of the left and right channels;
[0065] S4. By introducing a compensation factor into the ideal estimation results of the background noise interference signals of the left and right channels, the distortion of the background noise interference signal caused by the equalizer is corrected to obtain an estimation result of the compensated background noise interference signal;
[0066] S5. Calculate the posterior signal-to-noise ratio based on the estimated results of the dual-channel acoustic signal of the target area and the compensated background noise interference signal; calculate the prior signal-to-noise ratio using a decision-directed method based on the posterior signal-to-noise ratio;
[0067] S6. constructing a gain function based on the prior signal-to-noise ratio, and multiplying the gain function by the left and right channel frequency domain signals respectively to obtain enhanced left and right channel frequency domain signals;
[0068] S7. Perform inverse fast Fourier transform on the frequency domain signals of the left and right channels respectively and use a Hanning window to smoothly splice them to obtain reconstructed time domain signals, that is, obtain the sound signals of the left and right channels of the golden pomfret.
[0069] In step S1, the dual-channel acoustic signal of the golden pomfret breeding area is a time domain form of a mixture of left and right channels. In the time domain, the signal can be understood as the superposition of the target signal and the background noise. is the left channel mixed signal, is the right channel mixed signal, which can be modeled as:
[0070]
[0071]
[0072] in is the target signal, i.e. the sound signal of golden pomfret, is background noise. Next, the signal is frame-segmented (e.g., if each frame is 20ms, then the frame shift is 10ms), and a Hamming window is applied to each frame. A fast Fourier transform is performed on each frame to obtain the spectrum of each frame, converting the time domain signal into a frame-level frequency domain signal:
[0073]
[0074]
[0075] in for Frame Frequency left channel frequency domain signal (spectrum); for Frame Frequency right channel frequency domain signal (spectrum); represents the fast Fourier transform conversion; is the left channel acoustic signal; is the right channel acoustic signal.
[0076] In step S2, the expressions for training the left equalizer and the right equalizer respectively by the normalized least mean square method (NLMS) are:
[0077]
[0078]
[0079] in express Left equalizer under the frame; express Right equalizer under the frame; express Left equalizer under the frame; express Right equalizer under the frame; is a constant, ranging from 0.01 to 0.1; represents the L2 norm; for Left channel frequency domain signal under the frame; for The right channel frequency domain signal under the frame; the superscript T indicates the transpose of the matrix.
[0080] In this embodiment, white noise is used as a proxy for the sound signal of golden pomfret during training because it has full-band coverage characteristics, and the direction of the target signal source is simulated by the transfer function to ensure training accuracy. This step will find an optimal complex coefficient filter for each of the two channels. , so that one of the channel signals can be approximated to the other channel signal by multiplying this filter, that is, canceling the difference between the two.
[0081] In step S1, the input signal (the dual-channel acoustic signal of the golden pomfret breeding area) has been segmented according to the frame. In step S3, based on the left and right equalizers, the spectrum of the left and right channel frequency domain signals is aligned and the target signal is eliminated as much as possible. The expression for the ideal estimation result of the background noise interference signal of the left and right channels is:
[0082]
[0083]
[0084] in and They are Frame The ideal estimation result of the background noise interference signal of the left channel and Frame Frequency ideal estimation result of right channel background noise interference signal; for Frame Frequency: left channel frequency domain signal; for Frame Frequency: right channel frequency domain signal; and They are Frame Frequency left and right equalizers.
[0085] The purpose of step S3 is to align the spectrum of the left and right channels based on the EC psychoacoustic model, analyze the target signal components, eliminate the target signal as much as possible, retain the background noise interference to assess the degree of noise pollution, and determine whether this part of the soundscape tends to be noisy or quiet.
[0086] After calculating the corresponding noise interference situation for each signal segment, the distortion caused by the equalization filter on the background noise signal is corrected to improve the accuracy of the interference signal sound energy estimation. Therefore, a compensation factor is introduced in step S4, and its expression is:
[0087]
[0088] in for Frame Frequency compensation factor; Expressing hope; express Frame Frequency dual-channel frequency domain signal; Represents the ideal estimation result of the background noise interference signal of the left and right channels The complex conjugate of It means taking the square of the magnitude; Indicates the left channel; Indicates the right channel.
[0089] The expression for introducing the compensation factor into the ideal estimation result of the background noise interference signal of the left and right channels is:
[0090]
[0091]
[0092] in and After compensation Frame Frequency estimation results of left channel background noise interference signal and compensated Frame Frequency estimation result of right channel background noise interference signal; for Frame Frequency ideal estimation result of left channel background noise interference signal; for Frame Frequency is the ideal estimation result of the background noise interference signal in the right channel.
[0093] The calculation expressions of the posterior signal-to-noise ratio and the prior signal-to-noise ratio in step S5 are respectively:
[0094]
[0095]
[0096] in for Frame the posterior signal-to-noise ratio of the frequency; for Frame the a priori signal-to-noise ratio of the frequency; It is a dual-channel acoustic signal for the golden pomfret breeding area; for Frame The estimation result of the background noise interference signal of frequency is given by and composition; is the smoothing factor, which is usually set to 0.98; For the current Frame Frequency enhancement results; for Frame The estimation result of the background noise interference signal frequency; Indicates taking the maximum value; the calculation expression of the priori signal-to-noise ratio of the first frame is: .
[0097] In step S6, the Wiener filtering method is used, where the expression of the gain function constructed based on the posterior signal-to-noise ratio is:
[0098]
[0099] in for Frame Gain function of frequency.
[0100] The expression for multiplying the gain function by the left and right channel frequency domain signals is:
[0101]
[0102]
[0103] in for Frame The frequency domain signal of the left channel after frequency enhancement; for Frame The frequency domain signal of the right channel after frequency enhancement; for Frame Frequency: left channel frequency domain signal; for Frame Frequency: right channel frequency domain signal.
[0104] Step S6 is to calculate the gain of the frequency points of the frame using the signal-to-noise ratio, multiply the gain function by the energy of each frequency point of the left and right channel signals, enhance the amplitude spectrum of the golden pomfret sound signal, and retain its original phase information, so as to maintain the shared gain of the left and right channels. , avoid destroying spatial positioning cues (ITD / ILD), and achieve target signal enhancement and background noise interference suppression.
[0105] In the specific implementation process, step S5 and step S6 are performed alternately and frame by frame. Frame proceeds to step S5, and then Frame proceeds to step S6, then +1 frame to proceed to step S5, and then +1 frame goes to step S6, and so on.
[0106] After enhancing the sound signal of golden pomfret and suppressing the background noise signal, , Perform inverse fast Fourier transform and use Hanning window for smooth splicing. Use overlap-add (OLA) to superimpose each frame signal and reconstruct it into a continuous time domain signal:
[0107]
[0108] Final Output , The final sound signals of the left and right channels of the golden pomfret are obtained to realize the recognition of the sound of the fish school.
[0109] In this embodiment, time grouping may be introduced in step S1, and spectral grouping may be introduced in step S6 to specialize the algorithm performance:
[0110] After introducing time grouping, a sliding window (e.g., 30 frames) is used to calculate the average posterior signal-to-noise ratio of the local signal (of the dual-channel acoustic signal in the golden pomfret breeding area). Each window uses independent processing intensity to enhance response speed and reduce latency.
[0111] After the introduction of spectrum grouping, the frame-level spectrum of the input signal is divided into multiple sub-bands, each of which uses a different processing intensity, which has stronger recognition performance when dealing with non-stationary input signals with large frequency fluctuations.
[0112] Temporal grouping is used during frame segmentation and sliding window processing to preserve local temporal information while avoiding distortion caused by long-term signal analysis. Spectral grouping is used during gain calculation and signal enhancement to reduce noise fluctuations in high-frequency bands while effectively improving signal quality in low-frequency bands.
[0113] In summary, the present invention can suppress multiple noise sources and produce targeted gain effects by analyzing the noise situation through equalization processing and combining the prior signal-to-noise ratio of the signal. It also introduces specialized technologies in the time and frequency dimensions to make the method more flexible, solving the problems of difficulty in identifying fish sound signals in non-stationary environmental noise and severe distortion of target signals at low signal-to-noise ratios. At the same time, it also provides a more efficient and accurate unsupervised solution for situations where supervised training data cannot be collected.
Claims
1. A method for identifying the vocalizations of a school of golden pomfret in a cage based on a dual-channel time-varying Wiener filter, characterized in that: The following steps are involved: The dual-channel acoustic signals from the golden pomfret breeding area are collected and frame segmented. A Hamming window is added to each frame and converted into a frequency domain signal through fast Fourier transform to obtain a dual-channel frequency domain signal; the dual channels are the left channel and the right channel; Based on the dual-channel frequency domain signal, the left and right equalizers are trained separately by the normalized least mean square method to make the target signal parts in the left and right channels consistent. Based on the left and right equalizers, the left and right channel frequency domain signals are spectrally aligned and the target signal is eliminated as much as possible to obtain an ideal estimation result of the background noise interference signal of the left and right channels; By introducing a compensation factor into the ideal estimation results of the background noise interference signals of the left and right channels, the distortion caused by the equalizer on the background noise interference signal is corrected, and the estimated results of the compensated background noise interference signal are obtained; The posterior signal-to-noise ratio is calculated based on the estimation results of the dual-channel acoustic signal of the target area and the compensated background noise interference signal; the prior signal-to-noise ratio is calculated based on the posterior signal-to-noise ratio through a decision-guided method; A gain function is constructed based on the prior signal-to-noise ratio, and the gain function is multiplied by the frequency domain signals of the left and right channels respectively to obtain enhanced frequency domain signals of the left and right channels; The frequency domain signals of the left and right channels are respectively subjected to inverse fast Fourier transform and smoothed together using a Hanning window to obtain the reconstructed time domain signals, that is, the sound signals of the left and right channels of the golden pomfret.
2. The method for recognizing the sound of a school of golden pomfret in a cage based on dual-channel time-varying Wiener filtering according to claim 1, characterized in that: When performing frame segmentation, each frame is 20ms and the frame shift is 10ms.
3. The method for recognizing the sound of a school of golden pomfret in a cage based on dual-channel time-varying Wiener filtering according to claim 1, characterized in that: The expression converted to frequency domain signal by fast Fourier transform is: ; ; in for Frame Frequency: left channel frequency domain signal; for Frame Frequency: right channel frequency domain signal; represents the fast Fourier transform conversion; is the left channel acoustic signal; is the right channel acoustic signal.
4. The method for recognizing the sound of a school of golden pomfret in a cage based on dual-channel time-varying Wiener filtering according to claim 1, characterized in that: The expressions for training the left equalizer and the right equalizer respectively by the normalized least mean square method are: ; ; in express Left equalizer under the frame; express Right equalizer under the frame; express Left equalizer under the frame; express Right equalizer under the frame; is a constant; represents the L2 norm; for Left channel frequency domain signal under the frame; for The right channel frequency domain signal under the frame; the superscript T indicates the transpose of the matrix; During the training process, white noise signals were used as a proxy for the vocal signals of golden pomfret, and the direction of the target signal source was simulated through the transfer function to ensure training accuracy.
5. The method for recognizing the sound of a school of golden pomfret in a cage based on dual-channel time-varying Wiener filtering according to claim 1, characterized in that: Based on the left and right equalizers, the left and right channel frequency domain signals are spectrally aligned and the target signal is eliminated as much as possible. The expression for the ideal estimation result of the background noise interference signal of the left and right channels is: ; ; in and They are Frame The ideal estimation result of the background noise interference signal of the left channel and Frame Frequency ideal estimation result of right channel background noise interference signal; for Frame Frequency: left channel frequency domain signal; for Frame Frequency: right channel frequency domain signal; and They are Frame Frequency left and right equalizers.
6. The method for recognizing the sound of a school of golden pomfret in a cage based on dual-channel time-varying Wiener filtering according to claim 1, characterized in that: The expression of the compensation factor is: ; in for Frame Frequency compensation factor; Expressing hope; express Frame Frequency dual-channel frequency domain signal; Represents the ideal estimation result of the background noise interference signal of the left and right channels The complex conjugate of It means taking the square of the magnitude; Indicates the left channel; Indicates the right channel.
7. The method for recognizing the sound of a school of golden pomfret in a cage based on dual-channel time-varying Wiener filtering according to claim 6, characterized in that: The expression for introducing the compensation factor into the ideal estimation result of the background noise interference signal of the left and right channels is: ; ; in and After compensation Frame Frequency estimation results of left channel background noise interference signal and compensated Frame Frequency estimation result of right channel background noise interference signal; for Frame Frequency ideal estimation result of left channel background noise interference signal; for Frame Frequency is the ideal estimation result of the background noise interference signal in the right channel.
8. The method for recognizing the sound of a school of golden pomfret in a cage based on dual-channel time-varying Wiener filtering according to claim 7, characterized in that: The calculation expressions of the posterior signal-to-noise ratio and the prior signal-to-noise ratio are: ; ; in for Frame the posterior signal-to-noise ratio of the frequency; for Frame the a priori signal-to-noise ratio of the frequency; It is a dual-channel acoustic signal for the golden pomfret breeding area; for Frame The estimation result of the background noise interference signal of frequency is given by and composition; is the smoothing factor; for Frame Frequency enhancement results; for Frame The estimation result of the background noise interference signal frequency; Indicates taking the maximum value; the calculation expression of the priori signal-to-noise ratio of the first frame is: .
9. The method for recognizing the sound of a school of golden pomfret in a cage based on dual-channel time-varying Wiener filtering according to claim 8, characterized in that: The expression of the gain function based on the prior signal-to-noise ratio is: ; in for Frame Gain function of frequency.
10. The method for recognizing the sound of a school of golden pomfret in a cage based on dual-channel time-varying Wiener filtering according to claim 9, characterized in that: The expression for multiplying the gain function by the left and right channel frequency domain signals is: ; ; in for Frame The frequency domain signal of the left channel after frequency enhancement; for Frame The frequency domain signal of the right channel after frequency enhancement; for Frame Frequency: left channel frequency domain signal; for Frame Frequency: right channel frequency domain signal.
Citation Information
Patent Citations
Speech enhancement method for speech recognition in noise environment
CN108831495A
Statistical characteristic analysis method based on sound signal characteristics of large-scale net cage culture fish school
CN115394313A