Environmental sound classification and noise reduction method and system for intelligent Bluetooth hearing-aid earphone
By building a scene feature vector library and dynamic sound source correction technology, smart Bluetooth hearing aid headphones can accurately distinguish target speech from noise in complex environments, solving the problem of unnatural listening in traditional methods and achieving more efficient noise reduction and speech enhancement effects.
Patent Information
- Application Number
- CN202510973876.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-14
AI Technical Summary
传统智能蓝牙助听耳机在复杂环境中难以区分目标语音与背景噪声,导致对话者声音被误抑制、听感不自然,且无法动态修正声源定位,影响语音增强效果。
By collecting ambient sound signals, building a scene feature vector library, extracting Mel spectrum and auditory perception features, and combining dynamic attention mechanism and multi-layer perceptron, sound source direction correction and cross-modal interaction feature calculation are performed to generate a noise-reduced audio signal.
It achieves precise enhancement of target speech and effective suppression of background noise in complex environments, improving the noise reduction effect and user experience of hearing aid headphones. It is particularly suitable for scenarios with multiple people in conversation and strong reverberation.
Smart Images

Figure CN120786233A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hearing aid equipment, and in particular to a method and system for classifying and reducing ambient sound noise in an intelligent Bluetooth hearing aid headset. Background Art
[0002] Smart Bluetooth hearing aid earphones are wearable devices that integrate ambient sound acquisition, speech enhancement, and wireless audio transmission. They are widely used in everyday situations such as phone calls, music playback, and hearing assistance. These devices typically utilize multi-microphone arrays and digital signal processing to identify and suppress ambient noise, thereby improving user hearing clarity and comfort in complex acoustic environments. Traditional noise reduction methods primarily include fixed filtering based on statistical models (such as spectral subtraction and Wiener filtering), adaptive filtering, and directional beamforming. While these methods can reduce background noise to a certain extent, they still have significant limitations in practical applications. For example, in typical noisy environments such as restaurants and streets, traditional methods struggle to effectively distinguish target speech from background noise. Consequently, while suppressing ambient noise, they also attenuate the speaker's voice, resulting in muffled speech, an unnatural sounding experience, and even reduced communication efficiency. Furthermore, existing beamforming technologies often employ fixed directional strategies and are unable to dynamically adjust the focus direction based on actual sound field changes. This leads to inaccurate localization and reduced enhancement when there are numerous interfering sound sources or when the speaker's position changes. Therefore, there is an urgent need for an intelligent noise reduction method that can combine environmental sound type recognition and dynamic correction of sound source positioning to achieve accurate enhancement of target speech and effective suppression of background noise in complex environments, thereby improving the performance and user experience of Bluetooth hearing aids in real application scenarios. Summary of the Invention
[0003] In view of this, the present invention aims to provide a method and system for ambient sound classification and noise reduction for smart Bluetooth hearing aid headphones, so as to solve the problem that traditional methods have difficulty in distinguishing target speech from different types of background noise in complex environments, resulting in the interlocutor's voice being mistakenly suppressed, the listening experience being unnatural, and the inability to dynamically correct the sound source positioning, resulting in poor speech enhancement effect.
[0004] In order to achieve the above object, the present invention discloses a method for classifying and reducing ambient sound noise of a smart Bluetooth hearing aid headset, comprising the following steps: S1: Collect the original audio signals of background noise and real-time binaural sound field signals in typical scenes, and pre-process them respectively to obtain the background noise data and real-time sound field data of each typical scene; S2: Based on the background noise data, for each typical scene, we extract the Mel spectrum features and auditory perception features, calculate the fusion features, and then perform dimension compression to obtain the scene feature vector library; S3: Extracting live sound field features based on live sound field data, calculating preliminary classification features, and then extracting scene feature vectors from the scene feature vector library that match the live sound field. S4: Locate the direction of the sound source based on the live sound field data; calculate the sound source direction correction coefficient based on the live sound field characteristics and the scene feature vector in the scene feature vector library that matches the live sound field; and correct the sound source direction using the sound source direction correction coefficient to obtain the corrected sound source direction. S5: Based on the scene feature vectors in the scene feature vector library that match the actual sound field, the corrected sound source direction, and the actual sound field characteristics, the cross-modal interaction features are calculated. The noise suppression strength and sound source gain strength are then calculated, and the noise-reduced audio is generated in combination with the corrected sound source direction. S6: Generate a synchronized stereo audio signal based on the noise-reduced audio.
[0005] Furthermore, the step S1 further includes: S11: Using smart Bluetooth hearing aids, collect raw audio signals of background noise in typical scenarios, perform noise reduction filtering, and perform frame segmentation and windowing processing to obtain background noise data for each typical scenario. The typical scenarios include noisy restaurants, road traffic, pedestrian flow in alleys, open office areas, shopping malls and supermarkets, airport terminals, and outdoor parks. S12: Through smart Bluetooth hearing aid headphones, real-time binaural sound field signals are collected, and channel alignment and time domain synchronization processing are performed to obtain live sound field data.
[0006] Furthermore, the step S2 further includes: According to the Mel spectrum features and auditory perception features, one-dimensional convolution is performed respectively to obtain Mel attention features and auditory attention features, and then the splicing operation, Sigmoid function and weight matrix are combined to generate Mel attention weights and auditory attention weights; according to the Mel attention weights and auditory attention weights, the Mel attention features and auditory attention features are weightedly fused to generate fused features.
[0007] Furthermore, the step S2 further includes: S21: Based on the background noise data, the Mel spectrum features are extracted by combining short-time Fourier transform with Mel filter bank. The calculation method is: ; in, is the Mel spectrum feature, is the Mel filter bank, is the short-time Fourier transform, is the mean zeroing, A is the background noise data; S22: Based on the background noise data, the auditory perception features are extracted through the gammatone filter bank. The calculation method is: ; in, is the auditory perception feature, is the gammatone filter bank; S23: Calculate the fusion feature based on the Mel spectrum feature and auditory perception feature. The calculation method is: ; in, is the Mel attention feature, is a one-dimensional convolution, is the auditory attention feature, is the Mel attention weight, is the auditory attention weight, is the Sigmoid function, is the Mel attention weight matrix, is the auditory attention weight matrix, For splicing operations, It is a fusion feature; S24: Dimensionally compress the fused features to generate a scene feature vector. The calculation method is: ; in, is the scene feature vector, is the adaptive pooling layer, is global average pooling; S25: Using steps S21 to S24, traverse and process each typical scene including noisy restaurants, road traffic, alley crowds, open office areas, shopping malls and supermarkets, airport terminals, and outdoor parks to obtain a scene feature vector library V.
[0008] It should be further explained that, in step S2, the present invention extracts Mel spectrum features by combining short-time Fourier transform with Mel filter group, models the characteristics of frequency nonlinear distribution, and extracts auditory perception features by gammatone filter group to capture the nonlinear features in acoustic signals that are highly correlated with the response characteristics of the human auditory system; in the process of fusion feature calculation, the Mel attention weight and auditory attention weight are used to enable the model to adaptively enhance the feature dimensions that play a key role in scene classification according to the local characteristics of the input signal, while suppressing redundant information interference; compared with traditional single feature extraction methods (such as feature splicing that only relies on Mel spectrum coefficients or fixed weights), the present invention has significant advantages in feature expression ability and scene adaptability; traditional methods are limited by the rigid constraints of feature extraction methods, and are prone to feature confusion problems in complex noise scenes (such as noisy restaurants and road traffic noise), while the present invention dynamically allocates weights through the attention mechanism, which can effectively enhance the discriminability of target features; in addition, in dimensional compression In the first stage, the adaptive pooling layer is combined with the global average pooling, which not only retains the spatial distribution information of the feature map, but also realizes the generation of high-order abstract representation through global feature compression, further improving the generalization ability of the feature vector for complex scenes. In terms of adaptability to typical scenes (such as noisy restaurants, road traffic, and open office areas), the present invention can accurately capture the key differences in the acoustic characteristics of the scene through the collaborative design of multi-feature fusion and dynamic attention mechanism. For example, in the restaurant scene, the Mel spectrum feature can effectively characterize the energy distribution of the human voice frequency band, while the auditory perception feature is more sensitive to non-stationary noise such as the sound of tableware colliding and background whispers. In the road traffic scene, the fused feature can enhance the joint representation ability of low-frequency tire friction and high-frequency wind noise through the dynamic adjustment of attention weights. This multi-dimensional modeling method based on the scene feature vector library is significantly superior to the traditional single-modal feature extraction and fixed fusion strategy, providing highly robust prior knowledge support for the subsequent classification and sound source localization correction of various scenes.
[0009] Furthermore, the step S3 includes: S31: Based on the live sound field data, the live sound field features are obtained through wavelet-Mel joint spectrum analysis. The calculation method is: ; in, For the live sound field characteristics, is the Mel frequency scale mapping function, is discrete wavelet transform, is the live sound field data; S32: Based on the actual sound field characteristics, preliminary classification features are generated through one-dimensional convolution. The calculation method is: ; in, It is the preliminary classification feature; S33: Calculate the scene classification confidence vector based on the preliminary classification features and the scene feature vector library. The calculation method is: ; in, is the similarity between the preliminary classification feature and the i-th scene feature vector in the scene feature vector library, i is the scene index, is the cosine similarity, is the i-th scene feature vector in the scene feature vector library, is the attention weight of the i-th scene feature vector, is the Softmax function, is the scene similarity weight matrix, is the scene classification confidence vector, is the total number of scenes in the scene feature vector library, is one-hot encoding; S34: According to the scene classification confidence vector, a scene feature vector matching the live sound field in the scene feature vector library is selected. The calculation method is: ; in, To match the scene index, is the maximum index function, is the scene feature vector in the scene feature vector library that matches the live sound field, For the query function.
[0010] Furthermore, the step S4 includes: S41: Calculate the initial value of the sound source azimuth in the horizontal plane based on the actual sound field data, combined with the geometric layout parameters of the microphone array and the sound wave propagation model ; S42: According to the initial value of the sound source azimuth , the direction of the sound source is located by using a fixed weight beamforming algorithm, and the calculation method is: ; in, is the direction of the sound source, is the fixed beamforming weight of the jth microphone, j is the microphone index, is the initial value of the j-th microphone's azimuth angle at the sound source The frequency domain response in the direction, To take the absolute value; S43: Calculate the sound source direction correction coefficient based on the live sound field characteristics and the scene feature vector in the scene feature vector library that matches the live sound field. The calculation method is: ; in, is the sound source direction correction feature, is the ReLU function, Correct the first weight matrix for the sound source direction, is layer normalization, Correct the door for the sound source direction, is the hyperbolic tangent function, To correct the gate weight matrix, is the sound source direction correction coefficient, Modifying the second weight matrix for the sound source direction; S44: The sound source direction is corrected by combining the sound source direction correction coefficient, and the corrected sound source direction is generated through weighted fusion. The calculation method is: ; in, is the corrected sound source direction, It is a multi-layer perceptron.
[0011] It should be further explained that, in step S42, the present invention adopts a fixed-weight beamforming algorithm to locate the initial sound source direction and determine the target direction; however, due to the constraint of fixed weights, it is susceptible to interference sources in complex noise environments (such as the sound of tableware collision in restaurant scenes or low-frequency tire noise in road traffic); for this reason, step S43 introduces a dynamic correction mechanism, which generates sound source direction correction features by splicing the live sound field features and the scene feature vector library that matches the live sound field, and performs layer normalization and ReLU activation. The hyperbolic tangent function and the correction gate weight matrix are combined to generate a correction gating signal, and finally the sound source direction correction coefficient is output through the Softmax function; the coefficient is corrected in the initial sound source direction by combining the multi-layer perceptron through the element-by-element product and weighted fusion mechanism, which can effectively avoid non-target sound source interference in typical scenes such as restaurants and roads; Traditional methods rely on preset weights to enhance fixed directions, making them difficult to cope with dynamic changes in the sound field (such as movement of the interlocutor or sudden changes in the direction of the interference source). The present invention, by matching the scene feature vector library with the sound source direction correction coefficient generated by the attention mechanism, can dynamically adjust the direction correction strength according to the real-time sound field characteristics and scene type. For example, in a restaurant scene, the sound source direction correction coefficient will increase attention to the human voice frequency band and suppress the interference of the sound of tableware colliding; in a road traffic scene, it will prioritize enhancing low-frequency speech components and weakening the impact of tire friction noise. In addition, S44 combines the scene prior direction information output by the MLP through a weighted fusion mechanism. While retaining the computational efficiency of fixed-weight beamforming, it significantly improves the adaptability to complex sound field distributions. In particular, it exhibits higher positioning accuracy and stability in reverberant environments (such as open office areas) or multilingual scenes (such as shopping malls and supermarkets), providing highly robust directional prior support for subsequent noise suppression and speech enhancement.
[0012] Furthermore, the step S5 includes: S51: Generate cross-modal interaction features through dynamic attention fusion based on the scene feature vector library that matches the live sound field, the corrected sound source direction, and the live sound field features. The calculation method is: ; in, is the cross-modal interaction gate, is the cross-modal interaction feature, is the attention mechanism, is the dot product, is the cross-modal interaction weight matrix; S52: Calculate the noise suppression strength and sound source gain strength based on the cross-modal interaction characteristics. The calculation method is: ; in, is the noise suppression strength, is the noise suppression strength weight matrix, is the sound source gain intensity, is the sound source gain intensity weight matrix, is the Softplus function; S53: Generate the noise-reduced audio according to the noise suppression strength, the sound source gain strength, and the corrected sound source direction. The calculation method is: ; in, is the audio after noise reduction, is the minimum variance distortionless response beamforming filter, is a directional null filter.
[0013] It should be further explained that, compared with the traditional fixed-parameter noise suppression strategy, step S52 of the present invention generates noise suppression strength and sound source gain strength based on cross-modal interaction features, wherein the noise suppression strength controls the attenuation degree of the non-target frequency band through ReLU activation and Sigmoid function, while the sound source gain strength ensures smooth enhancement of the voice band through the Softplus function, thereby realizing dynamic weighting of noise and speech features; this dual-channel dynamic parameter generation method is more adaptable to changes in speech distribution in complex acoustic environments than traditional single filtering or fixed-ratio enhancement methods; In step S53, the present invention further combines the corrected sound source direction information, focuses on the target speech direction through minimum variance distortionless response beamforming, and uses a directional null filter to form a suppression beam in the non-target direction, thereby achieving effective suppression of spatial domain noise; in the final audio reconstruction process, the system performs frequency domain weighted fusion on the original audio according to the noise suppression intensity and the sound source gain intensity, thereby significantly reducing background interference while retaining the clarity of the target speech; this design is particularly suitable for typical scenarios with multi-person conversations, strong reverberation, or frequent changes in speaker positions (such as open office areas, shopping malls, supermarkets, etc.), while improving speech intelligibility, avoiding problems such as "speech being flattened" and "unnatural listening" brought about by traditional noise reduction algorithms.
[0014] Furthermore, the step S6 includes: S61: Generate compressed audio data through sub-band filtering and quantization coding according to the noise-reduced audio; S62: For the compressed audio data, the packet sending order is adjusted through an adaptive QoS scheduling algorithm to obtain a stable audio data stream; S63: For the stably transmitted audio data stream, time slice binding and clock calibration are performed to eliminate the phase difference of the binaural sound field and obtain a synchronized stereo audio signal.
[0015] The present invention also discloses an ambient sound classification and noise reduction system for an intelligent Bluetooth hearing aid headset, comprising: Audio data acquisition module: collects the original audio signals of background noise and real-time binaural sound field signals in typical scenes, and pre-processes them respectively to obtain the background noise data and real-time sound field data of each typical scene; Scene feature vector library construction module: Based on background noise data, for each typical scene, Mel spectrum features and auditory perception features are extracted, fusion features are calculated, and then dimension compression is performed to obtain a scene feature vector library; Scene matching module: extracts the live sound field features based on the live sound field data, calculates the preliminary classification features, and then extracts the scene feature vectors in the scene feature vector library that match the live sound field; Sound source direction correction module: locates the sound source direction according to the live sound field data, calculates the sound source direction correction coefficient and corrects the sound source direction to obtain the corrected sound source direction; Audio noise reduction module: This module calculates cross-modal interaction features based on the scene feature vectors in the scene feature vector library that match the live sound field, the corrected sound source direction, and the live sound field features. It then calculates the noise suppression strength and sound source gain strength, and combines this with the corrected sound source direction to generate the denoised audio. Audio generation module: Generates synchronized stereo audio signals based on the noise-reduced audio.
[0016] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention addresses the problem that traditional methods are difficult to cope with the different background noise characteristics in typical scenes such as restaurants and roads, resulting in insufficient clarity of the target speech. By constructing a scene feature vector library, the sound features in various application scenes are extracted and matched with the current sound field features to achieve dynamic correction of the sound source direction; based on the corrected sound source direction, the cross-modal interaction features, noise suppression strength and sound source gain strength are calculated, and the audio signal is differentially processed to gain the target speech and reduce the noise intensity. The direction correction and frequency domain weighting mechanisms are effectively integrated, which enhances the system's adaptability to various sound field changes and improves the noise reduction effect of smart Bluetooth hearing aid headphones in different environments.
[0017] (2) Based on the positioning of the initial sound source direction, the present invention introduces a dynamic correction mechanism to address the problem that traditional methods are easily affected by interference sources in complex noise environments. The sound source direction correction coefficient is calculated by splicing the actual sound field features and the scene feature vectors in the scene feature vector library that match the actual sound field. The correction coefficient is adjusted in the direction of the initial sound source by combining the multi-layer perceptron through the element-by-element product and weighted fusion mechanism, so that the system can effectively avoid interference from non-target sound sources in typical scenes such as restaurants and roads, and significantly improve the direction estimation accuracy and stability.
[0018] (3) Compared with the traditional fixed parameter noise suppression strategy, the present invention generates noise suppression strength and sound source gain strength respectively based on cross-modal interaction characteristics; this dual-channel dynamic parameter generation method is more environmentally adaptable than the traditional single filtering or fixed ratio enhancement method; in addition, the present invention further combines the corrected sound source direction, uses minimum variance distortionless response beamforming to focus on the target speech direction, and uses a directional null filter to form a suppression beam in the non-target direction to achieve spatial domain noise suppression; in the final audio reconstruction process, the system performs frequency domain weighted fusion on the original audio based on the noise suppression strength and sound source gain strength, while retaining the clarity of the target speech, significantly reducing background interference, which is particularly suitable for typical scenarios with multi-person conversations, strong reverberation or frequent changes in speaker positions. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A schematic flow chart of a method for classifying and reducing ambient sound in a smart Bluetooth hearing aid headset provided by the present invention; Figure 2 A schematic diagram of the algorithm flow for sound source direction correction provided by the present invention; Figure 3 This is the local heat map of each scene feature vector in the scene feature vector library provided by the present invention. DETAILED DESCRIPTION
[0020] The present invention will be further described below with reference to the accompanying drawings, but the present invention is not limited in any way. Any changes or substitutions made based on the teachings of the present invention fall within the scope of protection of the present invention.
[0021] Example 1: A method for classifying and reducing ambient sound noise in a smart Bluetooth hearing aid headset, such as Figure 1 As shown, the following steps are included: S1: Collect the original audio signals of background noise and real-time binaural sound field signals in typical scenes, and pre-process them separately to obtain background noise data and real-time sound field data of each environment, including: S11: Using smart Bluetooth hearing aids, collect raw audio signals of background noise in typical scenarios, perform noise reduction filtering, and perform frame segmentation and windowing processing to obtain background noise data for each typical scenario. The typical scenarios include noisy restaurants, road traffic, pedestrian flow in alleys, open office areas, shopping malls and supermarkets, airport terminals, and outdoor parks. S12: Through smart Bluetooth hearing aid headphones, real-time binaural sound field signals are collected, and channel alignment and time domain synchronization processing are performed to obtain live sound field data.
[0022] S2: Based on the background noise data, for each typical scene, we extract the Mel spectrum features and auditory perception features, calculate the fusion features, and then perform dimensionality compression to obtain the scene feature vector library, including: S21: Based on the background noise data, the Mel spectrum features are extracted by combining short-time Fourier transform with Mel filter bank. The calculation method is: ; in, is the Mel spectrum feature, is the Mel filter bank, is the short-time Fourier transform, is the mean zeroing, A is the background noise data; S22: Based on the background noise data, the auditory perception features are extracted through the gammatone filter bank. The calculation method is: ; in, is the auditory perception feature, is the gammatone filter bank; S23: Calculate the fusion feature based on the Mel spectrum feature and auditory perception feature. The calculation method is: ; in, is the Mel attention feature, is a one-dimensional convolution, is the auditory attention feature, is the Mel attention weight, is the auditory attention weight, is the Sigmoid function, is the Mel attention weight matrix, is the auditory attention weight matrix, For splicing operations, It is a fusion feature; S24: Dimensionally compress the fused features to generate a scene feature vector. The calculation method is: ; in, is the scene feature vector, is the adaptive pooling layer, is global average pooling; S25: Using steps S21 to S24, traverse and process each typical scene including noisy restaurants, road traffic, alley crowds, open office areas, shopping malls and supermarkets, airport terminals, and outdoor parks to obtain a scene feature vector library V.
[0023] like Figure 3 As shown, the scene feature vector library obtained after processing through steps S21 to S25 includes the following feature representations of typical scenes (taking a 128-dimensional feature vector as an example, the values of the first 15 dimensions are shown): For example, in a noisy restaurant scene, the values are: [0.65, 0.88, 0.72, 0.45, 0.82, 0.30, 0.78, 0.25, 0.68, 0.55, 0.40, 0.62, 0.85, 0.75, 0.70]; the second dimension (0.88) and the fifth dimension (0.82) show peaks, corresponding to the mid-frequency energy of the noisy background voice conversation; the 13th dimension (0.85) represents the transient characteristics of high-frequency tableware collisions; and the 7th dimension (0.78) reflects the continuous fluctuation of the ambient noise. Road traffic scene: [0.85, 0.65, 0.58, 0.20, 0.90, 0.40, 0.92, 0.60, 0.30, 0.25, 0.70, 0.75, 0.60, 0.40, 0.30]; the first dimension (0.85) and the fifth dimension (0.90) represent the low-frequency roar of the engine; the seventh dimension (0.92) shows the characteristics of violent volume fluctuations; This scene feature vector library can be used to classify environmental noise. By calculating the similarity between the input audio features and the scene features in the library, the automatic identification of the environment type can be achieved. Based on the classification results, the system can adaptively adopt the optimal noise reduction method. For example, when it is identified as a noisy restaurant scene, the focus is on suppressing the mid-frequency noise of the human voice corresponding to the 2nd and 5th dimensions and the high-frequency sound of tableware collision in the 13th dimension. When it is identified as a traffic environment, the low-frequency and transient noise are mainly eliminated by focusing on the 1st and 7th dimension features. This environmental perception method based on the feature vector library enables the noise reduction system to be precisely adjusted according to the acoustic characteristics of different scenes, and can achieve better voice enhancement effects and more natural noise suppression performance compared to general noise reduction solutions.
[0024] S3: Based on the live sound field data, extract the live sound field features and calculate the preliminary classification features. Then, combined with the scene feature vector library, extract the scene feature vector in the scene feature vector library that matches the live sound field, including: S31: Based on the live sound field data, the live sound field features are obtained through wavelet-Mel joint spectrum analysis. The calculation method is: ; in, For the live sound field characteristics, is the Mel frequency scale mapping function, is discrete wavelet transform, is the live sound field data; S32: Based on the actual sound field characteristics, preliminary classification features are generated through one-dimensional convolution. The calculation method is: ; in, It is the preliminary classification feature; S33: Calculate the scene classification confidence vector based on the preliminary classification features and the scene feature vector library. The calculation method is: ; in, is the similarity between the preliminary classification feature and the i-th scene feature vector in the scene feature vector library, i is the scene index, is the cosine similarity, is the i-th scene feature vector in the scene feature vector library, is the attention weight of the i-th scene feature vector, is the Softmax function, is the scene similarity weight matrix, is the scene classification confidence vector, is the total number of scenes in the scene feature vector library, is one-hot encoding; S34: According to the scene classification confidence vector, a scene feature vector matching the live sound field in the scene feature vector library is selected. The calculation method is: ; in, To match the scene index, is the maximum index function, is the scene feature vector in the scene feature vector library that matches the live sound field, For the query function.
[0025] S4: Locate the sound source direction based on the live sound field data; calculate the sound source direction correction coefficient based on the live sound field characteristics and the scene feature vector in the scene feature vector library that matches the live sound field; and correct the sound source direction using the sound source direction correction coefficient to obtain the corrected sound source direction, including: S41: Based on the live sound field data, combined with the geometric layout parameters of the microphone array and the sound wave propagation model, calculate the initial value of the sound source azimuth in the horizontal plane ; S42: According to the initial value of the sound source azimuth , the direction of the sound source is located by using a fixed weight beamforming algorithm, and the calculation method is: ; in, is the direction of the sound source, is the fixed beamforming weight of the jth microphone, j is the microphone index, is the initial value of the j-th microphone's azimuth angle at the sound source The frequency domain response in the direction, To take the absolute value; S43: Calculating a sound source direction correction coefficient based on the live sound field characteristics and the scene feature vector in the scene feature vector library that matches the live sound field; like Figure 2As shown, the calculation process of the sound source direction correction coefficient includes: generating a sound source direction correction feature based on the live sound field feature and the scene feature vector matching the live sound field in the scene feature vector library, combining the splicing operation, layer normalization, the first weight matrix of the sound source direction correction and the ReLU function; inputting the sound source direction correction feature into the hyperbolic tangent function, and generating a sound source direction correction gate in combination with the correction gate weight matrix; using the second weight matrix of the sound source direction correction to weight the element-by-element product of the two, and combining the Softmax function to generate a sound source direction correction coefficient; The specific calculation method is: ; in, is the sound source direction correction feature, is the ReLU function, Correct the first weight matrix for the sound source direction, is layer normalization, Correct the door for the sound source direction, is the hyperbolic tangent function, To correct the gate weight matrix, is the sound source direction correction coefficient, Modifying the second weight matrix for the sound source direction; S44: The sound source direction is corrected by combining the sound source direction correction coefficient, and the corrected sound source direction is generated through weighted fusion. The calculation method is: ; in, is the corrected sound source direction, It is a multi-layer perceptron.
[0026] S5: Based on the scene feature vectors in the scene feature vector library that match the actual sound field, the corrected sound source direction, and the actual sound field characteristics, the cross-modal interaction features are calculated. The noise suppression strength and sound source gain strength are then calculated. Combined with the corrected sound source direction, the noise-reduced audio is generated, including: S51: Generate cross-modal interaction features through dynamic attention fusion based on the scene feature vector library that matches the live sound field, the corrected sound source direction, and the live sound field features. The calculation method is: ; in, is the cross-modal interaction gate, is the cross-modal interaction feature, is the attention mechanism, is the dot product, is the cross-modal interaction weight matrix; S52: Calculate the noise suppression strength and sound source gain strength based on the cross-modal interaction characteristics. The calculation method is: ; in, is the noise suppression strength, is the noise suppression strength weight matrix, is the sound source gain intensity, is the sound source gain intensity weight matrix, is the Softplus function; S53: Generate the noise-reduced audio according to the noise suppression strength, the sound source gain strength, and the corrected sound source direction. The calculation method is: ; in, is the audio after noise reduction, is the minimum variance distortionless response beamforming filter, is a directional null filter.
[0027] S6: Generate a synchronized stereo audio signal based on the noise-reduced audio, including: S61: Generate compressed audio data through sub-band filtering and quantization coding according to the noise-reduced audio; S62: For the compressed audio data, the packet sending order is adjusted through an adaptive QoS scheduling algorithm to obtain a stable audio data stream; S63: For the stably transmitted audio data stream, time slice binding and clock calibration are performed to eliminate the phase difference of the binaural sound field and obtain a synchronized stereo audio signal.
[0028] Example 2: The present invention also discloses an ambient sound classification and noise reduction system for smart Bluetooth hearing aid headphones, comprising: Audio data acquisition module: collects the original audio signals of background noise and real-time binaural sound field signals in typical scenes, and pre-processes them respectively to obtain the background noise data and real-time sound field data of each typical scene; Scene feature vector library construction module: Based on background noise data, for each typical scene, Mel spectrum features and auditory perception features are extracted, fusion features are calculated, and then dimension compression is performed to obtain a scene feature vector library; Scene matching module: extracts the live sound field features based on the live sound field data, calculates the preliminary classification features, and then extracts the scene feature vectors in the scene feature vector library that match the live sound field; Sound source direction correction module: locates the sound source direction according to the live sound field data, calculates the sound source direction correction coefficient and corrects the sound source direction to obtain the corrected sound source direction; The audio noise reduction module: according to the scene feature vector in the scene feature vector library matched with the live sound field, the corrected sound source direction, the live sound field feature, the cross-modal interaction feature is calculated, and then the noise suppression strength and the sound source gain strength are calculated, and combined with the corrected sound source direction, the audio after noise reduction is generated; The audio generation module: according to the audio after noise reduction, a synchronous stereo audio signal is generated.
[0029] It should be noted that the above-mentioned embodiment number of the application is only for description, not representing the advantages and disadvantages of the embodiments. And the term "include", "contain" or any other variant in this paper is intended to cover non-exclusive inclusion, so that the process, device, article or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes elements inherent to such process, device, article or method. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0030] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment method can be realized by software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is the better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, including a plurality of instructions to make a terminal device (which can be a mobile phone, computer, server, or network device, etc.) execute the method described in the embodiments of the present application.
[0031] The above is only the preferred embodiment of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. A method for classifying and reducing ambient sound noise in a smart Bluetooth hearing aid headset, characterized in that: The following steps are involved: S1: Collect the original audio signals of background noise and real-time binaural sound field signals in typical scenes, and pre-process them respectively to obtain the background noise data and real-time sound field data of each typical scene; S2: Based on the background noise data, for each typical scene, we extract the Mel spectrum features and auditory perception features, calculate the fusion features, and then perform dimension compression to obtain the scene feature vector library; S3: Extracting live sound field features based on live sound field data, calculating preliminary classification features, and then extracting scene feature vectors from the scene feature vector library that match the live sound field. S4: Locate the direction of the sound source based on the live sound field data; calculate the sound source direction correction coefficient based on the live sound field characteristics and the scene feature vector in the scene feature vector library that matches the live sound field; and correct the sound source direction using the sound source direction correction coefficient to obtain the corrected sound source direction. S5: Based on the scene feature vectors in the scene feature vector library that match the actual sound field, the corrected sound source direction, and the actual sound field characteristics, the cross-modal interaction features are calculated. The noise suppression strength and sound source gain strength are then calculated, and the noise-reduced audio is generated in combination with the corrected sound source direction. S6: Generate a synchronized stereo audio signal based on the noise-reduced audio.
2. The method for ambient sound classification and noise reduction of smart Bluetooth hearing aid headphones according to claim 1, characterized in that: The S1 step includes: S11: Using smart Bluetooth hearing aids, collect raw audio signals of background noise in typical scenarios, perform noise reduction filtering, and perform frame segmentation and windowing processing to obtain background noise data for each typical scenario. The typical scenarios include noisy restaurants, road traffic, pedestrian flow in alleys, open office areas, shopping malls and supermarkets, airport terminals, and outdoor parks. S12: Through smart Bluetooth hearing aid headphones, real-time binaural sound field signals are collected, and channel alignment and time domain synchronization processing are performed to obtain live sound field data.
3. The method for ambient sound classification and noise reduction of smart Bluetooth hearing aid headphones according to claim 1, characterized in that: In step S2, the calculation process of the fusion feature includes: According to the Mel spectrum features and auditory perception features, one-dimensional convolution is performed respectively to obtain Mel attention features and auditory attention features, and then the splicing operation, Sigmoid function and weight matrix are combined to generate Mel attention weights and auditory attention weights; according to the Mel attention weights and auditory attention weights, the Mel attention features and auditory attention features are weightedly fused to generate fused features.
4. The method for ambient sound classification and noise reduction of smart Bluetooth hearing aid headphones according to claim 3, characterized in that: The S2 step includes: S21: Based on the background noise data, the Mel spectrum features are extracted by combining short-time Fourier transform with Mel filter bank. The calculation method is: ; in, is the Mel spectrum feature, is the Mel filter bank, is the short-time Fourier transform, is the mean zeroing, A is the background noise data; S22: Based on the background noise data, the auditory perception features are extracted through the gammatone filter bank. The calculation method is: ; in, is the auditory perception feature, is the gammatone filter bank; S23: Calculate the fusion feature based on the Mel spectrum feature and auditory perception feature. The calculation method is: ; in, is the Mel attention feature, is a one-dimensional convolution, is the auditory attention feature, is the Mel attention weight, is the auditory attention weight, is the Sigmoid function, is the Mel attention weight matrix, is the auditory attention weight matrix, For splicing operations, It is a fusion feature; S24: Dimensionally compress the fused features to generate a scene feature vector. The calculation method is: ; in, is the scene feature vector, is the adaptive pooling layer, is global average pooling; S25: Using steps S21 to S24, traverse and process each typical scene including noisy restaurants, road traffic, alley crowds, open office areas, shopping malls and supermarkets, airport terminals, and outdoor parks to obtain a scene feature vector library V.
5. The method for ambient sound classification and noise reduction of smart Bluetooth hearing aid headphones according to claim 4, characterized in that: The S3 step includes: S31: Based on the live sound field data, the live sound field features are obtained through wavelet-Mel joint spectrum analysis. The calculation method is: ; in, For the live sound field characteristics, is the Mel frequency scale mapping function, is discrete wavelet transform, is the live sound field data; S32: Based on the actual sound field characteristics, preliminary classification features are generated through one-dimensional convolution. The calculation method is: ; in, It is the preliminary classification feature; S33: Calculate the scene classification confidence vector based on the preliminary classification features and the scene feature vector library. The calculation method is: ; in, is the similarity between the preliminary classification feature and the i-th scene feature vector in the scene feature vector library, i is the scene index, is the cosine similarity, is the i-th scene feature vector in the scene feature vector library, is the attention weight of the i-th scene feature vector, is the Softmax function, is the scene similarity weight matrix, is the scene classification confidence vector, is the total number of scenes in the scene feature vector library, is one-hot encoding; S34: According to the scene classification confidence vector, a scene feature vector matching the live sound field in the scene feature vector library is selected. The calculation method is: ; in, To match the scene index, is the maximum index function, is the scene feature vector in the scene feature vector library that matches the live sound field, For the query function.
6. The method for ambient sound classification and noise reduction of smart Bluetooth hearing aid headphones according to claim 5, characterized in that: The S4 step comprises: S41: Based on the live sound field data, combined with the geometric layout parameters of the microphone array and the sound wave propagation model, calculate the initial value of the sound source azimuth in the horizontal plane ; S42: According to the initial value of the sound source azimuth , the direction of the sound source is located by using a fixed weight beamforming algorithm, and the calculation method is: ; in, is the direction of the sound source, is the fixed beamforming weight of the jth microphone, j is the microphone index, is the initial value of the j-th microphone's azimuth angle at the sound source The frequency domain response in the direction, To take the absolute value; S43: Calculate the sound source direction correction coefficient based on the live sound field characteristics and the scene feature vector in the scene feature vector library that matches the live sound field. The calculation method is: ; in, is the sound source direction correction feature, is the ReLU function, Correct the first weight matrix for the sound source direction, is layer normalization, Correct the door for the sound source direction, is the hyperbolic tangent function, To correct the gate weight matrix, is the sound source direction correction coefficient, Modifying the second weight matrix for the sound source direction; S44: The sound source direction is corrected by combining the sound source direction correction coefficient, and the corrected sound source direction is generated through weighted fusion. The calculation method is: ; in, is the corrected sound source direction, It is a multi-layer perceptron.
7. The method for ambient sound classification and noise reduction of smart Bluetooth hearing aid headphones according to claim 6, characterized in that: The step S5 comprises: S51: Generate cross-modal interaction features through dynamic attention fusion based on the scene feature vector library that matches the live sound field, the corrected sound source direction, and the live sound field features. The calculation method is: ; in, is the cross-modal interaction gate, is the cross-modal interaction feature, is the attention mechanism, is the dot product, is the cross-modal interaction weight matrix; S52: Calculate the noise suppression strength and sound source gain strength based on the cross-modal interaction characteristics. The calculation method is: ; in, is the noise suppression strength, is the noise suppression strength weight matrix, is the sound source gain intensity, is the sound source gain intensity weight matrix, is the Softplus function; S53: Generate the noise-reduced audio according to the noise suppression strength, the sound source gain strength, and the corrected sound source direction. The calculation method is: ; in, is the audio after noise reduction, is the minimum variance distortionless response beamforming filter, is a directional null filter.
8. The method for ambient sound classification and noise reduction of smart Bluetooth hearing aid headphones according to claim 7, characterized in that: The step S6 comprises: S61: Generate compressed audio data through sub-band filtering and quantization coding according to the noise-reduced audio; S62: For the compressed audio data, the packet sending order is adjusted through an adaptive QoS scheduling algorithm to obtain a stable audio data stream; S63: For the stably transmitted audio data stream, time slice binding and clock calibration are performed to eliminate the phase difference of the binaural sound field and obtain a synchronized stereo audio signal.
9. An ambient sound classification and noise reduction system for smart Bluetooth hearing aid headphones, characterized in that: include: Audio data acquisition module: collects the original audio signals of background noise and real-time binaural sound field signals in typical scenes, and pre-processes them respectively to obtain the background noise data and real-time sound field data of each typical scene; Scene feature vector library construction module: Based on background noise data, for each typical scene, Mel spectrum features and auditory perception features are extracted, fusion features are calculated, and then dimension compression is performed to obtain a scene feature vector library; Scene matching module: extracts the live sound field features based on the live sound field data, calculates the preliminary classification features, and then extracts the scene feature vectors in the scene feature vector library that match the live sound field; Sound source direction correction module: locates the sound source direction according to the live sound field data, calculates the sound source direction correction coefficient and corrects the sound source direction to obtain the corrected sound source direction; Audio noise reduction module: This module calculates cross-modal interaction features based on the scene feature vectors in the scene feature vector library that match the live sound field, the corrected sound source direction, and the live sound field features. It then calculates the noise suppression strength and sound source gain strength, and combines this with the corrected sound source direction to generate the denoised audio. Audio generation module: generates synchronized stereo audio signals based on the noise-reduced audio; To implement the ambient sound classification and noise reduction method of the smart Bluetooth hearing aid headset as described in any one of claims 1-8.
Citation Information
Cited By
Noise monitoring method and system based on voiceprint recognition
CN121191532A