A cockpit voice noise reduction method, device, equipment and storage medium

By extracting spectral features and performing interactive enhancement processing on in-vehicle voice signals, and combining a dual-decoder structure with sub-band and low-frequency features, the problem of noise suppression and voice preservation in dynamic environments of in-vehicle voice noise reduction technology is solved, thereby improving voice clarity and naturalness.

CN121237107BActive Publication Date: 2026-08-25CHONGQING WUTONG CAR LINK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511391711.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-08-25
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Existing in-vehicle voice noise reduction technologies are difficult to adapt to dynamically changing driving environments, and traditional methods tend to over-suppress speech components when removing background noise, leading to speech distortion and degraded sound quality, which affects user experience.

Method used

By acquiring the initial noisy speech signal in the cockpit, spectral features are extracted and equivalent rectangular bandwidth is converted. Subband features and low-frequency features are processed separately. Interactive enhancement processing and a dual decoder structure are adopted to generate fused spectral features. Subband masks and low-frequency masks are used for noise suppression.

Benefits of technology

While suppressing noise across the entire frequency band, it effectively preserves low-frequency speech characteristics, improves speech clarity and naturalness, and enhances the user experience of in-vehicle voice systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237107B_ABST
    Figure CN121237107B_ABST
Patent Text Reader

Abstract

The application provides a cockpit voice noise reduction method and device, equipment and a storage medium. The initial noisy voice signal in the cockpit is obtained, and the spectrum feature is extracted to obtain the full-band spectrum feature. The equivalent rectangular bandwidth conversion and low-frequency part interception are performed on the full-band spectrum feature to obtain the sub-band feature and the low-frequency feature. The interactive enhancement processing is performed on the sub-band feature and the low-frequency feature respectively to obtain the sub-band coding feature and the low-frequency coding feature, and the fusion spectrum feature is generated by fusion. The sub-band feature decoder and the low-frequency feature decoder are respectively input to obtain the sub-band mask and the low-frequency mask to determine the noise reduction spectrum feature based on the sub-band mask, the low-frequency mask and the full-band spectrum feature. The application processes the sub-band feature and the low-frequency feature in the full-band spectrum feature respectively, and adopts the interactive enhancement processing and the double-decoder structure to realize noise suppression, which can suppress the full-band noise and retain the low-frequency voice feature, and can improve the voice clarity and naturalness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cabin noise reduction technology, and in particular to a cabin voice noise reduction method, apparatus, device and storage medium. Background Technology

[0002] In in-vehicle voice communication and interaction systems, voice quality is severely affected by complex driving environments. Mechanical noises such as wind noise, road noise, and engine noise generated during vehicle operation, along with environmental interference from external sources like street noise and horns, plus background noise from occupants such as conversations and entertainment system sounds, collectively create a multi-source, high-noise acoustic environment. This not only significantly reduces the intelligibility of the voice signal but also hinders in-vehicle voice applications such as voice recognition and voice communication. To address the application challenges caused by these issues, noise suppression technology has become a key component of in-vehicle audio preprocessing systems. Its goal is to extract the purest possible human voice from noisy speech signals to improve voice quality and system robustness. Common noise reduction schemes include traditional signal processing methods such as spectral subtraction and Wiener filtering, as well as end-to-end speech enhancement models based on deep learning.

[0003] However, the aforementioned noise reduction methods rely on prior assumptions about the statistical characteristics of noise, making them difficult to adapt to dynamically changing driving environments. Furthermore, existing deep learning models often employ single feature processing methods, failing to achieve comprehensive collaborative processing. This leads to the over-suppression of useful components in speech while removing background noise, resulting not only in speech distortion and reduced sound quality but also the introduction of artificial effects such as musical noise. This severely impacts the naturalness and clarity of speech, thus limiting the user experience of in-vehicle voice systems. Summary of the Invention

[0004] The purpose of this application is to provide a cockpit voice noise reduction method, apparatus, device, and storage medium to solve the above-mentioned technical problems.

[0005] This application provides a cockpit voice noise reduction method, which includes: acquiring an initial noisy voice signal in the cockpit, and extracting spectral features from the initial noisy voice signal to obtain full-band spectral features; performing an equivalent rectangular bandwidth transformation on the full-band spectral features to obtain sub-band features, and truncating the low-frequency portion of the full-band spectral features to obtain low-frequency features, wherein the low-frequency truncation includes truncating spectral features below a preset low-frequency threshold from the full-band spectral features; performing interactive enhancement processing on the sub-band features and the low-frequency features respectively to obtain sub-band coded features and low-frequency coded features, and fusing the sub-band coded features and low-frequency coded features to generate fused spectral features; inputting the fused spectral features into a sub-band feature decoder and a low-frequency feature decoder respectively to obtain a sub-band mask and a low-frequency mask, and determining noise reduction spectral features based on the sub-band mask, the low-frequency mask, and the full-band spectral features.

[0006] In one embodiment of this application, the interactive enhancement processing of the sub-band features and the low-frequency features includes: inputting the sub-band features and the low-frequency features into a sub-band feature encoder; performing attention calculation using the low-frequency features as a query and the sub-band features as a key and value to obtain sub-band coding weights; and reconstructing the sub-band features using the sub-band coding weights to obtain sub-band coded features; and inputting the sub-band features and the low-frequency features into a low-frequency feature encoder; performing attention calculation using the sub-band features as a query and the low-frequency features as a key and value to obtain low-frequency coding weights; and reconstructing the low-frequency features using the low-frequency coding weights to obtain low-frequency coded features.

[0007] In one embodiment of this application, fusing the sub-band coding features and the low-frequency coding features to generate fused spectral features includes: concatenating the sub-band coding features and the low-frequency coding features to form concatenated coding features; inputting the concatenated coding features into a preset neural network for weight allocation to obtain coding fusion weights, wherein the preset neural network is used for weight allocation of feature fusion of the sub-band coding features and the low-frequency coding features; and performing a weighted summation of the sub-band coding features and the low-frequency coding features according to the coding fusion weights to obtain fused spectral features.

[0008] In one embodiment of this application, inputting the fused spectral features into a sub-band feature decoder and a low-frequency feature decoder to obtain a sub-band mask and a low-frequency mask includes: inputting the fused spectral features into the sub-band feature decoder for decoding processing, and outputting a sub-band mask corresponding to the full-band spectral features, wherein the sub-band mask is used to suppress noise in the full-band portion of the full-band spectral features; and inputting the fused spectral features into the low-frequency feature decoder for decoding processing, and outputting a low-frequency mask corresponding to the low-frequency portion of the full-band spectral features, wherein the low-frequency mask is used to suppress noise in the low-frequency portion of the full-band spectral features.

[0009] In one embodiment of this application, determining the noise reduction spectral features based on the sub-band mask, low-frequency mask, and full-band spectral features includes: multiplying the sub-band mask by the full-band spectral features to obtain a sub-band noise reduction result; multiplying the low-frequency mask by the low-frequency portion of the full-band spectral features to obtain a low-frequency noise reduction result; and fusing the sub-band noise reduction result with the low-frequency noise reduction result to obtain the noise reduction spectral features.

[0010] In one embodiment of this application, fusing the sub-band denoising result with the low-frequency denoising result to obtain denoised spectral features includes: predicting the confidence level of the sub-band denoising result to obtain a sub-band confidence score, and predicting the confidence level of the low-frequency denoising result to obtain a low-frequency confidence score; normalizing the sub-band confidence score and the low-frequency confidence score to obtain sub-band denoising weights and low-frequency denoising weights, and weighting and summing the sub-band denoising result and the low-frequency denoising result based on the sub-band denoising weights and the low-frequency denoising weights to obtain denoised spectral features.

[0011] In one embodiment of this application, after determining the noise reduction spectrum features based on the sub-band mask, low-frequency mask, and full-band spectrum features, the cockpit voice noise reduction method further includes: performing an inverse Fourier transform on the noise reduction spectrum features to obtain the time-domain signal to be processed, and performing overlapping and addition processing on the time-domain signal to be processed to obtain a time-domain voice signal.

[0012] This application embodiment also provides a cockpit voice noise reduction device, which includes: a signal feature extraction module, used to acquire an initial noisy voice signal in the cockpit, and extract spectral features from the initial noisy voice signal to obtain full-band spectral features; perform equivalent rectangular bandwidth transformation on the full-band spectral features to obtain sub-band features, and truncate the low-frequency portion of the full-band spectral features to obtain low-frequency features, wherein the low-frequency truncation includes truncating spectral features below a preset low-frequency threshold from the full-band spectral features; a spectral feature noise reduction module, used to perform interactive enhancement processing on the sub-band features and the low-frequency features respectively to obtain sub-band coded features and low-frequency coded features, and fuse the sub-band coded features and low-frequency coded features to generate fused spectral features; input the fused spectral features to the sub-band feature decoder and the low-frequency feature decoder respectively to obtain a sub-band mask and a low-frequency mask, so as to determine the noise reduction spectral features based on the sub-band mask, the low-frequency mask and the full-band spectral features.

[0013] This application also provides an electronic device, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the cockpit voice noise reduction method as described in any of the above embodiments.

[0014] This application also provides a computer-readable storage medium storing computer-readable instructions, which, when executed by a computer's processor, cause the computer to perform the cockpit voice noise reduction method as described in any of the above embodiments.

[0015] The beneficial effects of this invention are as follows: This application provides a cockpit speech noise reduction method, apparatus, device, and storage medium. It acquires an initial noisy speech signal from within the cockpit, extracts spectral features from the initial noisy speech signal to obtain full-band spectral features, performs an equivalent rectangular bandwidth transformation on the full-band spectral features to obtain sub-band features, and truncates the low-frequency portion of the full-band spectral features to obtain low-frequency features. It then performs interactive enhancement processing on the sub-band features and low-frequency features to obtain sub-band encoded features and low-frequency encoded features, and fuses the sub-band encoded features and low-frequency encoded features to generate fused spectral features. These fused spectral features are then input to a sub-band feature decoder and a low-frequency feature decoder to obtain a sub-band mask and a low-frequency mask, respectively. Noise reduction spectral features are then determined based on the sub-band mask, low-frequency mask, and full-band spectral features. This application achieves noise suppression by separately processing the sub-band features and low-frequency features within the full-band spectral features and employing interactive enhancement processing and a dual-decoder structure. This effectively preserves low-frequency speech features while suppressing full-band noise, thereby improving speech clarity and naturalness.

[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0018] Figure 1 This is a schematic diagram illustrating an exemplary system architecture as shown in an exemplary embodiment of this application;

[0019] Figure 2 This is a flowchart illustrating a cockpit voice noise reduction method in an exemplary embodiment of this application;

[0020] Figure 3 This is a flowchart illustrating a specific cockpit voice noise reduction method as an exemplary embodiment of this application;

[0021] Figure 4 This is an exemplary embodiment of the present application illustrating a specific ERB feature encoder execution flowchart;

[0022] Figure 5 This is an exemplary embodiment of the present application illustrating a specific ERB feature decoder execution flowchart;

[0023] Figure 6 This is an exemplary embodiment of the present application illustrating a specific low-frequency feature decoder execution flowchart;

[0024] Figure 7 This is a schematic diagram of a cockpit voice noise reduction device shown in an exemplary embodiment of this application;

[0025] Figure 8 This is a schematic diagram of the structure of a computer system for an electronic device, as illustrated in an exemplary embodiment of this application. Detailed Implementation

[0026] The embodiments of this application will be described below with reference to the accompanying drawings and specific examples. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be understood that the preferred embodiments are only for illustrating this application and are not intended to limit the scope of protection of this application.

[0027] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. Therefore, the illustrations only show the components related to this application and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0028] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the present application. However, it will be apparent to those skilled in the art that embodiments of the present application may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the present application.

[0029] The term "and / or" used in this application describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the related objects before and after it are in an "or" relationship.

[0030] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an exemplary system architecture as shown in an exemplary embodiment of this application.

[0031] Reference Figure 1 As shown, the system architecture may include a vehicle 110 and a computer device 120. The computer device 120 acquires the initial noisy speech signal from the cabin of the vehicle 110, extracts spectral features from the initial noisy speech signal to obtain full-band spectral features, performs equivalent rectangular bandwidth transformation on the full-band spectral features to obtain sub-band features, and truncates the low-frequency portion of the full-band spectral features to obtain low-frequency features. The sub-band features and low-frequency features are then interactively enhanced to obtain sub-band coded features and low-frequency coded features, respectively. These sub-band coded features and low-frequency coded features are then fused to generate fused spectral features. The fused spectral features are input to the sub-band feature decoder and the low-frequency feature decoder to obtain a sub-band mask and a low-frequency mask, respectively, to determine the noise reduction spectral features based on the sub-band mask, the low-frequency mask, and the full-band spectral features. The aforementioned vehicle 110 includes a cockpit with voice functionality, which includes at least a microphone to collect noisy speech within the cockpit; the aforementioned computer equipment 120 refers to a computing power support terminal device used to carry out the program implementation environment for the cockpit speech noise reduction method, including but not limited to tablet devices, microcomputers, embedded computers, and cloud servers.

[0032] In a schematic manner, computer device 120 acquires the initial noisy speech signal from the cabin of vehicle 110, extracts spectral features from the initial noisy speech signal to obtain full-band spectral features, performs equivalent rectangular bandwidth transformation on the full-band spectral features to obtain sub-band features, and truncates the low-frequency portion of the full-band spectral features to obtain low-frequency features. Interactive enhancement processing is performed on the sub-band features and low-frequency features respectively to obtain sub-band coded features and low-frequency coded features. The sub-band coded features and low-frequency coded features are then fused to generate fused spectral features. The fused spectral features are input to the sub-band feature decoder and the low-frequency feature decoder respectively to obtain the sub-band mask and the low-frequency mask. The noise reduction spectral features are then determined based on the sub-band mask, the low-frequency mask, and the full-band spectral features. This application achieves noise suppression by separately processing the sub-band features and low-frequency features in the full-band spectral features and employing interactive enhancement processing and a dual-decoder structure. While suppressing full-band noise, it effectively preserves low-frequency speech features, thereby improving speech clarity and naturalness.

[0033] Figure 2 This is a flowchart illustrating an exemplary embodiment of the present application of a cockpit voice noise reduction method, which can... Figure 1 It can be executed in the implementation environment described above, but it can also be implemented in other implementation environments. No specific limitations are imposed on the aforementioned implementation environments here. (See also...) Figure 2 As shown in the flowchart, the cockpit voice noise reduction method includes at least steps S210 to S240, which are described in detail below:

[0034] In step S210, the initial noisy speech signal in the cockpit is acquired, and the spectral features of the initial noisy speech signal are extracted to obtain the full-band spectral features.

[0035] In one embodiment of this application, spectral feature extraction of the initial noisy speech signal includes extracting the complete frequency domain of the speech signal through time-frequency transformation, including but not limited to using short-time Fourier transform or Mel filter banks. The initial noisy speech signal is a sound signal collected from inside the vehicle cabin, containing the sound of people speaking inside the cabin and various environmental noises, such as the static sound of doors opening and closing, and the dynamic sound of tire noise and wind noise. The final form is noisy speech after the mixture of various sounds. For example, if the initial noisy speech signal has 512 sampling points, after undergoing a 512-point FFT (Fast Fourier Transform), a spectral feature of length 257 will be obtained.

[0036] In step S220, the full-band spectrum features are converted to equivalent rectangular bandwidth to obtain sub-band features, and the low-frequency part of the full-band spectrum features is truncated to obtain low-frequency features.

[0037] In one embodiment of this application, the aforementioned low-frequency portion truncation includes truncating spectral features below a preset low-frequency threshold from the full-band spectral features.

[0038] In one embodiment of this application, the above-mentioned equivalent rectangular bandwidth conversion includes dividing the spectrum into sub-bands that conform to the characteristics of human hearing. In some feasible environments, the conversion can be performed by an ERB (Equivalent Rectangular Bandwidth) filter bank. The equivalent rectangular bandwidth conversion can reflect the differences in the sensitivity of the auditory system to noise in different frequency bands.

[0039] In a specific embodiment of this application, during the process of performing equivalent rectangular bandwidth conversion on full-band spectral features, the full-band spectral features are converted into 32-dimensional sub-band features. It is important to note that the width of the equivalent rectangular bandwidth feature needs to be predefined before calculating the sub-band features. For example, if the width of the sub-band feature ERB is W = [w1, w2, w3, ..., wn], where n = 32, then the calculation of the equivalent rectangular bandwidth feature conversion includes:

[0040] ERB(1)=(spec[1]+…+spec[w1]) / w1;

[0041] ERB(2)=(spec[w1+1]+spec[w1+2]+…+spec[w1+w2]) / w2;

[0042]

[0043] Here, spec refers to the full-band spectral characteristics.

[0044] In a specific embodiment of this application, during the process of low-frequency truncation of full-band spectral features, in some feasible environments, when the number of sampling points of the initial noisy speech signal is 512, the full-band spectral features include 257 frequency points, and the low-frequency truncation can take the first 48 dimensions of the 257 frequency points.

[0045] In step S230, the sub-band features and low-frequency features are interactively enhanced to obtain sub-band coded features and low-frequency coded features, and the sub-band coded features and low-frequency coded features are fused to generate fused spectral features.

[0046] In one embodiment of this application, subband features and low-frequency features are input into a subband feature encoder. Attention is calculated using low-frequency features as queries and subband features as keys and values ​​to obtain subband coding weights. Subband features are then reconstructed using these subband coding weights to obtain subband coding features.

[0047] In one embodiment of this application, the subband feature encoder includes a neural network module for processing subband features after equivalent rectangular bandwidth transformation, including but not limited to a multilayer perceptron or a convolutional neural network. It establishes the correlation between subband features and low-frequency features through an attention mechanism, including modifying the subband features by adding low-frequency feature information to the subband feature information. The aforementioned attention calculation assigns weights based on the similarity between features to capture potential correlations between cross-frequency band features. Feature reconstruction refers to weighted fusion of the original features based on attention weights, including fusion through matrix multiplication, to strengthen the parts highly correlated with the subband features.

[0048] Specifically, in the subband feature encoder, low-frequency features are used as query vectors and key vectors of subband features for similarity matching to generate subband encoding weights that reflect the degree of association between the two. The value vectors of subband features are then weighted and summed based on the subband encoding weights to retain the subband features while introducing global information from low-frequency features.

[0049] In one embodiment of this application, subband features and low-frequency features are input into a low-frequency feature encoder. Attention is calculated using subband features as queries and low-frequency features as keys and values ​​to obtain low-frequency coding weights. Low-frequency features are then reconstructed using these low-frequency coding weights to obtain low-frequency coding features.

[0050] In one embodiment of this application, the low-frequency feature encoder includes a neural network module for processing truncated low-frequency features. In some embodiments, it adopts a structure symmetrical to the sub-band feature encoder. Feature reconstruction refers to weighted fusion of the original features based on attention weights, including fusion through matrix multiplication, to strengthen the parts that are highly correlated with the low-frequency features.

[0051] Specifically, in the low-frequency feature encoder, the sub-band feature is used as a query vector to match the key vector of the low-frequency feature, generating a low-frequency encoding weight, and reconstructing the low-frequency feature based on the low-frequency encoding weight, so that the low-frequency feature can integrate the detailed information in the sub-band feature.

[0052] In one embodiment of this application, a bidirectional attention mechanism enables sub-band features and low-frequency features to complement each other across frequency bands during the encoding process, thereby enhancing the representational capabilities of their respective features. Specifically, this approach can enhance the noise resistance of low-frequency features while preserving high-frequency speech details, reduce speech distortion caused by fragmented frequency band feature processing, improve the intelligibility and naturalness of speech signals in complex noise scenarios, and suppress noise interference on specific frequency band features through cross-band feature interaction, thereby enhancing the coherence of speech components across different frequency bands.

[0053] In one embodiment of this application, fusing subband coding features and low-frequency coding features includes concatenating the subband coding features and low-frequency coding features to form concatenated coding features, inputting the concatenated coding features into a preset neural network for weight allocation to obtain coding fusion weights, the preset neural network being used for weight allocation of feature fusion of subband coding features and low-frequency coding features, and performing weighted summation of subband coding features and low-frequency coding features according to the coding fusion weights to obtain fused spectral features.

[0054] In one embodiment of this application, concatenating subband coding features with low-frequency coding features includes forming a composite feature vector by connecting subband coding features and low-frequency coding features along the feature dimension. The process can be based on a fully connected layer or an attention mechanism to achieve channel-dimensional concatenation, so as to preserve the high-frequency details of the subband features and the structural information of the low-frequency features.

[0055] In one embodiment of this application, the aforementioned preset neural network refers to a deep learning model used for feature fusion, which can dynamically allocate feature weights through the adaptive learning capability of the neural network. In a specific embodiment, it can be characterized as follows:

[0056] weight=torch.sigmoid(MLP(torch.cat([new_erb_feat,new_df_feat],dim=-1)))

[0057] Wherein, new_erb_feat represents the subband encoded feature output by the ERB encoder; new_df_feat represents the low-frequency encoded feature output by the low-frequency encoder; torch.cat(...) represents splicing the subband encoded feature and the low-frequency encoded feature; MLP(...) represents a deep learning model including a small multilayer perceptron, used to learn and assign weights to the two features; torch.sigmoid(...) represents compressing the output of the deep learning model to the [0,1] interval, finally obtaining a floating-point number weight.

[0058] If the weight is close to 1, it means that the subband coding features of the current frame are more important, such as high frequencies needing to be suppressed in noisy environments; if the weight is close to 0, it means that low frequency coding features are more critical.

[0059] During the fusion process, the encoding fusion weights are proportional parameters reflecting the importance of subbands and low-frequency features. They are obtained by normalizing the neural network output using the Softmax function. The feature fusion process can be characterized as follows:

[0060] output=weight*new_erb_feat+(1-weight)*new_df_feat

[0061] It is represented as a linear combination of sub-band coding features and low-frequency coding features according to weights, which is used to retain the auditory optimization characteristics of sub-band coding features, while injecting the temporal dynamic information of low-frequency coding features to avoid the rigidity problem of fixed ratio fusion.

[0062] In one embodiment of this application, during the feature fusion stage, the interactively enhanced sub-band coding features and low-frequency coding features are first concatenated along the channel dimension to form a concatenated coding feature containing full-band information. This concatenated coding feature is then fed into a pre-defined neural network model, which analyzes the correlation between the sub-band features and low-frequency features in the time-frequency domain through multi-layer nonlinear transformations, outputting fusion weight coefficients. Finally, the sub-band coding features and low-frequency coding features are fused according to dynamically allocated weights using a weighted summation method, generating a fused spectral feature that retains both high-frequency noise suppression capabilities and low-frequency voice protection characteristics. This process effectively balances the needs of high-frequency noise suppression and low-frequency voice protection, achieving more accurate frequency band feature fusion in complex vehicle noise scenarios. It avoids speech distortion or noise residue problems caused by traditional fixed fusion strategies, significantly improving the clarity and naturalness of denoised speech.

[0063] In step S240, the fused spectral features are input to the sub-band feature decoder and the low-frequency feature decoder respectively to obtain the sub-band mask and the low-frequency mask, so as to determine the noise reduction spectral features based on the sub-band mask, the low-frequency mask and the full-band spectral features.

[0064] In one embodiment of this application, the fused spectral features are input into a sub-band feature decoder for decoding, and a sub-band mask corresponding to the full-band spectral features is output. The sub-band mask is used to suppress noise in the full-band portion of the full-band spectral features.

[0065] The subband feature decoder is a spectrum decoding module built on a fully connected network or convolutional neural network. Its input is the fused spectral features, and its output is a mask coefficient matrix covering the entire frequency band. The subband mask is a two-dimensional matrix generated by the subband feature decoder, where each element corresponds to the gain coefficient at each frequency point of the full-band spectrum, used to attenuate noise components in the frequency domain. In some specific implementations, the subband feature decoder can decode the fused spectral features into a 257-dimensional mask.

[0066] In one embodiment of this application, the fused spectral features are input into a low-frequency feature decoder for decoding, and a low-frequency mask corresponding to the low-frequency part of the full-band spectral features is output. The low-frequency mask is used to suppress noise in the low-frequency part of the full-band spectral features.

[0067] The low-frequency feature decoder refers to a decoding module built on a neural network structure based on bandpass filtering constraints. Its output is a mask coefficient that operates only on the low-frequency band. Specifically, it refers to a one-dimensional vector generated by the low-frequency feature decoder, where the vector elements correspond to the gain coefficients of each frequency point in the low-frequency band, used to selectively suppress low-frequency noise. In some specific implementation environments, the low-frequency feature decoder can decode the fused spectral features into a 48-dimensional low-frequency mask.

[0068] Specifically, the fused spectral features are simultaneously input into two independently operating decoder modules. The sub-band feature decoder maps the fused features to a full-band mask through multi-layer nonlinear transformations. This mask is then multiplied point-by-point with the original noisy speech spectrum in the frequency domain, achieving global suppression of broadband noise. Meanwhile, the low-frequency feature decoder, during decoding, uses a frequency band constraint mechanism to reconstruct features only in low-frequency regions below a preset threshold, generating a mask that operates only on the low-frequency band. This mask is then multiplied with the low-frequency portion of the original spectrum, focusing on eliminating low-frequency interference. This approach effectively addresses the mixed interference characteristics of broadband and low-frequency noise in the vehicle cabin environment, maintaining high-frequency speech clarity while avoiding distortion from low-frequency speech formants, significantly improving the naturalness and intelligibility of the denoised speech.

[0069] In one embodiment of this application, when determining the noise reduction spectrum features based on the sub-band mask, low-frequency mask, and full-band spectrum features, the sub-band mask is multiplied by the full-band spectrum features to obtain the sub-band noise reduction result, and the low-frequency mask is multiplied by the low-frequency part of the full-band spectrum features to obtain the low-frequency noise reduction result. Then, the sub-band noise reduction result and the low-frequency noise reduction result are fused to obtain the noise reduction spectrum features.

[0070] The multiplication process involves performing a point-by-point multiplication operation between the mask matrix and the spectral feature matrix, specifically through matrix element multiplication. Specifically, after obtaining the sub-band mask and the low-frequency mask, the sub-band mask is first applied to the full-band spectral feature, suppressing noise components in the mid-to-high frequency regions through point-by-point multiplication while preserving the main body of the speech signal. The low-frequency mask is then applied to a pre-trimmed low-frequency portion of the full-band spectral feature, suppressing low-frequency noise such as engine vibration through the same operation. Finally, the processed sub-band noise reduction result and the low-frequency noise reduction result are fused in the frequency domain to form a noise-reduced spectral feature that balances noise suppression across the entire frequency band with speech integrity.

[0071] The process of fusing subband denoising results with low-frequency denoising results to obtain denoised spectral features includes: predicting the confidence level of subband denoising results to obtain subband confidence scores; predicting the confidence level of low-frequency denoising results to obtain low-frequency confidence scores; normalizing the subband confidence scores and low-frequency confidence scores to obtain subband denoising weights and low-frequency denoising weights; and then weighting and summing the subband denoising results and low-frequency denoising results based on the subband denoising weights and low-frequency denoising weights to obtain denoised spectral features.

[0072] In one embodiment of this application, confidence prediction includes evaluating the reliability of the denoising result through a computational model. This can be achieved by probabilistically predicting the degree of speech component retention in the denoising result based on a neural network model. Normalization refers to converting the confidence score into weight coefficients using a softmax function, which can dynamically allocate the fusion ratio between sub-band denoising results and low-frequency denoising results. Specifically, the confidence prediction module generates reliability indices for both types of results. This can be achieved by performing feature analysis on the denoising result using an evaluation network constructed with, but not limited to, convolutional layers and fully connected layers. Then, normalization converts the confidence score into weight values ​​with a sum of 1. The sub-band denoising result is then multiplied by the sub-band denoising weight, and the low-frequency denoising result is multiplied by the low-frequency denoising weight. These are then summed to form a fused denoised spectral feature.

[0073] In one embodiment of this application, after determining the noise reduction spectral features based on the sub-band mask, low-frequency mask and full-band spectral features, the noise reduction spectral features are subjected to inverse Fourier transform to obtain the time-domain signal to be processed, and the time-domain signal to be processed is subjected to overlapping and addition processing to obtain the time-domain speech signal.

[0074] In one embodiment of this application, the inverse Fourier transform refers to the process of converting the denoised spectral features represented in the frequency domain into a time domain signal, including the inverse fast Fourier transform algorithm, which can restore the denoised spectral features after frequency domain denoising to a playable time domain waveform; the overlapping addition processing refers to the operation of smoothing the inter-frame superposition of the time domain signals processed by frame segmentation. In the specific implementation process, adjacent frame signals are weighted and superimposed by a window function with a fixed step size to eliminate the signal discontinuity caused by frame segmentation and restore the complete time domain speech signal.

[0075] Figure 3This is a flowchart illustrating a specific cockpit voice noise reduction method according to an exemplary embodiment of this application. It should be noted that in the following specific embodiments, the noisy speech spectral characteristics are consistent with the full-band spectral characteristics in the above embodiments, the ERP noise reduction result is consistent with the sub-band noise reduction result in the above embodiments, the ERP feature encoder is consistent with the sub-band feature encoder in the above embodiments, the low-frequency noise reduction result is consistent with the low-frequency noise reduction result in the above embodiments, and the final noise reduction result is consistent with the noise reduction spectral characteristics in the above embodiments. (Refer to...) Figure 3 As shown in the flowchart of this specific cockpit speech noise reduction method, the noisy speech spectral features are subjected to ERP feature extraction to obtain sub-band features, and low-frequency features are extracted to obtain low-frequency features. The extracted sub-band features and low-frequency features are then input into the ERP feature encoder and low-frequency feature encoder, respectively. The ERP feature encoder outputs the sub-band encoded features, and the low-frequency feature encoder outputs the low-frequency encoded features. Figure 4 This is an exemplary embodiment of the present application illustrating a specific ERB feature encoder execution flowchart, with reference to... Figure 4 As shown, after receiving ERB feature input, the ERB feature encoder performs feature interaction enhancement through three Conv convolutional layers, and performs weighted calculation and reconstruction of features based on liner linear layers and sequence processing by GRU gated recurrent units to output sub-band encoded features. Similarly, the low-frequency feature encoder also includes Conv convolutional layers, liner linear layers and GRU gated recurrent units, but the low-frequency feature encoder has fewer Conv convolutional layers than the ERB feature encoder.

[0076] In one specific embodiment of this application, sub-band coding features and low-frequency coding features are input into a feature fusion module to generate fused spectral features. These fused spectral features are then input into the ERP feature decoder and the low-frequency feature decoder, respectively. The ERP feature decoder outputs the ERP noise reduction result. Figure 5 This is an exemplary embodiment of the present application illustrating a specific ERB feature decoder execution flowchart, with reference to... Figure 5 As shown, after the ERB feature decoder receives the merged features (fused spectral features), it undergoes a nonlinear transformation through three layers of Tconv transposed convolutions, and then a Conv convolutional layer maps the fused features into a full-band mask for output; the low-frequency feature decoder outputs the low-frequency noise reduction result, specifically, Figure 6 This is an exemplary embodiment of the present application illustrating a specific low-frequency feature decoder execution flowchart, with reference to... Figure 6As shown, after the low-frequency feature decoder receives the merged features (fused spectral features), the features are transformed by two Conv convolutional layers, then processed by a sequence based on a GRU gated recurrent unit, and finally linearly processed by a liner linear layer to output the low-frequency noise reduction result. The ERP noise reduction result and the low-frequency noise reduction result are then fused to obtain the final noise reduction result.

[0077] This application provides a cockpit speech noise reduction method, apparatus, device, and storage medium. It acquires an initial noisy speech signal from within the cockpit, extracts spectral features from the initial noisy speech signal to obtain full-band spectral features, performs an equivalent rectangular bandwidth transformation on the full-band spectral features to obtain sub-band features, and truncates the low-frequency portion of the full-band spectral features to obtain low-frequency features. It then performs interactive enhancement processing on the sub-band features and low-frequency features to obtain sub-band coded features and low-frequency coded features, respectively. Finally, it fuses the sub-band coded features and low-frequency coded features to generate fused spectral features. These fused spectral features are input to the sub-band feature decoder and the low-frequency feature decoder to obtain a sub-band mask and a low-frequency mask, respectively. Noise reduction spectral features are then determined based on the sub-band mask, low-frequency mask, and full-band spectral features. This application achieves noise suppression by separately processing the sub-band features and low-frequency features within the full-band spectral features and employing interactive enhancement processing and a dual-decoder structure. This effectively preserves low-frequency speech features while suppressing full-band noise, thereby improving speech clarity and naturalness.

[0078] The following describes an embodiment of the apparatus described in this application, which can be used to execute the cockpit voice noise reduction method described in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the cockpit voice noise reduction method described above in this application.

[0079] Figure 7 This is a schematic diagram illustrating a cockpit voice noise reduction device according to an exemplary embodiment of this application. The device can be applied to... Figure 2 The method implementation process shown can be based on the device Figure 1 The implementation environment shown can be applied to other exemplary implementation environments and specifically configured in other devices. This embodiment does not limit the implementation environment to which the device is applicable.

[0080] like Figure 7 As shown, the exemplary cockpit voice noise reduction device includes: a signal feature extraction module 701 and a spectrum feature noise reduction module 702.

[0081] The signal feature extraction module 701 is used to acquire the initial noisy speech signal in the cockpit and extract the spectral features of the initial noisy speech signal to obtain full-band spectral features; perform equivalent rectangular bandwidth transformation on the full-band spectral features to obtain sub-band features, and truncate the low-frequency part of the full-band spectral features to obtain low-frequency features. The low-frequency part truncation includes truncating the spectral features below a preset low-frequency threshold in the full-band spectral features; the spectral feature denoising module 702 is used to perform interactive enhancement processing on the sub-band features and low-frequency features respectively to obtain sub-band coded features and low-frequency coded features, and fuse the sub-band coded features and low-frequency coded features to generate fused spectral features; input the fused spectral features to the sub-band feature decoder and low-frequency feature decoder respectively to obtain sub-band mask and low-frequency mask, and determine the denoised spectral features based on the sub-band mask, low-frequency mask and full-band spectral features.

[0082] Embodiments of this application also provide an electronic device, including: one or more processors; and a storage device for storing one or more programs, which, when executed by one or more processors, cause the electronic device to implement the cockpit voice noise reduction method provided in the above embodiments.

[0083] Figure 8 This is a schematic diagram illustrating the structure of a computer system for an electronic device, as shown in an exemplary embodiment of this application. It should be noted that... Figure 8 The computer system 800 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0084] like Figure 8 As shown, the computer system 800 includes a Central Processing Unit (CPU) 801, which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, based on a program stored in Read-Only Memory (ROM) 802 or a program loaded from storage into Random Access Memory (RAM) 803. The RAM 803 also stores various programs and data required for system operation. The CPU 801, ROM 802, and RAM 803 are interconnected via a bus. An I / O interface 805 is also connected to the bus 804, where the I / O interface 805 refers to an input / output interface.

[0085] The following components are connected to I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section performs communication processing via a network such as the Internet. A drive is also connected to I / O interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 810 as needed so that computer programs read from it can be installed into storage section 808 as needed.

[0086] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 809, and / or installed from removable medium 811. When the computer program is executed by central processing unit (CPU) 801, it performs various functions defined in the system of this application.

[0087] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0088] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0089] In the corresponding figures of the above embodiments, connecting lines can represent the connection relationship between various components, indicating more constitutive signal paths and / or one or more ends of some lines having arrows to indicate the main information flow direction. Connecting lines are an identifier and are not a limitation on the scheme itself, but rather the use of these lines in conjunction with one or more exemplary embodiments helps to more easily connect circuits or logic units. Any signal represented (determined by design requirements or preferences) can actually include one or more signals that can be transmitted in any direction and can be implemented in any suitable type of signal scheme.

[0090] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0091] Another aspect of this application provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method as described above. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not assembled into the electronic device.

[0092] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the cockpit voice noise reduction method as described in any of the above embodiments.

[0093] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0094] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0095] This application can be used in a wide range of general-purpose or special-purpose computing system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc.

[0096] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0097] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.

Claims

1. A cockpit voice noise reduction method, characterized in that, The cockpit voice noise reduction method includes: Acquire the initial noisy speech signal in the cockpit, and extract the spectral features of the initial noisy speech signal to obtain the full-band spectral features; The full-band spectrum features are converted into equivalent rectangular bandwidth to obtain sub-band features, and the low-frequency portion of the full-band spectrum features is truncated to obtain low-frequency features. The low-frequency truncation includes truncating the spectrum features of the full-band spectrum features that are below a preset low-frequency threshold. The sub-band features and the low-frequency features are respectively subjected to interactive enhancement processing to obtain sub-band coding features and low-frequency coding features, and the sub-band coding features and low-frequency coding features are fused to generate fused spectral features; The fused spectral features are input to the sub-band feature decoder and the low-frequency feature decoder respectively to obtain the sub-band mask and the low-frequency mask, so as to determine the noise reduction spectral features based on the sub-band mask, the low-frequency mask and the full-band spectral features. The interactive enhancement processing of the sub-band features and the low-frequency features includes: inputting the sub-band features and the low-frequency features into a sub-band feature encoder; performing attention calculation using the low-frequency features as a query and the sub-band features as a key and value to obtain sub-band encoding weights; and reconstructing the sub-band features using the sub-band encoding weights to obtain sub-band encoded features; and inputting the sub-band features and the low-frequency features into a low-frequency feature encoder; performing attention calculation using the sub-band features as a query and the low-frequency features as a key and value to obtain low-frequency encoding weights; and reconstructing the low-frequency features using the low-frequency encoding weights to obtain low-frequency encoded features. The process of inputting the fused spectral features into a sub-band feature decoder and a low-frequency feature decoder to obtain a sub-band mask and a low-frequency mask includes: inputting the fused spectral features into the sub-band feature decoder for decoding, and outputting a sub-band mask corresponding to the full-band spectral features, wherein the sub-band mask is used to suppress noise in the full-band portion of the full-band spectral features; and inputting the fused spectral features into the low-frequency feature decoder for decoding, and outputting a low-frequency mask corresponding to the low-frequency portion of the full-band spectral features, wherein the low-frequency mask is used to suppress noise in the low-frequency portion of the full-band spectral features. Determining the noise reduction spectral features based on the sub-band mask, low-frequency mask, and full-band spectral features includes: multiplying the sub-band mask by the full-band spectral features to obtain a sub-band noise reduction result; multiplying the low-frequency mask by the low-frequency portion of the full-band spectral features to obtain a low-frequency noise reduction result; and fusing the sub-band noise reduction result with the low-frequency noise reduction result to obtain the noise reduction spectral features.

2. The cockpit voice noise reduction method according to claim 1, characterized in that, The sub-band coding features and low-frequency coding features are fused to generate fused spectral features, including: The sub-band coding features are concatenated with the low-frequency coding features to form concatenated coding features; The concatenated coding features are input into a preset neural network for weight allocation to obtain coding fusion weights. The preset neural network is used to allocate weights for feature fusion of sub-band coding features and low-frequency coding features. The sub-band coding features and the low-frequency coding features are weighted and summed according to the coding fusion weights to obtain the fused spectral features.

3. The cockpit voice noise reduction method according to claim 1, characterized in that, The sub-band noise reduction result is fused with the low-frequency noise reduction result to obtain the noise reduction spectral features, including: The confidence score of the sub-band noise reduction result is predicted, and the confidence score of the low-frequency noise reduction result is predicted, and the low-frequency confidence score is obtained. The sub-band confidence score and the low-frequency confidence score are normalized to obtain the sub-band denoising weight and the low-frequency denoising weight. The sub-band denoising result and the low-frequency denoising result are then weighted and summed based on the sub-band denoising weight and the low-frequency denoising weight to obtain the denoised spectral features.

4. The cockpit voice noise reduction method according to any one of claims 1-3, characterized in that, After determining the noise reduction spectral features based on the sub-band mask, low-frequency mask, and full-band spectral features, the cockpit voice noise reduction method further includes: The denoised spectral features are subjected to inverse Fourier transform to obtain the time-domain signal to be processed, and the time-domain signal to be processed is overlapped and added to obtain the time-domain speech signal.

5. A cockpit voice noise reduction device, characterized in that, For implementing the method as described in any one of claims 1-4, the cockpit voice noise reduction device comprises: The signal feature extraction module is used to acquire the initial noisy speech signal in the cockpit, and to extract the spectral features of the initial noisy speech signal to obtain full-band spectral features; to perform equivalent rectangular bandwidth transformation on the full-band spectral features to obtain sub-band features, and to truncate the low-frequency portion of the full-band spectral features to obtain low-frequency features, wherein the low-frequency truncation includes truncating the spectral features of the full-band spectral features that are below a preset low-frequency threshold; The spectral feature denoising module is used to perform interactive enhancement processing on the sub-band features and the low-frequency features respectively to obtain sub-band coded features and low-frequency coded features, and to fuse the sub-band coded features and low-frequency coded features to generate fused spectral features; the fused spectral features are input to the sub-band feature decoder and the low-frequency feature decoder respectively to obtain the sub-band mask and the low-frequency mask, so as to determine the denoising spectral features based on the sub-band mask, the low-frequency mask and the full-band spectral features.

6. An electronic device, characterized in that, It includes a processor, a memory, and a communication bus; the communication bus is used to connect the processor and the memory; the processor is used to execute a computer program stored in the memory to implement the cockpit voice noise reduction method as described in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, It stores a computer program that enables the computer to perform the cockpit voice noise reduction method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Audio noise reduction processing method and device, storage medium and electronic equipment

    CN116959476A

  • Voice enhancement methods and systems

    US20230317093A1