Voice signal processing method and device, electronic equipment and storage medium

By using a joint denoising and recognition model to process speech signals, the problem of noise interference in VHF communication is solved, achieving accurate denoising and recognition of speech signals and improving communication quality.

CN121565174APending Publication Date: 2026-02-24SHANGHAI MERCHANT SHIP DESIGN & RES INST
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511750898.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing technologies in VHF voice signal communication struggle to simultaneously address both acoustic environmental noise and channel noise, resulting in poor voice recognition accuracy and failing to meet communication accuracy requirements.

Method used

A joint denoising and recognition model is adopted, which combines a denoising sub-model and a speech recognition sub-model to acquire and process speech spectrum features and real-time channel quality features, thereby achieving accurate denoising and recognition and generating clear and accurate speech signals.

Benefits of technology

It improves the denoising accuracy and recognition accuracy of voice signals, enhances the clarity and accuracy of communication voice signals, and takes into account the impact of environmental noise and channel noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565174A_ABST
    Figure CN121565174A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a voice signal processing method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining and inputting a to-be-enhanced voice spectrum feature and a corresponding real-time channel quality feature to a pre-trained denoising recognition joint model, the denoising recognition joint model comprising a denoising sub-model and a voice recognition sub-model; based on the to-be-enhanced speech spectrum features and the corresponding real-time channel quality features, obtaining de-noised speech spectrum features corresponding to the to-be-enhanced speech spectrum features through a de-noising sub-model; obtaining a de-noised speech recognition text corresponding to the to-be-enhanced speech spectrum feature through a speech recognition sub-model based on the de-noised speech spectrum feature corresponding to the to-be-enhanced speech spectrum feature; and generating an enhanced voice signal corresponding to the to-be-enhanced voice spectrum feature based on the de-noised voice recognition text. According to the embodiment of the invention, the accuracy of denoising and enhancing the voice signal can be improved, and the accuracy of recognizing the voice signal is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a speech signal processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] In scenarios involving voice communication, such as maritime operations using Very High Frequency (VHF) voice signals, it is often necessary to convert the voice signal into text. For example, in cross-language communication, voice signals frequently contain acoustic environmental noise and channel noise. For instance, during VHF communication voice analog frequency modulation (FM), if the signal is weak, it can produce a strong static "hissing" sound and signal attenuation distortion. Existing technologies cannot simultaneously address both types of noise when denoising, resulting in either insufficient or excessive denoising. The voice signals resulting from insufficient or excessive denoising are not the optimal inputs for downstream voice device models to easily recognize, leading to poor voice recognition performance and failing to meet the accuracy requirements of voice communication. Summary of the Invention

[0003] This invention provides a speech signal processing method, apparatus, electronic device, and storage medium, which can improve the accuracy of speech signal denoising and enhancement, and improve the accuracy of speech signal recognition.

[0004] In a first aspect, embodiments of the present invention provide a speech signal processing method, comprising:

[0005] The speech spectrum features to be enhanced and the corresponding real-time channel quality features are acquired and input into a pre-trained denoising and recognition joint model, which includes a denoising sub-model and a speech recognition sub-model.

[0006] Based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features, the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced are obtained through a denoising sub-model.

[0007] Based on the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced, the denoised speech recognition text corresponding to the speech spectrum features to be enhanced is obtained through a speech recognition sub-model; and

[0008] Enhanced speech signals are generated based on the spectral features of the speech to be enhanced, derived from denoised speech recognition text.

[0009] Secondly, embodiments of the present invention provide a speech signal processing apparatus, comprising:

[0010] The feature acquisition and input module is used to acquire and input the speech spectrum features to be enhanced and the corresponding real-time channel quality features into the pre-trained denoising and recognition joint model, which includes a denoising sub-model and a speech recognition sub-model.

[0011] The denoising module is used to obtain the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced through a denoising sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features.

[0012] The recognition module is used to obtain the denoised speech recognition text corresponding to the denoised speech spectrum features corresponding to the denoised speech spectrum features to be enhanced through a speech recognition sub-model; and

[0013] The enhanced speech signal generation module is used to generate an enhanced speech signal corresponding to the spectral features of the speech to be enhanced based on the denoised speech recognition text.

[0014] Thirdly, embodiments of the present invention also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the speech signal processing method as described in any of the embodiments of the present invention.

[0015] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the speech signal processing method as described in any of the embodiments of the present invention.

[0016] This invention provides a speech signal processing method, apparatus, electronic device, and storage medium. By acquiring and inputting the spectral features of the speech to be enhanced and the corresponding real-time channel quality features into a joint denoising and recognition model including a denoising sub-model and a speech recognition sub-model, the denoising sub-model obtains the denoised speech spectral features corresponding to the spectral features of the speech to be enhanced based on the spectral features of the speech to be enhanced and the corresponding real-time channel quality features. Then, the speech recognition sub-model obtains the denoised speech recognition text corresponding to the denoised speech spectral features of the speech to be enhanced. This approach can take into account both environmental noise and channel noise in the speech signal, avoiding insufficient or excessive denoising, improving the accuracy of denoising, and enhancing the accuracy of speech signal recognition. Furthermore, this invention generates an enhanced speech signal corresponding to the spectral features of the speech to be enhanced based on the denoised speech recognition text, which can improve the clarity and accuracy of communication speech signals. Attached Figure Description

[0017] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic flowchart of a speech signal processing method provided in an embodiment of the present invention;

[0019] Figure 2 This is another schematic flowchart of the speech signal processing method provided in this embodiment of the invention;

[0020] Figure 3 This is another schematic flowchart of the speech signal processing method provided in this embodiment of the invention;

[0021] Figure 4 This is another schematic flowchart of the speech signal processing method provided in this embodiment of the invention;

[0022] Figure 5 This is another schematic flowchart of the speech signal processing method provided in this embodiment of the invention;

[0023] Figure 6 This is a schematic diagram of a voice communication device provided in an embodiment of the present invention;

[0024] Figure 7 This is a schematic diagram of a speech signal processing device provided in an embodiment of the present invention;

[0025] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0026] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0028] Figure 1 This is a schematic flowchart illustrating a speech signal processing method provided in an embodiment of the present invention. This embodiment is applicable to voice communication devices, and the method can be executed by the speech signal processing device provided in this embodiment. This device can be implemented using software and / or hardware. In a specific embodiment, the device can be integrated into an electronic device, such as a computer or server. The following embodiments will illustrate this using the integration of the device into an electronic device as an example. (Reference) Figure 1 The method may specifically include the following steps:

[0029] Step 101: Obtain and input the spectral features of the speech to be enhanced and the corresponding real-time channel quality features into the pre-trained denoising and recognition joint model. The denoising and recognition joint model includes a denoising sub-model and a speech recognition sub-model. This step enables accurate denoising and precise recognition of the original speech signal based on the spectral features of the speech to be enhanced and the real-time channel quality through the denoising and recognition joint model.

[0030] Specifically, the communication device used in the embodiments of the present invention can be a VHF communication device, or other types of communication devices, such as a HF communication device or a LOW-HEN communication device.

[0031] Specifically, the aforementioned speech spectral features to be enhanced can be Mel spectrograms, or other types of spectral feature signals, such as the signal spectrum obtained by performing a short-time Fourier transform on the corresponding original speech signal.

[0032] Specifically, the original speech signal corresponding to the speech spectral features to be enhanced can be a VHF speech signal, or other types of speech signals, such as 4G / 5G communication signals.

[0033] Specifically, the aforementioned original voice signal can be either the communication voice signal to be sent or the communication voice signal received.

[0034] Specifically, the aforementioned real-time channel quality characteristics can be understood as real-time characteristics that characterize the degree of electrostatic interference and the degree of signal attenuation and distortion in the transmission channel that transmits the original voice signal.

[0035] Optionally, the process of obtaining the real-time channel quality features corresponding to the speech spectrum features to be enhanced includes: obtaining the latest historical speech signal transmitted by the transmission channel corresponding to the speech spectrum features to be enhanced, and obtaining the speech quality index value of the latest historical speech signal; and constructing the real-time channel quality features corresponding to the speech spectrum features to be enhanced based on the speech quality index value of the latest historical speech signal.

[0036] Specifically, the aforementioned latest historical voice signal can be a single voice signal or multiple voice signals.

[0037] Specifically, the aforementioned latest historical voice signal can be the communication voice signal that needs to be sent when the voice signal was most recently sent, or it can be the communication voice signal that was most recently received.

[0038] Specifically, the aforementioned speech quality metrics may include the signal-to-noise ratio and / or high-frequency noise quantization metrics of the latest historical speech signals.

[0039] Specifically, the signal-to-noise ratio of the latest historical speech signal can be obtained through the following steps:

[0040] Voice Activity Detection (VAD) will retrieve the latest historical audio frames. Divided into a set of speech frames and pure noise frame set ;

[0041] Calculate the average power of the speech frame: ;

[0042] Calculate the average power of a frame with only noise: ;

[0043] Calculate the signal-to-noise ratio: .

[0044] Specifically, the high-frequency noise quantization index value of the latest historical voice signal can be obtained through the following steps:

[0045] Obtain the high-frequency band energy corresponding to the latest historical voice signal and low-frequency energy The high-frequency energy can be, for example, the energy of a speech signal in the range of 4kHz to 8kHz, and the low-frequency energy can be, for example, the energy of a speech signal in the range of 0.1kHz to 1kHz.

[0046] Calculate the high-frequency noise quantization index value: .

[0047] In a specific example, the process of constructing the real-time channel quality features corresponding to the spectral features of the speech to be enhanced based on the speech quality index values ​​of the latest historical speech signal includes: constructing a two-dimensional feature vector from the signal-to-noise ratio and high-frequency noise quantization index values. The above real-time channel quality characteristics were obtained.

[0048] Specifically, the aforementioned speech quality metrics may also include the Mel-Frequency Cepstral Coefficients (MFCC) vector and / or power variance.

[0049] Step 102: Based on the spectral features of the speech to be enhanced and the corresponding real-time channel quality features, obtain the denoised speech spectral features corresponding to the spectral features of the speech to be enhanced through a denoising sub-model. This step can accurately denoise the original speech signal based on the spectral features of the speech to be enhanced and the real-time channel quality, taking into account both environmental noise and channel noise in the speech signal, avoiding the problems of insufficient or excessive denoising.

[0050] Specifically, the aforementioned denoising sub-model can be a pre-trained encoder-decoder architecture network model based on the corresponding sample speech spectrum features to be enhanced, sample channel quality features, and sample enhanced speech spectrum features. The encoder-decoder architecture network model can be, for example, a Grouped Temporal Convolutional Recurrent Network (GTCRN) or a U-net network architecture model.

[0051] Specifically, the process of obtaining the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced through a denoising sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features can include:

[0052] The speech spectrum features to be enhanced and the corresponding real-time channel quality features are input into the denoising sub-model, so that the denoising sub-model can calculate and output the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features.

[0053] Specifically, the process of obtaining the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced through a denoising sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features may include: after fusing the speech spectrum features to be enhanced and the corresponding real-time channel quality features, outputting the denoised speech spectrum features based on the corresponding fused features through the denoising sub-model; or obtaining the intermediate audio features of the speech spectrum features to be enhanced through a part of the network of the denoising sub-model, fusing the real-time channel quality features and the intermediate audio features, and outputting the denoised speech spectrum features based on the corresponding fused features through the remaining part of the network of the denoising sub-model.

[0054] Step 103: Based on the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced, the speech recognition sub-model obtains the denoised speech recognition text corresponding to the speech spectrum features to be enhanced. This step enables the accurate recognition text corresponding to the original speech signal based on the accurately denoised speech spectrum features.

[0055] Specifically, the aforementioned speech recognition sub-model can be a speech recognition model in the existing technology, such as a "whisper" or "Functional Automatic Speech Recognition (FunASR)" model.

[0056] Specifically, the process of obtaining the denoised speech recognition text corresponding to the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced through the speech recognition sub-model may include: inputting the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced into the speech recognition sub-model, so that the speech recognition sub-model can calculate and output the denoised speech recognition text corresponding to the speech spectrum features to be enhanced based on the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced.

[0057] Step 104: Generate an enhanced speech signal corresponding to the spectral features of the speech to be enhanced based on the denoised speech recognition text. This step can generate a clear and accurate communication speech signal based on the accurately recognized text corresponding to the original speech signal, thereby improving the clarity and accuracy of the communication speech signal.

[0058] Understandably, firstly converting the spectral features of the denoised speech into text allows for the extraction of core content based on the semantic logic of the speech, completely ignoring meaningless signals corresponding to noise. Then, based on the pure text content, any residual noise can be eliminated at its source, resulting in output speech with a purity far exceeding that of traditional denoising methods, thus improving the clarity of communication speech signals. Furthermore, when the original speech is severely distorted by noise, traditional denoising methods may fail to recover key information obscured by the noise, while speech-to-text can correctly identify the semantics of the distorted speech, such as through inference from context, thereby improving the accuracy of communication speech signals.

[0059] Specifically, the process of generating the enhanced speech signal corresponding to the spectral features of the speech to be enhanced based on the denoised speech recognition text can include: directly converting the denoised speech recognition text into a speech signal using a text-to-speech tool to obtain the enhanced speech signal corresponding to the spectral features of the speech to be enhanced.

[0060] Optionally, the process of generating an enhanced speech signal corresponding to the spectral features of the speech to be enhanced based on the denoised speech recognition text includes: translating the denoised speech recognition text from the current language to the target language to obtain the target language text; and generating the target language speech signal based on the target language text to obtain the enhanced speech signal.

[0061] It is understandable that since the two parties in the communication may be using different languages, this embodiment of the invention translates the denoised speech recognition text from the current language to the target language to obtain the target language text, and then generates the target language speech signal based on the target language text, so that voice communication can break down language barriers and improve communication efficiency and accuracy.

[0062] Optional, such as Figure 2 As shown, before step 101, a pre-trained denoising and recognition joint model is obtained through the following steps:

[0063] Step 201: Connect the output of the denoising sub-model to be trained to the pre-trained speech recognition sub-model to obtain the denoising and recognition joint model to be trained.

[0064] Specifically, the output of the denoising sub-model to be trained can be connected to the input of the speech recognition sub-model to be trained to obtain a joint denoising and recognition model to be trained.

[0065] Step 202: Obtain and input the sample speech spectrum features and the corresponding real-time channel quality features into the denoising and recognition joint model to be trained.

[0066] Step 203: Based on the sample speech spectrum features and the corresponding real-time channel quality features, obtain the denoised speech spectrum features corresponding to the sample speech spectrum features through a denoising sub-model.

[0067] Step 204: Based on the denoised speech spectrum features corresponding to the sample speech spectrum features, obtain the denoised speech recognition text corresponding to the sample speech spectrum features through the speech recognition sub-model.

[0068] Step 205: Calculate the denoising loss based on the denoised speech spectrum features corresponding to the sample speech spectrum features, calculate the speech recognition loss based on the denoised speech recognition text corresponding to the sample speech spectrum features, and calculate the total loss of the denoising recognition joint model based on the denoising loss and the speech recognition loss.

[0069] Specifically, the process of calculating the denoising loss based on the denoised speech spectrum features corresponding to the sample speech spectrum features can be performed based on the following loss function:

[0070]

[0071] in, Indicates the noise reduction loss. The clean Mel spectrum corresponding to the spectral features of the sample speech. Let T be the denoised Mel spectrum corresponding to the spectral features of the sample speech, and let T and F represent the number of time frames and frequency intervals of the Mel spectrum, respectively.

[0072] Specifically, the process of calculating the speech recognition loss based on the denoised speech recognition text corresponding to the sample speech spectral features can be performed based on the following loss function:

[0073]

[0074] in, Indicates speech recognition loss, This represents the denoised speech recognition text corresponding to the spectral features of the sample speech. This represents the real text corresponding to the pre-determined spectral features of the sample speech.

[0075] Specifically, the process of calculating the total loss of the joint denoising and recognition model based on denoising loss and speech recognition loss can be performed using the following formula:

[0076]

[0077] in, λ represents the total loss, and λ represents the hyperparameter, which can be set based on empirical data or the results of multiple trials.

[0078] Step 206: Keep the parameters of the speech recognition sub-model unchanged, and adjust the model parameters of the denoising sub-model based on the total loss of the denoising and recognition joint model.

[0079] Understandably, current speech recognition models are relatively mature. Considering feasibility and convenience, this embodiment of the invention obtains a joint denoising and recognition model to be trained by connecting the output of the denoising sub-model to be trained to a pre-trained speech recognition sub-model, and calculates the total loss of the joint denoising and recognition model based on the denoising loss and speech recognition loss. This allows for joint training of the denoising sub-model and the speech recognition sub-model, avoiding the possibility of over-denoising leading to the loss of key features or under-denoising leading to the masking of key features, thereby reducing the performance of the speech recognition model. This facilitates efficient training to obtain a joint denoising and recognition model that can accurately denoise speech signals and improve the accuracy of speech signal recognition.

[0080] The speech signal processing method provided by the embodiments of the present invention will be further described below.

[0081] Optionally, the denoising sub-model includes an encoder, a feature adaptation layer, a bottleneck layer, a decoder, and an output layer.

[0082] Optional, such as Figure 3 As shown, that is Figure 1 Step 102 may include the following steps:

[0083] Step 1021: The real-time channel quality features are reconstructed into the feature tensor shape of the bottleneck layer through the feature adaptation layer to obtain the adapted channel quality features.

[0084] Specifically, the aforementioned feature adaptation layer can be a lightweight multi-layer perceptron (MLP).

[0085] Step 1022: Obtain the bottleneck layer speech features corresponding to the speech spectrum features to be enhanced through the encoder and the bottleneck layer.

[0086] Step 1023: The bottleneck layer output features are obtained by fusing the bottleneck layer speech features and the adapted channel quality features through the bottleneck layer.

[0087] Specifically, the process of fusing the bottleneck layer speech features and the adapted channel quality features through the bottleneck layer can include: splicing the bottleneck layer speech features and the adapted channel quality features together.

[0088] Specifically, the process of fusing the bottleneck layer speech features and the adapted channel quality features through the bottleneck layer can also be fused using other existing fusion methods, such as element-weighted summation.

[0089] Step 1024: Based on the bottleneck layer output features, obtain the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced through the decoder and output layer.

[0090] Specifically, the real-time channel quality features can be reconstructed into the feature tensor shapes of other network layers of the denoising sub-model, and the reconstructed real-time channel quality features can be fused with the speech features of the corresponding layers. Based on the fused features, the remaining network structure of the denoising sub-model can be used to output the denoised speech spectrum features.

[0091] The embodiments of the present invention can integrate channel quality features into speech features at the bottleneck layer with the lowest feature dimension and the highest semantic abstraction. This allows the channel quality features to be uniformly transmitted to each layer along the upsampling process of the decoder, maximizing their guiding role in the feature recovery of the entire network, while taking into account both fusion efficiency and feature consistency.

[0092] The speech signal processing method provided by the embodiments of the present invention will be further described below.

[0093] Optionally, the original speech signal corresponding to the aforementioned speech spectral features to be enhanced is the communication speech signal to be transmitted.

[0094] Optional, such as Figure 4 and Figure 5 As shown, the speech signal processing method provided in this embodiment of the invention may include the following steps:

[0095] Step 401: Acquire the original voice signal through a voice input device.

[0096] Specifically, when the original voice signal is the received communication voice signal, the original voice signal is obtained through the host output device interface of the voice communication device, that is, the speaker interface of the host.

[0097] Optionally, the voice input device includes a single-channel voice input device and a multi-channel voice input device.

[0098] Specifically, the aforementioned single-channel voice input device can be, for example, a VHF standard hand microphone for a VHF communication device, such as... Figure 6 As shown.

[0099] Specifically, the aforementioned multi-channel voice input device can be a VHF smart microphone that uses a multi-channel microphone array.

[0100] Specifically, the process of acquiring the original voice signal through the voice input device includes: acquiring the original voice signal by connecting the first audio input interface of a VHF standard hand microphone; or acquiring the original voice signal by connecting the second audio input interface of a VHF smart hand microphone.

[0101] Specifically, both VHF standard hand microphones and VHF smart hand microphones are equipped with a push-to-talk (PPT) button. After the user presses the corresponding PPT button, the original voice is recorded. The first and second audio input interfaces detect the PPT control signal in real time and acquire the original voice signal after detecting the PPT control signal.

[0102] Step 402: The original speech signal is directly recognized by the speech recognition model, and the original speech signal is determined to be a high-urgency signal based on the direct recognition result and high-urgency keywords.

[0103] Specifically, the aforementioned speech recognition model can be a speech device model in the existing technology.

[0104] Specifically, the aforementioned high-urgency keywords can be pre-set words that reflect the urgency level of the original voice, such as words like "emergency collision avoidance" and "MAYDAY".

[0105] Step 403: When the original speech signal is not a high-urgent signal, acquire and input the speech spectrum features to be enhanced and the corresponding real-time channel quality features into the pre-trained denoising and recognition joint model.

[0106] Optionally, when the original voice signal is a high-urgent signal, the original voice signal can be directly transmitted to the host of the voice communication device through the host input device interface of the voice communication device, so that the host of the voice communication device can directly send the original voice signal.

[0107] Specifically, the aforementioned input device interface can be understood as a microphone interface.

[0108] Specifically, when the original voice signal is the received communication voice signal, the original voice signal can be directly recognized by the voice recognition model. Based on the direct recognition result and high urgency keywords, it can be determined whether the original voice signal is a high urgency signal. If the original voice signal is a high urgency signal, the original voice can be played directly through the speaker.

[0109] Step 404: Based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features, obtain the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced through a denoising sub-model.

[0110] Step 405: Based on the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced, obtain the denoised speech recognition text corresponding to the speech spectrum features to be enhanced through the speech recognition sub-model.

[0111] Step 406: Generate an enhanced speech signal corresponding to the spectral features of the speech to be enhanced based on the denoised speech recognition text.

[0112] Step 407: The enhanced voice signal is transmitted to the host of the voice communication device through the host input device interface of the voice communication device, so as to send the enhanced voice signal through the host of the voice communication device, or to play the enhanced voice signal through the speaker.

[0113] Optional, such as Figure 5 As shown, before step 102, the speech signal processing method provided in this embodiment of the invention further includes: determining the source input device type of the speech spectral features to be enhanced.

[0114] Optional, such as Figure 5 As shown, the denoising sub-models provided in this embodiment of the invention include a single-channel denoising sub-model and a multi-channel denoising sub-model.

[0115] Specifically, the network structure of the single-channel noise reduction sub-model can be 2D U-Net, and the network structure of the multi-channel noise reduction sub-model can be 3D U-Net.

[0116] Optionally, step 102, namely the process of obtaining the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced through a denoising sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features, may include: when the source input device of the speech spectrum features to be enhanced is a single-channel speech input device, obtaining the speech spectrum features to be enhanced through a single-channel denoising sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features; when the source input device of the speech spectrum features to be enhanced is a multi-channel speech input device, obtaining the speech spectrum features to be enhanced through a multi-channel denoising sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features.

[0117] Specifically, such as Figure 5 As shown, when the source interface of the PPT control signal is the first audio interface, the speech spectrum features to be enhanced can be obtained through a single-channel noise reduction sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features; when the source interface of the PPT control signal is not the first audio interface, the speech spectrum features to be enhanced can be obtained through a multi-channel noise reduction sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features.

[0118] Understandably, VHF hosts are mandatory certification devices under the Global Maritime Distress and Safety System (GMDSS). Any attempt to "modify" the internal circuitry or software of a VHF host will immediately invalidate its GMDSS certification. This invention, by using the existing input and output interfaces of the VHF host as the signal input / output links, enables noise reduction and enhancement of the voice signal without modifying any hardware or software of the VHF host. It is plug-and-play, thus perfectly circumventing maritime certification issues.

[0119] In addition, this implementation is compatible with both single-channel VHF standard hand microphones and multi-channel VHF multi-channel hand microphones, which can meet users' different voice enhancement needs and improve user experience.

[0120] Figure 7 This is a structural diagram of a speech signal processing apparatus provided in an embodiment of the present invention. This apparatus is suitable for executing the speech signal processing method provided in the embodiment of the present invention and is integrated into a voice communication device. Figure 7 As shown, the device may specifically include:

[0121] The feature acquisition and input module 701 is used to acquire and input the spectral features of the speech to be enhanced and the corresponding real-time channel quality features into a pre-trained denoising and recognition joint model. The denoising and recognition joint model includes a denoising sub-model and a speech recognition sub-model. This module enables accurate denoising and precise recognition of the corresponding original speech signal based on the spectral features of the speech to be enhanced and the real-time channel quality through the denoising and recognition joint model.

[0122] Optionally, the feature acquisition and input module 701 can be specifically used to acquire the latest historical speech signal transmitted through the transmission channel corresponding to the speech spectral features to be enhanced, and to acquire the speech quality index value of the latest historical speech signal; and

[0123] Based on the speech quality index values ​​of the latest historical speech signals, construct the real-time channel quality features corresponding to the spectral features of the speech to be enhanced.

[0124] Optionally, the denoising sub-model includes an encoder, a feature adaptation layer, a bottleneck layer, a decoder, and an output layer.

[0125] Optionally, the voice signal processing provided in this embodiment of the invention further includes an input and emergency signal judgment module, which is used to acquire the original voice signal through a voice input device, directly recognize the original voice signal through a voice recognition model, and determine whether the original voice signal is a high emergency signal based on the direct recognition result and high urgency keywords.

[0126] Optionally, the voice input device includes a single-channel voice input device and a multi-channel voice input device.

[0127] Optionally, the denoising sub-model includes a single-channel denoising sub-model and a multi-channel denoising sub-model.

[0128] Optionally, the feature acquisition and input module 701 can be specifically used to, when the original speech signal is not a high-urgent signal, start the acquisition and input of the speech spectrum features to be enhanced and the corresponding real-time channel quality features into the pre-trained denoising and recognition joint model;

[0129] Optionally, the voice signal processing device provided in this embodiment of the invention further includes a direct transmission module, which is used to directly transmit the original voice signal to the host of the voice communication device through the host input device interface of the voice communication device when the original voice signal is a high emergency signal, so as to directly send the original voice signal through the host of the voice communication device.

[0130] The denoising module 702 is used to obtain the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced through a denoising sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features. This module can accurately denoise the original speech signal based on the speech spectrum features to be enhanced and the real-time channel quality, taking into account both environmental noise and channel noise in the speech signal, avoiding the problems of insufficient or excessive denoising.

[0131] Optionally, the speech signal processing provided in this embodiment of the invention further includes a source input type acquisition module, which is used to determine the source input device type of the speech spectrum features to be enhanced before obtaining the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced through a denoising sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features.

[0132] Optionally, the denoising module 702 can be specifically used to obtain the speech spectrum features to be enhanced through a single-channel denoising sub-model when the source input device of the speech spectrum features to be enhanced is a single-channel speech input device.

[0133] Optionally, the denoising module 702 can be specifically used to obtain the speech spectrum features to be enhanced through a multi-channel speech input device when the source input device of the speech spectrum features to be enhanced is a multi-channel speech input device.

[0134] Optionally, the denoising module 702 can be specifically used to reconstruct the real-time channel quality features into the feature tensor shape of the bottleneck layer through the feature adaptation layer to obtain the adapted channel quality features.

[0135] Obtain the bottleneck layer speech features corresponding to the speech spectrum features to be enhanced by the encoder and the bottleneck layer;

[0136] The bottleneck layer output features are obtained by fusing the bottleneck layer speech features and the adapted channel quality features; and

[0137] Based on the output features of the bottleneck layer, the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced are obtained through the decoder and output layer.

[0138] Optionally, the original speech signal corresponding to the speech spectral features to be enhanced is the communication speech signal to be transmitted.

[0139] The recognition module 703 is used to obtain the denoised speech recognition text corresponding to the denoised speech spectrum features corresponding to the denoised speech spectrum features to be enhanced through a speech recognition sub-model. This module can obtain the accurate recognition text corresponding to the original speech signal based on the accurately denoised speech spectrum features.

[0140] The enhanced speech signal generation module 704 is used to generate an enhanced speech signal corresponding to the spectral features of the speech to be enhanced based on the denoised speech recognition text. Combined with modules 701 to 703, this module can generate a clear and accurate communication speech signal based on the accurately recognized text corresponding to the original speech signal, thereby improving the clarity and accuracy of the communication speech signal.

[0141] Optionally, the enhanced speech signal generation module 704 can be specifically used to translate the denoised speech recognition text from the current language to the target language to obtain target language text; and to generate a target language speech signal based on the target language text to obtain an enhanced speech signal.

[0142] Optionally, the speech signal processing apparatus provided in this embodiment of the invention further includes: a denoising and recognition joint model acquisition module, used to, before acquiring and inputting the speech spectrum features to be enhanced and the corresponding real-time channel quality features into the pre-trained denoising and recognition joint model, connect the output of the denoising sub-model to be trained to the pre-trained speech recognition sub-model to obtain the denoising and recognition joint model to be trained; acquire and input the sample speech spectrum features and the corresponding real-time channel quality features into the denoising and recognition joint model to be trained; acquire the denoised speech spectrum features corresponding to the sample speech spectrum features through the denoising sub-model based on the sample speech spectrum features and the corresponding real-time channel quality features; acquire the denoised speech recognition text corresponding to the sample speech spectrum features through the speech recognition sub-model based on the denoised speech spectrum features corresponding to the sample speech spectrum features; calculate the denoising loss based on the denoised speech spectrum features corresponding to the sample speech spectrum features, calculate the speech recognition loss based on the denoised speech recognition text corresponding to the sample speech spectrum features, and calculate the total loss of the denoising and recognition joint model based on the denoising loss and the speech recognition loss; and keep the parameters of the speech recognition sub-model unchanged, and adjust the model parameters of the denoising sub-model based on the total loss of the denoising and recognition joint model.

[0143] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional modules is merely an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the functional modules described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0144] This invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the speech signal processing method provided in any of the above embodiments.

[0145] This invention also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the speech signal processing method provided in any of the above embodiments.

[0146] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the speech signal processing methods described in this invention.

[0147] The following is for reference. Figure 8 It shows a schematic diagram of the structure of a computer system 800 suitable for implementing an electronic device according to embodiments of the present invention. Figure 8The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0148] like Figure 8 As shown, the computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 802 or programs loaded from storage section 808 into random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the system 800. The CPU 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0149] The following components are connected to I / O interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to I / O interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 810 as needed so that computer programs read from it can be installed into storage section 808 as needed.

[0150] In particular, according to the embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 809, and / or installed from removable medium 811. When the computer program is executed by central processing unit (CPU) 801, it performs the functions defined above in the system of this invention.

[0151] It should be noted that the computer-readable medium shown in this invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0153] The modules and / or units described in the embodiments of the present invention can be implemented in software or hardware. The described modules and / or units can also be housed in a processor; for example, a processor may be described as including a feature acquisition and input module, a noise reduction module, a recognition module, and an enhanced speech signal generation module. The names of these modules do not necessarily limit the module itself.

[0154] In another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable medium carries one or more programs, which, when executed by the device, cause the device to include: acquiring and inputting the speech spectrum features to be enhanced and the corresponding real-time channel quality features into a pre-trained denoising and recognition joint model, the denoising and recognition joint model including a denoising sub-model and a speech recognition sub-model; acquiring denoised speech spectrum features corresponding to the speech spectrum features to be enhanced through the denoising sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features; acquiring denoised speech recognition text corresponding to the speech spectrum features to be enhanced through the speech recognition sub-model based on the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced; and generating an enhanced speech signal corresponding to the speech spectrum features to be enhanced based on the denoised speech recognition text.

[0155] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A speech signal processing method, applicable to voice communication equipment, characterized in that, include: The speech spectrum features to be enhanced and the corresponding real-time channel quality features are acquired and input into a pre-trained denoising and recognition joint model, which includes a denoising sub-model and a speech recognition sub-model. Based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features, the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced are obtained through a denoising sub-model. Based on the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced, the denoised speech recognition text corresponding to the speech spectrum features to be enhanced is obtained through a speech recognition sub-model. as well as Enhanced speech signals are generated based on the spectral features of the speech to be enhanced, derived from denoised speech recognition text.

2. The speech signal processing method according to claim 1, characterized in that, The acquisition of real-time channel quality features corresponding to the spectral features of the speech to be enhanced includes: Obtain the latest historical speech signal transmitted through the transmission channel corresponding to the speech spectral features to be enhanced, and obtain the speech quality index value of the latest historical speech signal; and Based on the speech quality index values ​​of the latest historical speech signals, construct the real-time channel quality features corresponding to the spectral features of the speech to be enhanced.

3. The speech signal processing method according to claim 2, characterized in that, The denoising sub-model includes an encoder, a feature adaptation layer, a bottleneck layer, a decoder, and an output layer; The process of obtaining the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced through a denoising sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features includes: The adapted channel quality features are obtained by reconstructing the real-time channel quality features into the feature tensor shape of the bottleneck layer through the feature adaptation layer. Obtain the bottleneck layer speech features corresponding to the speech spectrum features to be enhanced by the encoder and the bottleneck layer; The bottleneck layer output features are obtained by fusing the bottleneck layer speech features and the adapted channel quality features; and Based on the output features of the bottleneck layer, the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced are obtained through the decoder and output layer.

4. The speech signal processing method according to claim 1, characterized in that, The original speech signal corresponding to the speech spectrum features to be enhanced is the communication speech signal to be sent; Before acquiring and inputting the spectral features of the speech to be enhanced and the corresponding real-time channel quality features into the pre-trained denoising and recognition joint model, the method further includes: The raw voice signal is acquired through a voice input device; The original speech signal is directly identified by a speech recognition model, and the original speech signal is judged to be a high-urgency signal based on the direct recognition results and high-urgency keywords. The step of acquiring and inputting the spectral features of the speech to be enhanced and the corresponding real-time channel quality features into a pre-trained denoising and recognition joint model includes: When the original speech signal is not a high-urgent signal, the process of acquiring and inputting the speech spectrum features to be enhanced and the corresponding real-time channel quality features into the pre-trained denoising and recognition joint model is initiated. The method further includes: when the original voice signal is a high emergency signal, directly transmitting the original voice signal to the host of the voice communication device through the host input device interface of the voice communication device, so as to directly send the original voice signal through the host of the voice communication device.

5. The speech signal processing method according to claim 4, characterized in that, The voice input device includes a single-channel voice input device and a multi-channel voice input device; The denoising sub-model includes a single-channel denoising sub-model and a multi-channel denoising sub-model; Before obtaining the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced through a denoising sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features, the method further includes: Determine the type of the source input device for the speech spectral features to be enhanced; The process of obtaining the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced through a denoising sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features includes: When the source input device for the speech spectrum features to be enhanced is a single-channel speech input device, the speech spectrum features to be enhanced are obtained through a single-channel noise reduction sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features. When the source input device for the speech spectrum features to be enhanced is a multi-channel speech input device, the speech spectrum features to be enhanced are obtained through a multi-channel noise reduction sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features.

6. The speech signal processing method according to claim 1, characterized in that, The process of generating enhanced speech signals corresponding to the spectral features of the speech to be enhanced based on denoised speech recognition text includes: Translate the denoised speech recognition text from the current language to the target language to obtain the target language text; and The enhanced speech signal is obtained by generating a target language speech signal based on the target language text.

7. The speech signal processing method according to claim 1, characterized in that, Before acquiring and inputting the spectral features of the speech to be enhanced and the corresponding real-time channel quality features into the pre-trained denoising and recognition joint model, the method includes: The output of the denoising sub-model to be trained is connected to the pre-trained speech recognition sub-model to obtain the denoising and recognition joint model to be trained. The sample speech spectrum features and corresponding real-time channel quality features are acquired and input into the denoising and recognition joint model to be trained; Based on the sample speech spectrum features and the corresponding real-time channel quality features, the denoised speech spectrum features corresponding to the sample speech spectrum features are obtained through a denoising sub-model. Based on the denoised speech spectrum features corresponding to the sample speech spectrum features, the denoised speech recognition text corresponding to the sample speech spectrum features is obtained through a speech recognition sub-model. The denoising loss is calculated based on the denoised speech spectrum features corresponding to the sample speech spectrum features; the speech recognition loss is calculated based on the denoised speech recognition text corresponding to the sample speech spectrum features; and the total loss of the joint denoising and recognition model is calculated based on the denoising loss and the speech recognition loss. Keep the parameters of the speech recognition sub-model unchanged, and adjust the model parameters of the denoising sub-model based on the total loss of the denoising and recognition joint model.

8. A voice signal processing device, integrated into a voice communication device, characterized in that, include: The feature acquisition and input module is used to acquire and input the speech spectrum features to be enhanced and the corresponding real-time channel quality features into the pre-trained denoising and recognition joint model, which includes a denoising sub-model and a speech recognition sub-model. The denoising module is used to obtain the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced through a denoising sub-model based on the speech spectrum features to be enhanced and the corresponding real-time channel quality features. The recognition module is used to obtain the denoised speech recognition text corresponding to the denoised speech spectrum features corresponding to the speech spectrum features to be enhanced through the speech recognition sub-model. as well as The enhanced speech signal generation module is used to generate an enhanced speech signal corresponding to the spectral features of the speech to be enhanced based on the denoised speech recognition text.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the speech signal processing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the speech signal processing method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Speech recognition model training method and device, equipment and storage medium

    CN113178192A

  • Underwater acoustic signal noise reduction and recognition combined training method, system and equipment and medium

    CN117765966A

  • Signal noise reduction method and device and electronic equipment

    CN119420424A

  • Air traffic control instruction end-to-end speech recognition method under high noise condition

    CN119559940A

  • Airport control decision support system and method based on semantic recognition of controller instruction

    US20230177969A1