Voice transmission method, device and electronic equipment

By identifying and denoising the voice signal, subtitle information and noise reduction voice information are obtained, and cross-encoding is performed to solve the problem of packet loss during audio signal transmission in real-time communication, and the integrity and comprehensibility of the voice signal are achieved.

CN114999455BActive Publication Date: 2025-05-16BEIJING INTENGINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210447695.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-26
Publication Date
2025-05-16
Estimated Expiration
2042-04-26

AI Technical Summary

Technical Problem

In real-time communication, audio signals are easily disturbed during transmission and lead to packet loss, resulting in poor quality and incoherence of audio signals at the receiving end, which are difficult to understand.

Method used

By performing speech recognition and noise reduction on the voice signal, subtitle information and noise reduction voice information are obtained respectively, and cross-encoded and sent to the target device.

Benefits of technology

Even in the case of packet loss, a large amount of information will not be lost. The target device can decode and restore complete subtitle information and noise-reducing voice information, improving the user's understanding of the voice signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114999455B_ABST
    Figure CN114999455B_ABST
Patent Text Reader

Abstract

The present invention discloses a voice transmission method, device and electronic device, the method comprising: acquiring a voice signal, and performing voice recognition and voice noise reduction on the voice signal respectively, to obtain subtitle information and noise reduction voice information corresponding to the voice signal; cross-coding the subtitle information and the noise reduction voice information to obtain cross-coded information; and sending the cross-coded information to a target device. The technical solution provided by the present invention, through the cross-coding method, is not easy to cause information loss even if the packet is lost, and the target device can decode through the cross-coded information to restore the complete subtitle information and noise reduction voice information. Through the combination of subtitle information and noise reduction voice information, the user's understanding of the voice signal is improved, avoiding the situation where the user does not understand due to the loss of the voice signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communications, and in particular to a voice transmission method, device and electronic equipment. Background Art

[0002] In the field of instant communication such as online conferences, the audio is usually mixed with various on-site noises. In order to improve the call quality, the audio needs to be noise-reduced first. However, when the noise-reduced audio is transmitted to the receiving end through the channel, packet loss may occur due to various interferences, making the audio signal quality at the receiving end poor, incoherent, or even difficult to understand. Therefore, how to improve the integrity of information transmission during instant communication is an urgent problem to be solved. Summary of the invention

[0003] In view of this, the embodiments of the present invention provide a voice transmission method, device and electronic device, thereby improving the integrity of information transmission.

[0004] According to the first aspect, an embodiment of the present invention provides a voice transmission method, which includes: acquiring a voice signal, and performing voice recognition and voice noise reduction on the voice signal respectively to obtain subtitle information and noise reduction voice information corresponding to the voice signal; cross-encoding the subtitle information and the noise reduction voice information to obtain cross-encoded information; and sending the cross-encoded information to a target device.

[0005] Optionally, performing speech recognition and speech noise reduction on the speech signal to obtain subtitle information and noise-reduced speech information corresponding to the speech signal includes: acquiring frequency domain features and high-dimensional features of the speech signal; detecting whether the speech signal contains real speech based on the high-dimensional features; if the speech signal contains real speech, performing noise reduction on the speech signal based on the frequency domain features and high-dimensional features of the speech signal to obtain the noise-reduced speech information; and performing speech recognition based on the high-dimensional features to obtain the subtitle information.

[0006] Optionally, obtaining the frequency domain features and high-dimensional features of the speech signal includes: converting the speech signal from the time domain to the frequency domain to obtain the frequency domain features; filtering the frequency domain features to obtain filtered features; inputting the filtered features into a first encoder, extracting high-dimensional information from the filtered features, and using the extracted high-dimensional information as high-dimensional features.

[0007] Optionally, detecting whether the speech signal contains real speech based on the high-dimensional features includes: inputting the high-dimensional features into a second encoder and obtaining deepened high-dimensional features output by the second encoder, the second encoder being used to extract high-dimensional information from the high-dimensional features; inputting the deepened high-dimensional features into an activity detection layer to output a detection result through the activity detection layer, and determining whether the speech signal contains real speech based on the degree of match between the detection result and a preset label.

[0008] Optionally, the denoising of the speech signal based on the frequency domain features and high dimensional features of the speech signal to obtain the denoised speech information includes: fusing the deepened high dimensional features with the frequency domain features to obtain fused features; inputting the fused features into a noise reduction decoder for decoding to obtain decoded frequency domain features; and converting the decoded frequency domain features from the frequency domain to the time domain to obtain the denoised speech information.

[0009] Optionally, performing speech recognition based on the high-dimensional features to obtain the subtitle information includes: further encoding the high-dimensional features through a third encoder, and then inputting the high-dimensional features into a recognition decoder to decode the further encoded high-dimensional features through the recognition decoder to obtain the subtitle information; wherein the parameters in the recognition decoder, the noise reduction decoder, the first encoder, the second encoder, the third encoder and the activity detection layer are determined by a speech processing model composed of the recognition decoder, the noise reduction decoder, the first encoder, the second encoder, the third encoder and the activity detection layer through joint training.

[0010] Optionally, the method further includes: if the speech signal does not contain real speech, outputting a silent signal as the noise reduction speech information.

[0011] According to the second aspect, an embodiment of the present invention provides a voice transmission device, which includes: a voice processing module, used to obtain a voice signal, and perform voice recognition and voice noise reduction on the voice signal respectively, to obtain subtitle information and noise reduction voice information corresponding to the voice signal; a cross-coding module, used to cross-code the subtitle information and the noise reduction voice information to obtain cross-coded information; and a sending module, used to send the cross-coded information to a target device.

[0012] According to the third aspect, an embodiment of the present invention provides an electronic device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the method described in the first aspect or any optional implementation manner of the first aspect by executing the computer instructions.

[0013] According to a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method described in the first aspect or any optional implementation manner of the first aspect.

[0014] The technical solution provided by this application has the following advantages:

[0015] The technical solution provided by the present application obtains the subtitle information and noise reduction voice information of the voice signal by identifying and reducing the noise of the voice signal. The subtitle information and the noise reduction voice information are then fused and encoded by cross-coding, and then the cross-coded information is sent to the target device. Through the cross-coding method, even if there is packet loss, a large amount of information will not be lost, so that the target device can decode the cross-coded information and restore the complete subtitle information and noise reduction voice information. By combining the subtitle information and the noise reduction voice information, the user's understanding of the voice signal is improved, avoiding the situation where the user does not understand due to the loss of the voice signal.

[0016] In addition, when performing speech recognition and speech denoising on a speech signal, a recognition encoder (including a first encoder and a third encoder connected in series) is used to extract high-dimensional features for the recognition task from the frequency domain features of the speech signal, and then an independent recognition decoding layer is used to decode the high-dimensional features to obtain the recognized subtitle information. In the speech denoising part, a second encoder is used to further encode the high-dimensional features output by the first encoder. On the one hand, the high-dimensional features are deepened and the effect of speech denoising is improved. On the other hand, the first encoder in the recognition task is used as a shared encoder, and secondary encoding is performed on this basis, so that the speech denoising part does not need to set up a more complex encoder, and the encoder needs to set fewer parameters, reducing the training complexity. Then the deepened high-dimensional features and the frequency domain features of the speech signal are fused to make the high-dimensional components in the original features of the speech signal more prominent, so that the denoised speech information generated by the decoder denoising is less lost. In addition, in the voice activity detection part, it is also based on the deepened high-dimensional features extracted by the second encoder to identify whether the current speech signal contains real speech, so as to determine whether the current voice activity is actually included. If the current voice signal does not include voice activity, no noise reduction processing is performed in the non-voice segment but silence is output. On the one hand, the end device enters low power consumption to reduce the amount of calculation, and on the other hand, the problem of residual noise left by the voice noise reduction model leading to poor listening experience is improved. Third, the encoder and decoder involved in the embodiment of the present invention are jointly trained as a whole. The obtained multi-task speech processing model changes the original voice noise reduction and speech recognition series structure, and jointly trains recognition and noise reduction, which effectively improves the current conference transcription system. The phenomenon of high recognition error rate of the speech recognition module caused by the common use of noise reduction speech as input. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present invention in any way. In the accompanying drawings:

[0018] Figure 1 A schematic diagram showing the steps of a voice transmission method in one embodiment of the present invention is shown;

[0019] Figure 2 A schematic flow chart of a voice transmission method in one embodiment of the present invention is shown;

[0020] Figure 3 Another schematic flow chart of a voice transmission method in one embodiment of the present invention is shown;

[0021] Figure 4 A schematic diagram of a service demand determination process of a voice transmission method in one embodiment of the present invention is shown;

[0022] Figure 5 A schematic diagram of the structure of a voice transmission device in one embodiment of the present invention is shown;

[0023] Figure 6 A schematic structural diagram of an electronic device in one embodiment of the present invention is shown. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0025] See also Figure 1 In one embodiment, a voice transmission method comprises the following steps:

[0026] Step S101: Acquire a speech signal, and perform speech recognition and speech noise reduction on the speech signal to obtain subtitle information and noise-reduced speech information corresponding to the speech signal.

[0027] Step S102: cross-encode the subtitle information and the noise reduction speech information to obtain cross-encoded information.

[0028] Step S103: Send the cross-coding information to the target device.

[0029] Specifically, in an embodiment of the present invention, after the current client collects the voice signal through a microphone or other sound receiving device, the voice signal is processed by voice recognition and voice noise reduction. In order to prevent the voice noise reduction from causing the voice to lose key information, in this embodiment, the voice recognition is performed on the original voice signal, rather than the voice signal after noise reduction. The subtitle information in the voice is identified by voice recognition technology, and the voice signal after noise reduction is obtained by voice noise reduction. Afterwards, both the noise reduction voice information and the subtitle voice information are sent to the target device, so that if the voice received by the target device is unclear due to network packet loss, the user can also be assisted in understanding the meaning of the voice based on the subtitle information, thereby improving the integrity of the voice and the comprehensibility of the voice signal.

[0030] Furthermore, in this embodiment, the identified subtitle information and noise reduction voice information are cross-coded, and the cross-coded information is sent to the target device. After the target device receives the cross-coded information, it decodes the cross-coded information to restore the noise reduction voice information and subtitle information. Through the cross-coding technology, the signal is distributed in multiple code words to improve the error correction capability, which can effectively alleviate the information loss caused by network packet loss. Further improve the integrity of voice signal transmission. For example: Assume that the error-free message obtained by splicing the subtitle information and the noise reduction voice information in the normal order is: aaaabbbbccccddddeeeeffffgggg. When cross-coding is not used, due to the sudden error caused by channel interference, the message received by the target device is: aaaabbbbccc____deeeeffffgggg. The code word c is changed by 1 bit, which can be corrected; the code word d is changed by 3 bits and cannot be correctly decoded. This embodiment uses cross-coding, and cross-arranges the subtitle information and the noise reduction voice information according to preset rules. The error-free message is encoded as abcdefgabcdefgabcdefgabcdefg. Assuming that a transmission error occurs, the message received by the target device is abcdefgabcd____bcdefgabcdefg, and the received message is decoded to obtain the message: aa_abbbbccccdddde_eef_ffg_gg; only one bit in each group of codewords is changed, so one bit of error correction code can be correctly decoded. Based on this, when the communication environment is harsh, the audio and subtitle information after noise reduction are cross-coded. On the one hand, the lost information can be restored by guessing, and on the other hand, the probability of the pronunciation and the corresponding subtitle being lost at the same time is greatly reduced. When one is lost, the other can also be used as a supplement to the information, thereby further improving the integrity of the voice signal transmission.

[0031] Specifically, in one embodiment, the above step S101 specifically includes the following steps:

[0032] Step 1: Obtain the frequency domain features and high-dimensional features of the speech signal.

[0033] Step 2: Detect whether the speech signal contains real speech based on high-dimensional features.

[0034] Step 3: If the speech signal contains real speech, the speech signal is denoised based on the frequency domain features and high-dimensional features of the speech signal to obtain denoised speech information.

[0035] Step 4: Perform speech recognition based on high-dimensional features to obtain subtitle information.

[0036] Specifically, Figure 2 As shown, in an embodiment of the present invention, the collected voice signal is first converted into a frequency domain feature. The frequency domain feature has a better feature expression than the time domain signal, so the general signal analysis starts from the frequency domain perspective. The embodiment of the present invention obtains the frequency domain feature of the original voice signal through framing, windowing and Fourier transform. Then the frequency domain feature is processed for high-dimensional information extraction, so as to extract the high-dimensional features in the frequency domain features that can better express the user's real voice. Since the high-dimensional features containing the real user's voice and the high-dimensional features not containing the real voice have a large difference in performance, the features containing the real user's voice have more high-frequency components, and the features containing only white noise have fewer high-frequency components. Therefore, it is identified whether the current voice signal contains the user's real voice information based on the high-dimensional features.

[0037] After that, the activity detection task is used to detect whether the current voice signal contains real voice information. In this embodiment, if the current voice signal does not contain the user's real voice information, no noise reduction and recognition are performed, and the current voice signal is directly muted without subtitle information. This keeps the current client device in a low power consumption state to avoid wasting resources. No noise reduction is performed in the non-voice segment, but silence is output instead, which improves the problem of residual noise left by the voice noise reduction model, resulting in poor listening experience.

[0038] If it is determined that the current voice signal contains the user's real voice information, the voice signal continues to be recognized and denoised. The process of user voice recognition is to directly input the high-dimensional features, or to further encode them and then input them into an independent recognition decoder to decode the high-dimensional features, thereby obtaining the corresponding subtitles. The recognition decoder can be composed of a neural network composed of multiple non-linear layers, including but not limited to FC, CNN, LSTM, Conformer, and Transformer neural networks. The specific setting method can refer to the prior art and will not be repeated here. In this embodiment, the process of denoising the voice signal is to first fuse the frequency domain features and the high-dimensional features, and then input the fused features into an independent noise reduction decoder for noise reduction to obtain noise-reduced voice information. The frequency domain features ensure that there are complete low-dimensional components in the noise-reduced voice information, and the high-dimensional features deepen the high-dimensional components of the voice signal. Noise reduction is performed based on the fused features generated based on the frequency domain features and the high-dimensional features, so as to obtain the noise-reduced user voice without information loss to the greatest extent.

[0039] Specifically, Figure 3 As shown, the above step 1 specifically includes the following steps:

[0040] Step 5: Convert the speech signal from time domain to frequency domain to obtain frequency domain features.

[0041] Step 6: Filter the frequency domain features to obtain the filtering features.

[0042] Step 7: Input the filter features into the first encoder, extract high-dimensional information from the filter features, and use the extracted high-dimensional information as high-dimensional features.

[0043] Specifically, in this embodiment, the frequency domain features of the original speech signal are first obtained by framing, windowing and Fourier transform, and then the frequency domain signal is filtered based on the filter bank technology, so as to further process the frequency domain features and obtain the filter features for neural network training. The filter bank technology used in this embodiment includes but is not limited to Mel filtering and logarithm, which converts the original frequency domain features into frequency domain features that are closer to human ear hearing. Then the filter features are input into the first encoder to extract high-dimensional features. The first encoder is composed of multiple nonlinear layers, including but not limited to FC, CNN, LSTM, Conformer, and Transformer neural network layers.

[0044] Specifically, in this embodiment, the above step 2 specifically includes the following steps:

[0045] Step 8: Input the high-dimensional features into the second encoder, and obtain the deepened high-dimensional features output by the second encoder. The second encoder is used to extract high-dimensional information from the high-dimensional features.

[0046] Step nine: input the deepened high-dimensional features into the activity detection layer to output the detection result through the activity detection layer, and determine whether the speech signal contains real speech based on the matching degree between the detection result and the preset label.

[0047] Specifically, during voice activity detection and voice noise reduction, if only high-dimensional features are input into the built noise reduction decoder or activity detection layer for noise reduction and detection, the final trained model is not deep enough, the activity monitoring is not accurate enough, the background noise is retained too much, and the noise reduction result affects the listening experience. To improve this situation, the prior art usually processes these two tasks separately through a voice activity detection module and a voice noise reduction module that are independent of voice recognition, but setting up an independent voice activity detection module and a voice noise reduction module will greatly increase the complexity of neural network model training. Unlike the prior art, in this solution, an additional second encoder is added, and the high-dimensional features output by the first encoder belonging to the recognition task branch are input into the second encoder, and the deepened high-dimensional features output by the second encoder are obtained. The second encoder is used to further extract high-dimensional information from the high-dimensional features. Based on this, it is only necessary to redeploy a second encoder with fewer layers and fewer parameters to deepen the output of the first encoder, so that the first encoder is shared simultaneously in the three parts of voice recognition, voice noise reduction, and voice activity detection, saving a lot of energy in the deployment of the coding layer and parameter training. Improve training and computing efficiency. Then, whether the speech signal contains real speech is determined based on the deepened high-dimensional features output by the second encoder, thereby improving the accuracy of real speech determination.

[0048] Specifically, in one embodiment, during the process of recognizing speech, the output of the first encoder is also input into the third encoder, and after the high-dimensional features are further encoded by the third encoder, they are input into the recognition decoder, so that the further encoded high-dimensional features are decoded by the recognition decoder to obtain the subtitle information. In this embodiment, the third encoder and the first encoder as a whole can be regarded as a recognition encoder. The first encoder is the first few layers in the recognition encoder, and the third encoder is the last few layers in the recognition encoder. The first encoder is shared in the speech denoising task and the speech activity detection task as a shared layer in the recognition encoder to reduce the training complexity of the speech denoising task and the speech activity detection task. The third encoder is a further encoding layer of the first encoder and is independently applied to the speech recognition task to further deepen the features encoded by the speech recognition task and distinguish them from the features of other tasks.

[0049] Specifically, in one embodiment, based on the above steps eight to nine, the above step three specifically includes the following steps:

[0050] Step 10: Fuse the deepened high-dimensional features with the frequency domain features to obtain fused features.

[0051] Step 11: Input the fused features into the noise reduction decoder for decoding to obtain the decoded frequency domain features.

[0052] Step 13: Convert the decoded frequency domain features from the frequency domain to the time domain to obtain the noise-reduced speech information.

[0053] Specifically, in this embodiment, the deepened high-dimensional features are fused with the original frequency domain features, thereby further improving the depth of the high-dimensional features under the premise that all voice information is complete, and then the fused features are decoded, and the frequency domain features of the clean voice are predicted through decoding. Then, an inverse Fourier transform is performed to obtain noise reduction voice information with better completeness. Before the frequency domain features are fused with the deepened high-dimensional features, the size of the frequency domain features is adjusted through a nonlinear layer so that the size of the frequency domain features matches the size of the deepened high-dimensional features, thereby facilitating the fusion of the two features. The fusion method includes but is not limited to matrix cross multiplication, weighted multiplication of corresponding elements, matrix addition, and the like. The specific setting method of the noise reduction decoder can refer to the prior art and will not be repeated here.

[0054] Specifically, in one embodiment, parameters in the recognition decoder, the noise reduction decoder, the first encoder, the second encoder, the third encoder and the activity detection layer are determined by joint training of a speech processing model composed of the recognition decoder, the noise reduction decoder, the first encoder, the second encoder, the third encoder and the activity detection layer.

[0055] Specifically, in this embodiment, in order to further improve the accuracy of the parameters of each neural network layer, during model training, the recognition decoder, denoising decoder, first encoder, second encoder, third encoder and activity detection layer are regarded as a whole to form a multi-task trained speech processing model, and the training data is input. Then, the difference between the labels of subtitle recognition, denoising labels and labels of speech activity detection and the output results is used respectively, and the parameters in the recognition decoder, denoising decoder, first encoder, second encoder and activity detection layer are adjusted at the same time, thereby further improving the accuracy of the model parameters. The multi-task speech processing model changes the original voice denoising and speech recognition serial structure, and jointly trains recognition and denoising, effectively improving the phenomenon that the speech recognition module has a high recognition error rate in the current conference transcription system due to the common use of denoised speech as input.

[0056] Take a specific training example as an example:

[0057] 1. Prepare training data and labels, and extract speech features of the training data.

[0058] Prepare training data and labels. In an embodiment of the present invention, the training of a multi-task speech processing model requires the following training data and labels: Use clean speech audio and noisy audio to perform data augmentation to obtain a training set that can be used for model pre-training. In an embodiment of the present invention, clean speech audio is used as a label for speech denoising task training, and the word label corresponding to the speech frame is used as a label for speech recognition task training. The speech start and end time points corresponding to the clean speech are binary encoded, and the speech segment is encoded as 1 and the silent segment is encoded as 0 as a label for speech activity detection; speech data augmentation refers to first adding reverberation to the clean audio and the noisy audio to obtain clean reverberation audio and noise reverberation audio, and then according to the specified signal-to-noise ratio range, respectively calculate the clean reverberation audio energy and the noise reverberation audio energy to obtain the signal-to-noise ratio coefficient, and then superimpose the corresponding proportion of noise reverberation audio on the clean reverberation audio to obtain a noisy frequency, and finally generate a noisy frequency with a random amplitude coefficient according to the specified amplitude range, that is, to obtain the augmented speech.

[0059] 2. Extract features of speech data.

[0060] When extracting the features of the training speech data, the frequency domain features of the clean speech and the noisy speech are obtained by using operations such as framing, windowing, and Fourier transform. After taking the absolute value of the obtained frequency domain features, Mel filtering is performed and the logarithm is taken to finally obtain the filter features for training.

[0061] 3. Train a multi-task speech processing model.

[0062] The speech features extracted in step 2 are used to train a multi-task speech processing model to obtain a recognition, noise reduction, and speech activity detection model. Among them, three training tasks that are related to each other are created. The first training task is a speech recognition task, so a speech recognition network is built. In an embodiment of the present invention, a recognition network based on a coding-decoding structure is adopted, which includes multiple layers of nonlinear layers and can be constructed by FC, CNN, LSTM, Conformer, Transformer, etc. The recognition encoder in the recognition network includes a first encoder and a third encoder, wherein the first encoder is used as a shared encoder. During multi-task learning training, the noisy speech features obtained in step 1 are input into the constructed speech recognition network, and characters are used as labels to calculate the cross entropy loss function or sequence loss function between the prediction results and the labels, and back-propagation training is performed. For the convenience of understanding in the following text, the loss function here is expressed as L recognition express.

[0063] The second training task is the speech denoising task, so a speech denoising network is built. In an embodiment of the present invention, the denoising network adopts an encoding-decoding structure, and the decoder includes multiple nonlinear layers, which can be constructed by FC, CNN, LSTM, Conformer, Transformer, etc. The encoder of the denoising model and the first few layers of the speech recognition encoder are shared layers, as shared encoding layers (i.e., the first encoder), and then the latter few layers (i.e., the second encoder) are added, but the denoising decoder is independent. During multi-task learning training, if only the filtering features of the noisy speech are input into the built denoising network, the denoising depth of the denoising model obtained by the final training is not deep enough, and the background noise is retained too much, affecting the sense of hearing. To improve this situation, in this scheme, the filtering features of the noisy speech are first used to extract high-dimensional features using the first encoder, and then the second encoder is used to further extract and deepen the high-dimensional features. After the noisy speech frequency domain features pass through a single layer of nonlinear layers, they are fused with the obtained deepened high-dimensional features, and then input into the denoising model decoder for decoding. Through decoding, the predicted frequency domain features are obtained. The frequency domain features of clean speech are used as labels, and the mean square error loss between them and the predicted frequency domain features inferred by the denoising network is calculated for back propagation training. For the convenience of understanding later, the mean square error loss here is expressed as L denois express.

[0064] The third task is the voice activity detection task, which involves building a voice activity detection network. In an embodiment of the present invention, the voice activity detection network includes multiple nonlinear layers and can be constructed from neural networks such as FC, CNN, LSTM, Conformer, and Transformer. Its encoder layer uses the first encoder and the second encoder of the denoising model, but the activity detection task has an independent activity detection layer for the final task of determining the real voice signal. During multi-task learning training, the noisy speech features obtained in step 1 are input into the established voice activity detection network, and the binary encoding of the speech start and end time points corresponding to the clean speech is used as the label to determine whether the current frame is speech or silence. If it is speech, it is encoded as 1, and if it is silence, it is encoded as 0. Calculate the cross entropy loss function between the prediction result and the label, and perform back propagation training. To facilitate understanding in the following text, the cross entropy loss function here is expressed as L vad express.

[0065] Steps 1, 2, and 3 complete the model structure of multi-task speech processing. When performing multi-task learning training, the noisy speech features obtained in step 1 are used as the input features of the multi-task speech processing model. The speech features of the characters and clean speech, and the binary codes of the speech start and end time points corresponding to the clean speech are used as the labels of the recognition task, the denoising task, and the speech activity detection task, respectively, to calculate L recognition , L denois , Lvad , and then calculate the total loss function of multi-task learning:

[0066] L=αL recognition +βL denoise +γL vad

[0067] Among them, α, β, and γ are all weighting factors with values ​​between 0 and 1, which are used to adjust the influence of different loss functions on model training.

[0068] Using the total loss function, through back propagation and gradient descent algorithms, a multi-task speech processing model for speech recognition, speech denoising, and speech activity detection is trained.

[0069] When performing multi-task learning training, you can also first use the general corpus to train the speech recognition model, then fix the parameters of the shared encoding layer, and use the shared encoding layer as the encoder of the speech denoising and voice activity detection models, respectively using L denois , L vad Train and adjust the remaining parameters of the speech denoising and voice activity detection models, and finally use the total loss function to fine-tune the parameters of the multi-task speech processing model.

[0070] 4. Use the trained multi-task speech processing model to process the collected speech signal to obtain subtitle information and noise-reduced speech information. Figure 4 As shown, in this embodiment, it can also be determined whether the subtitle information needs to be sent based on business needs. If the user does not want to send the subtitle information, the subtitle information is saved in the specified storage area of ​​the client device as a conference transcription; the target device at the receiving end does not need to do any processing. If the user needs to transmit the noise-reduced audio and the recognition results together, a cross-coding method is used to encode the audio and text and transmit them in real time. After the receiving end receives the file, it decodes it according to the specified format, calls the microphone module to play the noise-reduced audio, and converts the recognition results into a subtitle file, or saves the recognition results in the specified storage area of ​​the terminal device as a conference transcription.

[0071] Through the above steps, the technical solution provided by the present application obtains the subtitle information and noise reduction voice information of the voice signal by identifying and reducing the noise of the voice signal. Then, the subtitle information and the noise reduction voice information are fused and encoded by cross-coding, and then the cross-coded information is sent to the target device. Through the cross-coding method, even if there is packet loss, a large amount of information will not be lost, so that the target device can decode the cross-coded information and restore the complete subtitle information and noise reduction voice information. Through the combination of subtitle information and noise reduction voice information, the user's understanding of the voice signal is improved, avoiding the situation where the user does not understand due to the loss of the voice signal.

[0072] In addition, when performing speech recognition and speech denoising on a speech signal, a recognition encoder (including a first encoder and a third encoder connected in series) is first used to extract high-dimensional features for the recognition task from the frequency domain features of the speech signal, and then an independent recognition decoding layer is used to decode the high-dimensional features to obtain the recognized subtitle information. In the speech denoising part, a second encoder is used to further encode the high-dimensional features output by the first encoder. On the one hand, the high-dimensional features are deepened and the effect of speech denoising is improved. On the other hand, the first encoder in the recognition task is used as a shared encoder, and secondary encoding is performed on this basis, so that the speech denoising part does not need to set up a more complex encoder, and the encoder needs to set fewer parameters, reducing the training complexity. Then the deepened high-dimensional features and the frequency domain features of the speech signal are fused to make the high-dimensional components in the original features of the speech signal more prominent, so that the denoised speech information generated by the decoder denoising is less lost. In addition, in the voice activity detection part, it is also based on the extracted deepened high-dimensional features to identify whether the current speech signal contains real speech, thereby determining whether the current voice activity is actually included. If the current voice signal does not include voice activity, no noise reduction processing is performed in the non-voice segment but silence is output. On the one hand, the end device enters low power consumption to reduce the amount of calculation, and on the other hand, the problem of residual noise left by the voice noise reduction model leading to poor listening experience is improved. Third, the encoder and decoder involved in the embodiment of the present invention are jointly trained as a whole. The obtained multi-task speech processing model changes the original voice noise reduction and speech recognition series structure, and jointly trains recognition and noise reduction, which effectively improves the current conference transcription system. The phenomenon of high recognition error rate of the speech recognition module caused by the common use of noise reduction speech as input.

[0073] like Figure 5 As shown, this embodiment also provides a voice transmission device, which includes:

[0074] The speech processing module 101 is used to obtain a speech signal, and perform speech recognition and speech noise reduction on the speech signal to obtain subtitle information and noise-reduced speech information corresponding to the speech signal. For details, please refer to the relevant description of step S101 in the above method embodiment, which will not be repeated here.

[0075] The cross-coding module 102 is used to cross-code the subtitle information and the noise reduction speech information to obtain cross-coded information. For details, please refer to the relevant description of step S102 in the above method embodiment, which will not be repeated here.

[0076] The sending module 103 is used to send the cross-coding information to the target device. For details, please refer to the relevant description of step S103 in the above method embodiment, which will not be repeated here.

[0077] The voice transmission device provided in the embodiment of the present invention is used to execute the voice transmission method provided in the above embodiment. Its implementation method and principle are the same. For details, please refer to the relevant description of the above method embodiment, which will not be repeated here.

[0078] Through the collaborative cooperation of the above-mentioned components, the technical solution provided by this application obtains the subtitle information and noise reduction voice information of the voice signal by identifying and reducing the noise of the voice signal. Then, the subtitle information and the noise reduction voice information are fused and encoded by cross-coding, and then the cross-coded information is sent to the target device. Through the cross-coding method, even if there is packet loss, a large amount of information will not be lost, so that the target device can decode the cross-coded information and restore the complete subtitle information and noise reduction voice information. Through the combination of subtitle information and noise reduction voice information, the user's understanding of the voice signal is improved, avoiding the situation where the user does not understand due to the loss of the voice signal.

[0079] In addition, when performing speech recognition and speech denoising on a speech signal, the high-dimensional features for the recognition task are first extracted from the frequency domain features of the speech signal by the recognition encoder (including the first encoder and the third encoder), and then the high-dimensional features are decoded using an independent recognition decoding layer to obtain the recognized subtitle information. In the speech denoising part, the high-dimensional features output by the first encoder are further encoded by the second encoder. On the one hand, the high-dimensional features are deepened and the effect of speech denoising is improved. On the other hand, the first encoder in the recognition task is used as a shared encoder, and secondary encoding is performed on this basis, so that the speech denoising part does not need to set up a more complex encoder, and the encoder needs to set fewer parameters, reducing the training complexity. Then the deepened high-dimensional features and the frequency domain features of the speech signal are fused to make the high-dimensional components in the original features of the speech signal more prominent, so that the denoised speech information generated by the decoder denoising is less lost. In addition, in the voice activity detection part, it is also based on the extracted deepened high-dimensional features to identify whether the current speech signal contains real speech, thereby determining whether the current voice activity is actually included. If the current voice signal does not include voice activity, no noise reduction processing is performed in the non-voice segment but silence is output. On the one hand, the end device enters low power consumption to reduce the amount of calculation, and on the other hand, the problem of residual noise left by the voice noise reduction model leading to poor listening experience is improved. Third, the encoder and decoder involved in the embodiment of the present invention are jointly trained as a whole. The obtained multi-task speech processing model changes the original voice noise reduction and speech recognition series structure, and jointly trains recognition and noise reduction, which effectively improves the current conference transcription system. The phenomenon of high recognition error rate of the speech recognition module caused by the common use of noise reduction speech as input.

[0080] Figure 6An electronic device according to an embodiment of the present invention is shown, the device includes a processor 901 and a memory 902, which can be connected via a bus or other means. Figure 6 The example of connecting through bus is taken in the following.

[0081] The processor 901 may be a central processing unit (CPU). The processor 901 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.

[0082] The memory 902 is a non-transitory computer-readable storage medium that can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as program instructions / modules corresponding to the methods in the above method embodiments. The processor 901 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory 902, that is, implementing the methods in the above method embodiments.

[0083] The memory 902 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required by at least one function; the data storage area may store data created by the processor 901, etc. In addition, the memory 902 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 902 may optionally include a memory remotely arranged relative to the processor 901, and these remote memories may be connected to the processor 901 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0084] One or more modules are stored in the memory 902 , and when executed by the processor 901 , the method in the above method embodiment is executed.

[0085] The specific details of the above electronic device can be understood by referring to the corresponding descriptions and effects in the above method embodiments, and will not be repeated here.

[0086] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the implemented program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above-mentioned types of memory.

[0087] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A voice transmission method, characterized in that: The method comprises: Acquire speech signals; Convert the speech signal from time domain to frequency domain to obtain frequency domain features; Filtering the frequency domain features to obtain filtering features; Inputting the filter feature into the first encoder, extracting high-dimensional information from the filter feature, and using the extracted high-dimensional information as the high-dimensional feature; Detect whether the speech signal contains real speech based on high-dimensional features; If the speech signal contains real speech, the speech signal is denoised based on the frequency domain features and high-dimensional features of the speech signal to obtain denoised speech information; Perform speech recognition based on high-dimensional features to obtain subtitle information; Cross-coding the subtitle information and the noise reduction speech information to obtain cross-coded information; The cross-coded information is sent to a target device.

2. The method according to claim 1, characterized in that The detecting whether the speech signal contains real speech based on the high-dimensional feature comprises: Inputting the high-dimensional features into a second encoder, and obtaining a deepened high-dimensional feature output by the second encoder, wherein the second encoder is used to extract high-dimensional information from the high-dimensional features; The deepened high-dimensional features are input into an activity detection layer to output a detection result through the activity detection layer, and whether the speech signal contains real speech is determined based on the matching degree between the detection result and a preset label.

3. The method according to claim 2, characterized in that The step of performing noise reduction on the speech signal based on the frequency domain features and high dimensional features of the speech signal to obtain the noise-reduced speech information includes: Fusing the deepened high-dimensional feature with the frequency domain feature to obtain a fused feature; Inputting the fused features into a noise reduction decoder for decoding to obtain decoded frequency domain features; The decoded frequency domain features are converted from the frequency domain to the time domain to obtain the noise reduction speech information.

4. The method according to claim 3, characterized in that The performing speech recognition based on the high-dimensional features to obtain the subtitle information includes: After further encoding the high-dimensional features through a third encoder, the high-dimensional features are input into a recognition decoder, so that the further encoded high-dimensional features are decoded by the recognition decoder to obtain the subtitle information; Among them, the parameters in the recognition decoder, the noise reduction decoder, the first encoder, the second encoder, the third encoder and the activity detection layer are determined by joint training of a speech processing model composed of the recognition decoder, the noise reduction decoder, the first encoder, the second encoder, the third encoder and the activity detection layer.

5. The method according to claim 1, characterized in that The method further comprises: If the speech signal does not contain real speech, a silent signal is output as the noise reduction speech information.

6. A voice transmission device, characterized in that: The device comprises: The speech processing module is used to obtain a speech signal; convert the speech signal from the time domain to the frequency domain to obtain frequency domain features; filter the frequency domain features to obtain filter features; input the filter features into a first encoder, extract high-dimensional information from the filter features, and use the extracted high-dimensional information as high-dimensional features; detect whether the speech signal contains real speech based on the high-dimensional features; if the speech signal contains real speech, denoise the speech signal based on the frequency domain features and high-dimensional features of the speech signal to obtain denoised speech information; perform speech recognition based on the high-dimensional features to obtain subtitle information; A cross-coding module, used for cross-coding the subtitle information and the noise reduction speech information to obtain cross-coded information; The sending module is used to send the cross-coding information to the target device.

7. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 5 by executing the computer instructions.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Combined file format for digital multimedia broadcasting (DMB) content, method and apparatus for handling dmb content of this format

    CN101611630A

  • Voice recognition method, device and system and storage medium

    CN110648655A