Audio processing method and device, audio model training method and device, electronic device, and computer-readable storage medium

By training a model to use direct sound and early reflection audio as target audio and suppressing late reflection audio, the problem of speech reverberation in cloud meetings and cloud classrooms is solved, improving the naturalness and clarity of the audio.

CN113963686BActive Publication Date: 2025-11-11ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110932183.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-13
Publication Date
2025-11-11
Estimated Expiration
2041-08-13

AI Technical Summary

Technical Problem

In cloud meetings and cloud classrooms, the audio emitted by users is reverberated due to reflections in the room, which reduces the intelligibility of the speech and affects the listening experience.

Method used

The training model uses direct sound and early reflection audio as target audio to generate reverberant training audio, suppresses late reflection audio, and employs a deep neural network model for audio processing.

Benefits of technology

It effectively protects the original target audio, ensuring the naturalness and clarity of the processed audio and enhancing the listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113963686B_ABST
    Figure CN113963686B_ABST
Patent Text Reader

Abstract

The application discloses an audio processing method and device, an audio model training method and device, an electronic device and a computer readable storage medium. The method comprises: obtaining to-be-processed audio; extracting a feature vector of the to-be-processed audio; and calculating the feature vector using a predetermined model obtained by training reverberation training audio generated based on predetermined sampling audio to obtain processed audio. In the embodiment of the application, the model is trained by using audio generated by direct sound and early reflection audio as target audio for training in model training, and the model thus trained is used to process mixed audio in actual use, so that the original target audio can be effectively protected and the naturalness and clarity of the processed audio can be ensured by selecting early reflection sound instead of direct sound as the model training and recovery target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to an audio processing method and apparatus, an audio model training method and apparatus, an electronic device, and a computer-readable storage medium. Background Technology

[0002] With the development of cloud technology, more and more users are choosing to use cloud services for online meetings, discussions, and classroom teaching. However, these cloud meetings and online classes require capturing the user's voice near their own terminal and transmitting it over the internet to the participants for playback. Since users typically conduct these meetings or classes in a room, it's inevitable that the audio captured by the user's terminal will be a mixture of direct audio from the user's voice, reflected audio from objects like walls (which have reflected once or twice), and late-stage reverberation after multiple reflections. This mixed audio severely reduces the intelligibility of the user's voice, significantly impacting the listening experience for other users. Summary of the Invention

[0003] This application provides an audio processing method and apparatus, an audio model training method and apparatus, an electronic device, and a computer-readable storage medium to address the shortcomings of existing technologies in producing unnatural reverberation audio processing effects.

[0004] To achieve the above objectives, embodiments of this application provide an audio processing method, including:

[0005] Obtain the audio to be processed;

[0006] Extract the feature vector of the audio to be processed;

[0007] The feature vectors are calculated using a predetermined model trained on reverberant training audio generated from predetermined sampled audio to obtain processed audio.

[0008] This application also provides an audio model training method, including:

[0009] Use a predetermined algorithm to generate reverberant training audio for a predetermined sample of audio.

[0010] Generate training target audio based on at least a portion of the predetermined sampled audio and the reverberated training audio;

[0011] The predetermined model is trained using the reverberant training audio as input and the training target audio as validation data.

[0012] This application also provides a method for processing conference audio, including:

[0013] The audio of speeches sent by the participating terminals is acquired through an audio acquisition device.

[0014] Extract the feature vector of the spoken audio;

[0015] The feature vector is calculated using a predetermined model trained on reverberant training audio generated from predetermined sampled audio to obtain processed audio;

[0016] The processed audio is then sent to other participating terminals in the conference.

[0017] This application also provides a classroom audio processing method, including:

[0018] The teaching audio sent by the teacher during the lecture was obtained by audio acquisition devices installed in the classroom;

[0019] Extract the feature vector of the teaching audio;

[0020] The feature vector is calculated using a predetermined model trained on reverberant training audio generated from predetermined sampled audio to obtain processed audio;

[0021] The processed audio is sent over the network to terminals that listen to classroom lectures via the network.

[0022] This application also provides an audio processing apparatus, including:

[0023] The acquisition module is used to acquire the audio to be processed;

[0024] The extraction module is used to extract the feature vector of the audio to be processed;

[0025] The processing module is used to calculate the feature vector using a predetermined model trained on reverberant training audio generated based on predetermined sampled audio to obtain processed audio.

[0026] This application also provides an audio model training device, including:

[0027] The first generation module is used to generate reverberant training audio for a predetermined sample of audio using a predetermined algorithm;

[0028] The second generation module is used to generate training target audio based on at least a portion of the predetermined sampled audio and the reverberant training audio.

[0029] A training module is used to train a predetermined model using the reverberant training audio as input and the training target audio as validation data.

[0030] This application also provides an electronic device, including:

[0031] Memory, used to store programs;

[0032] A processor is configured to run the program stored in the memory, wherein the program executes the audio processing method or audio model training method provided in the embodiments of this application.

[0033] This application also provides a computer-readable storage medium storing a computer program executable by a processor, wherein the program, when executed by the processor, implements the audio processing method or audio model training method provided in this application.

[0034] This application also provides a computer program product, comprising: a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to perform steps in the audio processing method or audio model training method provided in this application.

[0035] The audio processing method and apparatus, audio model training method and apparatus, electronic device, and computer-readable storage medium provided in this application training method and apparatus use audio generated by direct sound and early reflection sound as the target audio for training the model, and use the model trained in this way to process mixed audio in actual use. Therefore, by selecting early reflection sound instead of direct sound as the target for model training and recovery, the original target audio can be effectively protected, and the naturalness and clarity of the processed audio can be guaranteed.

[0036] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0037] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0038] Figure 1 This is a schematic diagram illustrating an application scenario of the audio processing solution provided in the embodiments of this application;

[0039] Figure 2A flowchart of an embodiment of the audio processing method provided in this application;

[0040] Figure 3 A flowchart of an embodiment of the audio processing method provided in this application;

[0041] Figure 4a A schematic diagram of the structure of an embodiment of the audio processing apparatus provided in this application;

[0042] Figure 4b A schematic diagram of an embodiment of the audio model training device provided in this application;

[0043] Figure 5 A schematic diagram of the structure of an embodiment of the electronic device provided in this application. Detailed Implementation

[0044] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0045] Example 1

[0046] The solution provided in this application can be applied to any system with audio data processing capabilities, such as a server system including chips with audio processing functions, etc. Figure 1 This is a schematic diagram illustrating an application scenario of the audio processing solution provided in the embodiments of this application. Figure 1 The scenario shown is merely one example of the applicable technical solutions of this application.

[0047] With the development of cloud technology, more and more users are choosing to use cloud services for online meetings, discussions, and classroom teaching. However, these cloud meetings and online classes require capturing the user's voice near their own terminal and transmitting it over the internet to the participants for playback. Since users typically conduct these meetings or classes in a room, the audio captured by the user's terminal inevitably contains a mixture of direct audio (the user's voice directly transmitted to the capture device), reflected audio (the user's voice after one or two reflections from objects like walls), and late-stage reverberation after multiple reflections. This mixed audio severely reduces the intelligibility of the user's voice, significantly impacting the listening experience for other users. Therefore, a reverberation suppression technique for the captured audio is needed.

[0048] For example, in such Figure 1 In the audio acquisition scenario illustrated, such as a classroom, when a teacher speaks from the podium on the left side of the classroom, their voice can propagate to the right and be captured by the audio acquisition device on the right. Reverberation occurs during this acquisition process. In the field of audio processing technology, reverberation refers to the process of sound source attenuating in space after it stops emitting sound; the continued existence of sound within space is called reverberation. Therefore, in scenarios like... Figure 1 In the classroom sound acquisition scenario shown, the audio signal ultimately acquired by the acquisition device can be considered to consist of three parts: Audio 1, the direct speech of the teacher, which arrives directly at the acquisition device without any reflection; Audio 2, the early reflected audio of the teacher's speech, which arrives at the acquisition device after one or two reflections, for example, through the classroom walls; and Audio 3, the late reflected audio of the teacher's speech, which arrives at the acquisition device after three or more reflections, for example, through three or more side walls of the classroom. Therefore, the final acquired audio signal is actually a mixture of Audio 1 (direct speech), Audio 2 (early reflected audio), and Audio 3 (late reflected audio). Typically, only the late reflected audio significantly affects the clarity of the speech in the acquired mixed audio, while the early reflected audio can actually enhance the energy of the speech when the direct sound intensity is weak. Therefore, in the field of mixed audio processing, it is usually necessary to suppress the late reflected audio in the acquired mixed audio.

[0049] To address this, existing technologies have proposed signal processing-based solutions. For example, in single-channel audio pickup scenarios, Wiener gain is calculated by estimating the energy of late reverberation in mixed audio using a pre-assumed reverberation statistical model. However, this approach is not ideal for suppressing reverberation. In multi-channel audio pickup scenarios, the Weighted Prediction Error (WPE) method has been proposed. This method first estimates the reverberation tail of the audio signal and then subtracts this estimated reverberation tail from the audio signal to obtain a maximum likelihood estimate of the weak reverberation signal. However, this method does not significantly improve the listening experience when microphone data is limited.

[0050] In addition, a reverberation suppression algorithm based on a deep learning model has been proposed in the existing technology. This algorithm uses the direct audio in the mixed audio as the processing target during training. However, since such a model is not smooth enough in time for the degree of reverberation suppression, the processed audio will have obvious energy fluctuations and sound unnatural.

[0051] Therefore, in this embodiment of the application, when training the model, for example, speech audio with a sampling frequency of 16k / 48k can be prepared as sampled speech data, and the positions of the sound source and the acquisition device can be randomly set in a simulated randomly generated room. For example, it can be as follows: Figure 1 The diagram shows the sound source positioned in the center of the left side of the room, and the acquisition device positioned in the center of the right side. Room impact response (RIP) data can then be generated using, for example, the IMAGE method proposed by Allen and Berkley in 1979. This RIP data can be used to describe the reverberation characteristics of the simulated room. The model can then be configured and initialized according to actual needs. For example, in embodiments of this application, the model can use a linear transformation model, a DFSMN (Deep Feedforward Sequential Memory Network) model, and a deep neural network model with nonlinear activation functions. After initializing the model parameters, mixed audio can be obtained by convolutional operations using sampled speech data and RIP data, and the early reflection audio from the sampled speech and RIP data, i.e. Figure 1 Audio 2 in the dataset is used to obtain the training target audio through convolution operations. In this embodiment, the early reflection audio can be selected from the RIP data at 50ms or 100ms after the direct sound.

[0052] Then, a Short Time Fourier Transform (STFT) algorithm can be used to extract energy feature vectors from the mixed audio obtained by convolution operations based on sampled speech data and RIP data. In this embodiment, the extracted energy features can be filter bank feature vectors. Therefore, the extracted feature vectors can be input into the model and processed through the model's forward computation to obtain, for example, time-frequency masking data. Thus, this time-frequency masking data can then be combined with the previously obtained audio based on sampled speech and RIP data. Figure 1 The loss function for the target audio obtained by the audio-2 convolution shown can be calculated, for example, by calculating the mean square error between the time-frequency masking data output by the model and the masking data of the target audio. Then, the gradient backpropagation algorithm can be used to adjust the model parameters based on the calculation result and recalculate the model until the loss function of the current round no longer decreases significantly compared with the loss function calculated in the previous round. This indicates that the model has converged, that is, the training of the model has been completed.

[0053] Subsequently, in this embodiment, the trained model can be used to perform calculations on the acquired mixed audio. Similar to the training process described above, filter bank feature vectors can be extracted from the mixed audio acquired by the acquisition device, and the extracted feature vectors can be input into the trained model for processing. For example, the time-domain signal with suppressed reverberation can be obtained by multiplying the masking data output by the model with the time-frequency spectrum of the acquired mixed audio, followed by an inverse Fourier transform.

[0054] Therefore, the audio processing scheme provided in this application uses audio generated by direct sound and early reflection audio as the target audio for training the model during model training, and uses the model trained in this way to process the mixed audio in actual use. Therefore, by selecting early reflection audio instead of direct sound as the target for model training and recovery, the original target audio can be effectively protected, ensuring the naturalness and clarity of the processed audio.

[0055] The above embodiments illustrate the technical principles and exemplary application framework of the embodiments of this application. The specific technical solutions of the embodiments of this application will be further described in detail below through multiple embodiments.

[0056] Example 2

[0057] Figure 2 This is a flowchart of an embodiment of the audio processing method provided in this application. The subject executing this method can be various terminals or server devices with audio processing capabilities, or it can be a device or chip integrated into these devices. Figure 2 As shown, the audio processing method may include the following steps:

[0058] S201, Obtain the audio to be processed.

[0059] In this embodiment, the speech emitted by the speech source can be acquired in the same space as the speech source. In other words, as the speech propagates in the acquisition space, a portion of the speech that propagates along the line connecting the speech source and the acquisition device can be directly acquired by the acquisition device, for example... Figure 1 The audio shown is directly transmitted to audio 1, while another portion of the speech can propagate in other directions and via, for example, audio 1. Figure 1 The wall shown in the image reflects light only once or twice before reaching the data acquisition device, for example, as... Figure 1 The early reflected audio 2 shown in the diagram has a final portion that undergoes multiple reflections before reaching the acquisition device, for example, as... Figure 1 The late reflection audio 3 shown is thus obtained in step S201. Therefore, the final audio to be processed consists of such direct audio 1, early reflection audio 2, and late reflection audio 3.

[0060] S202, extract the feature vector of the audio to be processed.

[0061] After obtaining the audio to be processed, which is a mixture of direct audio, early reflection audio, and late reflection audio, in step S201, feature vectors can be extracted from the audio to be processed in step S202. For example, filter bank feature vectors can be extracted.

[0062] S203, use a predetermined model to calculate the feature vector to obtain the processed audio.

[0063] In step S203, the feature vector extracted in step S202 can be input into a predetermined model for processing. For example, such a model can be a linear transformation model, a DFSMN (Deep Feedforward Sequential Memory Network) model, or a deep neural network model with a nonlinear activation function. For instance, in this embodiment, in step S203, a predetermined model trained on reverberation training audio generated from predetermined sampled audio can be used to calculate the feature vector extracted in step S202 to obtain the processed audio.

[0064] For example, the predetermined model in the embodiments of this application may be trained by the following training method according to the embodiments of this application.

[0065] S204, using a predetermined algorithm to generate reverberant training audio for predetermined sampled audio.

[0066] In this embodiment of the application, when training the model, for example, speech audio with a sampling frequency of 16k / 48k can be prepared as sampled speech data, and the positions of the sound source and the acquisition device can be randomly set in a simulated randomly generated room. For example, it can be as follows: Figure 1 The diagram shows the sound source positioned in the center of the left side of the room, and the acquisition device positioned in the center of the right side. Various pre-defined algorithms can then be used to generate reverberation training audio from the sampled audio. For example, the IMAGE method can be used to generate Room Impulse Response (RIP) data, which can be used to describe the reverberation characteristics of the simulated room. The model can then be configured and initialized according to actual needs. For example, in embodiments of this application, the model can use a linear transformation model, a DFSMN (Deep Feedforward Sequential Memory Network) model, and a deep neural network model with nonlinear activation functions. After initializing the model parameters, the sampled speech data and RIP data can be used to obtain a mixed sound through convolution operations.

[0067] S205, generate training target audio based on at least a portion of the predetermined sampled audio and the reverberated training audio.

[0068] In this embodiment, the number of sampled voices used in step S204 and the early reflection audio from the RIP data obtained in step S204 can be used, for example... Figure 1 The audio 2 in the data is used to obtain the training target audio through convolution operation. In this embodiment of the application, the early reflection audio used in step S205 can be selected from the early reflection audio 50ms or 100ms after the direct sound in the RIP data.

[0069] S206, Use reverberant training audio as input and training target audio as validation data to train a predetermined model.

[0070] In step S206, the reverberant training audio obtained in step S205 can be used as input to the model, and the training target audio obtained in step S205 can be used as validation data to train the model used in step S203.

[0071] For example, a Short Time Fourier Transform (STFT) algorithm can be used to extract energy feature vectors from the mixed audio obtained by convolution operations based on sampled speech data and RIP data. In this embodiment, the extracted energy features can be filter bank feature vectors. Therefore, the extracted feature vectors can be input into the model and processed through the model's forward computation to obtain, for example, time-frequency masking data. Thus, this time-frequency masking data can then be combined with the previously obtained audio based on sampled speech and RIP data. Figure 1 The loss function for the target audio obtained by the audio-2 convolution shown can be calculated, for example, by calculating the mean square error between the time-frequency masking data output by the model and the masking data of the target audio. Then, the gradient backpropagation algorithm can be used to adjust the model parameters based on the calculation result and recalculate the model until the loss function of the current round no longer decreases significantly compared with the loss function calculated in the previous round. This indicates that the model has converged, that is, the training of the model has been completed.

[0072] Therefore, the audio processing scheme provided in this application uses audio generated by direct sound and early reflection audio as the target audio for training the model during model training, and uses the model trained in this way to process the mixed audio in actual use. Therefore, by selecting early reflection audio instead of direct sound as the target for model training and recovery, the original target audio can be effectively protected, ensuring the naturalness and clarity of the processed audio.

[0073] Example 3

[0074] Figure 3 This is a flowchart of an embodiment of the audio processing method provided in this application. The subject executing this method can be various terminals or server devices with audio processing capabilities, or it can be a device or chip integrated into these devices. Figure 3 As shown, the audio processing method may include the following steps:

[0075] S301, Obtain the audio to be processed.

[0076] In this embodiment, the speech emitted by the speech source can be acquired in the same space as the speech source. In other words, as the speech propagates in the acquisition space, a portion of the speech that propagates along the line connecting the speech source and the acquisition device can be directly acquired by the acquisition device, for example... Figure 1 The audio shown is directly transmitted to audio 1, while another portion of the speech can propagate in other directions and via, for example, audio 1. Figure 1 The wall shown in the image reflects light only once or twice before reaching the data acquisition device, for example, as... Figure 1 The early reflected audio 2 shown in the diagram has a final portion that undergoes multiple reflections before reaching the acquisition device, for example, as... Figure 1 The late reflection audio 3 shown is thus obtained in step S301. Therefore, the final audio to be processed consists of such direct audio 1, early reflection audio 2, and late reflection audio 3.

[0077] S302, extract the feature vector of the audio to be processed.

[0078] After obtaining the audio to be processed, which is a mixture of direct audio, early reflection audio, and late reflection audio, in step S301, feature vectors can be extracted from the audio to be processed in step S302. For example, filter bank feature vectors can be extracted.

[0079] S303 uses a predetermined model to perform forward computation on the feature vectors to obtain masking data.

[0080] S304, multiply the masking data with the time spectrum of the audio to be processed and perform an inverse Fourier transform to obtain the processed audio.

[0081] In step S303, the feature vector extracted in step S302 can be input into a predetermined model to perform forward calculation on the feature vector and obtain, for example, time-frequency masking data. Then, in step S304, the masking data obtained in this way can be multiplied with the time spectrum of the audio to be processed obtained in step S301, and then an inverse Fourier transform can be performed to obtain the processed time-domain signal. This time-domain signal can be used as the processed audio for speech recognition or playback and other processing.

[0082] Specifically, in the embodiments of this application, such a model can be a linear transformation model, a DFSMN (Deep Feedforward Sequential Memory Network) model, or a deep neural network model with a nonlinear activation function. For example, in the embodiments of this application, in step S303, a predetermined model obtained by training based on reverberation training audio generated from predetermined sampled audio can be used to perform forward calculation on the feature vector extracted in step S302 to obtain masking data.

[0083] For example, the predetermined model in the embodiments of this application may be trained by the following training method according to the embodiments of this application.

[0084] S305 uses pre-sampled audio and pre-sampled room impact response data to perform convolution calculations to obtain reverberation training audio.

[0085] In this embodiment of the application, when training the model, for example, speech audio with a sampling frequency of 16k / 48k can be prepared as sampled speech data, and the positions of the sound source and the acquisition device can be randomly set in a simulated randomly generated room. For example, it can be as follows: Figure 1 The diagram shows the sound source positioned in the center of the left side of the room, and the acquisition device positioned in the center of the right side. Various pre-defined algorithms can then be used to generate reverberant training audio from the sampled audio. For example, the IMAGE method can be used to generate Room Impulse Response (RIP) data, which can be used to describe the reverberation characteristics of the simulated room. The model can then be configured and initialized according to actual needs. For example, in this embodiment, the model can use a linear transformation model, a DFSMN (Deep Feedforward Sequential Memory Network) model, or a deep neural network model with nonlinear activation functions. After initializing the model parameters, the sampled speech data and RIP data can be used to obtain a mixed sound through convolution operations.

[0086] S306 uses pre-sampled audio and early reflection audio to perform convolution calculations to obtain the training target audio.

[0087] In this embodiment, the sampled speech data used in step S305 and the early reflection audio from the RIP data obtained in step S305 can be used, for example... Figure 1The audio 2 in the dataset is used to obtain the training target audio through convolution operations. In this embodiment, the early reflection audio used in step S306 can be selected from the early reflection audio 50ms or 100ms after the direct sound in the RIP data.

[0088] S307, using reverberant training audio as input and training target audio as validation data to train a predetermined model.

[0089] In step S307, the reverberant training audio obtained in step S305 can be used as input to the model, and the training target audio obtained in step S306 can be used as validation data to train the model used in step S303.

[0090] For example, a Short Time Fourier Transform (STFT) algorithm can be used to extract energy feature vectors from the mixed audio obtained by convolution operations based on sampled speech data and RIP data. In this embodiment, the extracted energy features can be filter bank feature vectors. Therefore, the extracted feature vectors can be input into the model and processed through the model's forward computation to obtain, for example, time-frequency masking data. Thus, this time-frequency masking data can then be combined with the previously obtained audio based on sampled speech and RIP data. Figure 1 The loss function for the target audio obtained by the audio-2 convolution shown can be calculated, for example, by calculating the mean square error between the time-frequency masking data output by the model and the masking data of the target audio. Then, the gradient backpropagation algorithm can be used to adjust the model parameters based on the calculation result and recalculate the model until the loss function of the current round no longer decreases significantly compared with the loss function calculated in the previous round. This indicates that the model has converged, that is, the training of the model has been completed.

[0091] Therefore, the audio processing scheme provided in this application uses audio generated by direct sound and early reflection audio as the target audio for training the model during model training, and uses the model trained in this way to process the mixed audio in actual use. Therefore, by selecting early reflection audio instead of direct sound as the target for model training and recovery, the original target audio can be effectively protected, ensuring the naturalness and clarity of the processed audio.

[0092] Example 4

[0093] Figure 4a This is a schematic diagram of an embodiment of the audio processing apparatus provided in this application, which can be used to perform... Figure 2 or Figure 3 The audio processing method shown is as follows. Figure 4aAs shown, the audio processing device may include: an acquisition module 41, an extraction module 42, and a processing module 43.

[0094] The acquisition module 41 can be used to acquire the audio to be processed.

[0095] In this embodiment, the acquisition module 41 can acquire the speech emitted by the speech source in the same space as the speech source. In other words, as the speech emitted by the speech source propagates in the acquisition space, a portion of the speech that propagates along the line connecting the speech source and the acquisition module 41 can be directly acquired by the acquisition module 41, for example... Figure 1 The audio shown is directly transmitted to audio 1, while another portion of the speech can propagate in other directions and via, for example, audio 1. Figure 1 The wall reflection shown reaches the acquisition module 41 after one or two reflections, for example, as Figure 1 As shown in the early reflected audio 2, a portion of the audio will undergo multiple reflections before reaching the acquisition module 41, for example... Figure 1 The late reflection audio 3 is shown in the figure. Therefore, the acquisition module 41 finally acquires the audio to be processed, which consists of such direct audio 1, early reflection audio 2 and late reflection audio 3.

[0096] The extraction module 42 can be used to extract the feature vector of the audio to be processed.

[0097] After the acquisition module 41 acquires the audio to be processed, which is a mixture of direct audio, early reflection audio, and late reflection audio, the extraction module 42 can first extract feature vectors from the audio to be processed. For example, filter bank feature vectors can be extracted.

[0098] The processing module 43 can be used to calculate the feature vector using a predetermined model to obtain the processed audio. For example, in this embodiment, the processing module 43 can use a predetermined model trained on reverberant training audio generated based on predetermined sampled audio to calculate the feature vector extracted by the extraction module 42 to obtain the processed audio.

[0099] The processing module 43 can input the feature vector extracted by the extraction module 42 into a predetermined model for processing. For example, the processing module 43 can input the feature vector extracted by the extraction module 42 into a predetermined model to perform forward computation on the feature vector and obtain, for example, time-frequency masking data. Then, the masking data obtained in this way can be multiplied with the time spectrum of the audio to be processed obtained by the acquisition module 41, and then an inverse Fourier transform can be performed to obtain the processed time-domain signal. This time-domain signal can be used as the processed audio for speech recognition or playback. For example, such a model can be a linear transformation model, a DFSMN (Deep Feedforward Sequential Memory Network) model, or a deep neural network model with a nonlinear activation function. For example, the predetermined model in this embodiment can be via, for example, Figure 4b The audio model training apparatus shown in this application is trained according to an embodiment of the present application.

[0100] Figure 4b A schematic diagram of an embodiment of the audio model training device provided in this application.

[0101] For example, an audio model training apparatus according to an embodiment of this application may include: a first generation module 44, a second generation module 45, and a training module 46.

[0102] The first generation module 44 can be used to generate reverberant training audio for predetermined sampled audio using a predetermined algorithm. In this embodiment, when the audio model training device trains the model, for example, it can first prepare speech audio with a sampling frequency of 16k / 48k as sampled speech data, and the audio model training device can randomly set the positions of the sound source and the acquisition device in a simulated randomly generated room. For example, it can be as follows: Figure 1 The diagram shows the sound source positioned in the middle left of the room, and the acquisition device positioned in the middle right. The first generation module 44 can then use various predetermined algorithms to generate reverberant training audio from the sampled audio. For example, the IMAGE method can be used to generate Room Impulse Response (RIP) data, which can be used to describe the reverberation characteristics of the simulated room. The model can then be configured and initialized according to actual needs. For example, in embodiments of this application, the model can use a linear transformation model, a DFSMN (Deep Feedforward Sequential Memory Network) model, or a deep neural network model with nonlinear activation functions. After model parameter initialization, the first generation module 44 can obtain mixed audio through convolution operations using the sampled speech data and RIP data.

[0103] The second generation module 45 can be used to generate training target audio based on at least a portion of the predetermined sampled audio and the reverberant training audio.

[0104] In this embodiment, the second generation module 45 can use the same number of sampled speech points and early reflection audio from the RIP data, for example... Figure 1 The audio 2 in the data is used to obtain the training target audio through convolution operation. In this embodiment, the early reflection audio used by the second generation module 45 can be selected from the early reflection audio 50ms or 100ms after the direct sound in the RIP data.

[0105] The training module 46 can be used to train the predetermined model used by the processing module 43 using reverberant training audio as input and training target audio as validation data.

[0106] For example, a Short Time Fourier Transform (STFT) algorithm can be used to extract energy feature vectors from the mixed audio obtained by convolution operations based on sampled speech data and RIP data. In this embodiment, the extracted energy features can be filter bank feature vectors. Therefore, the extracted feature vectors can be input into the model and processed through the model's forward computation to obtain, for example, time-frequency masking data. Thus, this time-frequency masking data can then be combined with the previously obtained audio based on sampled speech and RIP data. Figure 1 The loss function for the target audio obtained by the audio-2 convolution shown can be calculated, for example, by calculating the mean square error between the time-frequency masking data output by the model and the masking data of the target audio. Then, the gradient backpropagation algorithm can be used to adjust the model parameters based on the calculation result and recalculate the model until the loss function of the current round no longer decreases significantly compared with the loss function calculated in the previous round. This indicates that the model has converged, that is, the training of the model has been completed.

[0107] Therefore, the audio processing apparatus provided in this application training method uses audio generated from direct sound and early reflections as the target audio for training the model, and uses the trained model to process the mixed audio in actual use. Thus, by selecting early reflections instead of direct sound as the target for model training and recovery, the original target audio can be effectively protected, ensuring the naturalness and clarity of the processed audio.

[0108] Example 5

[0109] The above describes the internal functions and structure of a data processing device, which can be implemented as an electronic device. Figure 5 A schematic diagram illustrating the structure of an embodiment of the electronic device provided in this application. (See attached diagram.) Figure 5 As shown, the electronic device includes a memory 51 and a processor 52.

[0110] Memory 51 is used to store programs. In addition to the programs described above, memory 51 can also be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device, contact data, phonebook data, messages, pictures, videos, etc.

[0111] The memory 51 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0112] Processor 52 is not limited to a central processing unit (CPU), but may also be a graphics processing unit (GPU), a field-programmable gate array (FPGA), an embedded neural network processor (NPU), or an artificial intelligence (AI) chip. Processor 52 is coupled to memory 51 and executes the program stored in memory 51 to perform the audio processing methods of embodiments two and three described above.

[0113] Furthermore, such as Figure 5 As shown, the electronic device may also include other components such as a communication component 53, a power supply component 54, an audio component 55, and a display 56. Figure 5 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 5 The components shown.

[0114] Communication component 53 is configured to facilitate wired or wireless communication between electronic devices and other devices. The electronic devices can access wireless networks based on communication standards, such as WiFi, 3G, 4G, or 5G, or combinations thereof. In one exemplary embodiment, communication component 53 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 53 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0115] Power supply component 54 provides power to various components of the electronic device. Power supply component 54 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device.

[0116] Audio component 55 is configured to output and / or input audio signals. For example, audio component 55 includes a microphone (MIC) configured to receive external audio signals when the electronic device is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 51 or transmitted via communication component 53. In some embodiments, audio component 55 also includes a speaker for outputting audio signals.

[0117] Display 56 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation.

[0118] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An audio model training method, comprising: A reverberant training audio is generated using a predetermined algorithm for a predetermined sampled audio; wherein the reverberant training audio includes room impact response data generated using a mirror sound source model method and the predetermined sampled audio; Training target audio is generated based on at least a portion of the predetermined sampled audio and the reverberant training audio, wherein at least a portion of the reverberant training audio is an early reflection audio of the predetermined sampled audio within a predetermined time period; The predetermined model is trained using the reverberant training audio as input and the training target audio as validation data.

2. The audio model training method according to claim 1, wherein, The process of generating reverberant training audio for predetermined sampled audio using a predetermined algorithm includes: The reverberation training audio is obtained by convolution calculation using the predetermined sampled audio and predetermined room impact response data.

3. The audio model training method according to claim 1, wherein generating the training target audio based on at least a portion of the predetermined sampled audio and the reverberant training audio comprises: The training target audio is obtained by performing convolution calculations using the predetermined sampled audio and the early reflection audio.

4. The audio model training method according to claim 1, wherein, The step of training the predetermined model using the reverberant training audio as input and the training target audio as validation data further includes: Calculate the loss function based on the output data of the predetermined model and the validation data; The parameters of the predetermined model are adjusted according to the loss function; The model training has converged based on the difference between the loss function value and the loss function value obtained in the previous training round.

5. The audio model training method according to claim 4, wherein, The step of calculating the loss function based on the output data of the predetermined model and the validation data includes: Calculate the mean square error between the output mask and the ideal mask, and The step of adjusting the parameters of the predetermined model according to the loss function includes: The parameters are adjusted using a gradient backpropagation algorithm based on the mean square error.

6. An audio processing method, comprising: Obtain the audio to be processed; Extract the feature vector of the audio to be processed; The feature vector is calculated using a predetermined model trained on a reverberant training audio generated based on a predetermined sampled audio and a training target audio to obtain processed audio. The reverberant training audio includes room impact response data generated using a mirror source model method and the predetermined sampled audio. The training target audio is generated based on at least a portion of the predetermined sampled audio and the reverberant training audio. At least a portion of the reverberant training audio is the early reflection audio of the predetermined sampled audio within a predetermined time.

7. The audio processing method according to claim 6, wherein, The step of using a predetermined model to calculate the feature vector to obtain the processed audio includes: The predetermined model is used to perform forward computation on the feature vector to obtain masking data; The masking data is multiplied by the time spectrum of the audio to be processed, and then an inverse Fourier transform is performed to obtain the processed audio.

8. A method for processing conference audio, comprising: The audio of speeches sent by the participating terminals is acquired through an audio acquisition device. Extract the feature vector of the spoken audio; The feature vector is calculated using a predetermined model trained on a reverberant training audio generated based on a predetermined sampled audio and a training target audio to obtain processed audio. The reverberant training audio includes room impact response data generated using a mirror source model method and the predetermined sampled audio. The training target audio is generated based on at least a portion of the predetermined sampled audio and the reverberant training audio. At least a portion of the reverberant training audio is the early reflection audio of the predetermined sampled audio within a predetermined time. The processed audio is then sent to other participating terminals in the conference.

9. A classroom audio processing method, comprising: The teaching audio sent by the teacher during the lecture was obtained by audio acquisition devices installed in the classroom; Extract the feature vector of the teaching audio; The feature vector is calculated using a predetermined model trained on a reverberant training audio generated based on a predetermined sampled audio and a training target audio to obtain processed audio. The reverberant training audio includes room impact response data generated using a mirror source model method and the predetermined sampled audio. The training target audio is generated based on at least a portion of the predetermined sampled audio and the reverberant training audio. At least a portion of the reverberant training audio is the early reflection audio of the predetermined sampled audio within a predetermined time. The processed audio is sent over the network to terminals that listen to classroom lectures via the network.

10. An electronic device, comprising: Memory, used to store programs; A processor is configured to run the program stored in the memory to perform the audio model training method as described in any one of claims 1-5 or the audio processing method as described in any one of claims 6-7.

11. A computer-readable storage medium having a computer program stored thereon that can be executed by a processor, wherein, When the program is executed by the processor, it implements the audio model training method as described in any one of claims 1-5 or the audio processing method as described in any one of claims 6-7.

12. A computer program product, wherein, include: A computer-readable storage medium storing computer instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Method for generating reverberation audio signals and training method of audio processing model

    CN112652290A