Method and apparatus for training audio processing model, storage medium, and electronic device

By pre-training and joint training of the branch network of the audio processing model, the problems of complexity and high power consumption of audio processing processes in remote meetings are solved, and efficient audio signal processing in low signal-to-noise ratio environments are achieved, reducing system power consumption and improving model adaptability and accuracy.

WO2025152852A1PCT designated stage expired Publication Date: 2025-07-24JINGDONG CITY BEIJING DIGITS TECH CO LTD +1

Patent Information

Application Number
PCT/CN2025/071596
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-16
Filing Date
2025-01-09
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

In the prior art, the audio processing flow in remote meetings is long, resulting in increased system power consumption, and it is impossible to effectively solve the complexity and power consumption problems of tasks such as echo cancellation, voice enhancement, and voice endpoint detection.

Method used

By obtaining a training sample set containing the first sample set and the second sample set, the first branch network of the audio processing model is pre-trained for performing echo cancellation and speech enhancement tasks, and the second branch network is pre-trained for performing speech endpoint detection tasks, and conducting joint training to form a comprehensive model.

Benefits of technology

It improves the accuracy of model processing in low signal-to-noise ratio environments, reduces system complexity and power consumption, can adapt to a large range of delays, covers a large range of signal-to-noise ratio environments in multiple scenarios, and improves the effect of voice endpoint detection at low signal-to-noise ratio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025071596_24072025_PF_FP_ABST
    Figure CN2025071596_24072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of artificial intelligence, and provides a method for training an audio processing model, an apparatus for training an audio processing model, a computer storage medium, and an electronic device. The method for training an audio processing model comprises: acquiring a training sample set; using a first sample set to pre-train a first branch network of an audio processing model to be trained, so as to obtain a pre-trained first branch network, and using a second sample set to pre-train a second branch network of said audio processing model, so as to obtain a pre-trained second branch network; and using the training sample set to jointly train the pre-trained first branch network and the pre-trained second branch network, so as to obtain a trained audio processing model, wherein the first branch network is used for executing echo cancellation and speech enhancement tasks, and the second branch network is used for executing a voice activity detection task. In the present disclosure, multiple audio processing tasks can be executed by means of one model, thereby reducing system power.
Need to check novelty before this filing date? Find Prior Art

Description

Audio processing model training method and device, storage medium, and electronic device

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] The present disclosure claims priority to Chinese patent application number CN 202410063886.8 filed on January 16, 2024, entitled “Training method and device for audio processing model, storage medium, and electronic device”, the entire contents of which are incorporated by reference into the present disclosure. Technical Field

[0003] The present disclosure relates to the field of artificial intelligence technology, and in particular to a training method for an audio processing model, a training device for an audio processing model, a computer storage medium, and an electronic device. Background Art

[0004] With technological advancements, remote conferencing has become a crucial aspect for improving work efficiency, reducing travel costs, and enabling remote meetings anytime, anywhere. The voice interaction experience in this scenario relies heavily on acoustic signal processing, primarily including upstream features like acoustic echo cancellation (AEC) and speech enhancement (SE). If backup of meeting content is required, downstream features like voice activity detection (VAD) and automatic speech recognition (ASR) are also required.

[0005] In the related art, the above-mentioned multiple signal processing tasks are generally performed through multiple separate modules. However, this solution will result in a longer audio processing flow, thereby increasing system power consumption.

[0006] In view of this, there is an urgent need in this field to develop a new training method and device for audio processing models.

[0007] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this disclosure. Summary of the Invention

[0008] The purpose of the present disclosure is to provide a training method for an audio processing model, a training device for an audio processing model, a computer storage medium and an electronic device, thereby overcoming, at least to a certain extent, the technical problem of high system power consumption caused by the limitations of related technologies.

[0009] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0010] According to a first aspect of the present disclosure, a method for training an audio processing model is provided, comprising:

[0011] Acquire a training sample set; the training sample set includes a first sample set and a second sample set;

[0012] Pre-training a first branch network of an audio processing model to be trained using the first sample set to obtain a pre-trained first branch network, and pre-training a second branch network of the audio processing model to be trained using the second sample set to obtain a pre-trained second branch network;

[0013] Using the training sample set, jointly train the pre-trained first branch network and the pre-trained second branch network to obtain the trained audio processing model;

[0014] The first branch network is used to perform echo cancellation and speech enhancement tasks, and the second branch network is used to perform speech endpoint detection tasks.

[0015] According to a second aspect of the present disclosure, there is provided a training device for an audio processing model, comprising:

[0016] A training sample set acquisition module, configured to acquire a training sample set; the training sample set includes a first sample set and a second sample set;

[0017] a pre-training module, configured to pre-train a first branch network of the audio processing model to be trained using the first sample set to obtain a pre-trained first branch network, and to pre-train a second branch network of the audio processing model to be trained using the second sample set to obtain a pre-trained second branch network;

[0018] a joint training module, configured to jointly train the pre-trained first branch network and the pre-trained second branch network using the training sample set to obtain the trained audio processing model;

[0019] The first branch network is used to perform echo cancellation and speech enhancement tasks, and the second branch network is used to perform speech endpoint detection tasks.

[0020] According to a third aspect of the present disclosure, a computer storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the training method of the audio processing model described in the first aspect is implemented.

[0021] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the training method of the audio processing model described in the first aspect above by executing the executable instructions.

[0022] As can be seen from the above technical solutions, the audio processing model training method, audio processing model training device, computer storage medium, and electronic device in the exemplary embodiments of the present disclosure have at least the following advantages and positive effects:

[0023] In the technical solutions provided by some embodiments of the present disclosure, on the one hand, the present disclosure obtains a training sample set including a first sample set and a second sample set, uses the first sample set to pre-train the first branch network of the audio processing model to be trained, and obtains a pre-trained first branch network (for performing echo cancellation and speech enhancement tasks), and uses the second sample set to pre-train the second branch network of the audio processing model to be trained (for performing speech endpoint detection tasks) to obtain a pre-trained second branch network, and uses the training sample set to jointly train the pre-trained first branch network and the pre-trained second branch network to obtain a trained audio processing model. On the one hand, by collecting rich training sample sets, it is possible to basically cover a wide range of signal-to-noise ratio environments in multiple scenarios through as rich data simulation as possible, thereby improving the model processing accuracy in low signal-to-noise ratio environments; on the other hand, one model can solve three audio signal processing tasks, reducing the mutual constraints brought about by the parameter adjustment of each sub-process in the related technology, greatly reducing the complexity of the system, and reducing the system power consumption.

[0024] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0026] FIG1 is a schematic diagram showing a flow chart of a method for training an audio processing model in an embodiment of the present disclosure;

[0027] FIG2 is a schematic diagram showing a process of obtaining a first sample set in an embodiment of the present disclosure;

[0028] FIG3 is a schematic diagram showing a flow chart of how to obtain a third sample set in an embodiment of the present disclosure;

[0029] FIG4 is a schematic diagram showing a flow chart of how to perform analog processing on an audio signal transmitted from a far end to a near end to obtain an analog audio signal in an embodiment of the present disclosure;

[0030] FIG5 is a flow chart showing how to mix an analog audio signal with a far-end standard audio signal to obtain an audio signal collected at the near end in an embodiment of the present disclosure;

[0031] FIG6 is a schematic diagram showing how to obtain a third sample set in an embodiment of the present disclosure;

[0032] FIG7 is a schematic diagram showing a flow chart of how to obtain a fourth sample set in an embodiment of the present disclosure;

[0033] FIG8 is a schematic diagram showing how to obtain the mic signal and the farend signal in the fourth sample set in an embodiment of the present disclosure;

[0034] FIG9 is a schematic diagram showing a flow chart of how to use the first sample set to pre-train the first branch network of the audio processing model to be trained to obtain the pre-trained first branch network in an embodiment of the present disclosure;

[0035] FIG10 is a flow chart showing how, in an embodiment of the present disclosure, the audio signal collected by the near end and the audio signal transmitted from the far end to the near end are converted and synthesized through the first branch network to obtain an output signal;

[0036] FIG11 is a flow chart showing how to determine a first loss value according to the degree of signal difference between an output signal and a far-end standard audio signal in an embodiment of the present disclosure;

[0037] FIG12 is a flow chart showing how to use the second sample set to pre-train the second branch network of the audio processing model to be trained to obtain the pre-trained second branch network in an embodiment of the present disclosure;

[0038] FIG13 is a flow chart showing how to jointly train a pre-trained first branch network and a pre-trained second branch network using a training sample set to obtain a trained audio processing model in an embodiment of the present disclosure;

[0039] FIG14 is a schematic diagram showing the model structure of a trained audio processing model in an embodiment of the present disclosure;

[0040] FIG15 is a schematic diagram showing the structure of a training device for an audio processing model in an exemplary embodiment of the present disclosure;

[0041] FIG16 is a schematic structural diagram of an electronic device in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0042] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the present disclosure will be more comprehensive and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or that other methods, components, devices, steps, etc. may be employed. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0043] The terms "a", "an", "the" and "said" are used in this specification to indicate the presence of one or more elements / components / etc.; the terms "including" and "having" are used to express open-ended inclusion and mean that additional elements / components / etc. may exist in addition to the listed elements / components / etc.; the terms "first" and "second" etc. are used only as labels and are not intended to limit the quantity of their objects.

[0044] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the drawings represent identical or similar parts, and thus repeated descriptions thereof will be omitted. Some of the blocks shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically separate entities.

[0045] In a phone call or teleconference, there's a near end and a far end. The near end is where any user is located, and the far end is where the other participants in the call are located. Acoustic echo cancellation (AEC) prevents far-end participants in a teleconference from hearing an echo of their own voice.

[0046] AEC can remove the far-end signal reflection component from the audio signal collected by the near-end microphone (called the mic signal) based on the audio signal transmitted from the far-end to the near-end (called the farend signal), and then transmit it to the far-end, so that the far-end user cannot hear his or her own echo. The main difficulties are:

[0047] 1. The speed and accuracy of the time delay alignment between the mic signal and the farend signal;

[0048] 2. Estimate the performance of the linear filter of the echo through the farend signal;

[0049] 3. Accuracy of residual echo estimation (non-linear echo cancellation).

[0050] To address the above difficulties, there are two main commonly used AEC algorithms:

[0051] The first is to use traditional algorithms to solve all three difficulties. For example, delay estimation uses the NLMS algorithm in the time domain or the GCC-PHAT algorithm in the frequency domain, and linear echo cancellation uses Kalman filtering. However, there is no unified method for residual echo estimation. Each manufacturer will perform corresponding nonlinear processing based on the distortion and reverberation level of the actual equipment.

[0052] The second approach maintains the solutions to the first two challenges, while addressing the third challenge through a neural network, allowing the network to learn the degree of nonlinear distortion. Consequently, the accuracy of delay estimation affects the effectiveness of linear echo cancellation, while the performance of the linear filter also affects the suppression of residual echo.

[0053] SE enhances (also known as noise reduction) noisy speech signals, filtering out ambient noise so that far-end users can hear the near-end speaker's voice more clearly. The main challenge is the diverse nature of noise, and speech at low signal-to-noise ratios can be severely suppressed by noise. Previous SE algorithms primarily used adaptive filtering to estimate noise components. This traditional approach relies heavily on empirical parameter settings, significantly degrading performance in complex scenarios, such as those with low signal-to-noise ratios.

[0054] As can be seen, most related technologies for AEC and SE tasks require delay estimation and an adaptive process, which can lead to significant distortion at the beginning of a meeting. Furthermore, the accuracy of delay estimation affects the effectiveness of linear echo cancellation, while the performance of the linear filter affects the suppression of residual echoes. These factors can lead to poor generalization capabilities when the scenario (such as meeting environment or noisy environment) changes.

[0055] VAD is the process of automatically identifying the start and end points of speech within a continuous speech signal, thereby distinguishing between speech segments and non-speech segments. Previous VAD methods used pre-set thresholds based on statistical characteristics of speech signals, such as energy and zero-crossing rate. Consequently, in low signal-to-noise ratio environments, false positives were common and the return delay for speech endpoints was significant.

[0056] In addition, in general, since the related art generally performs the above-mentioned multiple audio processing tasks separately, the processing flow is relatively long, and a relatively long processing flow will inevitably increase the power consumption of the system.

[0057] In an embodiment of the present disclosure, a training method for an audio processing model is first provided, which overcomes the defect of high system power consumption in the related art at least to a certain extent.

[0058] FIG1 shows a flow chart of a method for training an audio processing model in an embodiment of the present disclosure. The executing entity of the method for training an audio processing model may be a server that trains the audio processing model.

[0059] 1 , a method for training an audio processing model according to an embodiment of the present disclosure includes the following steps:

[0060] Step S110, obtaining a training sample set; the training sample set includes a first sample set and a second sample set;

[0061] Step S120: pre-training a first branch network of the audio processing model to be trained using the first sample set to obtain a pre-trained first branch network, and pre-training a second branch network of the audio processing model to be trained using the second sample set to obtain a pre-trained second branch network; wherein the first branch network is used to perform echo cancellation and speech enhancement tasks, and the second branch network is used to perform speech endpoint detection tasks;

[0062] Step S130: jointly train the pre-trained first branch network and the pre-trained second branch network using the training sample set to obtain a trained audio processing model.

[0063] In the technical solution provided by the embodiment shown in FIG1 , the present disclosure obtains a training sample set including a first sample set and a second sample set, uses the first sample set to pre-train the first branch network of the audio processing model to be trained, and obtains a pre-trained first branch network (for performing echo cancellation and speech enhancement tasks), and uses the second sample set to pre-train the second branch network of the audio processing model to be trained (for performing speech endpoint detection tasks), and obtains a pre-trained second branch network. The pre-trained first branch network and the pre-trained second branch network are jointly trained using the training sample set to obtain a trained audio processing model. On the one hand, by collecting a rich training sample set, it is possible to basically cover a wide range of signal-to-noise ratio environments of multiple scenarios through as rich data simulation as possible, thereby improving the model processing accuracy in low signal-to-noise ratio environments; on the other hand, one model can solve three audio signal processing tasks, reducing the mutual constraints brought about by the parameter adjustment of each sub-process in the related technology, greatly reducing the complexity of the system, and reducing the system power consumption. The specific implementation process of each step in FIG1 is described in detail below:

[0064] In step S110 , a training sample set is obtained.

[0065] In an exemplary embodiment of the present disclosure, the training sample set may include a first sample set and a second sample set. The first sample set may be used to train AEC and SE tasks, and the second sample set may be used to train VAD tasks.

[0066] Among them, the above-mentioned first sample set may include audio signals transmitted from the far end to the near end (that is, audio signals sent from the far end and transmitted to the near end through the network, hereinafter referred to as farend signals), far-end standard audio signals (which may be actual audio signals sent from the far end, hereinafter referred to as target signals), and audio signals collected by the near end (for example: audio signals collected by the near-end microphone, hereinafter referred to as mic signals).

[0067] The second sample set includes a far-end standard audio signal and a real speech recognition label corresponding to the far-end standard audio signal. The real speech recognition label is used to indicate whether each frame of the far-end standard audio signal is a speech signal. For example, the far-end standard audio signal and its speech text may be forcibly aligned to obtain the real speech label corresponding to each frame of the far-end standard audio signal.

[0068] The following first describes a specific implementation of how to obtain the first sample set in conjunction with FIG2 . Referring to FIG2 , FIG2 shows a flowchart of how to obtain the first sample set in an embodiment of the present disclosure, including steps S201 to S203:

[0069] In step S201 , a third sample set and a fourth sample set are obtained.

[0070] In this step, the third sample set may be a sample set for training the AEC task, and the fourth sample set may be a sample set for training the SE task.

[0071] The third sample set may include multiple training samples, and each training sample includes an audio signal transmitted from the far end to the near end, a far end standard audio signal, and an audio signal collected from the near end.

[0072] The fourth sample set may include multiple training samples, each of which includes an audio signal transmitted from the far end to the near end, a far end standard audio signal, and an audio signal collected from the near end.

[0073] It should be noted that the third sample set and the fourth sample set contain the same data type, but the data are generated in different ways.

[0074] The specific method of generating the third sample set will be described below with reference to FIG3 . Referring to FIG3 , FIG3 shows a flowchart of how to obtain the third sample set in an embodiment of the present disclosure, including steps S301 to S303:

[0075] In step S301, a first clean audio and a second clean audio are randomly selected from a preset massive clean audio sample, the first clean audio is used as the audio signal transmitted from the far end to the near end, and the second clean audio is used as the far end standard audio signal.

[0076] In this step, the first clean audio and the second clean audio can be randomly selected from a preset massive amount of clean audio samples. The clean audio can be the audio after noise removal processing. Then, the first clean audio can be used as the audio signal (farend signal) transmitted from the far end to the near end, and the second clean audio can be used as the far-end standard audio signal (target signal).

[0077] In step S302, the audio signal transmitted from the far end to the near end is analogized to obtain an analog audio signal.

[0078] In this step, the audio signal transmitted from the far end to the near end can be analogized to obtain an analog audio signal. Specifically, reference can be made to FIG4 , which shows a flow chart of how to analogize the audio signal transmitted from the far end to the near end to obtain an analog audio signal in an embodiment of the present disclosure, including steps S401 to S404:

[0079] In step S401, a random delay within a preset time range is added to an audio signal transmitted from a far end to a near end to obtain a first transformed audio signal.

[0080] In this step, a random delay within a preset duration range can be added to the audio signal transmitted from the far-end to the near-end (farend signal) to obtain a first transformed audio signal. For example, the preset duration range can be 0-1 second and can be set according to actual circumstances. This disclosure does not impose any specific limitations on this.

[0081] By adding the random delay, the network transmission time required for the far-end signal to be transmitted to the near-end can be simulated.

[0082] In step S402, random noise is added to the first transformed audio signal to obtain a second transformed audio signal.

[0083] In this step, it should be noted that the present disclosure can pre-collect multiple preset noises associated with preset scenarios to pre-configure a noise library. For example, the preset scenario can be a remote conference scenario, and the preset noises can include over 200 types of noises, such as keyboard sounds, fan sounds, and mobile phone ringtones. These noises can be customized based on actual circumstances and are not specifically limited in this disclosure.

[0084] Therefore, in this step, noise may be randomly selected from the noise library to add the randomly selected noise to the first transformed audio signal to obtain a second transformed audio signal.

[0085] In step S403, nonlinear disturbance is added to the second transformed audio signal to obtain a third transformed audio signal.

[0086] In this step, nonlinear perturbations may be added to the second transformed audio signal to obtain a third transformed audio signal. By adding nonlinear perturbations, it is possible to simulate different degrees of distortion in the audio signal caused by different equipment quality and installation environments.

[0087] In step S404, random reverberation is added to the third transformed audio signal to obtain an analog audio signal.

[0088] In this step, random reverberation may be added to the third transformed audio signal to simulate room environments of different sizes to obtain a simulated audio signal.

[0089] After the analog audio signal is obtained, referring to FIG. 3 , in step S303 , the analog audio signal and the far-end standard audio signal are mixed to obtain the near-end collected audio signal.

[0090] In this step, the analog audio signal can be mixed with the far-end standard audio signal to obtain the near-end collected audio signal (mic signal). Specifically, referring to FIG5 , FIG5 shows a flow chart of how to mix the analog audio signal with the far-end standard audio signal to obtain the near-end collected audio signal in an embodiment of the present disclosure, including steps S501 to S502:

[0091] In step S501 , random noise is added to a near-end standard audio signal to obtain a target audio signal.

[0092] In this step, random noise may be added to the near-end standard audio signal to obtain a target audio signal. The random noise may be randomly selected from the pre-configured noise library.

[0093] In step S502, the analog audio signal and the target audio signal are mixed to obtain an audio signal collected by the near end.

[0094] In this step, the analog audio signal and the target audio signal may be mixed to obtain an audio signal (mic signal) collected at the near end.

[0095] Next, referring to FIG6 , FIG6 is a schematic diagram showing how to obtain the third sample set in an embodiment of the present disclosure, as shown in FIG6 :

[0096] After randomly selecting a first clean audio and a second clean audio, the first clean audio may be used as an audio signal (farend signal) transmitted from the far end to the near end, and the second clean audio may be used as a far-end standard audio signal (target signal);

[0097] Next, a random delay of 0-1s, random noise, nonlinear perturbation, and random reverberation may be added to the first clean audio in sequence to obtain an analog audio signal.

[0098] and, adding random noise to the second clean audio to obtain a target audio signal;

[0099] Afterwards, the analog audio signal and the target audio signal may be mixed to obtain an audio signal collected at the near end (mic signal).

[0100] The following is an explanation of a specific implementation of how the fourth sample set is obtained in the present application with reference to FIG7 . FIG7 shows a flowchart of how the fourth sample set is obtained in an embodiment of the present disclosure, including steps S701 to S703:

[0101] In step S701, a first clean audio signal and a second clean audio signal are randomly selected from a preset massive amount of clean audio samples, and the second clean audio signal is used as a remote standard audio signal.

[0102] In this step, the first clean audio and the second clean audio may be randomly selected from the preset massive clean audio samples, and the second clean audio may be used as the remote standard audio signal (target signal).

[0103] In step S702, random noise is added to the first clean audio to obtain an audio signal collected at the near end.

[0104] In this step, reference may be made to the explanation of step S402 above to randomly select noise from the pre-configured noise library. Then, noise processing may be performed on the first clean audio based on the randomly selected noise to obtain the near-end collected audio signal (mic signal).

[0105] In step S703, a non-all-zero silence signal having the same length as the audio signal collected at the near end is used as the audio signal transmitted from the far end to the near end.

[0106] In this step, a non-zero silence signal of the same length as the audio signal collected by the near-end in step S702 can be added as the audio signal transmitted from the far-end to the near-end. The non-zero silence signal is a silence signal corresponding to a signal sequence consisting of 0s and 1s, i.e., an audio signal in which no sound is audible within the human hearing range.

[0107] Referring to FIG8 , FIG8 is a schematic diagram showing how to obtain the mic signal and the farend signal in the fourth sample set in an embodiment of the present disclosure, as shown in FIG8 :

[0108] The first clean audio and the second clean audio may be randomly selected, and the second clean audio is used as the remote standard audio signal (target signal);

[0109] Randomly add noise to the first clean audio to obtain an audio signal (mic signal) collected by the near end;

[0110] A non-all-zero silent signal of the same length as the audio signal collected at the near end is supplemented as the audio signal (farend signal) transmitted from the far end to the near end.

[0111] After obtaining the third sample set and the fourth sample set, you can continue to refer to Figure 2. In step S202, according to the preset proportion relationship, select the first target number of training samples from the third sample set, and select the second target number of training samples from the fourth sample set; the ratio of the first target number to the second target number meets the preset ratio condition.

[0112] In this step, the preset ratio can be 1:3, for example. For example, when the first sample set needs to include 4,000 training samples, 1,000 training samples can be randomly selected from the third sample set, and 3,000 training samples can be selected from the fourth sample set. It should be noted that the preset ratio is the optimal ratio obtained through multiple experiments in the present disclosure. By extracting training samples from the third and fourth sample sets according to this ratio, the performance of the subsequent model can be optimized.

[0113] In step S203 , the first target number of training samples and the second target number of training samples are mixed to obtain a first sample set.

[0114] In this step, after selecting the first target number of training samples from the third sample set and the second target number of training samples from the fourth sample set, the first target number of training samples and the second target number of training samples can be mixed to obtain the first sample set.

[0115] After obtaining the first sample set and the second sample set, you can then refer to Figure 1. In step S120, the first branch network of the audio processing model to be trained is pre-trained using the first sample set to obtain the pre-trained first branch network, and the second branch network of the audio processing model to be trained is pre-trained using the second sample set to obtain the pre-trained second branch network.

[0116] In this step, pre-training refers to the process of training a model on a large dataset before fine-tuning it on a smaller, task-specific dataset.

[0117] Specifically, referring to FIG9 , FIG9 shows a flowchart of how to use a first sample set to pre-train a first branch network of an audio processing model to be trained to obtain a pre-trained first branch network in an embodiment of the present disclosure, including steps S901 to S903:

[0118] In step S901, signal conversion and synthesis processing are performed on the audio signal collected by the near end and the audio signal transmitted from the far end to the near end through the first branch network to obtain an output signal.

[0119] In this step, the first branch network can be used to perform echo cancellation and speech enhancement tasks. The first branch network may include a Fourier transform unit, an encoding unit, a long-range dependency capture unit, and a decoding unit. A residual jump connection structure is formed between the decoding unit and the encoding unit. Referring to FIG10 , FIG10 shows how, in an embodiment of the present disclosure, the first branch network is used to perform signal conversion and synthesis processing on the audio signal collected by the near end and the audio signal transmitted from the far end to the near end to obtain an output signal, including steps S1001 to S1004:

[0120] In step S1001, a Fourier transform unit performs Fourier transform on an audio signal transmitted from a far end to a near end to obtain a first transformed audio signal.

[0121] In this step, illustratively, the Fourier transform unit may perform a short-time Fourier transform (STFT) on the audio signal transmitted from the far end to the near end to obtain a first transformed audio signal.

[0122] Among them, short-time Fourier transform (STFT, short-time Fourier transform, or short-term Fourier transform) is a mathematical transformation related to Fourier transform, which is used to determine the frequency and phase of the sine wave in a local area of ​​a time-varying signal.

[0123] After obtaining the first transformed audio signal, the process proceeds to step S1002 , where an encoding unit encodes the first transformed audio signal to obtain encoding information.

[0124] In this step, the encoding unit may be an encoder network composed of multiple two-dimensional convolutions, through which local information of the first transformed audio signal can be extracted and feature resolution can be reduced to obtain encoding information.

[0125] In step S1003, the long-range dependency capturing unit performs long-range dependency capturing on the coding features to obtain long-range dependency information.

[0126] In this step, the long-range dependency capture unit can be a GRU (Gate Recurrent Unit, recurrent neural network) / LSTM network (Long Short Term Memory, long short-term memory network), which can enhance modeling capabilities by learning long-range information, thereby obtaining long-range dependency information.

[0127] In step S1004, the residual information transmitted by the decoding unit and the encoding unit is decoded to obtain an output signal.

[0128] In this step, the decoding unit may use transposed convolution to decode the long-range dependency information and the residual information transmitted by the encoding unit to obtain an output signal of the original size.

[0129] After the output signal is obtained, referring to FIG. 9 , in step S902 , a first loss value is determined according to the degree of signal difference between the output signal and the far-end standard audio signal.

[0130] In this step, reference may be made to FIG11 , which illustrates a flow chart showing how to determine a first loss value based on the degree of signal difference between the output signal and the remote standard audio signal in an embodiment of the present disclosure, including steps S1101 to S1103:

[0131] In step S1101 , a time domain loss is determined based on the degree of signal difference between the output signal and the far-end standard audio signal in the time domain.

[0132] In this step, the output signal may be subjected to an inverse short-time Fourier transform (ISTFT) to obtain a time domain signal. Then, the time domain loss may be determined based on a preset time domain loss calculation formula, the strength of the time domain signal, and the strength of the remote standard audio signal.

[0133] For example, the above-mentioned preset time domain loss calculation formula can refer to the following formula 1:

[0134]

[0135] Among them, the above P s It can represent the strength of the above time domain signal. The above P n Indicates the strength of the far-end standard audio signal.

[0136] In step S1102 , the frequency domain loss is determined according to the degree of signal difference between the output signal and the far-end standard audio signal in the frequency domain.

[0137] In this step, the frequency domain loss can be determined based on the ratio between the frequency domain conversion result of the time domain signal and the frequency domain conversion result of the far-end standard audio signal. For example, the time domain signal can be subjected to a short-time Fourier transform again to obtain its frequency domain conversion result, wherein the frequency domain conversion result can include a real part and an imaginary part. Simultaneously, the far-end standard audio signal is also subjected to a short-time Fourier transform to obtain its frequency domain conversion result, which also includes a real part and an imaginary part. Furthermore, the frequency domain loss can be determined by comparing the two frequency domain conversion results (i.e., comparing the real part with the real part and the imaginary part with the imaginary part).

[0138] In step S1103, a first loss value is determined according to the time domain loss and the frequency domain loss.

[0139] In this step, illustratively, the first loss value may be determined according to the cumulative value of the time domain loss and the frequency domain loss.

[0140] After determining the first loss value, referring to FIG. 9 , in step S903 , the first branch network is iteratively trained according to the first loss value to obtain a pre-trained first branch network.

[0141] In this step, the first branch network can be iteratively trained according to the above-mentioned first loss value until the above-mentioned first loss value meets the preset convergence condition, thereby obtaining a pre-trained first branch network.

[0142] The following describes a specific implementation of how to obtain a pre-trained second branch network (for performing speech endpoint detection tasks) in the present disclosure, in conjunction with Figure 12. Referring to Figure 12, Figure 12 illustrates a flow chart of how to pre-train the second branch network of the audio processing model to be trained using a second sample set to obtain the pre-trained second branch network in an embodiment of the present disclosure, including steps S1201 to S1203:

[0143] In step S1201, speech recognition is performed on the remote standard audio signal through the second branch network to obtain a speech recognition label corresponding to each frame signal of the remote standard audio signal.

[0144] In this step, the second branch network may be used to perform speech recognition on the remote standard audio signal in the second sample set to obtain a speech recognition label (i.e., speech or non-speech) corresponding to each frame signal of the remote standard speech signal.

[0145] In step S1202, a second loss value is determined according to the degree of difference between the speech recognition label and the true speech recognition label.

[0146] In this step, a second loss value may be determined based on the degree of difference between the speech recognition label of each frame signal and the true speech recognition label of each frame signal. The second loss value may be a cross-entropy loss.

[0147] The cross-entropy loss only considers the loss of positive samples. It compares the model's predictions with the actual speech recognition labels of the data. As the predictions become more accurate, the cross-entropy value decreases. If the predictions are completely correct, the cross-entropy value is 0. Therefore, when training a split model, cross-entropy can be used as a loss function.

[0148] In step S1203, the second branch network is iteratively trained according to the second loss value to obtain a pre-trained second branch network.

[0149] In this step, the second branch network can be iteratively trained based on the second loss value until the second loss value meets a preset convergence condition, thereby obtaining a pre-trained second branch network.

[0150] After obtaining the pre-trained first branch network and the pre-trained second branch network, we can continue with FIG1 , and in step S130 , jointly train the pre-trained first branch network and the pre-trained second branch network using the training sample set to obtain a trained audio processing model.

[0151] In this step, joint training means that there are multiple subtasks in the model, and we can train these subtasks together.

[0152] Specifically, reference may be made to FIG13 , which illustrates a flowchart of how to jointly train a pre-trained first branch network and a pre-trained second branch network using a training sample set to obtain a trained audio processing model in an embodiment of the present disclosure, including steps S1301 to S1305:

[0153] In step S1301, Fourier transform is performed on the audio signal transmitted from the far end to the near end to obtain a target audio signal.

[0154] In this step, a short-time Fourier transform may be performed on the audio signal transmitted from the far end to the near end (farend signal) to obtain a target audio signal.

[0155] In step S1302 , a masking operation is performed on the target audio signal and the output signal to obtain a masked audio signal.

[0156] In this step, a mask operation can be performed on the target audio signal and the output signal of the first branch network to obtain a masked audio signal. Masking is a common operation in deep learning. Simply put, it is equivalent to adding a mask to the original tensor to block or select specific elements. Therefore, it is often used to construct tensor filters.

[0157] Through masking operation, the near-end speech in the output signal can be retained while other components are masked out, thereby obtaining a frequency domain representation of the pure near-end speech.

[0158] In step S1303, speech recognition is performed on the masked audio signal through the pre-trained second branch network to obtain a speech recognition label corresponding to each frame signal in the masked audio signal.

[0159] In this step, the pre-trained second-branch network performs speech recognition on the masked audio signal, obtaining a speech recognition label for each frame of the masked audio signal. By leveraging the powerful modeling capabilities of neural networks and directly evaluating each audio frame, and by modifying the network and training objectives, more accurate speech recognition can be achieved even at low signal-to-noise ratios.

[0160] In step S1304, a third loss value is determined according to the degree of difference between the speech recognition label and the true speech recognition label.

[0161] In this step, the third loss value can be determined based on the degree of difference between the speech recognition label of each frame signal output in step S1303 and the real speech recognition label corresponding to each frame signal of the above-mentioned remote standard audio signal (target signal). The third loss value can be a CE loss value.

[0162] In step S1305, the pre-trained first branch network and the pre-trained second branch network are iteratively trained according to the third loss value to obtain a trained audio processing model.

[0163] In this step, the pre-trained first branch network and the pre-trained second branch network can be iteratively trained according to the third loss value until the third loss value meets the preset convergence condition to obtain a trained audio processing model.

[0164] For example, referring to FIG14 , FIG14 shows a schematic diagram of the model structure of the trained audio processing model in an embodiment of the present disclosure, as shown in FIG14 :

[0165] The first branch network includes a Fourier transform unit, an encoding unit, a long-range dependency capture unit, and a decoding unit, and a residual jump connection structure is used between the decoding unit and the encoding unit;

[0166] The first branch network can perform signal conversion and synthesis processing on the audio signal collected by the near end (mic signal) and the audio signal transmitted from the far end to the near end (farend signal) to obtain an output signal;

[0167] The second branch network (VAD classifier) ​​can perform speech recognition on the output signal to obtain its corresponding speech recognition label.

[0168] It should be noted that after obtaining the above-mentioned trained audio processing model, in the application stage of the model, exemplarily, the present disclosure can obtain the signal to be processed (the signal to be processed may include the real-time audio signal collected by the near end and the real-time audio signal transmitted from the far end to the near end), and then, the signal to be processed can be input into the trained audio processing model. In an optional embodiment, the first branch network of the above-mentioned audio processing model can perform echo cancellation and speech enhancement processing on the above-mentioned signal to be processed to obtain a processed audio signal. In another optional embodiment, after the first branch network outputs the processed audio signal, if the user turns on the VAD function, the second branch network of the audio processing model can perform speech endpoint detection processing on the above-mentioned audio signal, thereby outputting the processed audio signal and the speech recognition label corresponding to each frame signal of the processed audio signal, so as to facilitate the downstream ASR and other related tasks.

[0169] In addition, it should be noted that the present disclosure also conducted model tests on the AEC task and SE task of the above-trained audio processing model. Referring to Table 1 and Table 2, Table 1 shows the test results for the AEC task, and Table 2 shows the test results for the above-mentioned SE task:

[0170] Table 1

[0171] Table 2

[0172] As shown in Table 1, we choose AECMOS to evaluate AEC tasks. AEC tasks are divided into three scenarios: far-end single-talk, near-end single-talk, and two-talk. AECMOS indicators are used to evaluate AEC tasks:

[0173] AECMOS is used to evaluate the mean opinion score (Acoustic Echo Cancellation Mean Opinion Score) of AEC tasks. A higher score indicates better speech quality.

[0174] AEC-Challenge-blind_test-clean and AEC-Challenge-blind_test-noisy are the public test sets of the AEC-challenge competition.

[0175] As shown in Table 2, we choose SDR, PESQ and STOI indicators to evaluate SE tasks. The explanation of each indicator is as follows:

[0176] AECMOS: Evaluates the mean opinion score (Acoustic Echo Cancellation Mean Opinion Score) of the AEC task. A higher score indicates better speech quality.

[0177] SDR: Source to Distortion Ratio (SDR) indicates the overall distortion of the signal. A higher score means more of the original speech information is retained.

[0178] PESQ: Perceptual Evaluation of Speech Quality (PESQ), with a value range of -0.5 to 4.5. A higher score indicates better auditory quality of the tested speech.

[0179] STOI: Short-Time Objective Intelligibility (STOI) measures the percentage of each word in speech that is understood or not understood. A higher score indicates that the speech is more understandable.

[0180] By comparing the traditional open-source AEC method of Webrtc, the DPCRN single-task AEC model with the same parameters, and the single-task SE model, it can be found that the model in this disclosure is better than the single-task model in AEC tasks, and its performance in SE tasks is almost the same as that of the single-task model.

[0181] Based on the above technical solutions, the present disclosure has at least the following technical effects:

[0182] First, the model in this disclosure can adapt to a wide range of delays (0-1000 milliseconds), reducing the need for separate delay estimation modules.

[0183] Second, through as rich data simulation as possible, we can basically cover a wide range of signal-to-noise ratio environments in multiple scenarios;

[0184] Third, one model can handle three audio signal processing tasks, reducing the mutual constraints caused by parameter adjustments of various sub-processes in related technologies, greatly reducing system complexity and power consumption;

[0185] Fourth, compared with the prior art solution of placing speech enhancement before speech endpoint detection (although this solution can input clean audio after speech enhancement into VAD, which is more helpful for VAD to distinguish speech, it will be subject to the effect of speech enhancement), the present disclosure couples the speech enhancement task and the speech endpoint detection task together during the training stage, allowing the speech enhancement task to assist the judgment of the speech endpoint detection classifier, thereby making it easier to improve the speech endpoint detection effect under low signal-to-noise ratio conditions.

[0186] Fifth, if the meeting content needs to be recorded, the present disclosure also supports turning on the VAD function, which automatically records the text content of the meeting voice through the downstream ASR module, making it convenient to summarize the meeting minutes after the meeting.

[0187] The present disclosure further provides a training device for an audio processing model. FIG15 shows a schematic structural diagram of the training device for an audio processing model in an exemplary embodiment of the present disclosure. As shown in FIG15 , the training device 1500 for an audio processing model may include a training sample set acquisition module 1510, a pre-training module 1520, and a joint training module 1530. Specifically:

[0188] The training sample set acquisition module 1510 is configured to acquire a training sample set; the training sample set includes a first sample set and a second sample set;

[0189] a pre-training module 1520 configured to pre-train a first branch network of the audio processing model to be trained using the first sample set to obtain a pre-trained first branch network, and to pre-train a second branch network of the audio processing model to be trained using the second sample set to obtain a pre-trained second branch network;

[0190] The joint training module 1530 is used to jointly train the pre-trained first branch network and the pre-trained second branch network using the training sample set to obtain the trained audio processing model; wherein, the first branch network is used to perform echo cancellation and speech enhancement tasks, and the second branch network is used to perform speech endpoint detection tasks.

[0191] In an exemplary embodiment of the present disclosure, the first sample set is obtained by:

[0192] Obtaining a third sample set and a fourth sample set; the third sample set or the fourth sample set includes a plurality of training samples, each training sample in the third sample set or the fourth sample set includes an audio signal transmitted from a far-end to a near-end, a far-end standard audio signal, and an audio signal collected from a near-end;

[0193] selecting a first target number of training samples from the third sample set and a second target number of training samples from the fourth sample set according to a preset ratio; wherein the ratio of the first target number to the second target number satisfies a preset ratio condition;

[0194] The first target number of training samples and the second target number of training samples are mixed to obtain the first sample set.

[0195] In an exemplary embodiment of the present disclosure, the third sample set is obtained by:

[0196] Randomly selecting a first clean audio and a second clean audio from a preset massive amount of clean audio samples, using the first clean audio as the audio signal transmitted from the far-end to the near-end, and using the second clean audio as the far-end standard audio signal;

[0197] Performing analog processing on the audio signal transmitted from the far end to the near end to obtain an analog audio signal;

[0198] The analog audio signal and the far-end standard audio signal are mixed to obtain the near-end collected audio signal.

[0199] In an exemplary embodiment of the present disclosure, the training sample set acquisition module 1510 performs analog processing on the audio signal transmitted from the far end to the near end to obtain an analog audio signal, including:

[0200] Adding a random delay within a preset time range to the audio signal transmitted from the far end to the near end to obtain a first transformed audio signal;

[0201] adding random noise to the first transformed audio signal to obtain a second transformed audio signal;

[0202] adding a nonlinear perturbation to the second transformed audio signal to obtain a third transformed audio signal;

[0203] Add random reverberation to the third transformed audio signal to obtain the analog audio signal.

[0204] In an exemplary embodiment of the present disclosure, the training sample set acquisition module 1510 mixes the analog audio signal and the far-end standard audio signal to obtain the near-end collected audio signal, including:

[0205] Adding random noise to the near-end standard audio signal to obtain a target audio signal;

[0206] The analog audio signal and the target audio signal are mixed to obtain the audio signal collected by the near end.

[0207] In an exemplary embodiment of the present disclosure, the fourth sample set is obtained by:

[0208] Randomly selecting a first clean audio signal and a second clean audio signal from a preset massive amount of clean audio samples, and using the second clean audio signal as the remote standard audio signal;

[0209] Randomly adding noise to the first clean audio to obtain the audio signal collected by the near end;

[0210] Using a non-all-zero silence signal having the same length as the audio signal collected by the near end as the audio signal transmitted from the far end to the near end;

[0211] The non-all-zero silence is a silence signal corresponding to a signal sequence consisting of 0s and 1s.

[0212] In an exemplary embodiment of the present disclosure, the first sample set includes an audio signal transmitted from the far end to the near end, a far end standard audio signal, and an audio signal collected by the near end;

[0213] The pre-training module 1520 pre-trains the first branch network of the audio processing model to be trained using the first sample set to obtain the pre-trained first branch network, including:

[0214] Performing signal conversion and synthesis processing on the audio signal collected by the near end and the audio signal transmitted from the far end to the near end through the first branch network to obtain an output signal;

[0215] determining a first loss value according to a degree of signal difference between the output signal and the far-end standard audio signal;

[0216] The first branch network is iteratively trained according to the first loss value to obtain the pre-trained first branch network.

[0217] In an exemplary embodiment of the present disclosure, the first branch network includes a Fourier transform unit, an encoding unit, a long-range dependency capture unit, and a decoding unit, and a residual jump connection structure is formed between the decoding unit and the encoding unit;

[0218] The pre-training module 1520 performs signal conversion and synthesis processing on the audio signal collected by the near end and the audio signal transmitted from the far end to the near end through the first branch network to obtain an output signal, including:

[0219] Performing Fourier transform on the audio signal transmitted from the far end to the near end by the Fourier transform unit to obtain a first transformed audio signal;

[0220] performing encoding processing on the first transformed audio signal by the encoding unit to obtain encoding information;

[0221] Capturing long-range dependency on the coding feature by the long-range dependency capturing unit to obtain long-range dependency information;

[0222] The residual information transmitted by the decoding unit and the encoding unit is decoded to obtain the output signal.

[0223] In an exemplary embodiment of the present disclosure, the pre-training module 1520 determines the first loss value according to the degree of signal difference between the output signal and the far-end standard audio signal, including:

[0224] determining a time domain loss according to a degree of signal difference between the output signal and the remote standard audio signal in the time domain;

[0225] determining a frequency domain loss according to a degree of signal difference between the output signal and the remote standard audio signal in the frequency domain;

[0226] The first loss value is determined according to the time domain loss and the frequency domain loss.

[0227] In an exemplary embodiment of the present disclosure, the pre-training module 1520 determines the time domain loss based on the degree of signal difference between the output signal and the far-end standard audio signal in the time domain, including:

[0228] Performing inverse Fourier transform on the output signal to obtain a time domain signal;

[0229] The time domain loss is determined based on a preset time domain loss calculation formula, the strength of the time domain signal, and the strength of the far-end standard audio signal.

[0230] In an exemplary embodiment of the present disclosure, the pre-training module 1520 determines the frequency domain loss based on the degree of signal difference between the output signal and the remote standard audio signal in the frequency domain, including:

[0231] The frequency domain loss is determined according to a ratio between a frequency domain conversion result of the time domain signal and a frequency domain conversion result of the remote standard audio signal.

[0232] In an exemplary embodiment of the present disclosure, the second sample set includes the far-end standard audio signal and a real speech recognition label corresponding to the far-end standard audio signal, wherein the real speech recognition label is used to indicate whether each frame signal of the far-end standard audio signal is a speech signal;

[0233] The pre-training module 1520 pre-trains the second branch network of the audio processing model to be trained using the second sample set to obtain a pre-trained second branch network, including:

[0234] Performing speech recognition on the remote standard audio signal through the second branch network to obtain a speech recognition label corresponding to each frame signal of the remote standard audio signal;

[0235] determining a second loss value according to a degree of difference between the speech recognition label and the true speech recognition label;

[0236] The second branch network is iteratively trained according to the second loss value to obtain the pre-trained second branch network.

[0237] In an exemplary embodiment of the present disclosure, the joint training module 1530 uses the training sample set to jointly train the pre-trained first branch network and the pre-trained second branch network to obtain a trained audio processing model, including:

[0238] Performing Fourier transform on the audio signal transmitted from the far end to the near end to obtain a target audio signal;

[0239] performing a mask operation on the target audio signal and the output signal to obtain a masked audio signal;

[0240] Performing speech recognition on the masked audio signal through the pre-trained second branch network to obtain a speech recognition label corresponding to each frame signal in the masked audio signal;

[0241] determining a third loss value according to a degree of difference between the speech recognition label and the true speech recognition label;

[0242] The pre-trained first branch network and the pre-trained second branch network are iteratively trained according to the third loss value to obtain a trained audio processing model.

[0243] The specific details of each module in the above-mentioned audio processing model training device have been described in detail in the corresponding audio processing model training method, so they will not be repeated here.

[0244] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0245] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0246] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0247] The present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device.

[0248] Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device.

[0249] Computer-readable storage media can transmit, propagate, or transfer programs for use by or in conjunction with an instruction execution system, apparatus, or device. Program code contained on a computer-readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wireline, optical cable, RF, or any suitable combination thereof.

[0250] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiments.

[0251] In addition, an electronic device capable of implementing the above method is also provided in an embodiment of the present disclosure.

[0252] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods, or program products. Therefore, various aspects of the present disclosure may be implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits," "modules," or "systems."

[0253] The electronic device 1600 according to this embodiment of the present disclosure is described below with reference to Figure 16. The electronic device 1600 shown in Figure 16 is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0254] As shown in FIG16 , electronic device 1600 is implemented as a general-purpose computing device. Components of electronic device 1600 may include, but are not limited to, the aforementioned at least one processing unit 1610, the aforementioned at least one storage unit 1620, a bus 1630 connecting various system components (including storage unit 1620 and processing unit 1610), and a display unit 1640.

[0255] The storage unit stores a program code, and the program code can be executed by the processing unit 1610, so that the processing unit 1610 performs the steps described in the "Exemplary Method" section of the present specification according to various exemplary embodiments of the present disclosure. For example, the processing unit 1610 can perform as shown in Figure 1: step S110, obtaining a training sample set; the training sample set includes a first sample set and a second sample set; step S120, using the first sample set to pre-train the first branch network of the audio processing model to be trained to obtain a pre-trained first branch network, and using the second sample set to pre-train the second branch network of the audio processing model to be trained to obtain a pre-trained second branch network; wherein the first branch network is used to perform echo cancellation and speech enhancement tasks, and the second branch network is used to perform speech endpoint detection tasks; step S130, using the training sample set to jointly train the pre-trained first branch network and the pre-trained second branch network to obtain a trained audio processing model.

[0256] The storage unit 1620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 16201 and / or a cache memory unit 16202 , and may further include a read-only memory unit (ROM) 16203 .

[0257] The storage unit 1620 may also include a program / utility 16204 having a set (at least one) of program modules 16205, such program modules 16205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0258] Bus 1630 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0259] Electronic device 1600 can also communicate with one or more external devices 1700 (e.g., a keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with electronic device 1600, and / or any device that enables electronic device 1600 to communicate with one or more other computing devices (e.g., a router, modem, etc.). Such communication can occur via input / output (I / O) interface 1650. Furthermore, electronic device 1600 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via network adapter 1660. As shown, network adapter 1660 communicates with other modules of electronic device 1600 via bus 1630. It should be understood that, although not shown, other hardware and / or software modules can be used in conjunction with electronic device 1600, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0260] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow from the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.

Claims

1. A training method for an audio processing model, wherein, Including: Obtaining a training sample set; The training sample set includes a first sample set and a second sample set; Using the first sample set to pre-train the first branch network of the audio processing model to be trained, obtaining a pre-trained first branch network, and using the second sample set to pre-train the second branch network of the audio processing model to be trained, obtaining a pre-trained second branch network; Using the training sample set to jointly train the pre-trained first branch network and the pre-trained second branch network, obtaining the trained audio processing model; Wherein, the first branch network is used to perform echo cancellation and speech enhancement tasks, and the second branch network is used to perform voice activity detection tasks.

2. The method according to claim 1, wherein, The first sample set is obtained by the following method: Obtaining a third sample set and a fourth sample set; multiple training samples are included in the third sample set or the fourth sample set, and each training sample in the third sample set or the fourth sample set includes an audio signal transmitted from the far end to the near end, a far-end standard audio signal, and an audio signal collected at the near end; Selecting a first target number of training samples from the third sample set and a second target number of training samples from the fourth sample set according to a preset proportional relationship; The ratio of the first target number to the second target number satisfies a preset ratio condition; Mixing the first target number of training samples and the second target number of training samples to obtain the first sample set.

3. The method according to claim 2, wherein The third sample set is obtained by the following method: Randomly selecting a first clean audio and a second clean audio from a preset large number of clean audio samples, using the first clean audio as the audio signal transmitted from the far end to the near end, and using the second clean audio as the far-end standard audio signal; Performing analog processing on the audio signal transmitted from the far end to the near end to obtain an analog audio signal; Mixing the analog audio signal and the far-end standard audio signal to obtain the audio signal collected at the near end.

4. The method according to claim 3, wherein The performing analog processing on the audio signal transmitted from the far end to the near end to obtain an analog audio signal includes: Adding a random time delay within a preset time range to the audio signal transmitted from the far end to the near end to obtain a first transformed audio signal; Adding random noise to the first transformed audio signal to obtain a second transformed audio signal; Adding a non-linear perturbation to the second transformed audio signal to obtain a third transformed audio signal; Adding random reverberation to the third transformed audio signal to obtain the analog audio signal.

5. The method according to claim 3, wherein The mixing the analog audio signal and the far-end standard audio signal to obtain the audio signal collected at the near end includes: Adding random noise to the near-end standard audio signal to obtain a target audio signal; Mixing the analog audio signal and the target audio signal to obtain the audio signal collected at the near end.

6. The method according to claim 2, wherein The fourth sample set is obtained by the following method: Randomly selecting a first clean audio and a second clean audio from a preset large number of clean audio samples, and using the second clean audio as the far-end standard audio signal; Add random noise to the first clean audio to obtain the audio signal collected proximally; Use a non-all-zero silence signal with the same length as the audio signal collected proximally as the audio signal transmitted distally to proximally; The non-all-zero silence is the silence signal corresponding to the signal sequence composed of 0 and 1.

7. The method according to claim 1, wherein The first sample set includes the audio signal transmitted distally to proximally, the distal standard audio signal, and the audio signal collected proximally; Pre-training the first branch network of the audio processing model to be trained using the first sample set to obtain a pre-trained first branch network, including: Performing signal transformation and synthesis processing on the audio signal collected proximally and the audio signal transmitted distally to proximally through the first branch network to obtain an output signal; Determine a first loss value according to the signal difference degree between the output signal and the distal standard audio signal; Iteratively train the first branch network according to the first loss value to obtain the pre-trained first branch network.

8. The method according to claim 7, wherein, The first branch network includes a Fourier transform unit, an encoding unit, a long-range dependence capture unit, and a decoding unit, and there is a residual skip connection structure between the decoding unit and the encoding unit; The performing signal transformation and synthesis processing on the audio signal collected proximally and the audio signal transmitted distally to proximally through the first branch network to obtain an output signal includes: Performing Fourier transform on the audio signal transmitted distally to proximally through the Fourier transform unit to obtain a first transformed audio signal; Performing encoding processing on the first transformed audio signal through the encoding unit to obtain encoded information; Performing long-range dependence capture on the encoded feature through the long-range dependence capture unit to obtain long-range dependence information; Performing decoding processing on the residual information passed by the decoding unit and the encoding unit to obtain the output signal.

9. The method according to claim 7, wherein The determining a first loss value according to the signal difference degree between the output signal and the distal standard audio signal includes: Determine the time-domain loss according to the signal difference degree between the output signal and the distal standard audio signal in the time domain; Determine the frequency-domain loss according to the signal difference degree between the output signal and the distal standard audio signal in the frequency domain; Determine the first loss value according to the time-domain loss and the frequency-domain loss.

10. The method according to claim 9, wherein The determining the time-domain loss according to the signal difference degree between the output signal and the distal standard audio signal in the time domain includes: Performing inverse Fourier transform on the output signal to obtain a time-domain signal; Determine the time-domain loss based on a preset time-domain loss calculation formula, the intensity of the time-domain signal, and the intensity of the distal standard audio signal.

11. The method according to claim 9, wherein, The determining the frequency-domain loss according to the signal difference degree between the output signal and the distal standard audio signal in the frequency domain includes: Determine the frequency-domain loss according to the ratio between the frequency-domain conversion result of the time-domain signal and the frequency-domain conversion result of the distal standard audio signal.

12. The method according to claim 7, wherein, The second sample set contains the remote standard audio signal and the true speech recognition label corresponding to the remote standard audio signal, and the true speech recognition label is used to characterize whether each frame signal of the remote standard audio signal is a speech signal; The pre-training of the second branch network of the audio processing model to be trained by using the second sample set to obtain a pre-trained second branch network includes: Performing speech recognition on the remote standard audio signal through the second branch network to obtain a speech recognition label corresponding to each frame signal of the remote standard audio signal; Determining a second loss value according to the difference degree between the speech recognition label and the true speech recognition label; Performing iterative training on the second branch network according to the second loss value to obtain the pre-trained second branch network.

13. The method according to claim 12, wherein, The joint training of the pre-trained first branch network and the pre-trained second branch network by using the training sample set to obtain a trained audio processing model includes: Performing Fourier transform on the audio signal remotely transmitted to the proximal end to obtain a target audio signal; Performing a masking operation on the target audio signal and the output signal to obtain a masked audio signal; Performing speech recognition on the masked audio signal through the pre-trained second branch network to obtain a speech recognition label corresponding to each frame signal in the masked audio signal; Determining a third loss value according to the difference degree between the speech recognition label and the true speech recognition label; Performing iterative training on the pre-trained first branch network and the pre-trained second branch network according to the third loss value to obtain a trained audio processing model.

14. A training device for an audio processing model, wherein, Including: A training sample set acquisition module for acquiring a training sample set; The training sample set includes a first sample set and a second sample set; A pre-training module for pre-training the first branch network of the audio processing model to be trained by using the first sample set to obtain a pre-trained first branch network, and pre-training the second branch network of the audio processing model to be trained by using the second sample set to obtain a pre-trained second branch network; A joint training module for jointly training the pre-trained first branch network and the pre-trained second branch network by using the training sample set to obtain the trained audio processing model; Wherein, the first branch network is used to perform echo cancellation and speech enhancement tasks, and the second branch network is used to perform speech endpoint detection tasks.

15. A computer storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the training method of the audio processing model according to any one of claims 1 to 13.

16. An electronic device, wherein, Including: A processor; And A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the training method of the audio processing model according to any one of claims 1 to 13 by executing the executable instructions.

Citation Information

Patent Citations

  • Voice activity detection method combined with voice enhancement

    CN113113049A

  • Speech enhancement model training method, speech enhancement model recognition method, electronic equipment and storage medium

    CN114283795A

  • Model training method, echo cancellation method, system and device and storage medium

    CN114530160A

  • Counterfeit voice detection method and device combining time domain and frequency domain, equipment and medium

    CN116092503A

  • Echo cancellation model training method and device, equipment and storage medium

    CN117219107A

Cited By

  • Voice endpoint detection method, related device, equipment and medium

    CN121438875A