Speech enhancement model training method and device, storage medium and equipment

By generating and processing room impulse responses and clean speech data in training samples and using control curves to preserve early reverberation, the problems of speech quality damage and data alignment in speech enhancement model training are solved, achieving more efficient speech enhancement effects.

CN116189698BActive Publication Date: 2025-09-12GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111427538.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-25
Publication Date
2025-09-12
Estimated Expiration
2041-11-25

AI Technical Summary

Technical Problem

Existing speech enhancement models are prone to voice quality damage during training, and improper reverberation processing can lead to the model being unable to be trained or having poor results.

Method used

By obtaining the room impulse response and clean speech data, training samples are generated. The room impulse response is processed using a control curve to retain early reverberation and reduce late reverberation. Target data is generated through convolution to train the speech enhancement model.

Benefits of technology

It effectively reduces distortion after signal processing, solves the problem of aligning training data and target data, preserves early reverberation, and improves voice quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189698B_ABST
    Figure CN116189698B_ABST
Patent Text Reader

Abstract

The present application provides a method and apparatus, storage medium, and equipment for training a speech enhancement model, relating to the field of speech enhancement technology. The method comprises: obtaining N groups of training samples, wherein the i-th group of training samples comprises: i-th training data and i-th target data; training a speech enhancement model using the N groups of training samples; obtaining the i-th group of training samples comprises: obtaining the i-th room impulse response and i-th clean speech data, processing the i-th room impulse response and the i-th clean speech data to obtain the i-th training data in the i-th group of training samples; determining the i-th control curve based on the i-th room impulse response, multiplying the i-th control curve with the i-th room impulse response to obtain the i'th room impulse response; and convolving the i-th clean speech data with the i'th room impulse response to obtain the i-th target data in the i-th group of training samples. The present application can reduce distortion after signal processing and solve the problem of aligning training data and target data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of speech enhancement technology, and in particular to a method and device for training a speech enhancement model, a readable storage medium, and an electronic device. Background Art

[0002] The goal of speech enhancement is to improve speech quality through various algorithms, extracting the purest possible speech signal from a signal containing interfering noise. Commonly used speech enhancement algorithms include: spectral subtraction-based speech enhancement algorithms, wavelet analysis-based speech enhancement algorithms, Kalman filtering-based speech enhancement algorithms, signal subspace-based enhancement methods, speech enhancement model training methods based on auditory masking effects, speech enhancement model training methods based on independent component analysis, and speech enhancement model training methods based on neural networks.

[0003] Training a speech enhancement model using a neural network inevitably causes significant damage to the speech signal, resulting in a decrease in speech quality. Furthermore, when using a clean speech signal and room impulse response to generate training data, the target signal can lead the input signal, ultimately rendering the model unfeasible and impossible to train.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention

[0005] The purpose of the present disclosure is to provide a training method and device for a speech enhancement model, a readable storage medium and an electronic device, which can at least to some extent overcome the shortcomings of the related art in which early reverberation is not retained during dereverberation and the model is relatively complex, resulting in significant damage to the speech.

[0006] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0007] According to a first aspect of the present disclosure, a method for training a speech enhancement model is provided, the method comprising: obtaining N groups of training samples, wherein the i-th group of training samples comprises: i-th training data and i-th target data, wherein N is a positive integer and i is a positive integer not greater than N; training a speech enhancement model using the N groups of training samples; wherein obtaining the i-th group of training samples comprises: obtaining an i-th room impulse response and i-th clean speech data, processing the i-th room impulse response and the i-th clean speech data to obtain the i-th training data in the i-th group of training samples; determining an i-th control curve based on the i-th room impulse response, multiplying the i-th control curve by the i-th room impulse response to obtain an i'th room impulse response, wherein i' is a positive integer not greater than N; and convolving the i-th clean speech data with the i'th room impulse response to obtain the i-th target data in the i-th group of training samples.

[0008] In one embodiment of the present disclosure, the i-th room impulse response includes multiple sampling points; the i-th control curve includes multiple control values, and the number of the control values ​​is the same as the number of sampling points in the i-th room impulse response; when the sampling point value at the tail of the i'th room impulse response is zero or the absolute value is very small, the tail truncation processing can be selected.

[0009] In one embodiment of the present disclosure, determining the i-th control curve based on the i-th room impulse response includes: determining an absolute value of each sampling point in the i-th room impulse response, wherein the absolute values ​​include multiple maximum values ​​with equal values; determining, in the i-th room impulse response, a sampling point corresponding to a first maximum value among the absolute values ​​as a peak position point; and determining a control value in the i-th control curve corresponding to the peak position point of the i-th room impulse response as a main control value of the i-th control curve.

[0010] In one embodiment of the present disclosure, the method for determining the i-th control curve based on the i-th room impulse response further includes: adjusting a control value of the i-th control curve by a parameter to determine the i-th control curve; wherein, other control values ​​except the main control value are not greater than the main control value.

[0011] In one embodiment of the present disclosure, the processing of the i-th room impulse response and the i-th clean speech data to obtain the i-th training data in the i-th group of training samples includes: convolving the i-th clean speech data with the i-th room impulse response to obtain the i-th training data in the i-th group of training samples.

[0012] In one embodiment of the present disclosure, the processing of the i-th room impulse response and the i-th clean speech data to obtain the i-th training data in the i-th group of training samples includes: convolving the i-th clean speech data with the i-th room impulse response and adding the convolution operation with noise data to obtain the i-th training data in the i-th group of training samples.

[0013] In one embodiment of the present disclosure, the processing of the i-th room impulse response and the i-th clean speech data to obtain the i-th training data in the i-th group of training samples includes: adding the i-th clean speech data to the noise data, and convolving the data with the i-th room impulse response to obtain the i-th training data in the i-th group of training samples.

[0014] According to a second aspect of the present disclosure, a training device for a speech enhancement model is provided, the device comprising: an acquisition module for acquiring N groups of training samples, wherein the i-th group of training samples comprises i-th training data and i-th target data, wherein N is a positive integer and i is a positive integer not greater than N; a training module for training the speech enhancement model using the N groups of training samples; wherein the acquisition module is specifically configured to: acquire an i-th room impulse response and i-th clean speech data, process the i-th room impulse response and the i-th clean speech data to obtain the i-th training data; determine an i-th control curve based on the i-th room impulse response, multiply the i-th control curve by the i-th room impulse response to obtain an i'th room impulse response, wherein i' is a positive integer not greater than N; and convolve the i-th clean speech data with the i'th room impulse response to obtain the i-th target data in the i-th group of training samples.

[0015] According to a third aspect of the present disclosure, a terminal is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the training method of the speech enhancement model of the first aspect when executing the computer program.

[0016] According to a fourth aspect of the present disclosure, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the training method of the speech enhancement model of the first aspect is implemented.

[0017] The speech enhancement model training method and device, readable storage medium, and electronic device provided by the embodiments of the present disclosure have the following technical effects:

[0018] In the training process of the speech enhancement model provided in the embodiments of the present disclosure, N groups of training samples are obtained, where the i-th group of training samples includes: i-th training data and i-th target data; the speech enhancement model is trained using the N groups of training samples; obtaining the i-th group of training samples includes: obtaining the i-th room impulse response and i-th clean speech data, processing the i-th room impulse response and the i-th clean speech data to obtain the i-th training data in the i-th group of training samples; determining the i-th control curve based on the i-th room impulse response, multiplying the i-th control curve with the i-th room impulse response to obtain the i'th room impulse response, where i' is a positive integer not greater than N; and convolving the i-th clean speech data with the i'th room impulse response to obtain the i-th target data in the i-th group of training samples. The present application can reduce distortion after signal processing, solve the alignment problem of training data and target data, and preserve early reverberation.

[0019] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0021] Figure 1 The present invention schematically shows a flow chart of a method for training a speech enhancement model provided by an embodiment of the present disclosure;

[0022] Figure 2 shows a schematic diagram of a speech enhancement model;

[0023] Figure 3 A schematic diagram of the control curve is shown;

[0024] Figure 4 Schematic diagram showing the impulse response of the i-th room;

[0025] Figure 5 Schematic diagram showing the impulse response of the i'th room;

[0026] Figure 6 The structure of a training device for a speech enhancement model provided by an embodiment of the present disclosure is schematically shown;

[0027] Figure 7 A block diagram of an electronic device provided by an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0028] In order to make the objectives, technical solutions and advantages of the present disclosure more clear, the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings.

[0029] In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0030] In the description of the present disclosure, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and cannot be understood as indicating or implying relative importance. For those skilled in the art, the specific meanings of the above terms in the present disclosure can be understood according to specific circumstances. In addition, in the description of the present disclosure, unless otherwise specified, "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship.

[0031] Below, each step of the method for training the speech enhancement model in this example implementation will be described in more detail with reference to the accompanying drawings and examples.

[0032] in, Figure 1 The flowchart of the training method of the speech enhancement model in an exemplary embodiment of the present disclosure is schematically shown. Figure 1 , the method comprises the following steps:

[0033] S101, obtaining the i-th group of training samples, obtaining a total of N groups of training samples, wherein the i-th group of training samples includes: i-th training data and i-th target data, wherein N is a positive integer, and i is a positive integer not greater than N.

[0034] S102: Train a speech enhancement model using N groups of training samples.

[0035] S11, wherein obtaining the i-th group of training samples includes: obtaining the i-th room impulse response and the i-th clean speech data, and processing the i-th room impulse response and the i-th clean speech data to obtain the i-th training data in the i-th group of training samples.

[0036] S12, determining an i-th control curve based on the i-th room impulse response, multiplying the i-th control curve by the i-th room impulse response to obtain an i'th room impulse response, where i' is a positive integer not greater than N; convolving the i-th clean speech data with the i'th room impulse response to obtain an i-th target data in the i-th group of training samples.

[0037] The multiplication of the i-th control curve and the i-th room impulse response refers to multiplying the control value of each control point in the i-th control curve by the corresponding sampling point in the i-th room impulse response. Obviously, for control points in the i-th control curve with a control value of 1, the multiplication operation can be omitted. Similarly, for sampling points in the i-th room impulse response for which no operation is performed or no operation is required, the control value of the corresponding control point can be considered to be 1. In implementation, control points with a control value of 1 can be omitted, that is, the number of points in the i-th control curve and the number of samples in the i-th room impulse response are not necessarily strictly equal, but are only equal in a general sense.

[0038] According to the convolution theorem, when the tail sampling point value of the impulse response of room i' is zero or very small, tail truncation can be performed. After tail truncation, the number of sampling points is reduced accordingly. Obviously, when the impulse response of room i' is obtained by tail truncation of the impulse response of room i, it can be regarded as setting the control value of the corresponding tail control point in the i-th control curve to 0.

[0039] exist Figure 1 During the training process of the speech enhancement model provided by the illustrated embodiment, N groups of training samples are obtained, and the i-th group of training samples includes: i-th training data and i-th target data, where N is a positive integer and i is a positive integer not greater than N; the speech enhancement model is trained using the N groups of training samples; obtaining the i-th group of training samples includes: obtaining the i-th room impulse response and i-th clean speech data, processing the i-th room impulse response and the i-th clean speech data to obtain the i-th training data in the i-th group of training samples; determining the i-th control curve based on the i-th room impulse response, multiplying the i-th control curve with the i-th room impulse response to obtain the i'th room impulse response, where i' is a positive integer not greater than N; and convolving the i-th clean speech data with the i'th room impulse response to obtain the i-th target data in the i-th group of training samples. The present application can reduce distortion after signal processing and solve the alignment problem of training data and target data, and can retain a certain degree of early reverberation as needed.

[0040] Figure 2 The following is a schematic diagram of the speech enhancement model. Figure 2 right Figure 1 The specific implementation methods of each step included in the embodiment shown are described in detail. This embodiment provides a training method for a speech enhancement model, and the specific implementation method is as follows:

[0041] In S101, the i-th group of training samples is obtained to obtain N groups of training samples, wherein the i-th group of training samples includes: i-th training data and i-th target data, wherein N is a positive integer, and i is a positive integer not greater than N.

[0042] In S102, a speech enhancement model is trained using N groups of training samples.

[0043] In S11, obtaining the i-th group of training samples includes: obtaining the i-th room impulse response and the i-th clean speech data, and processing the i-th room impulse response and the i-th clean speech data to obtain the i-th training data in the i-th group of training samples.

[0044] In an exemplary embodiment, as Figure 2 As shown, the i-th clean speech data and the i-th room impulse response are randomly obtained from the sample library, and convolution processing is performed on the i-th clean speech data and the i-th room impulse response to obtain the i-th training data. The i-th training data and the i-th target data obtained in the following embodiment will be used together as the i-th group of training samples.

[0045] In S12, an i-th control curve is determined based on the i-th room impulse response, and the i-th control curve is multiplied by the i-th room impulse response to obtain an i'th room impulse response, where i' is a positive integer not greater than N; the i-th clean speech data is convolved with the i'th room impulse response to obtain an i-th target data in the i-th group of training samples.

[0046] In an exemplary embodiment, first, the i-th group of training samples is obtained, the i-th group of training samples includes the i-th training data and the i-th target data, and multiple i-th groups of training samples form N groups of training samples. The method for obtaining the i-th target data is as follows: Figure 2 As shown, after obtaining the i-th clean speech data and the i-th room impulse response, S12 is executed. The specific execution method is referred to the following embodiment:

[0047] In an exemplary embodiment, after obtaining the i-th clean speech data and the i-th room impulse response, the first sampling point with the largest absolute value in the i-th room impulse response is determined, i.e., the peak position point. A corresponding peak point must also exist in the i-th control curve. Therefore, the control value point in the i-th control curve corresponding to the peak position point of the i-th room impulse response is recorded as the main control point p, and the control value corresponding to the main control point p is recorded as the main control value.

[0048] Reverberation consists of early and late reflections. During reverberation removal, early reflections enhance the speech signal and should therefore be retained. Therefore, parameters can be used to control the duration and intensity of reverberation in the room impulse response, selectively preserving different stages of reverberation.

[0049] In an exemplary embodiment, an i-th control curve can be generated using preset parameters and multiplied by the values ​​of each sampling point in the i-th room impulse response to generate an i'th room impulse response containing only direct sound and / or only early reverberation. Because sound propagation requires a certain amount of time, and the intensity of direct sound and early reflections is greater than that of late reflections, the amplitude of the sampling points of the i'th room impulse response is initially zero or very small, then rapidly increases, reaching a maximum amplitude during the direct sound and / or early reflection period, and then gradually decreases during the late reflection period. Therefore, the control values ​​of the i-th control curve can also vary according to this pattern. The aforementioned main control point p is located near the direct sound or early reflections, and the control values ​​in the i-th control curve are no greater than the main control point p.

[0050] In an exemplary embodiment, the control values ​​in the aforementioned i-th control curve can be adjusted using parameters to control the amplitude, location, and duration of the direct sound, early reverberation, and late reverberation in the impulse response of the i'th room. During adjustment, the control values ​​in the i-th control curve can be adjusted as needed, meaning that the shape of the i-th control curve can be arbitrarily varied. This is not a limitation in this embodiment, and only two or three examples are provided below for illustrative purposes. This does not represent all feasible solutions. For example, a bell-shaped curve similar to a normal distribution or a Gaussian distribution can also be used.

[0051] Figure 3 In an exemplary embodiment, since the amplitude of the sampling point of the impulse response of the i'th room is zero or very small at the beginning, it can be ignored. For this part, the control value can be set to 0, as shown in FIG. Figure 3 The left dotted line portion is shown. The amplitude of the impulse response of the i'th room before the main control point p shows a rapid upward trend. When controlling this section, the i'th control curve can be recorded as m, and a section of m can be recorded as m1. The control value parameters can be adjusted. For example, m1 can be preset as a section of an exponentially rising curve, such as Figure 3 As shown. After that, the impulse response of the i'th room enters the early stage of direct sound and early reflection sound. The section of the control curve m is recorded as m2. The amplitude of this section of the impulse response of the i'th room can be kept unchanged, that is, the control values ​​of the control points of the m2 section are all set to 1. At this time, m2 is a straight line, as shown Figure 3 shown.

[0052] In an exemplary embodiment, after the main control point p, the i-th control curve is used to control the early reverberation and the late reverberation. The section of the i-th control curve that controls the early reverberation is denoted as m3, and the amplitude of m3 is not changed. That is, the control values ​​of the control points in the m3 section are all set to 1. At this time, m3 is a straight line, as shown in FIG. Figure 3As shown. The curve for controlling the late reverberation in the i-th control curve is recorded as m4. The control parameters can also be adjusted. For example, m4 can be preset as a curve of exponential decay, as shown in Figure 3 As shown. The duration of m3 can be set to the length of the early reverberation that needs to be retained, and the duration of m4 can be set to the required reverberation time, such as T60. After m4, corresponding to the part where the amplitude of the impulse response of the i'th room gradually decays to 0, the corresponding control value parameter in the control curve can be directly set to 0, as shown Figure 3 According to the above embodiment, the following is generated: Figure 3 The curve shown.

[0053] In an exemplary embodiment, the parameters of segment m1 can also be set to all zeros (or the duration of segment m1 can be set to 0), or m1 can be changed to a linearly varying line. The parameters of segment m4 can also be set to all zeros (or the duration of segment m4 can be set to 0), or m4 can be changed to a linearly decaying line. Segments m2 and m3 can also be set to linearly varying lines or exponential curves.

[0054] In the exemplary embodiment, in addition to the examples above, the shapes and lengths of m1, m2, m3, and m4 (i.e., the control of the amplitude and duration of the different stages of the impulse response of the i'th room) can be adjusted according to actual conditions and there are no fixed rules. However, to ensure the effectiveness of speech enhancement, the control value corresponding to the main control point p must remain the maximum value in the i'th control curve, and the control values ​​of the remaining control points must not exceed the main control value. Furthermore, the segmentation of the i'th control curve is not limited to m1, m2, m3, and m4; the number of segments can be increased or decreased arbitrarily, without limitation.

[0055] In an exemplary embodiment, after generating the i-th control curve, the control value corresponding to the control point in the i-th control curve is multiplied by the value of the corresponding sampling point in the i-th room impulse response to obtain the controlled i'th room impulse response.

[0056] In an exemplary embodiment, after obtaining the impulse response for the i'th room, samples with zero or very small amplitudes at the tail end can be deleted and truncated. This process reduces the number of samples in the impulse response for the i'th room. According to the convolution theorem, this process does not affect the convolution result and can save storage and computing resources.

[0057] Figure 4 shows a schematic diagram of the impulse response of the ith room, Figure 5 Schematic diagram of the impulse response of the i'th room is shown. Figure 4The impulse response of the ith room shown in FIG. 1 has a long reverberation duration and a certain amount of noise at the head and tail. By using the method described in the above embodiment, the late reverberation in the impulse response of the ith room can be removed while retaining the early reverberation. After processing the impulse response of the ith room using the ith control curve, the impulse response of the ith room is obtained as shown in FIG. Figure 5 As shown in FIG, only the direct sound and early reverberation are retained in the impulse response of the i'th room, and the noise at the head and tail of the impulse response of the i'th room is removed.

[0058] In an exemplary embodiment, the i-th clean speech data is convolved with the i'th room impulse response to obtain the i-th target data. The i-th target data is the target label data during the training of the speech enhancement model.

[0059] In an exemplary embodiment, the i-th clean speech data and the i-th room impulse response are convolved to obtain the i-th training data in the i-th group of training samples, which together with the i-th target data serve as the i-th group of training samples of the neural network and are input into the neural network, such as Figure 2 As shown in FIG. , the speech enhancement model is continuously trained using the i-th group of training samples until the output of the model can achieve excellent speech enhancement results.

[0060] In an exemplary embodiment, the i-th training data input to the neural network may include not only reverberant speech with a room impulse response, but also noise. For example, in the above embodiment, the i-th training data in the i-th group of training samples may also include: convolving the i-th clean speech data with the i-th room impulse response and adding the resultant noise data to obtain the i-th training data; or adding the i-th clean speech data with the noise data and convolving the resultant noise data with the i-th room impulse response to obtain the i-th training data. Whether or not to add noise depends on whether the model requires noise reduction capabilities. Furthermore, to enrich the training samples, the i-th clean speech data and noise data may also be randomly scaled in amplitude.

[0061] In an exemplary embodiment, when removing reverberation and noise, there is no limitation on the signal enhancement method used. For example, it can be any one of the following methods: ideal binary mask (IBM), ideal ratio mask (IRM), ideal amplitude mask (IAM), phase-shifting mask (PSM), complex ideal ratio mask (CIRM), etc.

[0062] In an exemplary embodiment, the above-mentioned neural network can be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory neural network (LSTM), etc., without limitation herein.

[0063] The following are embodiments of the apparatus disclosed herein, which can be used to implement the method embodiments disclosed herein. For details not disclosed in the apparatus embodiments disclosed herein, please refer to the method embodiments disclosed herein.

[0064] in, Figure 6 FIG1 shows a structural diagram of a training device for a speech enhancement model according to an exemplary embodiment of the present disclosure. Figure 6 The training device of the speech enhancement model shown in the figure can be implemented as all or part of the terminal through software, hardware or a combination of both, and can also be integrated into the server as an independent module.

[0065] In an exemplary embodiment, the speech enhancement model training device 600 includes: an acquisition module 601 and a training module 602, wherein:

[0066] An acquisition module 601 is configured to acquire N groups of training samples, wherein the i-th group of training samples includes the i-th training data and the i-th target data, wherein N is a positive integer and i is a positive integer not greater than N;

[0067] A training module 602 is configured to train a speech enhancement model using N sets of training samples.

[0068] The acquisition module 601 is specifically configured to: acquire an i-th room impulse response and an i-th clean speech data, process the i-th room impulse response and the i-th clean speech data to obtain an i-th training data in the i-th group of training samples; determine an i-th control curve based on the i-th room impulse response, and multiply the i-th control curve by the i-th room impulse response to obtain an i'th room impulse response, where i' is a positive integer not greater than N; and convolve the i-th clean speech data with the i'th room impulse response to obtain an i-th target data in the i-th group of training samples.

[0069] It should be noted that the data synchronization device provided in the above embodiment only uses the division of the above functional modules as an example to illustrate the training method of the speech enhancement model. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the speech enhancement model training device provided in the above embodiment and the speech enhancement model training method embodiment belong to the same concept. Therefore, for details not disclosed in the embodiment of the device disclosed in the present invention, please refer to the embodiment of the speech enhancement model training method disclosed in the present invention, and no further details will be given here.

[0070] The serial numbers of the above-mentioned embodiments of the present disclosure are for description only and do not represent the advantages or disadvantages of the embodiments.

[0071] The present disclosure also provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the aforementioned methods. The readable storage medium may include, but is not limited to, any type of disk, including a floppy disk, an optical disk, a DVD, a CD-ROM, a microdrive, a magneto-optical disk, a ROM, a RAM, an EPROM, an EEPROM, a DRAM, a VRAM, a flash memory device, a magnetic card or an optical card, a nanosystem (including a molecular memory IC), or any type of medium or device suitable for storing instructions and / or data.

[0072] An embodiment of the present disclosure further provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above-mentioned method embodiments when executing the program.

[0073] Figure 7 Schematically shows a structural diagram of an electronic device according to an exemplary embodiment of the present disclosure. Figure 7 As shown, the electronic device 700 includes a processor 701 and a memory 702 .

[0074] In the embodiment of the present disclosure, the processor 701 is the control center of the computer system, which can be the processor of a physical machine or the processor of a virtual machine. The processor 701 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 701 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 701 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state.

[0075] In the embodiment of the present disclosure, the processor 701 is specifically configured to:

[0076] Obtain N groups of training samples, wherein the i-th group of training samples includes: i-th training data and i-th target data, wherein N is a positive integer and i is a positive integer not greater than N; train a speech enhancement model using the N groups of training samples; wherein obtaining the i-th group of training samples includes: obtaining an i-th room impulse response and i-th clean speech data, processing the i-th room impulse response and the i-th clean speech data to obtain i-th training data in the i-th group of training samples; determining an i-th control curve based on the i-th room impulse response, multiplying the i-th control curve by the i-th room impulse response to obtain an i'-th room impulse response, wherein i' is a positive integer not greater than N; and convolving the i-th clean speech data with the i'-th room impulse response to obtain i-th target data in the i-th group of training samples.

[0077] Furthermore, in one embodiment of the present disclosure, the above-mentioned i-th room impulse response includes multiple sampling points; the above-mentioned i-th control curve includes multiple control values, and the number of the above-mentioned control values ​​is the same as the number of sampling points of the above-mentioned i-th room impulse response; when the sampling point value at the tail of the above-mentioned i'th room impulse response is zero or the absolute value is very small, the tail truncation processing can be selected.

[0078] Optionally, in the above-mentioned i-th control curve, determining the i-th control curve based on the above-mentioned i-th room impulse response includes: determining the absolute value of each sampling point in the above-mentioned i-th room impulse response, wherein the above-mentioned absolute value includes multiple maximum values ​​with equal values; in the above-mentioned i-th room impulse response, determining the sampling point corresponding to the first maximum value in the above-mentioned absolute value as the peak position point; and determining the control value in the above-mentioned i-th control curve corresponding to the peak position point of the above-mentioned i-room impulse response as the main control value of the above-mentioned i-control curve.

[0079] Optionally, the above-mentioned method of determining the i-th control curve based on the i-th room impulse response further includes: adjusting the control value of the i-th control curve by parameters to determine the i-th control curve; wherein, except for the above-mentioned main control point, the control values ​​corresponding to the other control points are not greater than the above-mentioned main control value.

[0080] Optionally, the processing of the i-th room impulse response and the i-th clean speech data to obtain the i-th training data in the i-th group of training samples includes: convolving the i-th clean speech data with the i-th room impulse response to obtain the i-th training data in the i-th group of training samples.

[0081] Optionally, the processing of the i-th room impulse response and the i-th clean speech data to obtain the i-th training data in the i-th group of training samples includes: convolving the i-th clean speech data with the i-th room impulse response and adding the convolution operation with noise data to obtain the i-th training data in the i-th group of training samples.

[0082] Optionally, the processing of the i-th room impulse response and the i-th clean speech data to obtain the i-th training data in the i-th group of training samples includes: adding the i-th clean speech data to the noise data, and convolving the data with the i-th room impulse response to obtain the i-th training data in the i-th group of training samples.

[0083] The memory 702 may include one or more readable storage media, which may be non-transitory. The memory 702 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments of the present disclosure, the non-transitory readable storage medium in the memory 702 is used to store at least one instruction, which is used to be executed by the processor 701 to implement the method in the embodiment of the present disclosure.

[0084] In some embodiments, the electronic device 700 further includes a peripheral device interface 703 and at least one peripheral device. The processor 701, memory 702, and peripheral device interface 703 may be connected via a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 703 via a bus, signal lines, or circuit boards. Specifically, the peripheral device includes at least one of a display screen 704, a camera 707, and an audio circuit 706.

[0085] The peripheral device interface 703 can be used to connect at least one input / output (I / O)-related peripheral device to the processor 701 and the memory 702. In some embodiments of the present disclosure, the processor 701, the memory 702, and the peripheral device interface 703 are integrated on the same chip or circuit board; in some other embodiments of the present disclosure, any one or two of the processor 701, the memory 702, and the peripheral device interface 703 can be implemented on separate chips or circuit boards. This is not specifically limited in the present disclosure.

[0086] The display screen 704 is used to display a user interface (UI). The UI may include graphics, text, icons, videos, or any combination thereof. When the display screen 704 is a touch screen display, the display screen 704 is also capable of collecting touch signals on or above the surface of the display screen 704. The touch signals can be input as control signals to the processor 701 for processing. In this case, the display screen 704 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments of the present disclosure, there can be one display screen 704, provided on the front panel of the electronic device 700; in other embodiments of the present disclosure, there can be at least two display screens 704, provided on different surfaces of the electronic device 700 or in a foldable design; in still other embodiments of the present disclosure, the display screen 704 can be a flexible display, provided on a curved or foldable surface of the electronic device 700. The display screen 704 can also be provided in a non-rectangular, irregular shape, i.e., a special-shaped screen. The display screen 704 can be made of materials such as liquid crystal display (LCD) and organic light-emitting diode (OLED).

[0087] The camera 707 is used to capture images or videos. Optionally, the camera 707 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the electronic device, and the rear camera is arranged on the back of the electronic device. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and virtual reality (VR) shooting function or other fusion shooting functions. In some embodiments of the present disclosure, the camera 707 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0088] Audio circuit 706 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals for input to processor 701 for processing. For the purpose of stereo sound collection or noise reduction, multiple microphones may be provided, respectively, at different locations within electronic device 700. The microphone may also be an array microphone or an omnidirectional microphone.

[0089] Power supply 707 is used to power the various components of electronic device 700. Power supply 707 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 707 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0090] The electronic device structure block diagram shown in the embodiment of the present disclosure does not constitute a limitation on the electronic device 700. The electronic device 700 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0091] In the present disclosure, the terms "first", "second", etc. are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or order; the term "plurality" refers to two or more, unless otherwise expressly defined. The terms "installed", "connected", "connected", "fixed", etc. should be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; "connected" can be a direct connection or an indirect connection through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present disclosure can be understood according to the specific circumstances.

[0092] In the description of the present disclosure, it is necessary to understand that the terms "upper" and "lower" etc. indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or unit referred to must have a specific direction, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on the present disclosure.

[0093] The above descriptions are merely specific embodiments of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any modifications or substitutions that can be readily conceived by a person skilled in the art within the technical scope disclosed herein are intended to be covered by the scope of protection of the present disclosure. Therefore, equivalent modifications made according to the claims of the present disclosure are still within the scope of protection of the present disclosure.

Claims

1. A method for training a speech enhancement model, characterized in that: include: Obtain N groups of training samples, wherein the i-th group of training samples includes: i-th training data and i-th target data, wherein N is a positive integer, and i is a positive integer not greater than N; Training a speech enhancement model using the N groups of training samples; The step of obtaining the i-th group of training samples includes: Obtaining an i-th room impulse response and an i-th clean speech data, and processing the i-th room impulse response and the i-th clean speech data to obtain an i-th training data in the i-th group of training samples; An i-th control curve is determined based on the i-th room impulse response, and an i'-th room impulse response is obtained by multiplying the i-th control curve by the i-th room impulse response, where i' is a positive integer not greater than N. The i-th clean speech data is convolved with the i'-th room impulse response to obtain i-th target data in the i-th group of training samples.

2. The method for training a speech enhancement model according to claim 1, wherein: The impulse response of the i-th room includes a plurality of sampling points; The i-th control curve includes a plurality of control values, and the number of the control values ​​is the same as the number of sampling points of the i-th room impulse response; When the sampling point value of the tail of the impulse response of the i'th room is zero or the absolute value is very small, a tail truncation process may be selected.

3. The method for training a speech enhancement model according to claim 2, wherein: The determining the i-th control curve according to the i-th room impulse response includes: Determining the absolute value of each sampling point in the impulse response of the i-th room, wherein the absolute value includes a plurality of maximum values ​​having equal values; In the impulse response of the i-th room, determining the sampling point corresponding to the first maximum value in the absolute value as the peak position point; A control value corresponding to a peak position point of the impulse response of the i-th room in the i-th control curve is determined as a main control value of the i-th control curve.

4. The method for training a speech enhancement model according to claim 3, wherein: The method further comprises determining the i-th control curve according to the i-th room impulse response: The control value of the i-th control curve is adjusted by parameters to determine the i-th control curve; wherein, other control values ​​except the main control value are not greater than the main control value.

5. The method for training a speech enhancement model according to claim 1, wherein: The processing of the i-th room impulse response and the i-th clean speech data to obtain the i-th training data in the i-th group of training samples includes: The i-th clean speech data is convolved with the i-th room impulse response to obtain the i-th training data in the i-th group of training samples.

6. The method for training a speech enhancement model according to claim 5, wherein: The processing of the i-th room impulse response and the i-th clean speech data to obtain the i-th training data in the i-th group of training samples includes: The i-th clean speech data is convolved with the i-th room impulse response, and the convolution operation is performed with the noise data to obtain the i-th training data in the i-th group of training samples.

7. The method for training a speech enhancement model according to claim 5, wherein: The processing of the i-th room impulse response and the i-th clean speech data to obtain the i-th training data in the i-th group of training samples includes: The i-th clean speech data and the noise data are added and convolved with the i-th room impulse response to obtain the i-th training data in the i-th group of training samples.

8. A training device for a speech enhancement model, characterized in that: include: An acquisition module is configured to: acquire N groups of training samples, wherein the i-th group of training samples includes: the i-th training data and the i-th target data, wherein N is a positive integer and i is a positive integer not greater than N; A training module, configured to: train a speech enhancement model using the N groups of training samples; The acquisition module is specifically configured to: acquire an i-th room impulse response and an i-th clean speech data, process the i-th room impulse response and the i-th clean speech data to obtain an i-th training data in the i-th group of training samples; determine an i-th control curve based on the i-th room impulse response, multiply the i-th control curve by the i-th room impulse response to obtain an i'-th room impulse response, where i' is a positive integer not greater than N; and convolve the i-th clean speech data with the i'-th room impulse response to obtain an i-th target data in the i-th group of training samples.

9. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the training method of the speech enhancement model according to any one of claims 1 to 7 is implemented.

10. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the training method of the speech enhancement model according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Voice data enhancing method and system

    CN107481731A

  • Speech enhancement method and device, electronic equipment and storage medium

    CN112700786A