Voice distortion recovery method and device based on discrete domain, equipment and medium

Through the discrete domain-based speech distortion recovery method, the global characteristics of distorted speech are acquired and converted into discrete characterization sequences, and the problems of low versatility and high complexity of speech distortion recovery methods in the prior art are solved, and the speech distortion recovery effect with high versatility and low complexity are achieved.

CN120164474APending Publication Date: 2025-06-17UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510428850.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the prior art, speech distortion recovery methods are mainly used in continuous domains, limiting their universality, and requiring auxiliary text information or semantic information, increasing complexity.

Method used

A discrete domain-based speech distortion recovery method is proposed. By obtaining the global characteristics of distorted speech, inputting the discrete characterization prediction module of the speech distortion recovery model, determining the discrete characterization sequence of distorted speech, and inputting it into the speech decoding module to convert it into the recovery speech.

Benefits of technology

This method converts speech distortion recovery tasks into classification problems, can handle multiple types of speech distortion, is highly versatile, does not require auxiliary text or semantic information, and reduces complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164474A_ABST
    Figure CN120164474A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice distortion recovery method and device based on a discrete domain, equipment and a medium, and relates to the technical field of signal processing. The method comprises the following steps: acquiring distorted voice; inputting the distorted voice into a global feature extraction module of the voice distortion recovery model to determine global features of the distorted voice; inputting the global features of the distorted voice into a discrete representation prediction module of a voice distortion recovery model to determine a discrete representation sequence of the distorted voice; and inputting the discrete representation sequence of the distorted voice into a voice decoding module of a voice distortion recovery model, and converting the discrete representation sequence of the distorted voice into recovered voice. Therefore, according to the method, a voice distortion recovery task is converted into a classification problem, various types of voice distortion (including mixed distortion, unconventional distortion and the like) can be processed, the universality is high, voice distortion recovery can be executed without auxiliary text or semantic information, and the complexity is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of signal processing, and in particular, to a method, apparatus, device, and medium for speech distortion recovery based on the discrete domain. Background Art

[0002] In real life, speech signals are inevitably affected by various types of distortions during generation or transmission, such as noise, reverberation, phase distortion, compression distortion, etc., resulting in a decline in the quality of speech signals. Therefore, speech distortion recovery methods have emerged, and their goal is to recover high-quality original speech signals from these distorted speech signals. High-quality original speech signals can better serve application scenarios such as communication, speech recognition, and speech synthesis, and have important practical value.

[0003] In related technologies, with the rapid development of deep learning technology, distorted speech signals are usually input into a trained speech distortion recovery model to directly obtain the waveform or features of the original speech signal. However, these methods are mainly applied to speech distortion recovery in the continuous domain, which limits the generality of speech distortion recovery. Moreover, continuous-domain speech distortion recovery methods usually require auxiliary text information or semantic information to assist the speech distortion recovery process, which greatly increases the complexity of speech distortion recovery.

[0004] Therefore, how to obtain a speech distortion recovery method based on the discrete domain with lower complexity and higher generality has become an urgent technical problem to be solved. Summary of the Invention

[0005] Based on the above problems, the present application provides a method, apparatus, device, and medium for speech distortion recovery based on the discrete domain, which can improve the generality of speech distortion recovery and reduce the complexity of speech distortion recovery.

[0006] The embodiments of the present application disclose the following technical solutions:

[0007] In a first aspect, the present application discloses a method for speech distortion recovery based on the discrete domain, the method comprising:

[0008] Obtaining distorted speech, where the distorted speech is a speech signal generated by the original speech being affected by distortion during generation or transmission;

[0009] Determining the global features of the distorted speech by inputting the distorted speech into the global feature extraction module of the speech distortion recovery model;

[0010] Determining the discrete representation sequence of the distorted speech by inputting the global features of the distorted speech into the discrete representation prediction module of the speech distortion recovery model;

[0011] By inputting the discrete representation sequence of the distorted speech into the speech decoding module of the speech distortion recovery model, the discrete representation sequence of the distorted speech is converted into a recovered speech.

[0012] Optionally, the global feature extraction module includes an encoder, a residual vector quantizer, and a feature processor connected in sequence, where the residual vector quantizer includes N vector quantizers;

[0013] The determining the global feature of the distorted speech by inputting the distorted speech into the global feature extraction module of the speech distortion recovery model includes:

[0014] By inputting the distorted speech into the encoder, a downsampled feature vector is obtained;

[0015] By inputting the downsampled feature vector into the N vector quantizers, N quantization results respectively generated by the N vector quantizers are obtained;

[0016] By inputting the concatenated N quantization results into the feature processor, the global feature of the distorted speech is obtained.

[0017] Optionally, the discrete representation prediction module includes N discrete representation predictors;

[0018] The determining the discrete representation sequence of the distorted speech by inputting the global feature of the distorted speech into the discrete representation prediction module of the speech distortion recovery model includes:

[0019] By looking up a random codebook for a random discrete representation sequence, an initial feature vector is generated;

[0020] By inputting the global feature of the distorted speech and the initial feature vector into the first discrete representation predictor, a first raw discrete representation output by the first discrete representation predictor is obtained;

[0021] By inputting the global feature of the distorted speech and the sum of the raw feature vectors generated by looking up the codebooks of the corresponding vector quantizers for the n - 1 raw discrete representations output by the previous n - 1 discrete representation predictors into the nth discrete representation predictor, the nth raw discrete representation output by the nth discrete representation predictor is obtained, where 2 ≤ n ≤ N;

[0022] Calculating the classification probability distribution for the first raw discrete representation and the nth raw discrete representation through a softmax layer;

[0023] Determining the raw discrete representation with the maximum classification probability as the discrete representation sequence.

[0024] Optionally, converting the discrete representation sequence of the distorted speech into a restored speech by inputting the discrete representation sequence of the distorted speech into the speech decoding module of the speech distortion restoration model includes:

[0025] Obtaining continuous feature vectors by looking up codebooks respectively corresponding to the discrete representation sequences of the distorted speech;

[0026] After adding the continuous feature vectors, inputting them into the speech decoding module of the speech distortion restoration model to obtain a restored speech.

[0027] Optionally, the training method of the speech distortion restoration model is specifically as follows:

[0028] Obtaining a target discrete representation sequence by inputting training speech into an encoder and N vector quantizers;

[0029] Looking up the corresponding codebook according to the target discrete representation sequence to obtain a target quantization result;

[0030] Obtaining a training probability distribution by inputting the target quantization result into a discrete representation predictor;

[0031] Obtaining a target probability distribution by performing One-Hot encoding on the target discrete representation sequence;

[0032] Training the speech distortion restoration model by minimizing the distance between the target probability distribution and the training probability distribution through a cross-entropy loss function.

[0033] In a second aspect, the present application discloses a speech distortion restoration device based on the discrete domain. The device includes: a speech acquisition module, a feature determination module, a sequence determination module, and a speech conversion module;

[0034] The speech acquisition module is configured to acquire distorted speech, where the distorted speech is a speech signal generated by an original speech affected by distortion during generation or transmission;

[0035] The feature determination module is configured to determine the global feature of the distorted speech by inputting the distorted speech into the global feature extraction module of the speech distortion restoration model;

[0036] The sequence determination module is configured to determine the discrete representation sequence of the distorted speech by inputting the global feature of the distorted speech into the discrete representation prediction module of the speech distortion restoration model;

[0037] The speech conversion module is configured to convert the discrete representation sequence of the distorted speech into a restored speech by inputting the discrete representation sequence of the distorted speech into the speech decoding module of the speech distortion restoration model.

[0038] Optionally, the global feature extraction module includes an encoder, a residual vector quantizer, and a feature processor connected in sequence, where the residual vector quantizer includes N vector quantizers;

[0039] The feature determination module includes: a first determination module, a second determination module, and a third determination module;

[0040] The first determination module is configured to obtain a downsampled feature vector by inputting the distorted speech into the encoder;

[0041] The second determination module is configured to obtain N quantization results respectively generated by the N vector quantizers by inputting the downsampled feature vector into the N vector quantizers;

[0042] The third determination module is configured to obtain the global feature of the distorted speech by inputting the concatenated N quantization results into the feature processor.

[0043] Optionally, the discrete representation prediction module includes N discrete representation predictors;

[0044] The sequence determination module includes: a third determination module, a fourth determination module, a fifth determination module, a sixth determination module, and a seventh determination module;

[0045] The third determination module is configured to generate an initial feature vector by looking up a random codebook for a random discrete representation sequence;

[0046] The fourth determination module is configured to obtain a first original discrete representation output by the first discrete representation predictor by inputting the global feature of the distorted speech and the initial feature vector into the first discrete representation predictor;

[0047] The fifth determination module is configured to input the global feature of the distorted speech and the sum of the original feature vectors generated by looking up the codebooks of the corresponding vector quantizers for the n - 1 original discrete representations output by the previous n - 1 discrete representation predictors into the nth discrete representation predictor to obtain the nth original discrete representation output by the nth discrete representation predictor, where 2 ≤ n ≤ N;

[0048] The sixth determination module is configured to calculate the classification probability distribution of the first original discrete representation and the nth original discrete representation through a softmax layer;

[0049] The seventh determination module is configured to determine the original discrete representation with the maximum classification probability as the discrete representation sequence.

[0050] Optionally, the speech conversion module includes: a first conversion module and a second conversion module;

[0051] The first conversion module is configured to obtain continuous feature vectors by looking up codebooks respectively corresponding to the discrete representation sequences of the distorted speech;

[0052] The second conversion module is configured to add the continuous feature vectors and then input them into the speech decoding module of the speech distortion recovery model to obtain the recovered speech.

[0053] Optionally, the training unit of the speech distortion recovery model specifically includes:

[0054] The first training unit is configured to obtain target discrete representation sequences by inputting training speech into an encoder and N vector quantizers;

[0055] The second training unit is configured to look up the corresponding codebooks according to the target discrete representation sequences to obtain target quantization results;

[0056] The third training unit is configured to obtain a training probability distribution by inputting the target quantization results into a discrete representation predictor;

[0057] The fourth training unit is configured to obtain a target probability distribution by performing One-Hot encoding on the target discrete representation sequences;

[0058] The fifth training unit is configured to train the speech distortion recovery model by minimizing the distance between the target probability distribution and the training probability distribution through a cross-entropy loss function.

[0059] In a third aspect, the present application discloses a speech distortion recovery device based on the discrete domain, and the device includes: a memory and a processor;

[0060] The memory is configured to store a program;

[0061] The processor is configured to execute the program to implement each step of the speech distortion recovery method based on the discrete domain as described in the first aspect.

[0062] In a fourth aspect, the present application discloses a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, each step of the speech distortion recovery method based on the discrete domain as described in the first aspect is implemented.

[0063] Compared with the prior art, the present application has the following beneficial effects:

[0064] The embodiments of the present application provide a method, apparatus, device and medium for speech distortion recovery based on the discrete domain. The method includes: obtaining distorted speech; determining the global features of the distorted speech by inputting the distorted speech into the global feature extraction module of the speech distortion recovery model; determining the discrete representation sequence of the distorted speech by inputting the global features of the distorted speech into the discrete representation prediction module of the speech distortion recovery model; and converting the discrete representation sequence of the distorted speech into recovered speech by inputting the discrete representation sequence of the distorted speech into the speech decoding module of the speech distortion recovery model. Thus, this method transforms the speech distortion recovery task into a classification problem, which can not only handle various types of speech distortions (including mixed distortions and unconventional distortions, etc.), has high versatility, but also can perform speech distortion recovery without auxiliary text or semantic information, with low complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0066] Figure 1 It is a flowchart of a method for speech distortion recovery based on the discrete domain provided by the embodiments of the present application;

[0067] Figure 2 It is a schematic diagram of a speech distortion recovery model provided by the embodiments of the present application;

[0068] Figure 3 It is a schematic diagram of a speech distortion recovery device based on the discrete domain provided by the embodiments of the present application;

[0069] Figure 4 It is a schematic diagram of a computer-readable medium provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] As described above, with the rapid development of deep learning technology, speech signals affected by distortion are usually input into a trained speech distortion recovery model to directly obtain the waveform or features of the original speech signal. However, these methods are mainly applied to speech distortion recovery in the continuous domain, which limits the versatility of speech distortion recovery. Moreover, the speech distortion recovery methods in the continuous domain usually require auxiliary text information or semantic information to assist the speech distortion recovery process, which greatly increases the complexity of speech distortion recovery.

[0071] After research, the inventors proposed a method, apparatus, device, and medium for speech distortion recovery based on the discrete domain. The method includes: obtaining distorted speech; determining the global features of the distorted speech by inputting the distorted speech into the global feature extraction module of the speech distortion recovery model; determining the discrete representation sequence of the distorted speech by inputting the global features of the distorted speech into the discrete representation prediction module of the speech distortion recovery model; and converting the discrete representation sequence of the distorted speech into recovered speech by inputting the discrete representation sequence of the distorted speech into the speech decoding module of the speech distortion recovery model. Thus, this method transforms the speech distortion recovery task into a classification problem, which can not only handle various types of speech distortions (including mixed distortions and unconventional distortions, etc.), has high versatility, but also can perform speech distortion recovery without auxiliary text or semantic information, with low complexity.

[0072] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of this application.

[0073] See Figure 1 , which is a flowchart of a method for speech distortion recovery based on the discrete domain provided by an embodiment of this application. The method includes:

[0074] S101: Obtain distorted speech y, where the distorted speech is a speech signal generated by the original speech being affected by distortion during generation or transmission.

[0075] Exemplarily, the original speech x can be a clear recording, and the distorted speech is a speech signal generated after the original speech is affected by distortions such as noise, reverberation, and compression during generation or transmission. Here, T represents the waveform length of the original speech (i.e., the number of sampling points).

[0076] S102: Determine the global feature G of the distorted speech y by inputting the distorted speech y into the global feature extraction module of the speech distortion recovery model (y) .

[0077] First, extract the global feature of the distorted speech y from the distorted speech y where C represents the dimension of the global feature (i.e., the length of each feature vector), and L represents the number of frames of the global feature (i.e., the number of time steps of each feature vector). It can be understood that the global feature G (y) captures the overall information of the distorted speech y (including overall structure, spectral features, time-frequency characteristics, etc.), rather than local details.

[0078] See Figure 2 , which is a schematic diagram of a voice distortion recovery model provided by an embodiment of this application. As can be seen from Figure 2 , the voice distortion recovery model includes a global feature extraction module, and the global feature extraction module consists of an encoder φ E , a Residual Vector Quantizer (RVQ), and a feature processor φ FP in sequence. Among them, the residual vector quantizer consists of N vector quantizers (VQ) φ Q1 ,..., φ QN .

[0079] Specifically, first, input the distorted speech y into the encoder φ E of the global feature extraction module of the voice distortion recovery model to obtain a downsampled feature vector . Among them, K is the dimension of the downsampled feature vector, and L is the number of frames of the downsampled feature vector. It can be understood that the role of the encoder φ E is to convert the distorted speech y into a higher-level abstract feature representation (i.e., the downsampled feature vector E (y) ), and these downsampled feature vectors E (y) can capture the key information of the distorted speech y.

[0080] Subsequently, input the downsampled feature vector E (y) into the residual vector quantizer RVQ of the global feature extraction module of the voice distortion recovery model. Through each vector quantizer VQ in the residual vector quantizer RVQ, respectively quantize the downsampled feature vector E (y) to obtain the quantization result output by each vector quantizer VQ

[0081] Finally, splice and reduce the dimension of the quantization result output by each vector quantizer VQ to generate an intermediate feature. Then, input this intermediate feature into the feature processor φ FP to obtain the global feature G (y) of the distorted speech. Among them, the feature processor φ FP consists of multiple Conformer blocks. The Conformer block is a module that combines a convolutional neural network and a self-attention mechanism, and can capture local and global features simultaneously.

[0082] It can be understood that the quantization process of VQ removes the time-invariant features (i.e., redundant information that does not change with time) in the distorted speech y, thereby reducing the difficulty of speech distortion recovery.

[0083] S103: By inputting the global feature G of the distorted speech (y) into the discrete representation prediction module of the speech distortion recovery model, determine the discrete representation sequence of the distorted speech

[0084] It is known from Figure 2 that the speech distortion recovery model includes a discrete representation prediction module. The discrete representation prediction module consists of N discrete representation predictors φ TP1 ,..., φ TPN and the backbone network of each discrete representation predictor consists of multiple Conformer blocks.

[0085] The discrete representation prediction module is used to perform the following operations:

[0086] In the first step, by looking up the random codebook for the random discrete representation sequence d0 = [d 0,1 ,..., d 0,L T generate the initial feature vector where each discrete representation belongs to the set {1, 2,..., M}, and M is the codebook size of the vector quantizer VQ, representing the number of categories that each discrete representation can choose.

[0087] In the second step, for the first discrete representation predictor φ TP1 : By inputting the global feature G of the distorted speech (y) and the initial feature vector into the first discrete representation predictor φ TP1 , so that the first discrete representation predictor φ TP1 outputs the first raw discrete representation of the first vector quantizer

[0088] For the nth discrete representation predictor φ TPn (2 ≤ n ≤ N): By inputting the global feature G of the distorted speech (y) , and, the discrete representations output by the previous n - 1 discrete representation predictors and the raw feature vectors generated by looking up the codebooks of the corresponding vector quantizers into the nth discrete representation predictor φ TPn , so that the nth discrete representation predictor φ TPn outputs the nth raw discrete representation of the nth vector quantizer

[0089] ​It can be seen from this that the output of the discrete representation predictor can be shown as in the following formula (1):

[0090]

[0091] Step 3: For each original discrete representation Calculate the classification probability distribution for each frame through the softmax layer And select the category with the highest possibility (i.e., the original discrete representation with the largest classification probability distribution) from N categories as the discrete representation sequence

[0092] S104: By inputting the discrete representation sequence of the distorted speech into the speech decoding module of the speech distortion recovery model, the discrete representation sequence of the distorted speech is converted into the recovered speech

[0093] It is known from Figure 2 that the speech distortion recovery model includes a speech decoding module, and the speech decoding module includes a decoder φ D .

[0094] Specifically, by looking up the codebooks corresponding to the discrete representation sequence respectively, continuous feature vectors are obtained Subsequently, after adding the continuous feature vectors as the input of the decoder φ D the recovered speech is obtained Among them, the acquisition of the recovered speech can be shown as in the following formula (2):

[0095]

[0096] It should be noted that the speech distortion recovery model in steps S101 - S104 has been pre-trained, and the teacher forcing training (TFT) strategy is adopted during the training process. The specific steps are as follows:

[0097] Step 1: By inputting the training speech into the encoder φ E and N vector quantizers φ Q1 ,..., φ QN the target discrete representation sequence is obtained where represents the target discrete representation obtained by the nth vector quantizer at the lth frame.

[0098] Step 2: According to the target discrete representation Search for the corresponding codebook to obtain the target quantization result

[0099] In the third step, input the target quantization result into the discrete representation predictor φ TP2 ,..., φ TPN for teacher-forced training to obtain the training probability distribution where represents the probability of each discrete representation category on the l-th frame of the n-th quantizer.

[0100] In the fourth step, perform One-Hot encoding on the target discrete representation sequence to obtain the target probability distribution

[0101] In the fifth step, minimize the distance between the target probability distribution and the training probability distribution by using the cross-entropy loss function to train the speech distortion recovery model. The formula of the cross-entropy loss function is specifically shown as formula (3) below:

[0102]

[0103] where represents the probability value corresponding to the target category in the prediction distribution .

[0104] In summary, the present application provides a speech distortion recovery method based on the discrete domain. The method includes: obtaining distorted speech; determining the global features of the distorted speech by inputting the distorted speech into the global feature extraction module of the speech distortion recovery model; determining the discrete representation sequence of the distorted speech by inputting the global features of the distorted speech into the discrete representation prediction module of the speech distortion recovery model; and converting the discrete representation sequence of the distorted speech into recovered speech by inputting the discrete representation sequence of the distorted speech into the speech decoding module of the speech distortion recovery model. Thus, the method transforms the speech distortion recovery task into a classification problem, which can not only handle various types of speech distortions (including mixed distortions and unconventional distortions, etc.), has high versatility, but also does not require auxiliary text or semantic information to perform speech distortion recovery, with low complexity.

[0105] Refer to Figure 3 , which is a schematic diagram of a speech distortion recovery device based on the discrete domain provided by an embodiment of the present application. The speech distortion recovery device 300 based on the discrete domain includes: a speech acquisition module 301, a feature determination module 302, a sequence determination module 303, and a speech conversion module 304;

[0106] A voice acquisition module 301, configured to acquire distorted voice, where the distorted voice is a voice signal generated by the original voice affected by distortion during generation or transmission;

[0107] A feature determination module 302, configured to determine the global feature of the distorted voice by inputting the distorted voice into the global feature extraction module of the voice distortion recovery model;

[0108] A sequence determination module 303, configured to determine the discrete representation sequence of the distorted voice by inputting the global feature of the distorted voice into the discrete representation prediction module of the voice distortion recovery model;

[0109] A voice conversion module 304, configured to convert the discrete representation sequence of the distorted voice into recovered voice by inputting the discrete representation sequence of the distorted voice into the voice decoding module of the voice distortion recovery model.

[0110] In a specific implementation, the global feature extraction module includes an encoder, a residual vector quantizer, and a feature processor connected in sequence, where the residual vector quantizer includes N vector quantizers;

[0111] The feature determination module 302 includes: a first determination module, a second determination module, and a third determination module;

[0112] The first determination module is configured to obtain a downsampled feature vector by inputting the distorted voice into the encoder;

[0113] The second determination module is configured to obtain N quantization results respectively generated by the N vector quantizers by inputting the downsampled feature vector into the N vector quantizers;

[0114] The third determination module is configured to obtain the global feature of the distorted voice by inputting the concatenated N quantization results into the feature processor.

[0115] In a specific implementation, the discrete representation prediction module includes N discrete representation predictors;

[0116] The sequence determination module 303 includes a third determination module, a fourth determination module, a fifth determination module, a sixth determination module, and a seventh determination module;

[0117] The third determination module is configured to generate an initial feature vector by looking up a random codebook for a random discrete representation sequence;

[0118] The fourth determination module is configured to obtain a first original discrete representation output by the first discrete representation predictor by inputting the global feature of the distorted voice and the initial feature vector into the first discrete representation predictor;

[0119] A fifth determination module, configured to input the sum of the global feature of the distorted speech and the original feature vectors generated by looking up the codebooks of the corresponding vector quantizers for the n-1 original discrete representations output by the first n-1 discrete representation predictors into the nth discrete representation predictor, so as to obtain the nth original discrete representation output by the nth discrete representation predictor, where 2≤n≤N;

[0120] A sixth determination module, configured to calculate the classification probability distribution for the first original discrete representation and the nth original discrete representation through a softmax layer;

[0121] A seventh determination module, configured to determine the original discrete representation with the maximum classification probability as the discrete representation sequence.

[0122] In a specific implementation manner, the speech conversion module 304 includes: a first conversion module and a second conversion module;

[0123] The first conversion module is configured to obtain continuous feature vectors by looking up the codebooks corresponding to the discrete representation sequences of the distorted speech respectively;

[0124] The second conversion module is configured to input the continuous feature vectors after addition into the speech decoding module of the speech distortion recovery model to obtain the recovered speech.

[0125] In a specific implementation manner, the training unit of the speech distortion recovery model specifically includes:

[0126] A first training unit, configured to input the training speech into an encoder and N vector quantizers to obtain a target discrete representation sequence;

[0127] A second training unit, configured to look up the corresponding codebooks according to the target discrete representation sequence to obtain a target quantization result;

[0128] A third training unit, configured to input the target quantization result into a discrete representation predictor to obtain a training probability distribution;

[0129] A fourth training unit, configured to perform One-Hot encoding on the target discrete representation sequence to obtain a target probability distribution;

[0130] A fifth training unit, configured to train the speech distortion recovery model by minimizing the distance between the target probability distribution and the training probability distribution through a cross-entropy loss function.

[0131] In summary, the present application provides a voice distortion recovery device based on the discrete domain. The device includes: obtaining distorted speech; determining the global features of the distorted speech by inputting the distorted speech into the global feature extraction module of the voice distortion recovery model; determining the discrete representation sequence of the distorted speech by inputting the global features of the distorted speech into the discrete representation prediction module of the voice distortion recovery model; and converting the discrete representation sequence of the distorted speech into recovered speech by inputting the discrete representation sequence of the distorted speech into the speech decoding module of the voice distortion recovery model. Thus, the device transforms the voice distortion recovery task into a classification problem, which can not only handle various types of voice distortions (including mixed distortions and unconventional distortions, etc.), has high versatility, but also can perform voice distortion recovery without auxiliary text or semantic information, with low complexity.

[0132] The embodiments of the present application also provide corresponding voice distortion recovery devices and computer-readable media based on the discrete domain for implementing the voice distortion recovery method provided by the embodiments of the present application.

[0133] Among them, the voice distortion recovery device based on the discrete domain includes a memory and a processor. The memory is used to store instructions or codes, and the processor is used to execute the instructions or codes so that the device executes a voice distortion recovery method according to any embodiment of the present application.

[0134] See Figure 4 , this figure is a schematic diagram of a computer-readable medium provided by an embodiment of the present application. A computer program 411 is stored on the computer-readable medium 400. When the computer program 411 is executed by a processor, it implements the steps of the above-mentioned Figure 1 voice distortion recovery method based on the discrete domain.

[0135] It should be noted that in the context of the present application, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0136] It should be noted that the machine-readable medium described above in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. And in the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0137] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.

[0138] Although the subject matter has been described in language specific to structural features and / or method logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms for implementing the claims.

[0139] Although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present application. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. On the contrary, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0140] The above description is only a preferred embodiment of the present application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present application that have similar functions.

Claims

1. A speech distortion restoration method based on discrete domain, characterized in that: The method comprises: Acquire distorted speech, where the distorted speech is a speech signal generated when the original speech is distorted during generation or transmission; Determining the global features of the distorted speech by inputting the distorted speech into a global feature extraction module of a speech distortion recovery model; Determining a discrete representation sequence of the distorted speech by inputting the global features of the distorted speech into a discrete representation prediction module of the speech distortion recovery model; The discrete representation sequence of the distorted speech is input into a speech decoding module of the speech distortion restoration model to convert the discrete representation sequence of the distorted speech into restored speech.

2. The method according to claim 1, characterized in that The global feature extraction module includes an encoder, a residual vector quantizer and a feature processor connected in sequence, wherein the residual vector quantizer includes N vector quantizers; The step of inputting the distorted speech into a global feature extraction module of a speech distortion recovery model to determine the global features of the distorted speech comprises: By inputting the distorted speech into the encoder, a downsampled feature vector is obtained; By inputting the downsampled feature vector into the N vector quantizers, N quantization results generated by the N vector quantizers are obtained respectively; The global features of the distorted speech are obtained by splicing the N quantization results and inputting them into the feature processor.

3. The method according to claim 1, characterized in that The discrete representation prediction module includes N discrete representation predictors; The step of inputting the global features of the distorted speech into a discrete representation prediction module of the speech distortion recovery model to determine a discrete representation sequence of the distorted speech comprises: Generate an initial feature vector by searching a random codebook for a random discrete representation sequence; By inputting the global feature of the distorted speech and the initial feature vector into the first discrete representation predictor, a first original discrete representation output by the first discrete representation predictor is obtained; The global features of the distorted speech and the sum of the original feature vectors generated by searching the codebook of the corresponding vector quantizer for the n-1 original discrete representations output by the first n-1 discrete representation predictors are input into the nth discrete representation predictor to obtain the nth original discrete representation output by the nth discrete representation predictor, wherein 2≤n≤N; The first original discrete representation and the nth original discrete representation are subjected to a softmax layer to calculate a classification probability distribution; Determine the original discrete representation with the maximum classification probability as the discrete representation sequence.

4. The method according to claim 1, characterized in that: The step of converting the discrete representation sequence of the distorted speech into the restored speech by inputting the discrete representation sequence of the distorted speech into the speech decoding module of the speech distortion restoration model comprises: Obtaining a continuous feature vector by searching for codebooks corresponding to discrete representation sequences of the distorted speech; After the continuous feature vectors are added, they are input into the speech decoding module of the speech distortion restoration model to obtain the restored speech.

5. The method according to any one of claims 1 to 4, characterized in that: The training method of the speech distortion recovery model is as follows: By inputting the training speech into an encoder and N vector quantizers, a target discrete representation sequence is obtained; Searching a corresponding codebook according to the target discrete representation sequence to obtain a target quantization result; Obtaining a training probability distribution by inputting the target quantization result into a discrete representation predictor; Obtaining a target probability distribution by One-Hot encoding the target discrete representation sequence; The speech distortion restoration model is trained by minimizing the distance between the target probability distribution and the training probability distribution through a cross entropy loss function.

6. A speech distortion restoration device based on discrete domain, characterized in that: The device comprises: a speech acquisition module, a feature determination module, a sequence determination module and a speech conversion module; The speech acquisition module is used to acquire distorted speech, where the distorted speech is a speech signal generated when the original speech is distorted during generation or transmission; The feature determination module is used to determine the global features of the distorted speech by inputting the distorted speech into the global feature extraction module of the speech distortion recovery model; The sequence determination module is used to determine the discrete representation sequence of the distorted speech by inputting the global features of the distorted speech into the discrete representation prediction module of the speech distortion recovery model; The speech conversion module is used to convert the discrete representation sequence of the distorted speech into the restored speech by inputting the discrete representation sequence of the distorted speech into the speech decoding module of the speech distortion restoration model.

7. The device according to claim 6, characterized in that The global feature extraction module includes an encoder, a residual vector quantizer and a feature processor connected in sequence, wherein the residual vector quantizer includes N vector quantizers; The feature determination module includes: a first determination module, a second determination module and a third determination module; The first determining module is used to obtain a downsampled feature vector by inputting the distorted speech into the encoder; The second determining module is used to obtain N quantization results respectively generated by the N vector quantizers by inputting the downsampled feature vector into the N vector quantizers; The third determination module is used to obtain the global features of the distorted speech by splicing the N quantization results and inputting them into the feature processor.

8. The device according to claim 6, characterized in that The discrete representation prediction module includes N discrete representation predictors; The sequence determination module includes a third determination module, a fourth determination module, a fifth determination module, a sixth determination module and a seventh determination module; The third determination module is used to generate an initial feature vector by searching a random codebook for a random discrete representation sequence; The fourth determination module is used to obtain a first original discrete representation output by the first discrete representation predictor by inputting the global feature of the distorted speech and the initial feature vector into the first discrete representation predictor; The fifth determination module is used to obtain an nth original discrete representation output by the nth discrete representation predictor by inputting the global feature of the distorted speech and the sum of the original feature vectors generated by searching the codebook of the corresponding vector quantizer for the n-1 original discrete representations output by the first n-1 discrete representation predictors into the nth discrete representation predictor, wherein 2≤n≤N; The sixth determination module is used to calculate the classification probability distribution of the first original discrete representation and the nth original discrete representation through a softmax layer; The seventh determination module is used to determine the original discrete representation with the maximum classification probability as the discrete representation sequence.

9. A speech distortion restoration device based on discrete domain, characterized in that: The device comprises: a memory and a processor; The memory is used to store programs; The processor is used to execute the program to implement the various steps of the discrete domain-based speech distortion restoration method as described in any one of claims 1 to 5.

10. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the discrete domain-based speech distortion restoration method according to any one of claims 1 to 5 is implemented.