Speech Quality Evaluation Method and Device
The proposed speech quality evaluation method decouples from specific speech enhancement and ASR backends, using multiple acoustic models to assess and optimize speech quality across diverse environments, ensuring effective ASR performance without system updates.
Patent Information
- Application Number
- CN202010965579.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-15
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-09-15
AI Technical Summary
The existing speech quality evaluation methods cannot reflect the impact of voice enhancement algorithm changes on the automatic speech recognition (ASR) backend, and cannot be flexibly applied to many different scenarios, especially when the ASR system expands from one device to another, or when third-party applications access, it is impossible to ensure that the voice signal quality meets the access conditions.
The speech quality evaluation model constructed using at least two acoustic models is used to determine the loss information through the decoding results of the decoder, determine whether the enhanced speech quality meets the standards, and decouple it from the specific speech enhancement algorithm and ASR backend, which is suitable for a variety of different scenarios.
It can reflect the impact of the voice enhancement algorithm on the ASR backend, reduce the workload when expanding the device, ensure that the voice signal quality meets the access conditions, and improve the accuracy and service effectiveness of voice recognition.
Smart Images

Figure CN114187921B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular, to a method and device for evaluating speech quality. Background Art
[0002] With the development of artificial intelligence technology, automatic speech recognition (ASR) technology has become increasingly important. ASR is a key technology for realizing human-computer interaction. A machine can identify and understand human speech through ASR technology, convert speech into text, or convert speech into commands that the machine can understand, and perform corresponding operations according to the commands.
[0003] Due to the existence of actual environmental noise and interference, before performing speech recognition, it is necessary to perform speech enhancement processing on the collected speech signal. Speech enhancement is to use a speech enhancement algorithm to extract as pure an original speech as possible from the noisy speech, that is, to perform noise reduction processing on the speech, so as to improve the speech quality and reduce the recognition error rate. Usually, a speech quality evaluation method is needed to evaluate whether the quality of the enhanced speech meets the requirements, so as to evaluate the quality of the speech enhancement algorithm.
[0004] The speech quality evaluation models used in current speech quality evaluation methods are all based on the processed speech signals, independent of a specific ASR backend, and cannot show the impact of changes in the speech enhancement algorithm on ASR. For example, when ASR is extended from one device to another, the speech enhancement algorithm changes, and the enhanced speech also changes. The metrics obtained by the original speech quality evaluation algorithm cannot reflect the impact of the change in the speech enhancement algorithm on the ASR backend. Summary of the Invention
[0005] This application proposes a method and device for evaluating speech quality, which can be decoupled from a specific speech enhancement algorithm and a specific ASR backend, can reflect the impact of the speech enhancement algorithm on the ASR backend, and can be flexibly applied to a variety of different scenarios.
[0006] In a first aspect, a method for evaluating speech quality is provided, specifically including: first, enhancing the speech signal collected by a device through a speech enhancement algorithm to obtain enhanced speech; then inputting the enhanced speech into a speech quality evaluation model to obtain loss information of the enhanced speech, where the speech quality evaluation model includes at least two acoustic models, and the above loss information is determined based on the decoding results of the enhanced speech by the decoders of at least two acoustic models as the ground truth; finally, based on the above loss information, determining whether the quality of the enhanced speech meets the standard.
[0007] It should be understood that the loss information can also be referred to as perplexity or other names, which are not limited in the embodiments of the present application. The voice quality evaluation model adopted in the embodiments of the present application does not prefer a specific voice enhancement algorithm, enabling the voice quality evaluation method to be decoupled from a specific voice enhancement algorithm and a specific ASR backend, and being flexibly applicable to a variety of different scenarios. The method of the embodiments of the present application can be used to test whether the quality of the voice signal accessed by a third party can meet the access conditions. If the voice quality meets the standard, it means that the voice signal meets the access conditions.
[0008] In combination with the first aspect, in some implementation manners of the first aspect, based on the loss information, determining whether the quality of the enhanced voice meets the standard includes: if the loss information is less than or equal to a threshold, determining that the quality of the enhanced voice meets the standard; or, if the loss information is greater than the threshold, determining that the quality of the enhanced voice does not meet the standard.
[0009] In another possible implementation manner, determining whether the quality of the enhanced voice meets the standard according to the loss information of the enhanced voice includes: if the loss information is less than the threshold, determining that the quality of the enhanced voice meets the standard; or, if the loss information is greater than or equal to the threshold, determining that the quality of the enhanced voice does not meet the standard.
[0010] In combination with the first aspect, in some implementation manners of the first aspect, after determining whether the quality of the enhanced voice meets the standard, the method further includes: if the quality of the enhanced voice meets the standard, accepting the voice enhancement algorithm; or, if the quality of the enhanced voice does not meet the standard, rejecting the voice enhancement algorithm.
[0011] In the embodiments of the present application, accepting a voice enhancement algorithm can be understood as the ASR system allowing the access of the voice enhancement algorithm, and rejecting a voice enhancement algorithm can be understood as the ASR system rejecting the access of the voice enhancement algorithm.
[0012] In combination with the first aspect, in some implementation manners of the first aspect, the above method further includes: in the case where the quality of the enhanced voice does not meet the standard, optimizing the voice enhancement algorithm.
[0013] In the embodiments of the present application, when it is determined that the quality of the enhanced voice does not meet the standard, the voice enhancement algorithm is improved and optimized, driving the enhanced voice obtained through the voice enhancement algorithm to approach the ideal far-field voice with low reverberation, low external noise, no human voice noise, and no echo in terms of acoustic features. When the ASR system is extended from one device to another new device, the voice enhancement algorithm can be optimized by the method of the embodiments of the present application, without updating the ASR system, reducing the workload.
[0014] In combination with the first aspect, in some implementations of the first aspect, the above at least two speech recognition models include all or part of the following models: a model based on a convolutional neural network (CNN) structure and a connectionist temporal classification (CTC) loss function; a model based on a Transformer structure and a transducer loss function; a listen, attend and spell (LAS) model based on a cross-entropy loss function; a hidden Markov model - deep neural network (HMM-DNN) model based on a cross-entropy loss function. The above models are all common acoustic models.
[0015] In the embodiments of the present application, the speech quality evaluation model may include, but is not limited to, each of the above-listed acoustic models. The modeling units of these acoustic models may include entries, sub-words, pinyin, syllables, phonemes, etc. By comprehensively using a variety of typical acoustic models, the speech quality evaluation method of the embodiments of the present application can be neutral to a specific acoustic model.
[0016] In combination with the first aspect, in some implementations of the first aspect, inputting the enhanced speech into the speech quality evaluation model to obtain loss information of the enhanced speech includes: inputting the enhanced speech into the above at least two acoustic models respectively to obtain at least two sub-loss information, where the at least two sub-loss information corresponds to the above at least two acoustic models; determining the loss information based on the at least two sub-loss information.
[0017] In the embodiments of the present application, the loss information of the enhanced speech can be determined according to the sub-loss information corresponding to each of the at least two acoustic models.
[0018] In combination with the first aspect, in some implementations of the first aspect, the loss information is obtained by performing weighted summation on the above at least two sub-loss information.
[0019] In combination with the first aspect, in some implementations of the first aspect, inputting the enhanced speech into the above at least two acoustic models respectively to obtain at least two sub-loss information includes: inputting the enhanced speech into the first acoustic model of the above at least two acoustic models to obtain the decoding result of the decoder in the first acoustic model; using the decoding result of the decoder in the first acoustic model as the ground truth value to calculate the first sub-loss information of the enhanced speech.
[0020] In combination with the first aspect, in some implementations of the first aspect, the method further includes: training the above at least two acoustic models based on the labeled corpus to obtain a speech quality evaluation model.
[0021] It should be understood that since the speech quality evaluation model includes at least two acoustic models, training the speech quality evaluation model is to train the at least two acoustic models. The speech corpus is the labeled corpus of speech signals, and the labeled text of a speech signal can be understood as the true value of the speech signal. Taking the speech signal as the input of an acoustic model, the predicted labeled text of the speech signal is output, and this predicted labeled text is used as the predicted value to compare with the true value of the speech signal, continuously training the parameters of the acoustic model so that the predicted value of the acoustic model approaches the true value, thereby completing the training of the acoustic model.
[0022] Combined with the first aspect, in some implementation manners of the first aspect, the labeled corpus may include an ideal near-field corpus and an ideal far-field corpus. Among them, the ideal near-field corpus corresponds to the near-field scenario, and the above-mentioned ideal far-field corpus corresponds to the far-field scenario.
[0023] In the near-field scenario, the ideal near-field corpus can be used as the training corpus of the above-mentioned acoustic model; in the far-field scenario, an ideal far-field corpus with low reverberation, low external noise, no human voice noise, and no echo can be used as the training corpus of the above-mentioned acoustic model. Finally, two different speech quality evaluation models are trained, and different speech quality evaluation models are selected for different scenarios to evaluate the speech quality, which is more targeted and has a higher correlation.
[0024] In a second aspect, a speech quality evaluation device is provided for performing the method in any of the possible implementation manners in the first aspect above. Specifically, the device includes a module for performing the method in any of the possible implementation manners in the first aspect above.
[0025] In a third aspect, another speech quality evaluation device is provided, including a processor, which is coupled to a memory and can be used to execute instructions in the memory to implement the method in any of the possible implementation manners in the first aspect above. Optionally, the device further includes a memory. Optionally, the device further includes a communication interface, and the processor is coupled to the communication interface.
[0026] In one implementation manner, the speech quality evaluation device is a data processing device. When the speech quality evaluation device is a data processing device, the communication interface can be a transceiver or an input / output interface.
[0027] In another implementation manner, the speech quality evaluation device is a chip configured in a server. When the speech quality evaluation device is a chip configured in a server, the communication interface can be an input / output interface.
[0028] Fourth aspect, a processor is provided, including: an input circuit, an output circuit, and a processing circuit. The processing circuit is configured to receive a signal through the input circuit and transmit a signal through the output circuit, so that the processor executes the method in any one of the possible implementation manners in the above first aspect.
[0029] In a specific implementation process, the above-mentioned processor may be a chip, the input circuit may be an input pin, the output circuit may be an output pin, and the processing circuit may be transistors, gate circuits, flip-flops, and various logic circuits, etc. The input signal received by the input circuit may be received and input by, for example but not limited to, a receiver. The signal output by the output circuit may be output to, for example but not limited to, a transmitter and transmitted by the transmitter. Moreover, the input circuit and the output circuit may be the same circuit, which serves as the input circuit and the output circuit at different times respectively. The embodiments of the present application do not limit the specific implementation manners of the processor and various circuits.
[0030] Fifth aspect, a processing device is provided, including a processor and a memory. The processor is configured to read instructions stored in the memory, and may receive a signal through a receiver and transmit a signal through a transmitter to execute the method in any one of the possible implementation manners in the above first aspect.
[0031] Optionally, there is one or more processors and one or more memories.
[0032] Optionally, the memory may be integrated with the processor or separately provided from the processor.
[0033] In a specific implementation process, the memory may be a non-transitory memory, such as a read only memory (ROM), which may be integrated with the processor on the same chip or separately provided on different chips. The embodiments of the present application do not limit the type of the memory and the setting manner of the memory and the processor.
[0034] It should be understood that relevant data interaction processes, such as sending indication information, may be a process of outputting indication information from the processor, and receiving capability information may be a process of the processor receiving input capability information. Specifically, the data output by the processing may be output to the transmitter, and the input data received by the processor may come from the receiver. Among them, the transmitter and the receiver may be collectively referred to as a transceiver.
[0035] The processing device in the above fifth aspect may be a chip, and the processor may be implemented by hardware or software. When implemented by hardware, the processor may be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor may be a general-purpose processor that is implemented by reading software code stored in a memory. The memory may be integrated in the processor or may be outside the processor and exist independently.
[0036] In a sixth aspect, a computer program product is provided. The computer program product includes: a computer program (which may also be referred to as code or instructions). When the computer program is run, it causes the computer to execute the method in any one of the possible implementation manners in the first aspect above.
[0037] In a seventh aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program (which may also be referred to as code or instructions). When it runs on a computer, it causes the computer to execute the method in any one of the possible implementation manners in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a schematic flowchart of the voice quality evaluation method in this embodiment;
[0039] Figure 2 is a schematic flowchart of the voice quality evaluation method in this embodiment;
[0040] Figure 3 is a training diagram of the voice quality evaluation model in this embodiment;
[0041] Figure 4 is a schematic block diagram of the voice quality evaluation device in this embodiment;
[0042] Figure 5 is a schematic block diagram of another voice quality evaluation device in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The technical solutions in the present application will be described below with reference to the accompanying drawings.
[0044] Artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning, and decision-making.
[0045] With the continuous development of artificial intelligence technology, natural language human-computer interaction systems that enable interaction between humans and machines through natural language have become increasingly important. For humans and machines to interact through natural language, the system needs to be able to recognize the specific meaning of human natural language, and thus speech recognition technology has emerged. Speech recognition is the process by which a machine converts speech signals into corresponding text or commands through recognition and understanding, and then performs corresponding operations.
[0046] For ease of understanding, first, the relevant terms involved in this application are described.
[0047] 1. Automatic Speech Recognition (ASR)
[0048] ASR is a technology in which a machine converts speech signals into corresponding text or commands through recognition and understanding. Speech recognition is a broad interdisciplinary field that has very close relationships with disciplines such as acoustics, phonetics, linguistics, information theory, pattern recognition theory, and neurobiology.
[0049] An ASR system can generally be divided into two modules: "front-end" and "back-end". Among them, the main functions of the ASR front-end module are noise reduction, feature extraction, and endpoint detection; the function of the ASR back-end module is to perform statistical pattern recognition, also known as "decoding", on the feature vectors of the user's speech using trained "acoustic models" and "language models" to obtain the text information contained therein.
[0050] The above-mentioned acoustic model can be obtained by training on speech data. The input is the feature vector, and the output is information on basic pronunciation units such as phonemes, syllables, pinyin, Chinese characters, etc. Through the acoustic model, the distance between the feature vector sequence of the speech and each pronunciation unit can be calculated. The acoustic model is an important part of the speech recognition system and determines the performance of the speech recognition system.
[0051] 2. Speech Enhancement (SE)
[0052] Speech enhancement refers to the technology of extracting useful speech signals from the noise background and suppressing and reducing noise interference when the speech signal is interfered with or even submerged by various noises. The main goal of speech enhancement is to extract as pure an original speech as possible from the noisy speech signal for subsequent recognition, thereby improving the speech quality and reducing the recognition error rate. Therefore, speech enhancement often serves ASR and belongs to ASR front-end technology. Specifically, speech enhancement technology usually uses a microphone array composed of multiple microphones in a far-field environment with significant external noise to utilize the redundancy and difference in time domain and space between the speech recorded by multiple microphones and the noise to improve the quality of the speech signal.
[0053] Speech enhancement is generally achieved through speech enhancement algorithms, which can specifically include processes such as echo cancellation, beamforming, beam tracking, noise suppression, reverberation cancellation, and automatic gain control. Common speech enhancement algorithms can be classified into the following categories: speech enhancement algorithms based on spectral subtraction, speech enhancement algorithms based on wavelet analysis, speech enhancement algorithms based on Kalman filtering, enhancement methods based on signal subspace, speech enhancement methods based on auditory masking effect, speech enhancement methods based on independent component analysis, speech enhancement methods based on neural networks, and other algorithms, which will not be listed one by one here.
[0054] 3. Speech Quality Evaluation Methods
[0055] The enhanced speech needs to use speech quality evaluation methods to evaluate whether the quality of the enhanced speech meets the requirements, so as to evaluate the quality of the speech enhancement algorithm. Speech quality evaluation methods can be divided into two categories: subjective evaluation and objective evaluation.
[0056] Subjective evaluation methods are based on people's subjective feelings, such as the mean opinion score (MOS) method. However, subjective evaluation methods have high requirements for evaluators, poor repeatability, and long cycles, and subjective evaluation methods have gradually been replaced by objective evaluation methods based on algorithms.
[0057] The evaluation indicators of typical objective evaluation methods can include: signal-to-distortion ratio (SDR), perceptual evaluation of speech quality (PESQ), short-time objective intelligibility (STOI), etc. These methods all require an ideal reference signal to evaluate the quality of the enhanced speech signal based on whether the enhanced speech signal is close to the reference signal in the time domain or frequency domain.
[0058] The speech quality evaluation models adopted by current speech quality evaluation methods are all based on the processed speech signals, unable to present the impact of changes in speech enhancement algorithms on ASR, and unable to be flexibly applied to a variety of different scenarios.
[0059] In a possible scenario, an ASR system can be extended from one device to another new device. For example, the ASR system is extended from a linear-array-based smart TV to a circular-array-based smart speaker. In this case, the voice enhancement algorithm at the ASR front end will be improved according to microphone arrays with different shapes and different numbers of microphones. The change in the voice enhancement algorithm can be reflected in the change in the acoustic features brought about by the change in the voice signal after voice enhancement processing. Since the original voice quality evaluation method does not take into account the features of ASR, and ASR is very sensitive to changes in acoustic features, even if the metrics obtained by the voice quality evaluation method improve, the ASR metrics may not necessarily improve.
[0060] In another possible scenario, the speech recognition ability of the ASR system can be opened to third-party applications in the form of a service. In this case, the ASR backend needs to use a voice quality evaluation method to test whether the quality of the voice signal accessed by the third-party application can meet the access conditions. Only when the quality of the voice signal meets the standard can the ASR backend ensure the accuracy of speech recognition and the effectiveness of the service. If the voice quality evaluation method does not consider the ASR backend, then for different third-party applications, even if they get the same voice quality score, the final metrics of the recognition accuracy may not be exactly the same, so it cannot reflect the impact of voice quality on ASR.
[0061] In view of this, the embodiments of the present application provide a new voice quality evaluation method and device. By using a voice quality evaluation model constructed by at least two acoustic models, the quality of the voice enhanced by the voice enhancement algorithm is evaluated, which reflects the impact of the voice enhancement algorithm on ASR and can be decoupled from a specific voice enhancement algorithm and a specific ASR backend, and is flexibly applicable to a variety of different scenarios.
[0062] The method of the embodiments of the present application is applicable to any electronic device capable of speech recognition. For example, smart speakers, smart TVs, computers, cars, telephones, mobile phones, tablets, laptops, handheld computers, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, personal digital assistants (PDAs), etc. equipped with ASR systems. The embodiments of the present application are not limited thereto.
[0063] Before introducing the voice quality evaluation method and device provided by the embodiments of the present application, the following points are explained first.
[0064] First, in the embodiments shown below, the terms and English abbreviations, such as the voice quality evaluation model, loss information, etc., are all exemplary examples given for convenience of description, and should not constitute any limitation to this application. This application does not exclude the possibility of defining other terms that can achieve the same or similar functions in existing or future protocols.
[0065] Second, in the embodiments shown below, the first, second, and various numerical numbers are only for the convenience of description and are not used to limit the scope of the embodiments of this application. For example, to distinguish different acoustic models, different sub-loss information, etc.
[0066] Third, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, and c can mean: a, or b, or c, or a and b, or a and c, or b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0067] The voice quality evaluation method and device provided by this application will be described in detail below with reference to the accompanying drawings.
[0068] Figure 1 It is a schematic flowchart of a voice quality evaluation method 100 provided for the embodiments of this application. As Figure 1 shown, the method 100 includes the following steps:
[0069] S101, Enhance the voice signal collected by the device through a voice enhancement algorithm to obtain the enhanced voice.
[0070] S102, Input the enhanced voice into the voice quality evaluation model to obtain the loss information of the enhanced voice. The voice quality evaluation model includes at least two acoustic models, and the loss information is determined based on the decoding results of the decoders of at least two acoustic models for the enhanced voice as the true values when calculating the loss.
[0071] S103, Based on the loss information, determine whether the quality of the enhanced voice meets the standard.
[0072] In an embodiment of the present application, an electronic device may collect a voice signal, perform voice enhancement processing on the collected voice signal through a voice enhancement algorithm to obtain enhanced voice. The electronic device may input the enhanced voice into a voice quality evaluation model to obtain loss information of the enhanced voice, and then based on the loss information, determine whether the quality of the enhanced voice meets the standard.
[0073] It should be understood that the voice quality evaluation model in the embodiment of the present application is composed of at least two acoustic models, so the voice quality evaluation model does not prefer a specific voice enhancement algorithm. Through each of the at least two acoustic models, taking the decoding result of the decoder of each acoustic model for the enhanced voice as the ground truth, loss information of the enhanced voice can be obtained. The loss information may also be referred to as perplexity or other names, and the embodiment of the present application does not limit this.
[0074] The voice quality evaluation model adopted in the embodiment of the present application does not prefer a specific voice enhancement algorithm, enabling the voice quality evaluation method to be decoupled from a specific voice enhancement algorithm and a specific ASR backend, and flexibly applicable to a variety of different scenarios. The method in the embodiment of the present application can be used to test whether the quality of a third-party accessed voice signal can meet the access conditions. If the voice quality meets the standard, it means that the voice signal meets the access conditions.
[0075] As an optional embodiment, the electronic device may determine whether the quality of the enhanced voice meets the standard by comparing the loss information with a predetermined threshold, that is, whether the enhanced voice meets the access conditions.
[0076] In a possible implementation manner, the above determining whether the quality of the enhanced voice meets the standard based on the loss information includes: if the loss information is less than or equal to the threshold, determining that the quality of the enhanced voice meets the standard; or, if the loss information is greater than the threshold, determining that the quality of the enhanced voice does not meet the standard.
[0077] In another possible implementation manner, the above determining whether the quality of the enhanced voice meets the standard according to the loss information of the enhanced voice includes: if the loss information is less than the threshold, determining that the quality of the enhanced voice meets the standard; or, if the loss information is greater than or equal to the threshold, determining that the quality of the enhanced voice does not meet the standard.
[0078] As an optional embodiment, after the above determining whether the quality of the enhanced voice meets the standard, the method further includes: if the quality of the enhanced voice meets the standard, the electronic device may accept the voice enhancement algorithm; or, if the quality of the enhanced voice does not meet the standard, the electronic device may reject the voice enhancement algorithm.
[0079] In the embodiments of the present application, accepting a voice enhancement algorithm can be understood as the ASR system of the electronic device allowing the access of the voice enhancement algorithm, and rejecting a voice enhancement algorithm can be understood as the ASR system of the electronic device rejecting the access of the voice enhancement algorithm.
[0080] As an optional embodiment, the above method further includes: optimizing the voice enhancement algorithm when the quality of the enhanced voice does not meet the standard.
[0081] Specifically, when the electronic device determines that the quality of the enhanced voice does not meet the standard, it can improve the voice enhancement algorithm, optimize the voice enhancement algorithm, and drive the enhanced voice obtained through the voice enhancement algorithm to approach the ideal far-field voice with low reverberation, low external noise, no human voice noise, and no echo in terms of acoustic features. Exemplarily, if the interference human voices in different directions are not eliminated cleanly, the beamforming algorithm in the voice enhancement algorithm can be focused on for optimization; if the amplitude of the voice signal is too small, the automatic gain control algorithm in the voice enhancement algorithm can be focused on for adjustment.
[0082] When the ASR system is extended from one device to another new device, the voice enhancement algorithm can be optimized by the method of the embodiments of the present application, without updating the ASR system, reducing the workload.
[0083] As an optional embodiment, the above voice quality evaluation model may include all or part of the following acoustic models: a model based on the convolutional neural network main structure (convolutional neural network, CNN) and the connectionist temporal classification (CTC) loss function; a model based on the transformer main structure and the Transducer loss function; a listen, attention and spell (LAS) model based on the cross entropy loss function; a hidden Markov model (HMM)-deep neural network (deep neural nrtwork, DNN) model based on the cross entropy.
[0084] In the embodiments of the present application, the voice quality evaluation model may include, but is not limited to, the various acoustic models listed above. The modeling units of these acoustic models may include pinyin, syllables, phonemes, etc. By comprehensively using a variety of typical acoustic models, the voice quality evaluation method of the embodiments of the present application can be neutral to specific acoustic models.
[0085] As an alternative embodiment, inputting the enhanced speech into the speech quality evaluation model to obtain the loss information of the enhanced speech includes: inputting the enhanced speech into at least two acoustic models respectively to obtain at least two sub-loss information, where the at least two sub-loss information corresponds to the at least two acoustic models; determining the loss information based on the at least two sub-loss information.
[0086] In the embodiment of the present application, the loss information of the enhanced speech can be determined according to the sub-loss information corresponding to each acoustic model in the at least two acoustic models.
[0087] As an alternative embodiment, the above loss information can be obtained by performing a weighted sum on the at least two sub-loss information.
[0088] Exemplarily, assume that the above speech quality evaluation model includes N acoustic models, the first acoustic model, the second acoustic model,..., the Nth acoustic model. Then each acoustic model corresponds to a weight. The first acoustic model corresponds to the weight w1, the second acoustic model corresponds to the weight w2, and so on. The Nth acoustic model corresponds to the weight w N , and the weights of each acoustic model can be preset, or different weights can be preset for different scenarios. The output results (i.e., sub-loss information) of each acoustic model are p1, p2,..., p N , then the loss information can be obtained through the following formula:
[0089] w1×p1 + w2×p2 +... + w N ×p N .
[0090] Taking the first acoustic model as an example below, the process of determining the first sub-loss information corresponding to the first acoustic model is described. It should be understood that the process of determining the sub-loss information of other acoustic models is similar to that of the first acoustic model and will not be elaborated here.
[0091] As an alternative embodiment, inputting the enhanced speech into the at least two acoustic models respectively to obtain at least two sub-loss information includes: inputting the enhanced speech into the first acoustic model among the at least two acoustic models to obtain the decoding result of the decoder in the first acoustic model; using the decoding result of the decoder in the first acoustic model as the true value to calculate the first sub-loss information of the enhanced speech.
[0092] Figure 2A schematic flowchart of the voice quality evaluation method according to an embodiment of the present application is shown. Assume that the ASR system is installed on a new electronic device (hereinafter referred to as the new device). To make the voice quality evaluation method according to the embodiment of the present application more applicable to the new device, the voice enhancement algorithm of the ASR system can be optimized through a test data set. Specifically, Figure 2 By using the test data set, voice signals are obtained, the voice signals are enhanced through the voice enhancement algorithm, and then the quality of the enhanced voice signals is evaluated through the voice quality evaluation method according to the embodiment of the present application. When the voice enhancement algorithm does not meet the access conditions, the voice enhancement algorithm is further optimized.
[0093] First, the process of constructing the test data set is introduced. Exemplarily, recording can be performed in a recording studio to collect standard microphone corpus, which is also referred to as "standard microphone corpus" in the embodiment of the present application; noise corpus is collected in typical scenarios, which can be a living room, a dining room, a car traveling at high speed, etc., and the embodiment of the present application does not limit this. The above standard microphone corpus and noise corpus can be collectively referred to as reference corpus. When the ASR system is installed on the new device, the application scenario of the new device can be determined first, for example, noise type, signal-to-noise ratio, human voice interference, room size and reverberation, distance between the microphone and the human, etc. Then, a suitable recording room is set up and a suitable environment is arranged, and then the above reference corpus is played through an artificial mouth to simulate the speech of a human and the propagation of noise, so as to form real multi-channel voice signals adapted to the corresponding scenario of the new device, including Figure 2 the original recording and noise recording shown. It should be understood that the above factors such as noise type, signal-to-noise ratio, and human voice interference can be controlled by time-domain addition synthesis of recorded multi-channel speakers.
[0094] Then, the above multi-channel voice signals are processed through the voice enhancement algorithm to obtain enhanced voice. The loss information of the enhanced voice is calculated by using the voice quality evaluation method according to the embodiment of the present application. If the loss information is less than the threshold, the voice enhancement algorithm is accepted, and it is determined that the voice enhancement algorithm meets the requirements and no optimization is required; if the loss information is greater than or equal to the threshold, the voice enhancement algorithm is rejected, and the voice enhancement algorithm is further optimized.
[0095] It should be understood that after obtaining the optimized voice enhancement algorithm, the above method can also be used to continue to determine whether the algorithm needs to be further optimized until the voice enhancement algorithm meets the requirements.
[0096] It should also be understood that, generally, an electronic device performs enhancement processing on a batch of audio through a voice enhancement algorithm, and then calculates the loss information of the batch of audio through a voice quality evaluation model respectively. If the mean value of the loss information of the batch of audio is less than a threshold, the electronic device can determine that the voice quality of the enhanced voice signal meets the standard; otherwise, the voice quality of the enhanced voice signal does not meet the standard, and the voice enhancement algorithm can be continuously improved.
[0097] Based on the existing voice quality evaluation model above, the usage process of the voice quality evaluation model is introduced. Before using the voice quality evaluation model, the voice quality evaluation model can also be trained. Next, the training process of the voice quality evaluation model according to the embodiments of the present application is introduced.
[0098] As an optional embodiment, the above method further includes: training the at least two acoustic models based on the labeled corpus to obtain a voice quality evaluation model.
[0099] Specifically, since the voice quality evaluation model includes at least two acoustic models, training the voice quality evaluation model is to train the at least two acoustic models. The training process is as Figure 3 shown. The voice corpus is the labeled corpus of voice signals. The labeled text of a voice signal can be understood as the true value of the voice signal. Taking the voice signal as the input of an acoustic model, outputting the predicted labeled text of the voice signal, comparing the predicted labeled text as the predicted value with the true value of the voice signal, and continuously training the parameters of the acoustic model so that the predicted value of the acoustic model approaches the true value, thereby completing the training of the acoustic model.
[0100] Optionally, the above labeled corpus may include an ideal near-field corpus and an ideal far-field corpus. Among them, the above ideal near-field corpus corresponds to the near-field scenario, and the above ideal far-field corpus corresponds to the far-field scenario. In a possible implementation manner, in the near-field scenario, the ideal near-field corpus can be used as the training corpus of the above acoustic model; in the far-field scenario, an ideal far-field corpus with low reverberation, low external noise, no human voice noise, and no echo can be used as the training corpus of the above acoustic model.
[0101] It should be understood that the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0102] In the above text, in combination with Figures 1 to 3 , the usage and training methods of the voice quality evaluation model according to the embodiments of the present application are described in detail. Next, in combination with Figures 4 to 5 , the device for voice quality evaluation according to the embodiments of the present application will be described in detail.
[0103] Figure 4 Fig. 2 shows a voice quality evaluation device 400 provided by an embodiment of the present application. The device 400 includes: an enhancement module 401, a processing module 402, and a judgment module 403.
[0104] The enhancement module 401 is configured to enhance the voice signal collected by the device through a voice enhancement algorithm to obtain enhanced voice. The processing module 402 is configured to input the enhanced voice into a voice quality evaluation model to obtain loss information of the enhanced voice. The voice quality evaluation model includes at least two acoustic models. The loss information is determined based on the decoding results of the decoders of at least two acoustic models for the enhanced voice as the ground truth. The judgment module 403 is configured to judge whether the quality of the enhanced voice meets the standard based on the loss information.
[0105] Optionally, the judgment module 403 determines that the quality of the enhanced voice meets the standard if the loss information is less than or equal to a threshold; or determines that the quality of the enhanced voice does not meet the standard if the loss information is greater than the threshold.
[0106] Optionally, the judgment module 403 is configured to accept the voice enhancement algorithm if the quality of the enhanced voice meets the standard; or reject the voice enhancement algorithm if the quality of the enhanced voice does not meet the standard.
[0107] Optionally, the processing module 402 is configured to optimize the voice enhancement algorithm when the quality of the enhanced voice does not meet the standard.
[0108] Optionally, the at least two acoustic models include all or part of the following acoustic models:
[0109] A model based on a convolutional neural network (CNN) structure and a connectionist temporal classification (CTC) loss function; a model based on a Transformer structure and a transducer loss function; a listen, attend and spell (LAS) model based on a cross-entropy loss function; a hidden Markov model - deep neural network (HMM-DNN) model based on a cross-entropy loss function.
[0110] Optionally, the processing module 402 is configured to input the enhanced voice into the at least two acoustic models respectively to obtain at least two sub-loss information, and the at least two sub-loss information corresponds to the at least two acoustic models; determine the loss information based on the at least two sub-loss information.
[0111] Optionally, the loss information is obtained by performing weighted summation on the at least two sub-loss information.
[0112] Optionally, the processing module 402 is configured to input the enhanced speech into a first acoustic model of the at least two acoustic models, and obtain a decoding result of a decoder in the first acoustic model; use the decoding result of the decoder in the first acoustic model as a ground truth, and calculate first sub-loss information of the enhanced speech.
[0113] Optionally, the apparatus 400 further includes: a training module 404, configured to train the at least two acoustic models based on the labeled corpus to obtain a speech quality evaluation model.
[0114] Optionally, the labeled corpus includes: an ideal near-field corpus and an ideal far-field corpus, where the ideal near-field corpus corresponds to a near-field scenario, and the ideal far-field corpus corresponds to a far-field scenario.
[0115] It should be understood that the apparatus 400 is embodied in the form of functional modules herein. The term "module" herein may refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor, or a group of processors, etc.) for executing one or more software or firmware programs, a memory, a combined logic circuit, and / or other suitable components that support the described functions. In an alternative example, those skilled in the art can understand that the apparatus 400 may specifically be the electronic device in the foregoing embodiments, or the functions of the electronic device in the foregoing embodiments may be integrated in the apparatus 400, and the apparatus 400 may be configured to execute each process and / or step corresponding to the electronic device in the foregoing method embodiments. To avoid repetition, details are not described herein again.
[0116] The apparatus 400 has the function of implementing the corresponding steps executed by the electronic device in the foregoing method; the foregoing function may be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the foregoing function.
[0117] In the embodiments of the present application, Figure 4 the apparatus 400 in may also be a chip or a chip system, for example: a system on chip (SoC).
[0118] Figure 5 Another speech quality evaluation apparatus 500 provided by an embodiment of the present application is shown. The apparatus 500 includes a processor 501, a transceiver 502, and a memory 503. Among them, the processor 501, the transceiver 502, and the memory 503 communicate with each other through an internal connection path. The memory 503 is configured to store instructions, and the processor 501 is configured to execute the instructions stored in the memory 503 to control the transceiver 502 to send signals and / or receive signals.
[0119] It should be understood that the device 500 may specifically be the electronic device in the above embodiments, or the functions of the electronic device in the above embodiments may be integrated in the device 500, and the device 500 may be used to execute each step and / or process corresponding to the electronic device in the above method embodiments. Optionally, the memory 503 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may further include a non-volatile random access memory. For example, the memory may also store information about the device type. The processor 501 may be used to execute the instructions stored in the memory, and when the processor executes the instructions, the processor may execute each step and / or process corresponding to the electronic device in the above method embodiments.
[0120] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0121] In the implementation process, each step of the above method may be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed and completed by the hardware processor, or executed and completed by a combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor executes the instructions in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0122] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0123] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments, and will not be described in detail here.
[0124] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0125] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0126] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0127] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0128] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for evaluating speech quality, characterized in that, Including: Enhancing the voice signal collected by the device through a voice enhancement algorithm to obtain enhanced voice; Inputting the enhanced voice into a voice quality evaluation model to obtain loss information of the enhanced voice, where the voice quality evaluation model includes at least two acoustic models, and the loss information is determined based on the decoding results of the enhanced voice by the decoders of the at least two acoustic models as the ground truth; Based on the loss information, determining whether the quality of the enhanced voice meets the standard and evaluating the voice enhancement algorithm.
2. The method according to claim 1, characterized in that The determining whether the quality of the enhanced voice meets the standard based on the loss information includes: If the loss information is less than or equal to a threshold, determining that the quality of the enhanced voice meets the standard; or, If the loss information is greater than the threshold, determining that the quality of the enhanced voice does not meet the standard.
3. The method according to claim 1 or 2, characterized in that, The evaluating the voice enhancement algorithm includes: If the quality of the enhanced voice meets the standard, accepting the voice enhancement algorithm; or, If the quality of the enhanced voice does not meet the standard, rejecting the voice enhancement algorithm.
4. The method according to claim 3, characterized in that, The method further includes: Optimizing the voice enhancement algorithm in the case where the quality of the enhanced voice does not meet the standard.
5. The method according to any one of claims 1 to 4, characterized in that The at least two acoustic models include all or part of the following acoustic models: A model based on a convolutional neural network (CNN) structure and a connectionist temporal classification (CTC) loss function; A model based on a Transformer structure and a transducer loss function; A listen, attend and spell (LAS) model based on a cross-entropy loss function; A hidden Markov model - deep neural network (HMM-DNN) model based on a cross-entropy loss function.
6. The method according to any one of claims 1 to 5, characterized in that The inputting the enhanced voice into a voice quality evaluation model to obtain loss information of the enhanced voice includes: Inputting the enhanced voice into the at least two acoustic models respectively to obtain at least two sub-loss information, where the at least two sub-loss information corresponds to the at least two acoustic models; Based on the at least two sub-loss information, determining the loss information.
7. The method according to claim 6, wherein The loss information is obtained by performing weighted summation on the at least two sub-loss information.
8. The method according to claim 6 or 7, characterized in that The inputting the enhanced voice into the at least two acoustic models respectively to obtain at least two sub-loss information includes: Inputting the enhanced voice into the first acoustic model of the at least two acoustic models to obtain the decoding result of the decoder in the first acoustic model; Using the decoding result of the decoder in the first acoustic model as the ground truth to calculate the first sub-loss information of the enhanced voice.
9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: Training the at least two acoustic models based on labeled corpus to obtain the voice quality evaluation model.
10. The method according to claim 9, wherein The labeled corpus includes: Ideal near-field corpus and ideal far-field corpus, where the ideal near-field corpus corresponds to the near-field scenario and the ideal far-field corpus corresponds to the far-field scenario.
11. A voice quality evaluation device, characterized in that, Including: An enhancement module for enhancing the voice signal collected by the device through a voice enhancement algorithm to obtain enhanced voice; A processing module, configured to input the enhanced speech into a speech quality evaluation model to obtain loss information of the enhanced speech, where the speech quality evaluation model includes at least two acoustic models, and the loss information is determined based on the decoding results of the enhanced speech by the decoders of the at least two acoustic models as the ground truth; A judgment module, configured to judge whether the quality of the enhanced speech meets the standard based on the loss information, and evaluate the speech enhancement algorithm.
12. The device according to claim 11, characterized in that, Specifically, the judgment module is configured to: If the loss information is less than or equal to a threshold, determine that the quality of the enhanced speech meets the standard; or, If the loss information is greater than the threshold, determine that the quality of the enhanced speech does not meet the standard.
13. The device according to claim 11 or 12, characterized in that, Specifically, the judgment module is configured to: If the quality of the enhanced speech meets the standard, accept the speech enhancement algorithm; or, If the quality of the enhanced speech does not meet the standard, reject the speech enhancement algorithm.
14. The device according to claim 13, characterized in that, Specifically, the processing module is configured to: Optimize the speech enhancement algorithm in the case that the quality of the enhanced speech does not meet the standard.
15. The device according to any one of claims 11 to 14, characterized in that, The at least two acoustic models include all or part of the following acoustic models: A model based on a convolutional neural network (CNN) structure and a connectionist temporal classification (CTC) loss function; A model based on a Transformer structure and a transducer loss function; A listen, attend and spell (LAS) model based on a cross-entropy loss function; A hidden Markov model - deep neural network (HMM-DNN) model based on a cross-entropy loss function.
16. The device according to any one of claims 11 to 15, characterized in that, Specifically, the processing module is configured to: Input the enhanced speech into the at least two acoustic models respectively to obtain at least two sub-loss information, and the at least two sub-loss information corresponds to the at least two acoustic models; Determine the loss information based on the at least two sub-loss information.
17. The device according to claim 16, characterized in that, The loss information is obtained by performing weighted summation on the at least two sub-loss information.
18. The device according to claim 16 or 17, characterized in that, Specifically, the processing module is configured to: Input the enhanced speech into a first acoustic model of the at least two acoustic models to obtain a decoding result of a decoder in the first acoustic model; Use the decoding result of the decoder in the first acoustic model as the ground truth, and calculate first sub-loss information of the enhanced speech.
19. The device according to any one of claims 11 to 18, characterized in that, The device further includes: A training module, configured to train the at least two acoustic models based on the labeled corpus to obtain the speech quality evaluation model.
20. The device according to claim 19, characterized in that, The labeled corpus includes: An ideal near-field corpus and an ideal far-field corpus, where the ideal near-field corpus corresponds to a near-field scenario, and the ideal far-field corpus corresponds to a far-field scenario.
21. A voice quality evaluation device, characterized in that, It includes: A processor, where the processor is coupled to a memory, and the memory is used to store a computer program. When the processor calls the computer program, the device executes the method according to any one of claims 1 to 10.
22. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program includes instructions for implementing the method according to any one of claims 1 to 10.
23. A computer program product, characterized in that, For storing a computer program which, when run, causes a computer to perform the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Voice evaluation method and device
CN104464757A