Method and apparatus for speech quality assessment, speech recognition quality prediction and improvement

By employing methods for speech quality assessment and speech recognition quality prediction, the decoupling problem between the speech preprocessing module and the recognition module was solved, enabling independent optimization and improved fault diagnosis efficiency, thereby enhancing the voice interaction experience.

CN116210050BActive Publication Date: 2026-07-24YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YINWANG INTELLIGENT TECHNOLOGIES CO LTD
Filing Date
2021-09-30
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

The lack of existing technologies in systems and standards for separating and independently optimizing the speech preprocessing module and the speech recognition module makes it difficult to assess the quality of speech signals, locate and resolve problems in the speech recognition system independently, and affect the efficiency of fault diagnosis.

Method used

This paper provides a speech quality assessment method that evaluates speech quality by acquiring semantic information related to test speech and predicts speech recognition quality using a speech recognition quality function. This decouples preprocessing from speech recognition and allows for independent tuning of each module.

Benefits of technology

It enables effective evaluation of speech signal quality and prediction of speech recognition quality, and can independently locate and improve problems in the speech recognition system, thereby improving fault diagnosis efficiency and voice interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116210050B_ABST
    Figure CN116210050B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of speech recognition, and provides a speech quality evaluation method, a speech recognition quality prediction method and a method for improving speech recognition quality. First, test speech is acquired, then the speech quality of the test speech is evaluated according to semantic related information of the test speech, so as to determine whether the pre-processing of the speech needs to be adjusted, and the speech recognition quality is predicted according to the evaluated speech quality, so as to determine whether the parameters of the speech recognition model need to be adjusted, and the corresponding evaluation result or prediction result is output, so that the pre-processing or the parameter adjustment of the speech recognition model can be performed according to the evaluation result or the prediction result. The pre-processing and the parameter adjustment of the speech recognition model in the speech recognition process can be decoupled through the application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and in particular to speech quality assessment methods and apparatus, speech recognition quality prediction methods and apparatus, methods and apparatus for improving speech recognition quality, vehicles, computer-readable storage media, and computer program products. Background Technology

[0002] In a speech recognition system, the speech signal is acquired by a sensor, enhanced by a speech preprocessing module, and then sent to the speech recognition module for voice wake-up and recognition. Therefore, the recognition effect of a speech recognition system mainly depends on two factors: the performance of the speech recognition module and the quality of the speech signal. The quality of the speech signal is constrained by the environment, the audio acquisition hardware, and the algorithm of the speech preprocessing module.

[0003] The industry commonly uses a method of integrated debugging of the speech recognition module and the speech preprocessing module to test speech recognition quality and optimize performance. However, there is a lack of available systems and standards that allow for the independent optimization of the speech preprocessing module and the speech recognition module. This prevents the evaluation of the speech signal quality output by the speech preprocessing module, thus hindering the provision of optimization baselines and feedback for the acquisition hardware, speech preprocessing module, and speech recognition module. This approach is not conducive to the independent location and resolution of speech signal quality and speech recognition module performance issues in actual speech recognition operations, and it can easily lead to difficulties in troubleshooting speech recognition system faults. Therefore, decoupling the speech recognition module from the speech preprocessing module to obtain speech quality calibration and feedback is of great significance for improving the efficiency of module fault diagnosis. Summary of the Invention

[0004] In view of the above, this application provides a speech quality assessment method and apparatus, a speech recognition quality prediction method and apparatus, a method and apparatus for improving speech recognition quality, a vehicle, a computer-readable storage medium, and a computer program product.

[0005] To achieve the above objectives, a first aspect of this application provides a speech quality assessment method, comprising: acquiring test speech; assessing the speech quality of the test speech based on semantically relevant information of the test speech; and outputting the speech quality assessment result.

[0006] As described above, speech quality assessment of test speech can be achieved. This assessment is based on semantically relevant information about the test speech, thus the assessed speech quality reflects semantic information. Therefore, this assessed speech quality can be used to predict the speech recognition quality of the speech recognition model for the test speech. Furthermore, the output assessment results can include relevant content related to the assessed speech quality, which users can refer to to improve their speech quality.

[0007] For example, when applied to vehicles, it can enable in-vehicle voice quality monitoring and feedback, allowing users to take appropriate measures to maintain or improve the in-vehicle voice interaction experience. In another possible implementation, the vehicle can also perform corresponding operations based on the output evaluation results to improve the in-vehicle voice interaction experience. These operations could include automatically closing windows, automatically reducing in-vehicle volume (such as the volume of music playing in the vehicle), or automatically optimizing parameters.

[0008] As one possible implementation of the first aspect, the evaluation results of the speech quality include one or more of the following: the quality of the test speech; factors affecting the quality of the test speech; and methods for adjusting the quality of the test speech.

[0009] As described above, the output content can be flexibly configured. For example, the quality of the output test speech can be quantified with specific parameters, qualitatively rated as excellent, good, or average, or combined with images and text. Factors affecting test speech quality include external noise from an open window, excessive tire or engine noise due to high speed, and loud music or other sounds inside the vehicle, allowing users to make adjustments as needed. Adjustments to test speech quality can include closing windows, reducing speed, lowering in-car volume, replacing faulty microphones, and optimizing pre-processing module parameters.

[0010] As one possible implementation of the first aspect, evaluating the speech quality of test speech based on semantically relevant information of test speech includes: obtaining a first feature vector of test speech, wherein the first feature vector includes a time-frequency feature vector of test speech; obtaining a second feature vector of test speech based on the first feature vector of test speech, wherein the second feature vector is semantically relevant to test speech; and evaluating the speech quality of test speech based on the second feature vector.

[0011] As described above, a first feature vector, including the time-frequency feature vector of the speech, can be obtained first. Then, a second feature vector, semantically related to the speech, can be obtained based on this first feature vector. The speech quality of the test speech can then be evaluated using the second feature vector. Since the first feature vector is related to the time-frequency feature vector, the preprocessing parameters are connected to the second feature vector. Therefore, the speech quality of the test speech evaluated by the second feature vector can be used as a reference for optimizing the preprocessing parameters.

[0012] As one possible implementation of the first aspect, evaluating the speech quality of test speech based on the second feature vector includes: evaluating the speech quality of test speech using a first evaluation index, the first evaluation index including: a concentration index of the center positions of each group of test speech, wherein different groups of test speech have different semantics, the center position of each group of test speech is the center position of the second feature vector of that group of test speech in the second feature space, and the second feature space is the space in which the second feature vector is located.

[0013] Based on the above, the speech quality of the test speech can be evaluated according to one of the optional evaluation metrics mentioned above, which is calculated based on the second feature vector.

[0014] As one possible implementation of the first aspect, evaluating the speech quality of test speech based on the second feature vector includes: evaluating the speech quality of test speech using a second evaluation index, the second evaluation index including: a dispersion index of the second feature vectors of each group of test speech in a second feature space, wherein different groups of test speech have different semantics, and the second feature space is the space where the second feature vectors are located.

[0015] Based on the above, the speech quality of the test speech can be evaluated according to one of the optional evaluation metrics, which is calculated based on the second feature vector.

[0016] As one possible implementation of the first aspect, evaluating the speech quality of test speech based on the second feature vector includes: evaluating the speech quality of test speech using a third evaluation index, the third evaluation index including: the similarity index between the center position of each group of test speech and the center position of each group of reference speech corresponding to the semantics; wherein, different groups of test speech have different semantics, the center position of each group of test speech is the center position of the second feature vector of that group of test speech in the second feature space, the second feature space is the space where the second feature vector is located, different groups of reference speech have different semantics, and the center position of each group of reference speech is the center position of the second feature vector of that group of reference speech in the second feature space.

[0017] Based on the above, the speech quality of the test speech can be evaluated according to one of the optional evaluation metrics, which is calculated based on the second feature vector.

[0018] As one possible implementation of the first aspect, obtaining a first feature vector of the test speech includes: acquiring consecutive frames contained in the test speech, wherein adjacent frames contain overlapping information; and obtaining multiple feature vectors including frequency domain features based on the consecutive frames, wherein the multiple feature vectors constitute the first feature vector.

[0019] Therefore, when obtaining the first feature vector, the relationship between frames can be incorporated into the calculation of the second feature vector by using adjacent frames with overlap.

[0020] To achieve the above objectives, a second aspect of this application provides a speech recognition quality prediction method, comprising: acquiring test speech; predicting the speech recognition quality of the test speech by a speech recognition model based on a speech recognition quality function, wherein the speech recognition quality function is used to indicate the relationship between speech recognition quality and speech quality, and the speech quality is evaluated according to the method of the first aspect or any possible implementation thereof; and outputting the prediction result of the speech recognition quality.

[0021] As described above, this method uses voice quality to predict the quality of speech recognition models, which users can then use to improve their speech recognition performance. For example, when applied to vehicles, it allows for the monitoring and feedback of in-vehicle speech recognition quality, enabling users to take appropriate measures to maintain a good in-vehicle voice interaction experience.

[0022] As one possible implementation of the second aspect, the output prediction result of speech recognition quality includes one or more of the following: the speech recognition quality of the speech recognition model; factors affecting the speech recognition quality of the speech recognition model; and methods for adjusting the speech recognition quality of the speech recognition model.

[0023] Therefore, the predicted results can be flexibly configured. For example, the quality of the speech recognition model can be provided as specific quantitative parameters, or as a qualitative assessment (excellent, good, moderate), or combined with images and text. For instance, the factors affecting the speech recognition quality of the output can be those provided in the first aspect, or performance factors of the speech recognition module. The method for adjusting the output to test speech quality can be the adjustment method provided in the first aspect, or it can be a parameter tuning suggestion for the speech recognition model.

[0024] As one possible implementation of the second aspect, the process of constructing the speech recognition quality function includes: obtaining multiple sets of degraded reference speech; obtaining speech recognition results of the multiple sets of degraded reference speech according to the speech recognition model, and using them as first statistical results; using the multiple sets of degraded reference speech as test speech respectively, obtaining speech quality evaluation results of the multiple sets of degraded reference speech according to the method of the first aspect or any possible implementation of the first aspect, and obtaining a first evaluation result; and obtaining the speech recognition quality function according to the functional relationship between the first statistical result and the first evaluation result.

[0025] The above describes one way to construct a speech recognition quality function using reference speech. Specifically, without introducing other reference speech, a method of degradation using reference speech is adopted, which can effectively reduce the amount of reference speech data.

[0026] To achieve the above objectives, a third aspect of this application provides a method for improving speech recognition quality, comprising: acquiring test speech; obtaining a speech quality evaluation result of the test speech according to the method of the first aspect or any possible implementation thereof; and outputting the speech quality evaluation result when the speech quality is lower than a preset first baseline.

[0027] From the above, the speech quality assessment result of the test speech can be obtained, and this result includes content related to speech quality, such as whether preprocessing parameters need to be adjusted. For example, when applied to vehicles, it can monitor and provide feedback on in-vehicle speech quality, allowing users to take appropriate measures to maintain the in-vehicle voice interaction experience. This process can be independent of the speech recognition process of the speech recognition model, achieving decoupling between preprocessing and speech recognition.

[0028] As a possible implementation of the third aspect, it further includes: when the speech quality is greater than or equal to the first baseline, predicting the speech recognition quality according to the method of the second aspect or any possible implementation of the second aspect; when the speech recognition quality is lower than the preset second baseline, executing the prediction result of the speech recognition quality.

[0029] As described above, it's possible to determine whether to adjust preprocessing parameters or the speech recognition model based on test audio, thus decoupling preprocessing and speech recognition. This facilitates independent problem localization and performance optimization for each module.

[0030] To achieve the above objectives, a fourth aspect of this application provides a speech quality assessment device, comprising:

[0031] The module is used to acquire test speech; the module is used to evaluate the speech quality of the test speech based on semantic information; and the module is used to output the speech quality evaluation results.

[0032] As described above, speech quality assessment of test speech can be achieved. This assessment is based on semantically relevant information about the test speech, thus reflecting semantic information in the assessed speech quality. This assessed speech quality can then be used to predict the speech recognition quality of the speech recognition model for the test speech. Furthermore, the output assessment results can include content related to the assessed speech quality, which users can use to improve their speech quality. For example, when applied to vehicles, it can enable in-vehicle speech quality monitoring and feedback, allowing users to take appropriate measures to maintain or improve the in-vehicle voice interaction experience. In one possible implementation, the vehicle can also perform corresponding operations based on the output assessment results to improve the in-vehicle voice interaction experience. These operations could include automatically closing windows, automatically reducing in-vehicle volume (such as the volume of music playing in the vehicle), or automatically optimizing parameters.

[0033] As one possible implementation of the fourth aspect, the evaluation result of the voice quality output by the output module includes one or more of the following: the quality of the test voice; factors affecting the quality of the test voice; and the method of adjusting the quality of the test voice.

[0034] As one possible implementation of the fourth aspect, the evaluation module is specifically used to: obtain a first feature vector of the test speech, wherein the first feature vector includes a time-frequency feature vector of the test speech; obtain a second feature vector of the test speech based on the first feature vector of the test speech, wherein the second feature vector is semantically related to the test speech; and evaluate the speech quality of the test speech based on the second feature vector.

[0035] As a possible implementation of the fourth aspect, when the evaluation module evaluates the speech quality of the test speech based on the second feature vector, it includes: evaluating the speech quality of the test speech using a first evaluation index, the first evaluation index including: the concentration index of the center position of each group of test speech, wherein different groups of test speech have different semantics, the center position of each group of test speech is the center position of the second feature vector of that group of test speech in the second feature space, and the second feature space is the space where the second feature vector is located.

[0036] As a possible implementation of the fourth aspect, when the evaluation module evaluates the speech quality of the test speech based on the second feature vector, it includes: evaluating the speech quality of the test speech using a second evaluation index, the second evaluation index including: the dispersion index of the second feature vectors of each group of test speech in the second feature space, wherein the test speech of different groups has different semantics, and the second feature space is the space where the second feature vectors are located.

[0037] As a possible implementation of the fourth aspect, when the evaluation module evaluates the speech quality of the test speech based on the second feature vector, it includes: evaluating the speech quality of the test speech using a third evaluation index, the third evaluation index including: the similarity index between the center position of each group of test speech and the center position of each group of reference speech corresponding to the semantics; wherein, different groups of test speech have different semantics, the center position of each group of test speech is the center position of the second feature vector of that group of test speech in the second feature space, the second feature space is the space where the second feature vector is located, different groups of reference speech have different semantics, and the center position of each group of reference speech is the center position of the second feature vector of that group of reference speech in the second feature space.

[0038] As a possible implementation of the fourth aspect, when the evaluation module obtains the first feature vector of the test speech, it includes: acquiring consecutive frames contained in the test speech, wherein adjacent frames contain overlapping information; and obtaining multiple feature vectors including frequency domain features based on the consecutive frames, wherein the multiple feature vectors constitute the first feature vector.

[0039] To achieve the above objectives, a fifth aspect of this application provides a speech recognition quality prediction apparatus, comprising: an acquisition module for acquiring test speech; a prediction module for predicting the speech recognition quality of the test speech by a speech recognition model based on a speech recognition quality function, wherein the speech recognition quality function indicates the relationship between speech recognition quality and speech quality, and the speech quality is evaluated according to the first aspect or any possible implementation thereof; and an output module for outputting the prediction result of the speech recognition quality.

[0040] As described above, this method uses voice quality to predict the quality of speech recognition models, which users can then use to improve their speech recognition performance. For example, when applied to vehicles, it allows for the monitoring and feedback of in-vehicle speech recognition quality, enabling users to take appropriate measures to maintain a good in-vehicle voice interaction experience.

[0041] As one possible implementation of the fifth aspect, the prediction result of the speech recognition quality output by the output module includes one or more of the following: the speech recognition quality of the speech recognition model; factors affecting the speech recognition quality of the speech recognition model; and methods for adjusting the speech recognition quality of the speech recognition model.

[0042] As one possible implementation of the fifth aspect, the process of constructing the speech recognition quality function includes: obtaining multiple sets of degraded reference speech; obtaining speech recognition results of the multiple sets of degraded reference speech according to the speech recognition model, and using them as first statistical results; using the multiple sets of degraded reference speech as test speech respectively, obtaining speech quality evaluation results of the multiple sets of degraded reference speech according to the method of the first aspect or any possible implementation of the first aspect, and obtaining a first evaluation result; and obtaining the speech recognition quality function according to the functional relationship between the first statistical result and the first evaluation result.

[0043] To achieve the above objectives, a sixth aspect of this application provides an apparatus for improving speech recognition quality, comprising: an acquisition module for acquiring test speech; an evaluation and prediction module for obtaining a speech quality evaluation result of the test speech according to the method of the first aspect or any possible implementation thereof; and an output module for outputting the speech quality evaluation result when the speech quality is lower than a preset first baseline.

[0044] From the above, the speech quality assessment result of the test speech can be obtained, and based on this result, speech quality-related content can be output, including whether preprocessing parameters need to be adjusted. For example, when applied to vehicles, it can monitor and provide feedback on in-vehicle speech quality, allowing users to take appropriate measures to maintain the in-vehicle voice interaction experience. This process can be independent of the speech recognition process of the speech recognition model, achieving decoupling between preprocessing and speech recognition.

[0045] As a possible implementation of the sixth aspect, the evaluation prediction module is further configured to predict the speech recognition quality according to the method of the second aspect or any possible implementation of the second aspect when the speech quality is greater than or equal to the first baseline; the output module is further configured to execute the output speech recognition quality prediction result when the speech recognition quality is lower than the preset second baseline.

[0046] To achieve the above objectives, a seventh aspect of this application provides a vehicle comprising: a sound acquisition device for acquiring a user's voice commands; a preprocessing device for preprocessing the voice commands; a voice recognition device for recognizing the preprocessed voice commands; and the devices described in the fourth, fifth, and sixth aspects and any possible embodiments thereof.

[0047] To achieve the above objectives, the eighth aspect of this application provides a computing device including one or more processors and one or more memories, the memories storing program instructions that, when executed by one or more processors, cause the one or more processors to implement the method of the first aspect and any possible implementation thereof.

[0048] To achieve the above objectives, the ninth aspect of this application provides a computer-readable storage medium having program instructions stored thereon, which, when executed by a computer, cause the computer to implement the method of the first aspect and any possible implementation thereof.

[0049] To achieve the above objectives, the tenth aspect of this application provides a computer program product including program instructions that, when executed by a computer, cause the computer to implement the method of the first aspect and any possible implementation thereof.

[0050] In summary, the embodiments of this application decouple the evaluation of the speech preprocessing process from the speech recognition model's prediction of speech recognition, enabling the separate localization of preprocessing and speech recognition issues, which facilitates independent problem localization and performance optimization of each module. Furthermore, due to this decoupling, the evaluation of test speech quality can be used to prompt users with corresponding actions, thereby improving the in-vehicle voice interaction experience; similarly, the prediction of speech recognition can be used to prompt users with corresponding actions, further improving the in-vehicle voice interaction experience.

[0051] These and other aspects of this application will become more apparent in the description of the following embodiments(s). Attached Figure Description

[0052] The following description, with reference to the accompanying drawings, further illustrates the various features of this application and the relationships between them. The drawings are exemplary; some features are not shown to scale, and some drawings may omit conventional features in the field of this application that are not essential to it, or additional features that are not essential to this application may be shown. The combination of features shown in the drawings is not intended to limit this application. Furthermore, throughout this specification, the same reference numerals refer to the same things. Specific descriptions of the drawings are as follows:

[0053] Figure 1 This is a structural diagram of an application scenario involved in an embodiment of this application;

[0054] Figure 2A A schematic flowchart illustrating the speech quality assessment method provided in this application embodiment;

[0055] Figure 2B A schematic flowchart illustrating the speech quality assessment method for test speech provided in this application embodiment;

[0056] Figure 3A A flowchart illustrating the speech recognition quality prediction method provided in this application embodiment;

[0057] Figure 3B A flowchart illustrating the speech recognition quality function construction method provided in this application embodiment;

[0058] Figure 4 A flowchart illustrating a method for improving speech recognition quality provided in an embodiment of this application;

[0059] Figure 5 A flowchart illustrating a specific implementation of the method for improving speech recognition quality provided in this application;

[0060] Figure 6A A schematic diagram of the speech quality assessment device provided in the embodiments of this application;

[0061] Figure 6B A schematic diagram of the speech recognition quality prediction device provided in the embodiments of this application;

[0062] Figure 6C A schematic diagram of a device for improving speech recognition quality provided in an embodiment of this application;

[0063] Figure 7A This is a schematic diagram of the vehicle structure provided in an embodiment of this application;

[0064] Figure 7B A schematic diagram of the vehicle cabin provided in an embodiment of this application;

[0065] Figure 8 A schematic diagram of a computing device provided in an embodiment of this application. Detailed Implementation

[0066] The technical solutions provided in this application will be further described below with reference to the accompanying drawings and embodiments. It should be understood that the system architecture and business scenarios provided in the embodiments of this application are mainly for illustrating possible implementations of the technical solutions of this application and should not be construed as the sole limitation on the technical solutions of this application. Those skilled in the art will recognize that the technical solutions provided in this application are equally applicable to similar technical problems as system architectures evolve and new business scenarios emerge.

[0067] It should be understood that the speech quality assessment schemes provided in the embodiments of this application include speech quality assessment methods and apparatus, speech recognition quality prediction methods and apparatus, methods and apparatus for improving speech recognition quality, computer-readable storage media, and computer program products. Since these technical solutions solve problems based on the same or similar principles, some repetitive details may not be repeated in the following descriptions of specific embodiments. However, it should be considered that these specific embodiments have mutual references and can be combined with each other.

[0068] When evaluating the quality of the tested speech, the Mean Opinion Score (MOS) can be used. This index, also known as the subjective speech quality index, can be used to evaluate the quality of the tested speech at multiple levels. The quality of the tested speech is obtained by averaging the scores from all listeners.

[0069] When using the average opinion score as an indicator to assess the quality of the tested speech, differences in auditory ability and subjective auditory experience among different testers will cause differences in scores; especially when only a single sentence is provided without context, the differences in tester scores are significant, which will lead to low objectivity of the assessment results of the tested speech quality and poor adaptability of the assessment method.

[0070] Objective evaluation metrics can also be used to assess the quality of the tested speech. These metrics include Signal-to-Noise Ratio (SNR), Perceptual Evaluation of Speech Quality (PESQ), Perceptual Objective Listening Quality Analysis (POLQA), and Short-Time Objective Intelligibility (STOI). The PESQ algorithm requires a noisy attenuated signal and an original reference signal. After level adjustment, input filter filtering, time alignment and compensation, and auditory transformation of the two speech signals to be compared, parameters of both signals are extracted, and their time-frequency characteristics are combined to obtain a PESQ score. This score is then mapped to a Subjective Mean Opinion Score (MOS). POLQA is the successor to PESQ, extending to handle higher bandwidth audio signals. STOI is one of the important metrics for measuring speech intelligibility, used to evaluate the intelligibility of noisy speech that has been masked in the time domain or subjected to a short-time Fourier transform and weighted in the frequency domain. STOI scores the clean speech by comparing it with the speech to be evaluated.

[0071] The method of evaluating the quality of the tested speech using the above-mentioned objective evaluation indicators focuses on the correlation between sound characteristics and subjective listening experience from an acoustic perspective. However, its relationship with machine-oriented speech recognition (i.e., speech recognition into text or semantics) performance is unclear, making it difficult to serve as an effective reference for speech recognition modules during the research, development, and optimization of speech recognition systems.

[0072] This application provides an improved speech quality assessment scheme, wherein the speech quality related to semantics of the test speech can be determined based on the time-frequency characteristics of the test speech, and the recognition quality of the test speech can be predicted based on the speech quality related to semantics of the test speech. Furthermore, by comparing the speech quality related to semantics of the test speech and the predicted speech recognition quality with a baseline, the problem affecting the speech recognition quality can be identified, thereby improving the speech recognition quality by locating or solving the problem.

[0073] The voice quality assessment scheme provided in this application can be applied to fields such as quality detection and evaluation in the voice recognition process. For example, when applied to a smart cockpit in a vehicle, it can be used to determine the current voice quality or voice recognition quality within the cockpit, and then provide corresponding prompts or perform corresponding actions, such as reducing the volume of music in the cockpit, closing the windows, or optimizing parameters to improve the voice quality. Similarly, when applied to smart terminals such as mobile phones and smart speakers, it can assess the voice quality or voice recognition quality of the terminal's current environment, and then provide corresponding prompts, such as indicating whether the microphone is blocked, prompting to activate the camera to combine with lip reading during voice recognition to improve accuracy, or performing permitted actions (e.g., setting the corresponding application's permissions to access the camera), such as activating the camera to combine with lip reading during voice recognition. Furthermore, when applied as a quality detection terminal, it can be used to test the voice recognition function of a product under test, such as testing the voice recognition quality of a vehicle, to facilitate the optimization of voice recognition-related parameters.

[0074] In some embodiments, the aforementioned vehicle, smart terminal, and product under test typically include a microphone and a processor. The microphone is used to collect user voice; the processor can preprocess the collected voice and perform speech recognition on the preprocessed voice to convert it into text, and can further recognize instructions based on the recognized text. In some embodiments, when applied to a vehicle or smart terminal, the vehicle or smart terminal may also have a human-machine interface (HMI) for providing the aforementioned prompts to the user via display or sound. In some embodiments, the processor can also perform parameter optimization for preprocessing or speech recognition based on user operations through the HMI. In some embodiments, the processor of the aforementioned vehicle may be an electronic device, specifically a processor of an in-vehicle processing device such as a vehicle infotainment system or vehicle-mounted computer, or a conventional chip processor such as a central processing unit (CPU) or microcontroller unit (MCU). In some embodiments, when applied as a testing terminal for quality inspection, the testing terminal may have a HMI to provide the aforementioned prompts to the user via display or sound.

[0075] In some embodiments, the voice quality assessment scheme provided in this application can be embedded in the aforementioned vehicle, smart terminal, or product under test, existing as a functional module. In some embodiments, when the voice quality assessment scheme provided in this application is applied to an independent quality testing terminal, the testing terminal can communicate with the device under test, such as the aforementioned vehicle, smart terminal, or product under test, via wired or wireless means to obtain the required data, such as pre-processed voice data. Based on this data, voice quality detection or prediction of voice recognition quality can be performed, and test result prompts can be provided. In some embodiments, the testing terminal can also feed back the test results to the device under test, allowing the device under test to perform parameter optimization and other operations based on the test results.

[0076] Below, we will further combine Figure 1 This paper provides a brief description of a scenario in which the voice quality assessment scheme provided in the embodiments of this application is applied to a vehicle.

[0077] Figure 1This illustration shows a scenario where the voice quality assessment scheme provided in this application is applied to a vehicle. It can be used for voice quality assessment and also for assessing the quality of voice recognition. In this scenario, the vehicle's cabin includes: a sound acquisition module 110, a preprocessing module 120, a voice recognition module 130, an evaluation and prediction module 140, and an output module 150. The preprocessing module 120, the voice recognition module 130, and the evaluation and prediction module 140 can be implemented by the same processor in the vehicle, or they can be implemented by three or more processors respectively.

[0078] The sound acquisition module 110 can be a microphone, used to acquire test speech spoken by the user. Each test speech corresponds to a reference speech, and both correspond to the same sentence content, such as the same voice command, and have the same semantics. During the stage of establishing the various models for speech quality assessment, the sound acquisition module 110 is also used to acquire reference speech. Reference speech refers to the speech used in training the speech recognition module 130, or in training the semantically related feature model described later. The semantically related feature model is used for speech quality assessment and will be further explained later.

[0079] The preprocessing module 120 is used to perform preprocessing on the acquired sound, such as pre-emphasis, framing, or windowing, to make the user's test speech contained in the sound easier to recognize. Pre-emphasis includes emphasizing the high-frequency components of the speech, removing the influence of lip radiation, and increasing the high-frequency resolution of the speech. Framing utilizes the short-time stationarity of the speech signal to divide the speech signal into individual speech frames for processing, with overlap between adjacent speech frames to ensure continuity. Windowing strengthens the speech waveform near the sampled area of ​​each speech frame while weakening the rest of the waveform, thus smoothing the speech. The parameters adjusted in the preprocessing include one or more of the following: the frequency band targeted by pre-emphasis processing, the degree of emphasis, the frame length and overlapping frame length in framing processing, and the degree of strengthening and weakening of certain waveforms in windowing processing.

[0080] The Automatic Speech Recognition (ASR) module 130 is used to recognize the sentence content of the pre-processed test speech. The ASR module can identify words in the test speech and convert them into computer-readable character sequences. After obtaining the speech recognition content, control commands can be further recognized based on this content, and the vehicle actuator can execute the control commands corresponding to the speech. In some embodiments, the process of recognizing control commands from speech recognition content can be based on keyword matching or on semantic recognition technology using neural networks. Parameter adjustment of the speech recognition module refers to adjusting the parameters of the speech recognition model of the speech recognition module, such as adjusting the parameters and hyperparameters of the neural network implementing the speech recognition model.

[0081] The evaluation and prediction module 140 is used to perform speech quality evaluation and can generate a semantically relevant speech quality evaluation result for the test speech. In some embodiments, it can also predict the speech recognition quality of the test speech by the speech recognition module based on the speech quality evaluation result. In some embodiments, the evaluation and prediction module is also described as including an evaluation module and a prediction module, which respectively implement the above-mentioned speech quality evaluation and speech recognition quality prediction.

[0082] The output module 150 is used to output information such as speech quality assessment results or speech recognition quality prediction results. The output content can be provided to the vehicle controller so that the vehicle can perform corresponding operations. The output content can also be provided to the user through the vehicle's human-machine interface. In some embodiments, the output information includes the quality of the test speech, the speech recognition quality of the speech recognition model, factors affecting the quality of the test speech, and methods for adjusting the quality of the test speech. In some embodiments, the human-machine interface may include a display screen (such as a liquid crystal display, head-up display, HUD, etc.) and speakers in the vehicle cabin to provide prompts to the user through visual or auditory means.

[0083] In some embodiments, the human-machine interface can be a central control screen. After receiving the aforementioned information output by the output module 150 through the central control screen, the user can adjust the parameters of the preprocessing module 120 or the voice recognition module 130, or control relevant actuators in the vehicle, such as opening and closing windows or controlling the playback volume of the in-vehicle audio playback device. The parameter adjustment interface provided by the human-machine interface can be displayed in a way that is easy for ordinary users to understand and adjust (such as a graphical display), or it can be displayed in a way that is geared towards professional maintenance personnel.

[0084] In other embodiments, the evaluation and prediction module 140 can also be deployed on a standalone test device or in the cloud, and the output module 150 can also be deployed on a standalone test device. The test device can be a dedicated test device or a smart terminal with corresponding software installed, such as a mobile phone, computer, or tablet. When the modules are deployed on a standalone test device or in the cloud, communication between the vehicle and the test device or cloud can be achieved based on communication technology.

[0085] The following describes various method embodiments of this application based on the accompanying drawings.

[0086] This application provides a method for evaluating speech quality, which can be used to evaluate speech quality based on a test speech set. Figure 2A The flowchart of an embodiment of a voice quality assessment method is shown, which is illustrated using a vehicle as an example, and includes steps S210 to S230.

[0087] S210: Obtain test audio.

[0088] In this embodiment, the test voice is acquired by a sound acquisition module installed in the vehicle cabin. The sound acquisition module can be a microphone, and in some embodiments, it can be multiple microphones installed in different locations in the vehicle cabin.

[0089] In some embodiments, this step may be performed during the testing or inspection of the vehicle, and the test voice may be broadcast by the tester.

[0090] In some embodiments, this step can be performed while the user is using the vehicle, such as while the vehicle is in motion or parked. In this case, the test voice can be spoken by the driver (i.e., the user). The test voice is matched with the vehicle's voice commands. Since the driver already knows the voice content used in the voice commands, the voice commands spoken by the driver can be used as the test voice.

[0091] In some embodiments, this step may be triggered when the vehicle is unable to accurately recognize the user's (e.g., the driver's) voice commands, and the voice commands already spoken or re-spoken by the user may be used as test voices. In some embodiments, after this step is triggered, the vehicle may also prompt the user through a human-machine interface, using visual, text, or voice means, that the voice quality assessment process of this embodiment has been entered, and may further guide the user to speak the corresponding test voices.

[0092] In some embodiments, a user can repeatedly broadcast a single voice command (such as a voice instruction) multiple times. The resulting multiple voice recordings are also referred to as multiple voice recordings corresponding to the same set of test voice recordings, or multiple samples corresponding to a corpus. Users can also repeatedly broadcast several voice commands separately to obtain multiple voice recordings corresponding to each of these sets of test voice recordings. For example, multiple broadcasts of the command "turn on the air conditioner" constitute one set of test voice recordings, while multiple broadcasts of the command "turn up the volume" constitute another set. Each set of test voice recordings corresponds to a single voice command, or a single semantic meaning, or the same sentence content.

[0093] S220: Evaluate the speech quality of the test speech based on semantically relevant information of the test speech.

[0094] In this embodiment, the voice quality is evaluated using a vehicle evaluation module. This evaluation module is implemented by the vehicle's processor, which is signal-connected to the sound acquisition module.

[0095] In this embodiment, the speech quality evaluation result of the test speech is related to a predetermined semantic, so that the speech quality evaluation result can not only be used to evaluate speech quality, but also to predict the speech recognition quality of the speech recognition module.

[0096] In this embodiment, the semantically relevant information of the test speech is a semantically relevant feature vector generated by a neural network for the test speech, and this feature vector is the output of any layer preceding the output layer of the neural network. In other embodiments, the feature vector may also be a cascade of outputs from multiple layers preceding the output layer of the neural network, and these multiple layers may be any two or more layers.

[0097] In some embodiments, the semantic information related to the test speech can be a one-dimensional vector composed of the outputs of the output layer of the aforementioned neural network. For example, the output of the output layer could be a one-dimensional vector composed of the confidence scores of each speech instruction (i.e., each semantic meaning). The semantic information will be described in further detail in step S223 below.

[0098] S230: Output the evaluation results of the speech quality of the test speech.

[0099] In this embodiment, the output can be sent to the vehicle's human-machine interface (HMI) to display the evaluation results to the user. The HMI may include a display screen, which can be a vehicle central control screen, a head-up display (HUD), an augmented reality HUD (AR-HUD), etc. The HMI may also include a speaker and input components, which can be a touchscreen integrated into the display screen or independent buttons. The prompts can be executed via the HMI using images, text, or voice.

[0100] In some embodiments, the evaluation results include content related to the speech quality of the test speech, which may include one or a combination of the following: the quality of the test speech, factors affecting the quality of the test speech, and methods for adjusting the quality of the test speech.

[0101] In some embodiments, factors affecting the quality of the test voice may include one or a combination of the following: the windows are open, which introduces external noise; the vehicle speed is too high, which makes tire noise or engine noise too loud; other sounds inside the vehicle are too loud, such as music played in the vehicle; the performance of the microphone inside the vehicle or the location of its deployment.

[0102] In some embodiments, adjusting the test voice quality may include one or a combination of the following: closing the windows, reducing vehicle speed, reducing in-vehicle volume, replacing the faulty microphone, and optimizing the parameters of the preprocessing module.

[0103] In some embodiments, such as Figure 2B The flowchart shown above indicates that step S220 includes steps S221 to S225.

[0104] S221: Obtain the first feature vector of the test speech. The first feature vector includes the time-frequency features of the test speech.

[0105] In this embodiment, multiple feature vectors, including frequency domain features, can be obtained from each consecutive frame of the test speech, and these multiple feature vectors constitute the first feature vector. In this embodiment, adjacent frames may have overlapping information. In this embodiment, the acquired test speech can be segmented into consecutive frames by the preprocessing process of the preprocessing module. In some embodiments, the preprocessing process further includes pre-emphasis and windowing. By including overlapping information in adjacent frames, the frame data can be made complete, and the parameter changes of the feature vectors can be made smoother. In some other embodiments, adjacent frames may not contain overlapping information, which can reduce the amount of data to be processed and thus improve the processing speed.

[0106] The first feature vector includes the time-frequency features of the test speech, hence it is also called the time-frequency feature vector map. This map is a two-dimensional graph with one dimension being the time coordinate and the other dimension being the frequency coordinate. The intensity of each pixel is the intensity of the corresponding frequency of each consecutive frame.

[0107] In some embodiments, frequency domain features include one or a combination of the following: Mel-Frequency Cepstral Coefficients (MFCC), Linear Predictive Cepstral Coefficients (LPCC), and spectrum.

[0108] S223: Obtain the second feature vector of the test speech based on the first feature vector of the test speech. The second feature vector is related to the semantics of the speech.

[0109] In some embodiments, a second feature vector is extracted based on a first feature vector using a semantically related feature model, wherein the semantically related feature model represents the relationship between semantics and time-frequency features, and therefore the extracted second feature vector is semantically related.

[0110] In some embodiments, the semantically relevant feature model is constructed based on a first feature vector of a reference speech and the semantics of that reference speech. In some embodiments, the semantically relevant feature model can be constructed based on a neural network, such as a fully connected neural network (FCNN), a recurrent neural network (RNN), or a convolutional neural network (CNN). This embodiment uses a CNN network. When constructing the semantically relevant feature model, it can be trained using a preprocessing module. For example, during training, the reference speech has semantic annotations; the reference speech is processed by a preprocessing module to obtain a first feature vector, which is then input into the semantically relevant feature model. The semantically relevant model is trained based on whether the semantics of its output converge. Gradient descent, adversarial networks, etc., can be used during training.

[0111] In some embodiments, the semantically relevant feature model constructed based on a neural network includes a multi-layer network, and the output of the semantically relevant feature model corresponds to the classification of each semantic. The second feature vector can be a feature vector output from any layer of the multi-layer network. Since the first feature vector serves as the input to the neural network, the second feature vector is a feature vector extracted based on the first feature vector. On the other hand, since each semantic corresponds to the output of the neural network (which is essentially a classification network, with each category corresponding to each semantic), the second feature vector can be understood as a semantically relevant feature vector.

[0112] In some embodiments, when the output of a layer preceding the output layer of the neural network is used as the second feature vector, the second feature vector can be a feature vector with more dimensions than the first feature vector. For example, when using a CNN network, if multiple convolutional kernels are used to obtain the second feature vector, the results of the operations of the multiple convolutional kernels constitute the multiple dimensions of the second feature vector.

[0113] In some embodiments, the second feature vector can also be a feature vector formed by the concatenation of output vectors from two or more layers of the neural network (similar to residual connections). In this way, the second feature vector can have not only low-level features but also high-level features. For example, the outputs of the second layer and the outputs of the fourth layer of the neural network are concatenated to form the second feature vector.

[0114] In some embodiments, when the output of the output layer of the neural network is used as the second feature vector, the second feature vector is a vector composed of the confidence scores corresponding to each semantic. For example, when the output of the neural network has 20 nodes, which correspond to 20 categories (i.e., used to identify one of 20 semantics), the second feature vector can be a one-dimensional vector with 20 parameters, the value of each parameter corresponding to the confidence score of each semantic.

[0115] The reference speech set includes several sets of reference speech, each set corresponding to the same semantic meaning. For example, the semantic meaning could be the meaning of commonly used vehicle commands such as "turn on the air conditioner," "turn off the air conditioner," "turn up the volume," and "turn down the volume." The reference speech can be understood as speech collected in a low-noise environment, or as standard speech.

[0116] S225: Evaluate the speech quality of the test speech based on the second feature vector of the test speech.

[0117] Since the second feature vector is semantically related, the evaluated speech quality is also semantically related. When the test speech is a set of multiple test speech samples, the speech quality of the test speech can be understood as the speech quality of the test speech set.

[0118] In some embodiments, the speech quality assessment results include the following three assessment metrics:

[0119] The first evaluation metric represents the degree of concentration of the center positions of the test speech in each group. Different groups of test speech have different semantics. The center position of each group of test speech is the center position of the second feature vector of that group of test speech in the second feature space, where the second feature space is the space containing the second feature vector. A group of test speech may include multiple test speech samples, and the center position of that group of test speech samples can be calculated based on the distribution of the second feature vectors of these multiple test speech samples.

[0120] In some embodiments, the concentration index may be Mahalanobis distance, Euclidean distance, or other distance indexes that can calculate the similarity between samples.

[0121] In some embodiments, the evaluation method of the first evaluation index can be as follows: First, for each group of test speech, calculate the distance between the center position of the test speech and the center positions of other groups of test speech, and obtain the minimum value of the distance. Then, take the average of the minimum distances obtained for each group of test speech as an index of the concentration of the center positions of each group of test speech. The calculation method of the first evaluation index D1 can be shown in the following formula (1):

[0122]

[0123] Where C represents the grouping type in the test speech set, i.e., the corpus types of the test speech set, μ tj μ represents the center position of the j-th test speech in the test speech set. ti ∑ represents the center position of the i-th group of test speech in the test speech set, and ∑ represents the covariance matrix of the second feature vector distribution of the j-th group of test speech in the test speech set.

[0124] The second evaluation metric is an index representing the degree of dispersion of the second feature vector of each group of test speech in the second feature space, wherein the test speech of different groups has different semantics, and the second feature space is the space in which the second feature vector is located.

[0125] In some embodiments, the evaluation method of the second evaluation index can be as follows: First, for each group of test speech, calculate the semi-principal axis length of the distribution of the second feature vector of the group of test speech in the space where the second feature vector is located. Then, take the average of the semi-principal axis lengths obtained for each group of test speech as the dispersion index of each group of test speech. The calculation method of the second evaluation index D2 can be shown in the following formula (2):

[0126]

[0127] Where C represents the grouping type in the test speech set, d represents the dimension of the second feature vector of each test speech group in its space, and f jk Let represent the length of the k-th semi-principal axis of the second feature vector of the j-th test speech set, and ∑ represent the covariance matrix of the distribution of the second feature vector of the j-th test speech set.

[0128] The third evaluation metric is the similarity between the center position of the test speech in each group and the center position of the corresponding reference speech in each group. The test speech in different groups has different semantics, and the center position of each group of test speech is the center position of the second feature vector of that group of test speech in the second feature space, where the second feature space is the space containing the second feature vector. Similarly, the reference speech in different groups has different semantics, and the center position of each group of reference speech is the center position of the second feature vector of that group of reference speech in the second feature space.

[0129] In some embodiments, the evaluation method of the third evaluation index can be as follows: First, in the space where the second feature vector is located, for each group of test speech, calculate the distance between the center position of the test speech and the center position of a group of reference speech, where the reference speech and the test speech have the same semantics, that is, they correspond to the same sentence content. Then, take the average of the distances obtained for each group of test speech as an index of the similarity between the test speech set and the reference speech set. Wherein, the center position of a group of test speech refers to the distribution center of the second feature vector of the test speech, and the center position of a group of reference speech refers to the distribution center of the second feature vector of the reference speech. The calculation method of the third evaluation index D3 can be shown in the following formula (3):

[0130]

[0131] Where C represents the grouping type in the test speech set, μ rj μ represents the center position of the j-th group of reference speech. tj ∑ represents the center position of the j-th test speech, and ∑ represents the covariance matrix of the second feature vector distribution of the j-th reference speech in the reference speech set.

[0132] In some embodiments, the voice quality assessment result includes one of the above-mentioned assessment indicators, or it may include a combination of any number of assessment indicators, and different assessment indicators may have different weights.

[0133] This application also provides a method for predicting speech recognition quality, which can predict the speech recognition quality of test speech based on the speech quality evaluation results of the test speech described above. For example... Figure 3AA flow diagram of an embodiment of a speech recognition quality prediction method is shown, which includes steps S310 to S340.

[0134] S310: Get test audio.

[0135] This step can be referred to in the description of S210 or its various embodiments above, and will not be repeated here.

[0136] S320: Evaluate the speech quality of the test speech based on semantically relevant information of the test speech.

[0137] This step can be referred to in the description of S220 or its various embodiments above, and will not be repeated here.

[0138] In some embodiments, the speech recognition quality prediction method and the aforementioned speech quality assessment method can be integrated into a single process. In this case, the content described in steps S310 and S320 can be directly derived from the results of steps S210 and S220, without the need to repeat the same content.

[0139] S330: Based on the evaluated test speech quality, predict the speech recognition quality of the speech recognition model for the test speech using a speech recognition quality function. The speech recognition quality function indicates the relationship between speech recognition quality and overall speech quality.

[0140] In some embodiments, the recognition results of the speech recognition model are used in constructing the speech recognition quality function; therefore, the predicted speech recognition quality can be used as a prediction of the recognition quality of the speech recognition model. The speech recognition model can be implemented by the aforementioned speech recognition module.

[0141] S340: Outputs the predicted results of speech recognition quality.

[0142] The output speech recognition quality prediction results include: the predicted speech recognition quality of the speech recognition module, the factors affecting the speech recognition quality of the speech recognition module, and the methods to adjust the speech recognition quality of the speech recognition module.

[0143] In some embodiments, factors affecting the speech recognition quality of the speech recognition module include: the window being open, excessive vehicle speed, excessively loud in-vehicle noise, microphone performance and deployment location, and the performance of the speech recognition model of the speech recognition module.

[0144] In some embodiments, adjusting the speech recognition quality of the speech recognition module may include one or a combination of the following: closing the window, reducing vehicle speed, reducing in-vehicle volume, replacing a faulty microphone or optimizing microphone deployment, tuning the parameters of the preprocessing module, and tuning the parameters of the speech recognition model of the speech recognition module.

[0145] In this embodiment, the prediction results can be output to the vehicle's human-machine interface to display the evaluation results to the user. Please refer to the description of step S230 above for details.

[0146] In some embodiments, the speech recognition quality function in step S330 above can be as follows: Figure 3B The process shown is used to construct the function, which includes steps S321 to S327.

[0147] S321: Degrade each reference speech in the reference speech set multiple times, and each degradation forms a set of degraded speech, thereby obtaining multiple sets of degraded reference speech.

[0148] This process involves degrading the reference speech to varying degrees to generate multiple sets of degraded speech. Each set of degraded speech at each degree, or each set of degraded speech in different ways, constitutes a degraded speech set. The degradation method can be speech scrambling. In some embodiments, the reference speech in the reference speech set can be scrambled according to the possible noise environment of the vehicle, such as adding background music interference, simulated tire noise, wind resistance noise interference, simulated interference from other vehicle horns outside the vehicle, etc.

[0149] S323: According to the method of the aforementioned speech quality assessment method embodiment, the speech quality of multiple sets of degraded reference speech is assessed to obtain a first assessment result.

[0150] Each set of degraded reference speech is used as test speech. The speech quality of each set of degraded reference speech is evaluated according to the speech quality evaluation method in the aforementioned embodiment. The speech quality of each set of degraded reference speech constitutes the first evaluation result.

[0151] S325: Based on the above speech recognition model, multiple sets of degraded reference speech are identified and statistically analyzed, and the statistical speech recognition results are used as the first statistical results.

[0152] Specifically, a speech recognition model is used to identify each set of degraded reference speech, generating a speech recognition result for each set of degraded reference speech. The speech recognition results of each set of degraded reference speech constitute the first statistical result.

[0153] S327: Obtain the speech recognition quality function based on the functional relationship between the first statistical result and the first evaluation result.

[0154] In some embodiments, a speech recognition quality function can be constructed based on machine learning. For example, a speech recognition quality function can be constructed by fitting a polynomial to the first statistical results and first evaluation results of each set of degraded reference speech. Alternatively, a speech recognition quality function can be constructed based on deep learning, for example, by training a neural network model.

[0155] In some embodiments, each evaluation index in the first evaluation result is used as a dependent variable to construct a speech recognition quality function. In other embodiments, one or more combined indicators from the first evaluation result are combined and then used as dependent variables to construct a speech recognition quality function. The evaluation indicators here are, for example, those shown in formulas (1) to (3) above.

[0156] This application also provides a method for improving speech recognition quality. Based on the speech quality evaluation results and speech recognition quality prediction results of the test speech, it can be determined whether parameter tuning of the preprocessing module or the speech recognition module is needed to improve speech recognition quality. Figure 4 A flow diagram of an embodiment of a method for improving speech recognition quality is shown, which includes steps S410 to S480.

[0157] S410: Get test audio.

[0158] This step can be referred to in the description of step S210 above or its various embodiments, and will not be described in detail here.

[0159] S420: Obtain the speech quality of the evaluation test speech.

[0160] This step can be referred to in the description of step S220 above or its various embodiments, and will not be described in detail here.

[0161] S430: Determine whether the speech quality is lower than a preset first baseline. If the speech quality is lower than the first baseline, proceed to step S440; otherwise, proceed to step S450.

[0162] The first baseline, also known as the indicator baseline, is used to determine the evaluation indicators of speech quality. In some embodiments, a first baseline is set for each evaluation indicator of speech quality; in other embodiments, the evaluation indicators of speech quality are combined into one or more combined indicators, and then a corresponding first baseline is set.

[0163] S440: Outputs the evaluation results of the speech quality of the test speech.

[0164] This step can be referred to in the description of step S230 above or its various embodiments, and will not be described in detail here.

[0165] After completing this step in some embodiments, you can return to step S410 or end the current process.

[0166] In some embodiments, considering the fault tolerance of the speech recognition model, step S450 or step S480 may be performed to continue speech recognition.

[0167] S450: Predicts speech recognition quality.

[0168] This step can be referred to in the description of step S330 above or its various embodiments, and will not be described in detail here.

[0169] S460: Determine whether the predicted speech recognition quality is lower than a preset second baseline. If the speech recognition quality is lower than the preset second baseline, proceed to step S470; otherwise, proceed to step S480.

[0170] The second baseline is used to judge the quality of speech recognition, that is, to judge the speech recognition accuracy of the semantic recognition model, and is also known as the accuracy baseline.

[0171] S470: Output the prediction result of the speech recognition quality.

[0172] This step can be referred to in the description of step S340 above or its various embodiments, and will not be described in detail here.

[0173] After completing this step in some embodiments, you can return to step S410 or end the current process.

[0174] In some embodiments, considering the fault tolerance of the speech recognition model, step S480 can be continued to perform speech recognition.

[0175] S480: The test speech is recognized by the speech recognition module.

[0176] In some embodiments, when the speech quality evaluated in step S430 is higher than the first baseline and the speech recognition quality predicted in step S460 is higher than the second baseline, the accuracy of the speech recognition quality is considered to be high, and the recognition result can be used for subsequent purposes, such as for vehicle control.

[0177] In some embodiments, when the speech quality evaluated in step S430 is lower than the first baseline, or the speech recognition quality predicted in step S460 is lower than the second baseline, the speech recognition result can be further prompted to the user for confirmation in order to determine whether to use the speech recognition result.

[0178] In some embodiments, the user can adjust the preprocessing parameters or speech recognition model parameters based on the speech quality evaluation result output in step S440 or the speech recognition quality prediction result output in step S470 to improve the speech recognition quality. In other embodiments, the device under test, such as a vehicle, can automatically adjust the preprocessing parameters or speech recognition model parameters based on the speech quality evaluation result output in step S440 or the speech recognition quality prediction result output in step S470.

[0179] Below, to facilitate a further understanding of the technical solutions of the above embodiments, a specific implementation method for improving speech recognition quality will be described. In this specific implementation method, the steps of a speech quality assessment method and a speech recognition quality prediction method are involved. The steps corresponding to these two parts can also be separated from the foregoing embodiments as specific implementation methods of the speech quality assessment method and the speech recognition quality prediction method. For the sake of simplicity, the specific implementation methods of these two parts will not be described in detail.

[0180] like Figure 5 The flowchart illustrating a specific implementation of a speech recognition quality prediction method includes the following steps:

[0181] S510: The vehicle acquires a test voice set based on a reference voice set through microphones installed in the cockpit.

[0182] The test voice set includes several groups. For example, the tester sits in the driver's seat of the vehicle and broadcasts the test voices of each group in turn. The semantics of each group can correspond to a commonly used command. Each test voice set includes 10 test voices with the same content broadcast by the tester.

[0183] The content of the test audio broadcast corresponds to the content of the reference audio set. In this example, the vehicle guides the tester to broadcast the test audio through a human-machine interface. For example, the screen can display the audio content of each line of the corresponding reference audio set and the number of times it needs to be broadcast, which the tester then broadcasts accordingly.

[0184] S515: The collected test speech is preprocessed by the preprocessing module on the vehicle, including extracting the time-frequency feature vector map of each test speech in the test speech set, i.e., the first feature vector.

[0185] S520: Using the semantic relevance feature model, obtain the semantic relevance feature vector of the test speech in the test speech set, i.e., the second feature vector, based on the time-frequency feature vector map.

[0186] S525: Evaluate the speech quality of the test speech set based on the semantic relevance feature vector of the test speech, which can be evaluated using one or more of the above formulas (1) to (3).

[0187] S530: Based on the evaluation results of the voice quality, determine whether the voice quality is lower than the set target baseline. If it is lower than the set target baseline, proceed to step S535; otherwise, proceed to step S540.

[0188] S535: Outputs the evaluation results of the speech quality of the test speech.

[0189] In this example, the evaluation results can be output to the human-machine interface (HMI), including content related to the speech quality of the test speech. Specifically, the speech quality-related content displayed on the HMI might include prompts to the user to optimize the parameters and algorithms in the preprocessing module. This optimization interface and parameters can be displayed graphically.

[0190] S540: Using a speech recognition quality function, based on the speech quality of the test speech set, predicts the speech recognition quality of the speech recognition module for the test speech set.

[0191] S545: Determine whether the predicted recognition quality is lower than the set accuracy baseline. If it is lower than the accuracy baseline, proceed to step S550; otherwise, proceed to step S555.

[0192] S550: Output the prediction result of the speech recognition quality.

[0193] In this example, the prediction results can be output to the human-computer interface, and the prediction results include content related to the quality of speech recognition.

[0194] In this example, the content related to speech recognition quality displayed on the human-machine interface includes prompts to the user to optimize the speech recognition model of the speech recognition module. This can be achieved by graphically displaying the adjustable interface and parameters.

[0195] S555: This indicates that the quality of the current speech evaluation and the predicted speech recognition quality meet the standards.

[0196] Furthermore, in this specific embodiment, after the user has fine-tuned the parameters according to the output of step S535 or step S550, the speech recognition quality function can be further optimized. Specifically, the steps in the speech recognition quality function construction method can be used to retrain the speech recognition quality function to optimize it. For example, after optimizing the parameters or algorithm in the preprocessing, the speech recognition quality function can be retrained again according to steps S323-S327; after optimizing the speech recognition model parameters, the speech recognition quality function can be retrained again according to steps S325-S327.

[0197] This application also provides corresponding apparatuses. For details regarding the beneficial effects of these apparatuses or the technical problems they solve, please refer to the descriptions in the methods corresponding to each apparatus, or to the descriptions in the invention summary; only brief descriptions are provided here. The various apparatuses in this embodiment can be used to implement the optional embodiments of the methods described above. The following describes the various apparatus embodiments of this application based on the figures.

[0198] The embodiments of this application provide a speech quality assessment apparatus, which can be used to implement various embodiments of the speech quality assessment method, such as... Figure 6A As shown, the device has an acquisition module 610, an evaluation module 620, and an output module 630.

[0199] The acquisition module 610 is used to acquire test audio. Specifically, it is used to perform the aforementioned step S210 and its various embodiments.

[0200] The evaluation module 620 is used to evaluate the speech quality of the test speech based on semantically relevant information. Specifically, it is used to perform the aforementioned step S220 and its various embodiments.

[0201] The output module 630 is used to output the evaluation results of the speech quality. Specifically, it is used to perform the aforementioned step S230 and its various embodiments.

[0202] In some embodiments, the speech quality evaluation result output by the output module 630 includes one or more of the following: the quality of the test speech; factors affecting the quality of the test speech; and methods for adjusting the quality of the test speech. For details, please refer to the description in step S230 above.

[0203] In some embodiments, the evaluation module 620 is specifically configured to: obtain a first feature vector of the test speech, wherein the first feature vector includes a time-frequency feature vector of the test speech; obtain a second feature vector of the test speech based on the first feature vector of the test speech, wherein the second feature vector is semantically related to the test speech; and evaluate the speech quality of the test speech based on the second feature vector.

[0204] In some embodiments, when the evaluation module 620 evaluates the speech quality of the test speech based on the second feature vector, it includes: evaluating the speech quality of the test speech using a first evaluation index, the first evaluation index including: a concentration index of the center positions of each group of test speech, wherein the test speech of different groups has different semantics, the center position of each group of test speech is the center position of the second feature vector of that group of test speech in a second feature space, and the second feature space is the space in which the second feature vector is located.

[0205] In some embodiments, when the evaluation module 620 evaluates the speech quality of the test speech based on the second feature vector, it includes: evaluating the speech quality of the test speech using a second evaluation index, the second evaluation index including: a dispersion index of the second feature vector of each group of test speech in a second feature space, wherein the test speech of different groups has different semantics, and the second feature space is the space where the second feature vector is located.

[0206] In some embodiments, when evaluating the speech quality of the test speech based on the second feature vector, the evaluation module 620 includes: evaluating the speech quality of the test speech using a third evaluation index, the third evaluation index including: a similarity index between the center position of each group of test speech and the center position of each group of reference speech corresponding to the semantics; wherein, the test speech of different groups has different semantics, the center position of each group of test speech is the center position of the second feature vector of that group of test speech in the second feature space, the second feature space is the space where the second feature vector is located, the reference speech of different groups has different semantics, and the center position of each group of reference speech is the center position of the second feature vector of that group of reference speech in the second feature space.

[0207] In some embodiments, when the evaluation module 620 obtains the first feature vector of the test speech, it includes: acquiring consecutive frames contained in the test speech, wherein adjacent frames contain overlapping information; and obtaining a plurality of feature vectors including frequency domain features based on the consecutive frames, wherein the plurality of feature vectors constitute the first feature vector.

[0208] Embodiments of this application also provide an apparatus for speech recognition quality prediction, which can be used to implement method embodiments for speech recognition quality prediction, such as... Figure 6B As shown, the device has an acquisition module 612, a prediction module 622, and an output module 632.

[0209] The acquisition module 612 is used to acquire test audio. Specifically, it is used to perform the above-described step S310 and its various embodiments.

[0210] The prediction module 622 is used to predict the speech recognition quality of the test speech by the speech recognition model based on a speech recognition quality function, wherein the speech recognition quality function represents the relationship between speech recognition quality and speech quality, and the speech quality is evaluated according to any possible embodiment of the aforementioned speech quality evaluation method. Specifically, it is used to perform the above steps S320-S330 and their various embodiments.

[0211] The output module 632 is used to output the prediction result of the speech recognition quality. Specifically, it is used to perform the above step S340 and its various embodiments.

[0212] In some embodiments, the prediction result of the speech recognition quality output by the output module 632 includes one or more of the following: the speech recognition quality of the speech recognition model; factors affecting the speech recognition quality of the speech recognition model; and methods for adjusting the speech recognition quality of the speech recognition model.

[0213] In some embodiments, the construction process of the speech recognition quality function includes: obtaining multiple sets of degraded reference speech; obtaining speech recognition results of the multiple sets of degraded reference speech according to the speech recognition model, and using them as first statistical results; using the multiple sets of degraded reference speech as test speech, obtaining speech quality evaluation results of the multiple sets of degraded reference speech according to any possible embodiment of the speech quality evaluation method, and obtaining a first evaluation result; and obtaining the speech recognition quality function according to the functional relationship between the first statistical result and the first evaluation result.

[0214] Embodiments of this application also provide an apparatus for improving speech recognition quality, which can be used to implement method embodiments for improving speech recognition quality, such as... Figure 6C As shown, the voice quality assessment device has an acquisition module 614, an assessment prediction module 624, and an output module 634.

[0215] The acquisition module 614 is used to acquire test audio. Specifically, it is used to perform the above-described step S410 and its various embodiments.

[0216] The evaluation and prediction module 624 is used to evaluate the speech quality of the test speech. Any possible embodiment of the aforementioned speech quality evaluation method can be used for the evaluation. Specifically, it is used to perform steps S420-S430 and their various embodiments.

[0217] The output module 634 is used to execute the evaluation result of the output speech quality when the speech quality is lower than a preset first baseline. Specifically, it is used to execute the above step S440 and its various embodiments.

[0218] In some embodiments, the evaluation and prediction module 624 is further configured to predict the speech recognition quality according to any possible embodiment of the speech recognition quality prediction method when the speech quality is greater than or equal to the first baseline. Specifically, it is configured to perform the above steps S450-S460 and their respective embodiments.

[0219] The output module 634 is further configured to output the prediction result of the speech recognition quality when the speech recognition quality is lower than a preset second baseline. Specifically, it is used to execute the above step S470 and its various embodiments.

[0220] Embodiments of this application also provide a vehicle, such as Figure 7A and Figure 7B As shown, the vehicle includes a sound acquisition device 710, a preprocessing device 720, a speech recognition device 730, and the aforementioned speech quality assessment device, speech recognition quality prediction device, or device for improving speech recognition quality.

[0221] The sound acquisition device 710 is used to acquire commands spoken by the driver based on the semantics of reference speech. Figure 7B The microphone in the middle can be the one inside the cockpit. Figure 7B The example shown is located at the central control screen 740, but it can also be located at one or more other locations such as the instrument panel above the steering wheel, the rearview mirror in the cabin, or the steering wheel.

[0222] The preprocessing unit 720 is used to preprocess the voice broadcast by the driver.

[0223] The voice recognition device 730 is used to recognize driver commands when the voice quality of the driver's command meets the requirements and the predicted voice recognition quality meets the requirements.

[0224] The aforementioned speech quality assessment device, speech recognition quality prediction device, or device for improving speech recognition quality is used for the aforementioned purposes, and users can perform corresponding operations based on it to improve speech quality and speech recognition quality.

[0225] exist Figure 7B The vehicle cabin shown also includes a central control screen 740, which serves as a human-machine interface. Users receive and display information from a voice quality assessment device, a voice recognition quality prediction device, or a device for improving voice recognition quality through the central control screen 740. The central control screen 740 also displays an interface for adjusting parameters, allowing users to easily adjust the aforementioned parameters by operating the central control screen 740.

[0226] exist Figure 7B In the vehicle cabin shown, the preprocessing device 720, the voice recognition device 730, and the voice quality assessment device, voice recognition quality prediction device, or device for improving voice recognition quality can be implemented by one or more processors in the vehicle. In this embodiment, they can be implemented by the processor of the in-vehicle infotainment system.

[0227] Figure 8 This is a schematic structural diagram of a computing device 800 provided in an embodiment of this application. The computing device 800 includes: a processor 810, a memory 820, and a communication interface 830.

[0228] It should be understood that Figure 8 The communication interface 830 in the computing device 800 shown can be used to communicate with other devices.

[0229] The processor 810 can be connected to the memory 820. The memory 820 can be used to store the program code and data. Therefore, the memory 820 can be a storage module inside the processor 810, an external storage module independent of the processor 810, or a component that includes both the internal storage module of the processor 810 and the external storage module independent of the processor 810.

[0230] When the computing device 800 is running, the processor 810 executes the computer execution instructions in the memory 820 to perform the operation steps of the above method.

[0231] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, is used to perform one or more of the schemes described in the various embodiments of this application.

[0232] In the above description, the labels of the steps, such as S110, S120, etc., do not necessarily mean that the steps will be executed in this way. The order of the steps can be interchanged or executed simultaneously if permitted.

[0233] The terms "first," "second," "third," and similar terms used in the specification and claims of this application are used only to distinguish similar objects and do not represent a specific ordering of objects. It is understood that, where permissible, a specific order or sequence may be interchanged so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0234] The term "comprising" as used in the specification and claims of this application should not be construed as limiting itself to what follows; it does not exclude other elements or steps. Therefore, it should be interpreted as specifying the presence of the mentioned feature, integral, step, or component, but does not exclude the presence or addition of one or more other features, integrals, steps, or components, or groups thereof. Thus, the expression "device comprising means A and B" should not be limited to a device consisting solely of components A and B.

[0235] The term "embodiment" as used in this specification means that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in at least one embodiment of this application. Therefore, the phrase "in an embodiment" appearing throughout this specification does not necessarily refer to the same embodiment, but may refer to the same embodiment. Furthermore, in one or more embodiments, particular features, structures, or characteristics can be combined in any suitable manner, as will be apparent to those skilled in the art from this disclosure.

[0236] Note that the above are merely embodiments of this application and the technical principles employed. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this application, all of which fall within the scope of protection of this application.

Claims

1. A method for assessing speech quality, characterized in that, include: Get the test audio; The speech quality of the test speech is evaluated based on the semantic information of the test speech, wherein the speech quality of the test speech is related to the performance of the machine-oriented speech preprocessing module and the performance of the speech recognition module. The speech preprocessing module is used to preprocess the speech signal collected by the sensor, and the speech recognition module is used to perform speech recognition on the speech signal output by the speech preprocessing module. When the speech quality is lower than a preset first baseline, the speech quality evaluation result is output. The first baseline is the evaluation index of the speech quality. The evaluation result is used to indicate that the preprocessing parameters of the speech preprocessing module should be adjusted so that the speech quality of the preprocessed speech signal is greater than or equal to the first baseline.

2. The method according to claim 1, characterized in that, The speech quality assessment results include one or more of the following: The quality of the test audio; Factors affecting the quality of the test audio; The method for adjusting the quality of the test audio.

3. The method according to claim 1, characterized in that, The step of evaluating the speech quality of the test speech based on semantically relevant information includes: Obtain a first feature vector of the test speech, wherein the first feature vector includes the time-frequency feature vector of the test speech; A second feature vector of the test speech is obtained based on a first feature vector of the test speech, wherein the second feature vector is semantically related to the test speech; The speech quality of the test speech is evaluated based on the second feature vector.

4. The method according to claim 3, characterized in that, The step of evaluating the speech quality of the test speech based on the second feature vector includes: evaluating the speech quality of the test speech using a first evaluation metric. The first evaluation index includes: the concentration index of the center position of the test speech in each group, wherein the test speech in different groups has different semantics, and the center position of the test speech in each group is the center position of the second feature vector of the test speech in the second feature space, the second feature space being the space where the second feature vector is located.

5. The method according to claim 3, characterized in that, The step of evaluating the speech quality of the test speech based on the second feature vector includes: evaluating the speech quality of the test speech using a second evaluation metric. The second evaluation metric includes: a dispersion index of the second feature vector of each group of test speech in the second feature space, wherein the test speech of different groups has different semantics, and the second feature space is the space where the second feature vector is located.

6. The method according to claim 3, characterized in that, The step of evaluating the speech quality of the test speech based on the second feature vector includes: evaluating the speech quality of the test speech using a third evaluation metric. The third evaluation index includes: the similarity index between the center position of each group of test speech and the center position of each group of reference speech corresponding to the semantics; wherein, the test speech of different groups has different semantics, the center position of each group of test speech is the center position of the second feature vector of the test speech in the second feature space, the second feature space is the space where the second feature vector is located, the reference speech of different groups has different semantics, and the center position of each group of reference speech is the center position of the second feature vector of the reference speech in the second feature space.

7. The method according to any one of claims 3 to 6, characterized in that, Obtaining the first feature vector of the test speech includes: Obtain consecutive frames contained in the test speech, wherein adjacent frames contain overlapping information; Multiple feature vectors, including frequency domain features, are obtained from the consecutive frames, and the multiple feature vectors constitute the first feature vector.

8. A method for predicting speech recognition quality, characterized in that, include: Acquire test speech, wherein the speech quality of the test speech is greater than or equal to a first baseline, and the first baseline is an evaluation index of the speech quality; The speech recognition quality of the test speech is predicted by the speech recognition model according to the speech recognition quality function, wherein the speech recognition quality function is used to indicate the relationship between speech recognition quality and speech quality, and the speech quality is evaluated by the method according to any one of claims 1-7; Output the predicted results of the speech recognition quality.

9. The method according to claim 8, characterized in that, The output of the predicted speech recognition quality includes one or more of the following: The speech recognition quality of the speech recognition model; Factors affecting the speech recognition quality of the aforementioned speech recognition model; The method for adjusting the speech recognition quality of the speech recognition model.

10. The method according to claim 9, characterized in that, The process of constructing the speech recognition quality function includes: Obtain multiple sets of degraded reference speech; The speech recognition results of multiple sets of the degraded reference speech are obtained according to the speech recognition model and used as the first statistical result; Using multiple sets of degraded reference speech as test speech, a first evaluation result is obtained based on the speech quality evaluation results of the multiple sets of degraded reference speech; The speech recognition quality function is obtained based on the functional relationship between the first statistical result and the first evaluation result.

11. A method for improving speech recognition quality, characterized in that, include: Get the test audio; The speech quality assessment result of the test speech is obtained by the method according to any one of claims 1-7; When the speech quality of the preprocessed speech signal is greater than or equal to the first baseline, the speech recognition quality is predicted according to the method of claim 8 or 9. When the speech recognition quality is lower than a preset second baseline, the predicted result of the speech recognition quality is output.

12. An apparatus for speech quality assessment, characterized in that, include: The acquisition module is used to acquire test audio. An evaluation module is used to evaluate the speech quality of the test speech based on semantic information related to the test speech. The speech quality of the test speech is related to the performance of the machine-oriented speech preprocessing module and the performance of the speech recognition module. The speech preprocessing module is used to preprocess the speech signal collected by the sensor, and the speech recognition module is used to perform speech recognition on the speech signal output by the speech preprocessing module. The output module is used to output an evaluation result of the speech quality when the speech quality is lower than a preset first baseline. The first baseline is an evaluation index of the speech quality. The evaluation result is used to indicate that the preprocessing parameters of the speech preprocessing module should be adjusted so that the speech quality of the preprocessed speech signal is greater than or equal to the first baseline.

13. The apparatus according to claim 12, characterized in that, The speech quality evaluation results output by the output module include one or more of the following: The quality of the test audio; Factors affecting the quality of the test audio; The method for adjusting the quality of the test audio.

14. The apparatus according to claim 12, characterized in that, The evaluation module is specifically used for: Obtain a first feature vector of the test speech, wherein the first feature vector includes the time-frequency feature vector of the test speech; A second feature vector of the test speech is obtained based on a first feature vector of the test speech, wherein the second feature vector is semantically related to the test speech; The speech quality of the test speech is evaluated based on the second feature vector.

15. The apparatus according to claim 14, characterized in that, When evaluating the speech quality of the test speech based on the second feature vector, the evaluation module includes: evaluating the speech quality of the test speech using a first evaluation metric. The first evaluation index includes: the concentration index of the center position of the test speech in each group, wherein the test speech in different groups has different semantics, and the center position of the test speech in each group is the center position of the second feature vector of the test speech in the second feature space, the second feature space being the space where the second feature vector is located.

16. The apparatus according to claim 14, characterized in that, When evaluating the speech quality of the test speech based on the second feature vector, the evaluation module includes: evaluating the speech quality of the test speech using a second evaluation metric. The second evaluation metric includes: a dispersion index of the second feature vector of each group of test speech in the second feature space, wherein the test speech of different groups has different semantics, and the second feature space is the space where the second feature vector is located.

17. The apparatus according to claim 14, characterized in that, When evaluating the speech quality of the test speech based on the second feature vector, the evaluation module includes: evaluating the speech quality of the test speech using a third evaluation metric. The third evaluation index includes: the similarity index between the center position of each group of test speech and the center position of each group of reference speech corresponding to the semantics; wherein, the test speech of different groups has different semantics, the center position of each group of test speech is the center position of the second feature vector of the test speech in the second feature space, the second feature space is the space where the second feature vector is located, the reference speech of different groups has different semantics, and the center position of each group of reference speech is the center position of the second feature vector of the reference speech in the second feature space.

18. The apparatus according to any one of claims 14 to 17, characterized in that, When obtaining the first feature vector of the test speech, the evaluation module includes: Obtain consecutive frames contained in the test speech, wherein adjacent frames contain overlapping information; Multiple feature vectors, including frequency domain features, are obtained from the consecutive frames, and the multiple feature vectors constitute the first feature vector.

19. A device for predicting the quality of speech recognition, characterized in that, include: An acquisition module is used to acquire test speech, wherein the speech quality of the test speech is greater than or equal to a first baseline, and the first baseline is an evaluation index of the speech quality. A prediction module is used to predict the speech recognition quality of the speech recognition model for the test speech based on a speech recognition quality function, wherein the speech recognition quality function is used to indicate the relationship between speech recognition quality and speech quality, and the speech quality is evaluated by the method according to any one of claims 1-7; The output module is used to output the prediction result of the speech recognition quality.

20. The apparatus according to claim 19, characterized in that, The prediction result of the speech recognition quality output by the output module includes one or more of the following: The speech recognition quality of the speech recognition model; Factors affecting the speech recognition quality of the aforementioned speech recognition model; The method for adjusting the speech recognition quality of the speech recognition model.

21. The apparatus according to claim 20, characterized in that, The process of constructing the speech recognition quality function includes: Obtain multiple sets of degraded reference speech; The speech recognition results of multiple sets of the degraded reference speech are obtained according to the speech recognition model and used as the first statistical result; Using multiple sets of degraded reference speech as test speech, a first evaluation result is obtained based on the speech quality evaluation results of the multiple sets of degraded reference speech; The speech recognition quality function is obtained based on the functional relationship between the first statistical result and the first evaluation result.

22. A device for improving speech recognition quality, characterized in that, include: The acquisition module is used to acquire test audio. An evaluation and prediction module is used to obtain a speech quality evaluation result of the test speech by the method according to any one of claims 1-7; The evaluation and prediction module is further configured to: predict the speech recognition quality according to the method of claim 8 or 9 when the speech quality of the preprocessed speech signal is greater than or equal to the first baseline. The output module is used to: when the speech recognition quality is lower than a preset second baseline, execute the output of the predicted result of the speech recognition quality.

23. A vehicle, characterized in that, include: A sound acquisition device used to collect users' voice commands; A preprocessing device is used to preprocess the sound of the voice command; A voice recognition device is used to recognize the pre-processed sound; The apparatus according to any one of claims 12 to 22.

24. A computing device, characterized in that, It includes one or more processors and one or more memories, the memories storing program instructions that, when executed by the one or more processors, cause the one or more processors to perform the method of any one of claims 1-11.

25. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by a computer, the computer causes the computer to perform the method according to any one of claims 1-11.

26. A computer program product, characterized in that, It includes program instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1-11.