Speech recognition model training method, speech recognition method and device

CN120544541APending Publication Date: 2025-08-26HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510618526.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

[0003]本申请提供一种语音识别模型训练方法、语音识别方法及装置,用以解决现有技术中嘈杂环境中的语音识别准确度低的缺陷,实现提高嘈杂环境中的语音识别准确度

Benefits of technology

[0014] The present application also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-mentioned speech recognition model training methods or any of the above-mentioned speech recognition methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544541A_ABST
    Figure CN120544541A_ABST
Patent Text Reader

Abstract

The invention discloses a voice recognition model training method and device and a voice recognition method and device, and relates to the technical field of smart home, and the voice recognition model training method comprises the steps: carrying out the voice recognition training of a generative pre-training Transform model based on a first noiseless voice data set, and obtaining a first voice recognition model; wherein the generative pre-training Transform model is obtained through pre-training of a large model; performing voice recognition training on the first voice recognition model based on the first noise-containing voice data set to obtain a second voice recognition model; and jointly training the second speech recognition model based on the noise suppression task and the speech recognition task by using a multi-task learning framework to obtain a third speech recognition model. According to the speech recognition model training method, the speech recognition method and the speech recognition device provided by the invention, noise suppression and speech enhancement are realized based on a large model, and the speech recognition accuracy in a noisy environment is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of smart home technology, and in particular to a speech recognition model training method, a speech recognition method and a speech recognition device. Background Art

[0002] Current speech recognition technology performs well in quiet environments, but its accuracy significantly decreases in noisy ones. This is primarily because traditional speech recognition models struggle to effectively filter out background noise. While existing speech enhancement methods have improved speech recognition to some extent, they remain insufficient to address the challenges of speech recognition in noisy environments. Summary of the Invention

[0003] The present application provides a speech recognition model training method, a speech recognition method and a device, which are used to solve the defect of low speech recognition accuracy in noisy environments in the prior art and to improve the speech recognition accuracy in noisy environments.

[0004] The present application provides a speech recognition model training method, comprising: performing speech recognition training on a generative pre-trained Transformer model based on a first noise-free speech dataset to obtain a first speech recognition model; wherein the generative pre-trained Transformer model is obtained through large model pre-training; performing speech recognition training on the first speech recognition model based on a first noisy speech dataset to obtain a second speech recognition model; and using a multi-task learning framework to jointly train the second speech recognition model based on a noise suppression task and a speech recognition task to obtain a third speech recognition model.

[0005] According to a speech recognition model training method provided in the present application, the speech recognition training of a generative pre-trained Transformer model based on a first noise-free speech dataset includes: using an unsupervised learning method to perform speech recognition training on the generative pre-trained Transformer model based on the first noise-free speech dataset; the speech recognition training of the first speech recognition model based on a first noisy speech dataset includes: using a supervised learning method to perform speech recognition training on the first speech recognition model based on the first noisy speech dataset.

[0006] According to a speech recognition model training method provided by the present application, before performing speech recognition training on a generative pre-trained Transformer model based on a first noise-free speech dataset, the method further includes: collecting multi-source speech data; performing data cleaning on the multi-source speech data to obtain a second noise-free speech dataset; wherein the data cleaning includes noise filtering; and merging the second noise-free speech dataset with an open source third noise-free speech dataset to obtain the first noise-free speech dataset.

[0007] According to a speech recognition model training method provided by the present application, before performing speech recognition training on the first speech recognition model based on the first noisy speech dataset, the method further includes: adding random noise to the second noise-free speech dataset to obtain a second noisy speech dataset; and merging the second noisy speech dataset with an open source third noisy speech dataset to obtain the first noisy speech dataset.

[0008] According to a speech recognition model training method provided in the present application, the loss function for obtaining the first speech recognition model, the second speech recognition model and the third speech recognition model through model training adopts a composite loss function; wherein, the composite loss function is expressed as a weighted sum of mean square error and cross entropy loss.

[0009] According to a speech recognition model training method provided in the present application, the generative pre-trained Transformer model adopts a quantized neural network model.

[0010] The present application also provides a speech recognition method, comprising: obtaining speech data to be speech recognized; inputting the speech data into a third speech recognition model, and outputting a speech recognition result of the speech data.

[0011] According to a speech recognition method provided by the present application, before inputting the speech data into the third speech recognition model, the method further includes: using a noise suppression network to remove noise from the speech data to be speech recognized.

[0012] The present application also provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute any of the above-mentioned speech recognition model training methods or any of the above-mentioned speech recognition methods through the computer program.

[0013] The present application also provides a computer-readable storage medium, which includes a stored program, wherein when the program is run, it executes and implements any of the above-mentioned speech recognition model training methods or any of the above-mentioned speech recognition methods.

[0014] The present application also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-mentioned speech recognition model training methods or any of the above-mentioned speech recognition methods.

[0015] The speech recognition model training method, speech recognition method and device provided in the present application are: a first speech recognition model is obtained by performing speech recognition training on a generative pre-trained Transformer model based on a first noise-free speech data set; the generative pre-trained Transformer model is obtained by pre-training a large model; the first speech recognition model is trained for speech recognition based on a first noisy speech data set to obtain a second speech recognition model; a multi-task learning framework is used to jointly train the second speech recognition model based on a noise suppression task and a speech recognition task to obtain a third speech recognition model; noise suppression and speech enhancement are achieved based on the large model, and the speech recognition accuracy in noisy environments is significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present application or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1 This is a hardware environment diagram of a speech recognition model training method according to an embodiment of the present application.

[0018] Figure 2 This is a flow chart of the speech recognition model training method provided in this application.

[0019] Figure 3 It is a flowchart of the speech recognition method provided by this application.

[0020] Figure 4 It is a structural diagram of the speech recognition model training device provided in this application.

[0021] Figure 5 It is a structural diagram of the speech recognition device provided in this application.

[0022] Figure 6 It is a structural diagram of the electronic device provided in this application. DETAILED DESCRIPTION

[0023] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0025] According to one aspect of the embodiment of the present application, a speech recognition model training method is provided. The speech recognition model training method is widely used in smart home (Smart Home), smart home, smart home device ecology, smart residential (Intelligence House) ecology and other whole-house intelligent digital control application scenarios. Optionally, in this embodiment, the above-mentioned speech recognition model training method can be applied to Figure 1 In the hardware environment shown in FIG. 1 , which is composed of a terminal device 102 and a server 104. Figure 1 As shown, the server 104 is connected to the terminal device 102 via a network, and can be used to provide services (such as application services, etc.) for the terminal or a client installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data computing services for the server 104.

[0026] The aforementioned network may include, but is not limited to, at least one of the following: a wired network and a wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, and a local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity) and Bluetooth. The terminal device 102 may include, but is not limited to, a PC, a mobile phone, a tablet computer, a smart air conditioner, a smart range hood, a smart refrigerator, a smart oven, a smart stove, a smart washing machine, a smart water heater, a smart washing machine, a smart dishwasher, a smart projector, a smart TV, a smart clothes drying rack, smart curtains, a smart audio / video system, a smart socket, a smart speaker, a smart fresh air system, smart kitchen and bathroom equipment, smart bathroom equipment, a smart robot vacuum, a smart window cleaning robot, a smart robot mop, a smart air purifier, a smart steamer, a smart microwave oven, a smart kitchen appliance, a smart purifier, a smart water dispenser, a smart door lock, and the like.

[0027] Figure 2 This is a flow chart of the speech recognition model training method provided by this application. Figure 2 As shown, the method includes: Step S1: Perform speech recognition training on a generative pre-trained Transformer model based on a first noise-free speech dataset to obtain a first speech recognition model; wherein the generative pre-trained Transformer model is obtained through large model pre-training.

[0028] A generative pre-trained Transformer model is trained for speech recognition based on a noise-free speech dataset to obtain basic speech recognition capabilities, resulting in a first speech recognition model. For purposes of distinction, the noise-free speech dataset used to train the first speech recognition model is referred to as the first noise-free speech dataset. The generative pre-trained Transformer model is referred to as the GPT model. The GPT model used can be HomeGPT or another GPT model.

[0029] Step S2: Perform speech recognition training on the first speech recognition model based on the first noisy speech data set to obtain a second speech recognition model.

[0030] The first speech recognition model is trained for speech recognition based on the noisy speech dataset to obtain a second speech recognition model. For purposes of differentiation, the noisy speech dataset used to train the first speech recognition model is referred to as the first noisy speech dataset. After training, the first noisy speech dataset can achieve a certain degree of ability to distinguish between normal speech data and noise during speech recognition.

[0031] Step S3: Using a multi-task learning framework, jointly train the second speech recognition model based on the noise suppression task and the speech recognition task to obtain a third speech recognition model.

[0032] Using the multi-task learning (MTL) framework, the second speech recognition model is further jointly trained based on the noise suppression and speech recognition tasks, sharing some network parameters to generate a third speech recognition model. This third speech recognition model possesses both noise suppression and speech recognition capabilities. When jointly training the second speech recognition model based on the noise suppression and speech recognition tasks, it can be trained based on the first noisy speech dataset.

[0033] The speech recognition model training method provided in the present application obtains a first speech recognition model by performing speech recognition training on a generative pre-trained Transformer model based on a first noise-free speech dataset. The generative pre-trained Transformer model is obtained by pre-training a large model. The first speech recognition model is trained for speech recognition based on a first noisy speech dataset to obtain a second speech recognition model. The second speech recognition model is jointly trained based on a noise suppression task and a speech recognition task using a multi-task learning framework to obtain a third speech recognition model. Noise suppression and speech enhancement are achieved based on the large model, significantly improving the accuracy of speech recognition in noisy environments.

[0034] According to a speech recognition model training method provided in the present application, the speech recognition training of a generative pre-trained Transformer model based on a first noise-free speech dataset includes: using an unsupervised learning method to train the generative pre-trained Transformer model for speech recognition based on the first noise-free speech dataset. The speech recognition training of the first speech recognition model based on a first noisy speech dataset includes: using a supervised learning method to train the first speech recognition model for speech recognition based on the first noisy speech dataset.

[0035] When performing speech recognition training on the generative pre-trained Transformer model based on the first noise-free speech dataset, unsupervised learning is used for training to capture the basic patterns and features in the speech data, so that the trained first speech recognition model has basic speech recognition capabilities.

[0036] When performing speech recognition training on the first speech recognition model based on the first noisy speech dataset, a supervised learning method is used for fine-tuning to adapt to different types of noise environments, so that the trained second speech recognition model can better distinguish between noise and normal speech.

[0037] The speech recognition model training method provided in this application uses an unsupervised learning method to perform speech recognition training on a generative pre-trained Transformer model based on a first noise-free speech dataset, and uses a supervised learning method to perform speech recognition training on a first speech recognition model based on a first noisy speech dataset, thereby further improving the accuracy of speech recognition.

[0038] According to a speech recognition model training method provided by the present application, before performing speech recognition training on a generative pre-trained Transformer model based on a first noise-free speech dataset, the method further includes: collecting multi-source speech data; performing data cleaning on the multi-source speech data to obtain a second noise-free speech dataset; wherein the data cleaning includes noise filtering; and merging the second noise-free speech dataset with an open source third noise-free speech dataset to obtain the first noise-free speech dataset.

[0039] Before performing speech recognition training on the generative pre-trained Transformer model based on the first noise-free speech dataset, it is first necessary to obtain the first noise-free speech dataset. The process of obtaining the first noise-free speech dataset may include: Based on the specific speech recognition application scenario, multi-source speech data can be collected to improve data quality and diversity, as well as the accuracy of speech recognition. For example, if speech recognition is applied in a home environment, multi-source speech data can be collected based on the home environment to better improve the accuracy of speech recognition. If speech recognition is not restricted to certain scenarios, multi-source speech data can be collected randomly. Microphone arrays, cameras, and other sensors can be used to collect user voice data and environmental noise data in real time. Among them, the microphone array is used to capture speech and noise signals from different directions, and the camera is used to provide visual information to assist speech recognition. Whether other sensors are used and the type of sensors used can be determined according to specific needs.

[0040] Data cleaning is performed on the multi-source speech data to obtain a second noise-free speech dataset. Data cleaning aims to remove invalid or low-quality data to ensure the accuracy and reliability of the dataset. Data cleaning includes noise filtering. Filters (such as low-pass and high-pass filters) can be used to remove background noise. Data cleaning can also include outlier detection and signal enhancement. Outlier detection identifies and removes anomalous data points, such as extreme values. Signal enhancement, such as gain control, improves signal quality.

[0041] The second noise-free speech dataset is labeled. Data labeling involves manually or automatically annotating the collected speech data to generate the labels required for training. This step typically requires specialized knowledge and significant human effort. Manual or automatic labeling can be used. Manual labeling involves experts or trained personnel manually labeling the speech data. Automatic labeling can utilize pre-trained models or automatic labeling tools for preliminary labeling, followed by manual verification.

[0042] Data enhancement can also be performed on the second noise-free speech dataset. This can be achieved through methods such as time-frequency transformation and signal mixing, expanding the dataset and improving the robustness of the model. Time-frequency transformation transforms time-frequency features through operations such as time shifting, frequency shifting, and scaling, increasing data diversity. Signal mixing combines different speech signals to simulate multi-user speech data, improving multi-user speech recognition capabilities.

[0043] Since the amount of collected second noise-free speech dataset is limited, to obtain large-scale data for speech recognition training, the second noise-free speech dataset is combined with the open-source third noise-free speech dataset to obtain the first noise-free speech dataset for speech recognition training of the generative pre-trained Transformer model. The open-source third noise-free speech dataset can be a large-scale noise-free speech dataset such as LibriSpeech or VoxCeleb.

[0044] The speech recognition model training method provided in the present application collects multi-source speech data and performs data cleaning on the multi-source speech data to obtain a second noise-free speech data set, wherein the data cleaning includes noise filtering, and the second noise-free speech data set is merged with an open source third noise-free speech data set to obtain a first noise-free speech data set, thereby improving the data quality of the first noise-free speech data set and further improving the accuracy of speech recognition.

[0045] According to a speech recognition model training method provided by the present application, before performing speech recognition training on the first speech recognition model based on the first noisy speech dataset, the method further includes: adding random noise to the second noise-free speech dataset to obtain a second noisy speech dataset; and merging the second noisy speech dataset with an open source third noisy speech dataset to obtain the first noisy speech dataset.

[0046] Before performing speech recognition training on the first speech recognition model based on the first noisy speech dataset, it is necessary to first obtain the first noisy speech dataset. The process of obtaining the first noisy speech dataset includes: Random noise was added to the second noise-free speech dataset to improve the model's adaptability to noisy environments, resulting in the second noisy speech dataset. Due to the limited size of the second noisy speech dataset, to obtain large-scale data for speech recognition training, the second noisy speech dataset was merged with the open-source third noisy speech dataset to obtain the first noisy speech dataset. The open-source third noisy speech dataset uses speech datasets containing various types of noise, such as CHiME or Aurora.

[0047] The speech recognition model training method provided in this application adds random noise to a second noise-free speech dataset to obtain a second noisy speech dataset, and merges the second noisy speech dataset with an open source third noisy speech dataset to obtain a first noisy speech dataset, thereby improving the data quality of the first noisy speech dataset and further improving the accuracy of speech recognition.

[0048] According to a speech recognition model training method provided in the present application, the loss function for obtaining the first speech recognition model, the second speech recognition model and the third speech recognition model through model training adopts a composite loss function; wherein, the composite loss function is expressed as a weighted sum of mean square error and cross entropy loss.

[0049] Common loss functions include mean square error (MSE) and cross-entropy loss. MSE is used for regression tasks and is suitable for speech signal reconstruction in noise suppression. Its advantage is that it is sensitive to outliers and helps optimize the model's ability to process details. Cross-entropy loss is used for classification tasks and is suitable for classification problems in speech recognition. Its advantage is that it effectively models probability distributions and is suitable for optimizing classification accuracy.

[0050] The composite loss function in this application combines MSE and cross-entropy loss to balance noise suppression and speech recognition accuracy, while optimizing multiple tasks and improving overall model performance. The loss function used to obtain the first, second, and third speech recognition models through model training uses a composite loss function; the composite loss function is expressed as a weighted sum of mean square error and cross-entropy loss.

[0051] The composite loss function is expressed as: in, represents the loss value of the composite loss function, represents the weight of the mean square error, represents the mean square error, represents the weight of the cross entropy loss, represents the cross entropy loss.

[0052] The speech recognition model training method provided in this application adopts a composite loss function as the loss function when obtaining the first speech recognition model, the second speech recognition model and the third speech recognition model through model training. The composite loss function is expressed as the weighted sum of the mean square error and the cross entropy loss, which further improves the accuracy of speech recognition.

[0053] According to a speech recognition model training method provided in the present application, the generative pre-trained Transformer model adopts a quantized neural network model.

[0054] To speed up speech recognition processing and ensure a smooth user experience, this application's generative pre-trained Transformer model uses a quantized neural network model. In addition, parallel computing technology and hardware acceleration (such as GPUs and TPUs) can be used to improve the system's overall processing capabilities.

[0055] The speech recognition model training method provided in this application improves the speech recognition processing efficiency and enhances the user experience by adopting a quantized neural network model through a generative pre-trained Transformer model.

[0056] Figure 3 This is a flow chart of the speech recognition method provided by this application. Figure 3 As shown, the method includes: Step 101: Acquire speech data to be subjected to speech recognition.

[0057] This application is an application of the third speech recognition model in speech recognition. The speech data to be speech recognized is obtained. The speech data to be speech recognized can be data collected by a microphone array and a camera in an actual application environment.

[0058] Step 102: Input the speech data into a third speech recognition model, and output a speech recognition result of the speech data.

[0059] The speech data is input into the trained third speech recognition model, and a speech recognition result of the speech data is output.

[0060] The speech recognition method provided in the present application improves the accuracy of speech recognition by acquiring speech data to be speech recognized, inputting the speech data into a third speech recognition model, and outputting a speech recognition result of the speech data.

[0061] According to a speech recognition method provided by the present application, before inputting the speech data into the third speech recognition model, the method further includes: using a noise suppression network to remove noise from the speech data to be speech recognized.

[0062] Before inputting the speech data into the third speech recognition model, the existing noise suppression network can be used to remove noise from the speech data to be recognized. This filtering out background noise in real time enhances the quality of the speech signal, improving the subsequent speech recognition performance of the third speech recognition model and ensuring that the speech recognition system can accurately understand user commands. The noise suppression network can use a low-latency algorithm to ensure that the noise suppression process does not affect the real-time performance of voice interaction.

[0063] The noise suppression network can be based on a convolutional neural network (CNN) or a U-Net architecture, specifically designed to separate background noise from speech signals. CNNs are used for feature extraction and can effectively separate background noise from speech signals. U-Nets utilize skip connections to pass high-resolution features to the decoder, enhancing detail processing capabilities.

[0064] The speech recognition method provided in the present application further improves the accuracy of speech recognition by using a noise suppression network to remove noise from the speech data to be speech recognized before inputting the speech data into the third speech recognition model.

[0065] The speech recognition method provided in this application significantly improves the accuracy of speech recognition in noisy environments through large-model noise suppression and speech enhancement technology; by adopting a quantized neural network model, it ensures rapid response of voice interaction and improves user satisfaction.

[0066] The speech recognition model training device provided in this application is described below. The speech recognition model training device described below and the speech recognition model training method described above can be referenced to each other.

[0067] Figure 4 This is a schematic diagram of the structure of the speech recognition model training device provided by this application. Figure 4 As shown, the device includes a first training module 10, a second training module 20 and a third training module 30, wherein: the first training module 10 is used to: perform speech recognition training on a generative pre-trained Transformer model based on a first noise-free speech data set to obtain a first speech recognition model; wherein the generative pre-trained Transformer model is obtained through large model pre-training; the second training module 20 is used to: perform speech recognition training on the first speech recognition model based on a first noisy speech data set to obtain a second speech recognition model; the third training module 30 is used to: use a multi-task learning framework to jointly train the second speech recognition model based on a noise suppression task and a speech recognition task to obtain a third speech recognition model.

[0068] The speech recognition model training method provided in the present application obtains a first speech recognition model by performing speech recognition training on a generative pre-trained Transformer model based on a first noise-free speech dataset. The generative pre-trained Transformer model is obtained by pre-training a large model. The first speech recognition model is trained for speech recognition based on a first noisy speech dataset to obtain a second speech recognition model. The second speech recognition model is jointly trained based on a noise suppression task and a speech recognition task using a multi-task learning framework to obtain a third speech recognition model. Noise suppression and speech enhancement are achieved based on the large model, significantly improving the accuracy of speech recognition in noisy environments.

[0069] Figure 5 This is a schematic diagram of the structure of the speech recognition device provided by this application. Figure 5 As shown, the device includes an acquisition module 100 and a speech recognition module 200, wherein: the acquisition module 100 is used to: acquire speech data to be speech recognized; the speech recognition module 200 is used to: input the speech data into a third speech recognition model and output the speech recognition result of the speech data.

[0070] The speech recognition method provided in the present application improves the accuracy of speech recognition by acquiring speech data to be speech recognized, inputting the speech data into a third speech recognition model, and outputting a speech recognition result of the speech data.

[0071] Figure 6 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 6 As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 may call logic instructions in the memory 630 to execute a speech recognition model training method, which includes: performing speech recognition training on a generative pre-trained Transformer model based on a first noise-free speech data set to obtain a first speech recognition model; wherein the generative pre-trained Transformer model is obtained by pre-training a large model; performing speech recognition training on the first speech recognition model based on a first noisy speech data set to obtain a second speech recognition model; using a multi-task learning framework, jointly training the second speech recognition model based on a noise suppression task and a speech recognition task to obtain a third speech recognition model; Or execute a speech recognition method, which includes: obtaining speech data to be speech recognized; inputting the speech data into a third speech recognition model, and outputting a speech recognition result of the speech data.

[0072] In addition, the logical instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0073] On the other hand, the present application also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the speech recognition model training method provided by the above methods, the method comprising: performing speech recognition training on a generative pre-trained Transformer model based on a first noise-free speech dataset to obtain a first speech recognition model; wherein the generative pre-trained Transformer model is obtained by pre-training a large model; performing speech recognition training on the first speech recognition model based on a first noisy speech dataset to obtain a second speech recognition model; using a multi-task learning framework, jointly training the second speech recognition model based on a noise suppression task and a speech recognition task to obtain a third speech recognition model; Or execute a speech recognition method, which includes: obtaining speech data to be speech recognized; inputting the speech data into a third speech recognition model, and outputting a speech recognition result of the speech data.

[0074] In another aspect, the present application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the speech recognition model training method provided by each of the above methods, the method comprising: performing speech recognition training on a generative pre-trained Transformer model based on a first noise-free speech dataset to obtain a first speech recognition model; wherein the generative pre-trained Transformer model is obtained by pre-training a large model; performing speech recognition training on the first speech recognition model based on a first noisy speech dataset to obtain a second speech recognition model; using a multi-task learning framework, jointly training the second speech recognition model based on a noise suppression task and a speech recognition task to obtain a third speech recognition model; Or execute a speech recognition method, which includes: obtaining speech data to be speech recognized; inputting the speech data into a third speech recognition model, and outputting a speech recognition result of the speech data.

[0075] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0076] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A speech recognition model training method, characterized in that: include: Performing speech recognition training on a generative pre-trained Transformer model based on a first noise-free speech dataset to obtain a first speech recognition model; wherein the generative pre-trained Transformer model is obtained by pre-training with a large model; Performing speech recognition training on the first speech recognition model based on the first noisy speech dataset to obtain a second speech recognition model; Using a multi-task learning framework, the second speech recognition model is jointly trained based on the noise suppression task and the speech recognition task to obtain a third speech recognition model.

2. The speech recognition model training method according to claim 1, characterized in that The speech recognition training of the generative pre-trained Transformer model based on the first noise-free speech dataset includes: using an unsupervised learning method to perform speech recognition training on the generative pre-trained Transformer model based on the first noise-free speech dataset; The performing speech recognition training on the first speech recognition model based on the first noisy speech dataset includes: using a supervised learning method to perform speech recognition training on the first speech recognition model based on the first noisy speech dataset.

3. The speech recognition model training method according to claim 1, characterized in that Before performing speech recognition training on the generative pre-trained Transformer model based on the first noise-free speech dataset, the method further includes: Collect multi-source voice data; Performing data cleaning on the multi-source speech data to obtain a second noise-free speech data set; wherein the data cleaning includes noise filtering; The second noise-free speech dataset is merged with an open-source third noise-free speech dataset to obtain the first noise-free speech dataset.

4. The speech recognition model training method according to claim 3, characterized in that: Before performing speech recognition training on the first speech recognition model based on the first noisy speech dataset, the method further includes: adding random noise to the second noise-free speech data set to obtain a second noisy speech data set; The second noisy speech dataset is merged with an open-source third noisy speech dataset to obtain the first noisy speech dataset.

5. The speech recognition model training method according to claim 1, characterized in that: The loss function used when obtaining the first speech recognition model, the second speech recognition model and the third speech recognition model through model training adopts a composite loss function; wherein, the composite loss function is expressed as a weighted sum of mean square error and cross entropy loss.

6. The speech recognition model training method according to claim 1, characterized in that: The generative pre-trained Transformer model uses a quantized neural network model.

7. A speech recognition method based on the speech recognition model training method according to any one of claims 1 to 6, characterized in that: include: Acquire voice data to be used for voice recognition; The speech data is input into a third speech recognition model, and a speech recognition result of the speech data is output.

8. The speech recognition method according to claim 7, characterized in that: Before inputting the speech data into the third speech recognition model, the method further includes: A noise suppression network is used to remove noise from the speech data to be subjected to speech recognition.

9. A speech recognition model training device, characterized in that: include: A first training module is configured to perform speech recognition training on a generative pre-trained Transformer model based on a first noise-free speech dataset to obtain a first speech recognition model; wherein the generative pre-trained Transformer model is obtained by pre-training with a large model; A second training module is configured to perform speech recognition training on the first speech recognition model based on a first noisy speech data set to obtain a second speech recognition model; The third training module is used to: use a multi-task learning framework to jointly train the second speech recognition model based on the noise suppression task and the speech recognition task to obtain a third speech recognition model.

10. A speech recognition device, characterized in that: include: An acquisition module is used to: acquire voice data to be subjected to voice recognition; The speech recognition module is used to: input the speech data into a third speech recognition model and output a speech recognition result of the speech data.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program, when running, executes the speech recognition model training method according to any one of claims 1 to 6 or the speech recognition method according to claim 7 or 8.

12. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the speech recognition model training method according to any one of claims 1 to 6 or the speech recognition method according to claim 7 or 8 through the computer program.

Citation Information

Patent Citations

  • Speech enhancement model training method and device and speech enhancement method and device

    CN113593594A

  • Model training method and device and device for model training

    CN113707134A

  • Speech recognition method fused with speech enhancement

    CN114495969A

  • Speech recognition model training method and device, electronic equipment and storage medium

    CN116504252A

  • Voice multitask-based model training method and voice multitask processing method

    CN119580745A