A method, device and storage medium for training a speech recognition model

By utilizing unlabeled speech data and the correction and annotation of multiple open-source models, the problem of high training cost of speech recognition models is solved, and efficient and low-cost speech recognition model training is achieved.

CN116524905BActive Publication Date: 2026-06-02XIAOVO TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAOVO TECH
Filing Date
2023-02-15
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Training speech recognition models is costly, and existing technologies require a large amount of manually labeled data, resulting in excessive time and financial costs.

Method used

Using unlabeled speech data, the recognition results are corrected through multiple open-source speech recognition models. Combined with preset text correction principles and error correction models, the labeling results are determined, and the training model is iteratively trained to meet preset evaluation conditions.

Benefits of technology

While ensuring recognition accuracy, it significantly reduced training costs and achieved efficient speech recognition model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116524905B_ABST
    Figure CN116524905B_ABST
Patent Text Reader

Abstract

The application discloses a speech recognition model training method and device, equipment and a storage medium. The method comprises the following steps: obtaining unlabeled speech data of a target field, inputting the unlabeled speech data into at least two open-source speech recognition models respectively, and obtaining recognition results matched with the open-source speech recognition models; correcting the recognition results according to a preset text correction principle, and obtaining correction results matched with the open-source speech recognition models; determining a labeling result of the unlabeled speech data according to the correction results; and training a to-be-trained speech recognition model according to the unlabeled speech data and the labeling result, so as to obtain a speech recognition model whose recognition result of the speech data of the target field meets a preset evaluation condition. The technical scheme solves the problem of high training cost of the speech recognition model, can train the speech recognition model by using the unlabeled speech data, and greatly reduces the training cost while ensuring the recognition accuracy of the speech recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech processing technology, and in particular to a training method, apparatus, device, and storage medium for a speech recognition model. Background Technology

[0002] Currently, speech recognition models are mainly obtained through supervised training using large-scale speech data and labeled data matched with the speech data.

[0003] However, labeling large-scale speech data requires high manual costs and is time-consuming. Therefore, obtaining speech recognition models using unlabeled speech data and training methods is of great significance. Summary of the Invention

[0004] This invention provides a training method, apparatus, device, and storage medium for a speech recognition model to solve the problem of high training costs for speech recognition models. It can utilize unlabeled speech data for training, thereby significantly reducing training costs while ensuring the accuracy of the speech recognition model.

[0005] According to one aspect of the present invention, a method for training a speech recognition model is provided, the method comprising:

[0006] Obtain unlabeled speech data in the target domain, input the unlabeled speech data into at least two open-source speech recognition models respectively, and obtain recognition results that match each open-source speech recognition model;

[0007] According to the preset text correction principles, each recognition result is corrected to obtain the corrected results of each open-source speech recognition model; wherein, the text correction principles are determined based on the text data of the target domain;

[0008] Based on the correction results of matching various open-source speech recognition models, the annotation results of unlabeled speech data are determined;

[0009] Based on the unlabeled speech data and the annotation results, the speech recognition model to be trained is trained to obtain a speech recognition model whose recognition results for speech data in the target domain meet the preset evaluation conditions.

[0010] According to another aspect of the present invention, a training apparatus for a speech recognition model is provided, the apparatus comprising:

[0011] The recognition result determination module is used to acquire unlabeled speech data in the target domain, input the unlabeled speech data into at least two open-source speech recognition models respectively, and obtain recognition results that match each open-source speech recognition model;

[0012] The correction result determination module is used to correct each recognition result according to a preset text correction principle to obtain the correction result of each open source speech recognition model; wherein, the text correction principle is determined based on text data in the target domain;

[0013] The annotation result determination module is used to determine the annotation results of unlabeled speech data based on the correction results of matching various open-source speech recognition models;

[0014] The recognition model training module is used to train the speech recognition model to be trained based on the unlabeled speech data and the annotation results, so as to obtain a speech recognition model whose recognition results for speech data in the target domain meet the preset evaluation conditions.

[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the training method of the speech recognition model according to any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the training method of the speech recognition model according to any embodiment of the present invention.

[0020] The technical solution of this invention involves inputting unlabeled speech data from the target domain into at least two open-source speech recognition models to obtain recognition results that match each open-source speech recognition model. Then, according to a preset text correction principle, each recognition result is corrected to obtain corrected results that match each open-source speech recognition model. Next, based on the corrected results matched by each open-source speech recognition model, the annotation result of the unlabeled speech data is determined. Finally, based on the unlabeled speech data and the annotation result, the speech recognition model to be trained is trained to obtain a speech recognition model whose recognition results for speech data in the target domain meet preset evaluation conditions. This technical solution solves the problem of high training costs for speech recognition models by utilizing unlabeled speech data for training, significantly reducing training costs while ensuring the accuracy of the speech recognition model.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of a training method for a speech recognition model according to Embodiment 1 of the present invention;

[0024] Figure 2 This is a flowchart of a training method for a speech recognition model according to Embodiment 2 of the present invention;

[0025] Figure 3 This is a schematic diagram of the process of determining the training dataset according to Embodiment 2 of the present invention;

[0026] Figure 4 This is a schematic diagram of the structure of a training device for a speech recognition model according to Embodiment 3 of the present invention;

[0027] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the training method of the speech recognition model in the embodiments of the present invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be used interchangeably where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices. The acquisition, storage, use, and processing of data in the technical solutions of this application all comply with the relevant provisions of national laws and regulations.

[0030] Example 1

[0031] Figure 1 This is a flowchart illustrating a training method for a speech recognition model according to Embodiment 1 of the present invention. This embodiment is applicable to speech recognition model training scenarios. The method can be executed by a speech recognition model training device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0032] S110. Obtain unlabeled speech data in the target domain, and input the unlabeled speech data into at least two open-source speech recognition models respectively to obtain recognition results that match each open-source speech recognition model.

[0033] This solution can be executed by electronic devices such as computers and servers. For specialized fields such as telephone customer service, healthcare, and finance, voice data typically contains a large number of specialized terms, and the content is highly relevant. Therefore, targeted speech recognition models are needed to recognize the voice data. During the voice data recognition process, the electronic device can read a pre-built voice dataset. The voice dataset can include large-scale unlabeled voice data within the target domain. The target domain can be a vertical field such as healthcare or finance. Unlabeled voice data can be voice data without matching text labels, meaning the voice data has not been manually labeled.

[0034] Electronic devices can simultaneously input unlabeled speech data into multiple open-source speech recognition models for speech recognition. The recognition results of these open-source speech recognition models for non-vertical domain speech data meet preset evaluation criteria. In other words, these open-source speech recognition models have a certain recognition capability for unlabeled speech data and can achieve a certain recognition accuracy for non-vertical domain speech data. For example, the recognition accuracy of open-source speech recognition models for non-vertical domain speech data can be greater than 80%. It should be noted that the multiple open-source speech recognition models can be different; they may differ in model structure, training process, and training data.

[0035] After performing speech recognition on unlabeled speech data, each open-source speech recognition model can obtain the recognition results for the unlabeled speech data.

[0036] S120. According to the preset text correction principle, each recognition result is corrected to obtain the corrected results of each open source speech recognition model.

[0037] Understandably, electronic devices can acquire textual data such as books, papers, and journals within a target domain and statistically analyze the specialized vocabulary within the text. Simultaneously, electronic devices can extract textual features from the textual data to train models for text detection and text correction. Based on the textual data, electronic devices can determine text correction principles and, in accordance with these principles, perform word replacement, text correction, and other operations on the recognition results, thereby obtaining the corrected results corresponding to each recognition outcome.

[0038] S130. Based on the correction results of matching various open-source speech recognition models, determine the labeling results of the unlabeled speech data.

[0039] Electronic devices can compare the correction results corresponding to various open-source speech recognition models and determine the annotation result of the unlabeled speech data based on the comparison results. Specifically, if all correction results are the same, the correction result is used as the annotation result of the unlabeled speech data. If at least one correction result differs from the others, the unlabeled speech data is determined to be unlabeled.

[0040] Electronic devices can also calculate the similarity between the corrected results of every two open-source speech recognition models and determine the annotation result of the unlabeled speech data based on each similarity. For example, given three open-source speech recognition models, model A, model B, and model C, and their corrected matching results A, B, and C respectively, if the similarity between result A and result B is 100%, the similarity between result A and result C is 90%, and the similarity between result B and result C is 90%, the electronic device can use result A or result B as the annotation result of the unlabeled speech data.

[0041] S140. Based on the unlabeled speech data and the annotation results, train the speech recognition model to be trained to obtain a speech recognition model whose recognition results for speech data in the target domain meet the preset evaluation conditions.

[0042] Unlabeled speech data and the annotation results of unlabeled speech data matching are used as the dataset for training a speech recognition model, which is then iteratively trained. Based on the recognition results of speech data in the target domain, the electronic device can output a speech recognition model that meets the evaluation criteria. These evaluation criteria can be determined based on evaluation metrics such as recognition accuracy, loss, precision, recall, and F1 score. Specifically, evaluation criteria may include one or more of the following: recognition accuracy greater than a preset accuracy threshold, loss less than a preset loss threshold, precision greater than a preset precision threshold, and recall less than a preset recall rate. The electronic device can divide the dataset of the speech recognition model to be trained into training and testing sets. The training set is used for iterative training of the speech recognition model, and the testing set is used to test the trained speech recognition model. Based on the test results, the electronic device can determine the recognition accuracy, loss, precision, recall, and F1 score of the speech recognition model. By verifying whether the evaluation criteria are met, the electronic device can obtain a speech recognition model that meets the requirements of speech recognition in the target domain.

[0043] This technical solution involves inputting unlabeled speech data from the target domain into at least two open-source speech recognition models to obtain recognition results that match each model. Then, according to preset text correction principles, each recognition result is corrected to obtain corrected results that match each open-source speech recognition model. Based on the corrected results from each open-source model, the annotation results for the unlabeled speech data are determined. Finally, the speech recognition model to be trained is trained using the unlabeled speech data and the annotation results to obtain a speech recognition model whose recognition results for speech data in the target domain meet preset evaluation conditions. This technical solution solves the problem of high training costs for speech recognition models by utilizing unlabeled speech data for training, significantly reducing training costs while maintaining the accuracy of the speech recognition model.

[0044] Example Two

[0045] Figure 2 The figure is a flowchart of a method for training a speech recognition model provided in Example Two of the present invention. This example is a refinement based on the above example. As Figure 2 shown, the method includes:

[0046] S210. Obtain unlabeled speech data in the target domain, and input the unlabeled speech data into at least two open-source speech recognition models respectively to obtain recognition results matching each open-source speech recognition model.

[0047] S220. According to a preset set of target phrases, perform target vocabulary replacement on each recognition result respectively to obtain replacement results matching each open-source speech recognition model.

[0048] On the basis of the above solution, the set of target phrases is determined according to the vocabulary frequency statistics results in the text data of the target domain.

[0049] It is easy to understand that the electronic device can, according to the word frequency statistics results of professional vocabulary in the text data, use professional vocabulary with a word frequency greater than a preset threshold as target vocabulary. The electronic device can form target phrases by combining the target word and the words associated with the target word. The words associated with the target word can be synonyms, homophones, etc. of the target word. For example, if the target word is 5G, the target phrases can include homophones of 5G such as Wuji, Wuji, and Wujing, and can also include synonyms such as fifth-generation communication and fifth-generation mobile communication technology.

[0050] The electronic device can compare each recognition result with each target phrase in the set of target phrases in turn, replace all the words involved in the target phrase in each recognition result with the target word, and output the replacement results of each recognition result.

[0051] This solution performs the same target vocabulary replacement on each recognition result, which can eliminate the differences in the recognition of professional vocabulary in the target domain by different open-source speech recognition models, and is beneficial to the accurate judgment of the annotation results.

[0052] S230. Based on a pre-trained error correction model, perform text error correction on each replacement result respectively to obtain correction results matching each open-source speech recognition model.

[0053] Electronic devices can also pre-train error correction models based on deep learning algorithms using text data from the target domain. These models can be used to locate text errors in the replacement results, such as missing words, extra words, and misspelled words. Based on the text errors identified by the error correction model, the electronic device can correct the text errors in the replacement results according to the error type, thereby obtaining the corrected results matched by various open-source speech recognition models.

[0054] Optionally, the error correction model is trained using text data from the target domain and error correction labels for the text data.

[0055] This approach can train an error correction model to correct textual errors in the replacement results, which helps to obtain accurate correction results.

[0056] S240. Determine whether the number of open-source speech recognition models with the same correction result meets the preset quantity requirement. If yes, execute S250; otherwise, execute S260.

[0057] Electronic devices can compare the correction results corresponding to various open-source speech recognition models and determine the labeling result of the unlabeled speech data based on the proportion of identical correction results in the total number of correction results. Specifically, if 8 out of 10 correction results are the same, then that correction result is considered the labeling result of the unlabeled speech data. If 2 out of 10 correction results are the same, then the unlabeled speech data is determined to be unlabeled.

[0058] S250, Determine that the unlabeled speech data has no labeling result.

[0059] S260. The correction result is used as the annotation result of the unlabeled speech data.

[0060] S270. Use the unlabeled speech data with annotation results as the input to the speech recognition model to be trained. Based on the output results of the speech recognition model to be trained and the annotation results of the unlabeled speech data matching, iteratively train the speech recognition model to be trained.

[0061] The electronic device can filter out unlabeled speech data with annotation results from all unlabeled speech data, and use the annotated unlabeled speech data and its annotation results as the training dataset for the speech recognition model to be trained. Based on the output of the speech recognition model to be trained and the annotation results matched with the unlabeled speech data, the single device can iteratively train the speech recognition model to be trained. It should be noted that the speech recognition model to be trained can be one of various open-source speech recognition models, or it can be a self-built speech recognition model that has not undergone any training.

[0062] Figure 3 This is a schematic diagram illustrating the process of determining the training dataset according to Embodiment 2 of the present invention. Figure 3 As shown, in a specific scheme, the electronic device simultaneously inputs unlabeled speech data into open-source speech recognition model A and open-source speech recognition model B, obtaining recognition result A and recognition result B respectively. Both recognition result A and recognition result B are then replaced with high-frequency words from the target domain, yielding replacement result A and replacement result B respectively. High-frequency words can be words in the text data of the target domain that appear more than a preset threshold number of times. A text correction model is used to correct replacement result A and replacement result B, resulting in corrected result A and corrected result B. By comparing whether corrected result A and corrected result B are the same, the electronic device can determine whether the unlabeled speech data can be labeled and whether it needs to be added to the training dataset of the speech recognition model to be trained. Specifically, if corrected result A and corrected result B are the same, the electronic device can use either corrected result A or corrected result B as the labeling result for the unlabeled speech data and add the unlabeled speech data and the labeling result to the dataset. If corrected result A and corrected result B are different, the electronic device can discard the unlabeled speech data and not add it to the dataset.

[0063] This technical solution involves inputting unlabeled speech data from the target domain into at least two open-source speech recognition models to obtain recognition results that match each model. Then, according to preset text correction principles, each recognition result is corrected to obtain corrected results that match each open-source speech recognition model. Based on the corrected results from each open-source model, the annotation results for the unlabeled speech data are determined. Finally, the speech recognition model to be trained is trained using the unlabeled speech data and the annotation results to obtain a speech recognition model whose recognition results for speech data in the target domain meet preset evaluation conditions. This technical solution solves the problem of high training costs for speech recognition models by utilizing unlabeled speech data for training, significantly reducing training costs while maintaining the accuracy of the speech recognition model.

[0064] Example 3

[0065] Figure 4 This is a schematic diagram of the structure of a training device for a speech recognition model provided in Embodiment 3 of the present invention. Figure 4 As shown, the device includes:

[0066] The recognition result determination module 310 is used to acquire unlabeled speech data in the target domain, input the unlabeled speech data into at least two open-source speech recognition models respectively, and obtain recognition results that match each open-source speech recognition model;

[0067] The correction result determination module 320 is used to correct each recognition result according to a preset text correction principle to obtain the correction result of each open source speech recognition model; wherein, the text correction principle is determined based on text data in the target domain;

[0068] The annotation result determination module 330 is used to determine the annotation results of unlabeled speech data based on the correction results of matching various open-source speech recognition models;

[0069] The recognition model training module 340 is used to train the speech recognition model to be trained based on the unlabeled speech data and the annotation results, so as to obtain a speech recognition model whose recognition results for speech data in the target domain meet the preset evaluation conditions.

[0070] In this solution, optionally, the target domain is a vertical domain; the open-source speech recognition model's recognition results for speech data from non-vertical domains meet preset evaluation conditions.

[0071] In one feasible approach, the text correction principles include target word replacement and text error correction;

[0072] The correction result determination module 320 includes:

[0073] The word replacement unit is used to replace the target words in each recognition result according to the preset target word set, so as to obtain the replacement results matched by each open source speech recognition model.

[0074] The text correction unit is used to perform text correction on each replacement result based on the pre-trained correction model, and obtain the corrected results of the matching of each open source speech recognition model.

[0075] Based on the above scheme, optionally, the target word set is determined according to the word frequency statistics results in the text data of the target domain.

[0076] In this embodiment, optionally, the error correction model is trained using text data from the target domain and error correction labels for the text data.

[0077] In a preferred embodiment, the annotation result determination module 330 is specifically used for:

[0078] If the number of open-source speech recognition models with the same correction result meets the preset requirement, then the correction result will be used as the annotation result of the unlabeled speech data.

[0079] If the number of open-source speech recognition models with the same correction result does not meet the preset requirement, then the unlabeled speech data is determined to have no labeled result.

[0080] Based on the above scheme, optionally, the recognition model training module 340 is specifically used for:

[0081] Unlabeled speech data with annotation results is used as input to the speech recognition model to be trained. The speech recognition model is iteratively trained based on the output of the speech recognition model to be trained and the annotation results matched with the unlabeled speech data.

[0082] The speech recognition model training device provided in the embodiments of the present invention can execute the speech recognition model training method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0083] Example 4

[0084] Figure 5 A schematic diagram of an electronic device 410 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0085] like Figure 5 As shown, the electronic device 410 includes at least one processor 411 and a memory, such as a read-only memory (ROM) 412 or a random access memory (RAM) 413, communicatively connected to the at least one processor 411. The memory stores computer programs executable by the at least one processor. The processor 411 can perform various appropriate actions and processes based on the computer program stored in the ROM 412 or loaded from storage unit 418 into the RAM 413. The RAM 413 may also store various programs and data required for the operation of the electronic device 410. The processor 411, ROM 412, and RAM 413 are interconnected via a bus 414. An input / output (I / O) interface 415 is also connected to the bus 414.

[0086] Multiple components in electronic device 410 are connected to I / O interface 415, including: input unit 416, such as keyboard, mouse, etc.; output unit 417, such as various types of displays, speakers, etc.; storage unit 418, such as disk, optical disk, etc.; and communication unit 419, such as network card, modem, wireless transceiver, etc. Communication unit 419 allows electronic device 410 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0087] Processor 411 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 411 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 411 performs the various methods and processes described above, such as the training methods for a speech recognition model.

[0088] In some embodiments, the speech recognition model training method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 418. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 410 via ROM 412 and / or communication unit 419. When the computer program is loaded into RAM 413 and executed by processor 411, one or more steps of the speech recognition model training method described above may be performed. Alternatively, in other embodiments, processor 411 may be configured to execute the speech recognition model training method by any other suitable means (e.g., by means of firmware).

[0089] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0090] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0091] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0092] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0093] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0094] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0095] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0096] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for training a speech recognition model, the method comprising: The method includes: Obtain unlabeled speech data in the target domain, input the unlabeled speech data into at least two open-source speech recognition models respectively, and obtain recognition results that match each open-source speech recognition model; Each recognition result is compared sequentially with each target word in the target word set. All words involved in the target word in each recognition result are replaced with the target word to obtain the replacement results matched by each open-source speech recognition model. The target word is a professional term in the target domain text data with a word frequency greater than a preset threshold. The target word set includes at least one target word. The target word consists of the target word and the related words of the target word. The related words of the target word include synonyms and homophones of the target word. Based on the pre-trained error correction model, text error correction is performed on each replacement result to obtain the corrected results of matching each open-source speech recognition model. Based on the correction results of matching various open-source speech recognition models, the annotation results of unlabeled speech data are determined; Based on the unlabeled speech data and the annotation results, the speech recognition model to be trained is trained to obtain a speech recognition model whose recognition results for speech data in the target domain meet the preset evaluation conditions.

2. The method according to claim 1, characterized in that, The target domain is a vertical domain; the open-source speech recognition model meets the preset evaluation conditions for speech data from non-vertical domains.

3. The method according to claim 1, characterized in that, The error correction model is trained using text data from the target domain and error correction labels for that text data.

4. The method according to claim 1, characterized in that, The step of determining the annotation results of unlabeled speech data based on the correction results of matching various open-source speech recognition models includes: If the number of open-source speech recognition models with the same correction result meets the preset requirement, then the correction result will be used as the annotation result of the unlabeled speech data. If the number of open-source speech recognition models with the same correction result does not meet the preset requirement, then the unlabeled speech data is determined to have no labeled result.

5. The method according to claim 1, characterized in that, The step of training the speech recognition model to be trained based on the unlabeled speech data and the annotation results includes: Unlabeled speech data with annotation results is used as input to the speech recognition model to be trained. The speech recognition model is iteratively trained based on the output of the speech recognition model to be trained and the annotation results matched with the unlabeled speech data.

6. A training device for a speech recognition model, characterized in that, include: The recognition result determination module is used to acquire unlabeled speech data in the target domain, input the unlabeled speech data into at least two open-source speech recognition models respectively, and obtain recognition results that match each open-source speech recognition model; The correction result determination module is used to correct each recognition result according to a preset text correction principle to obtain the correction result of each open source speech recognition model; wherein, the text correction principle is determined based on text data in the target domain; The annotation result determination module is used to determine the annotation results of unlabeled speech data based on the correction results of matching various open-source speech recognition models; The recognition model training module is used to train the speech recognition model to be trained based on the unlabeled speech data and the annotation results, so as to obtain a speech recognition model whose recognition results for speech data in the target domain meet the preset evaluation conditions. The correction result determination module includes: The vocabulary replacement unit is used to compare each recognition result with each target word group in the target word group set in turn, and replace all words involved in the target word group in each recognition result with the target word to obtain the replacement result matched by each open source speech recognition model; the target word is a professional word in the target domain text data with a word frequency greater than a preset threshold, the target word group set includes at least one target word group, the target word group is composed of the target word and the related words of the target word, and the related words of the target word include synonyms and homophones of the target word; The text correction unit is used to perform text correction on each replacement result based on the pre-trained correction model, and obtain the corrected results of the matching of each open source speech recognition model.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the training method of the speech recognition model according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the training method of the speech recognition model according to any one of claims 1-5.