Method and apparatus for speech recognition
By performing incremental training and hot word decoding optimization in the speech recognition model, the problem of unsatisfactory new keyword recognition performance in the new scenario is solved, and more efficient training and better model adaptability are achieved.
Patent Information
- Application Number
- CN202110380567.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-09
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-04-09
AI Technical Summary
The existing speech recognition model has poor performance in identifying new keywords in new scenarios, and it is difficult to apply to new domains based on the recognition ability of the old scenario during cold start, which is prone to problems of training data imbalance and overfitting.
By obtaining the keywords to be identified, building the original scene training set and test set, performing incremental training, adjusting the model round by round, and combining with the hot word decoder to optimize the recognition results, the keyword recall rate is improved.
It effectively reduces the training time and data volume, solves the overfitting problem of incremental learning and the limitations of hot word decoding, and improves the recognition performance of new keywords and the applicability of the model in the new domain.
Smart Images

Figure CN115206296B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and particularly to methods and apparatuses for speech recognition. Background Art
[0002] End-to-end deep neural networks have become a popular framework in the field of speech recognition. Compared with traditional speech recognition frameworks, they can simplify the model construction and training processes. In practical applications, in many scenarios, it is required that existing speech recognition models can not only recognize speech inputs in new scenarios but also maintain the recognition accuracy in the original scenarios. For example, a speech recognition model trained on an original dataset needs to enhance its recognition ability for new keywords. Or, for the cold start of a new speech recognition scenario, while inheriting the recognition ability of the old speech recognition model, the model needs to be adapted to the new domain based on a small new dataset. Since new keywords or new scenarios are usually not in the past training datasets, when directly using the old model for a new task, the recognition performance will be very unsatisfactory. To solve this problem, a feasible method is to retrain the speech recognition model by mixing the old and new scenario datasets. However, this method may encounter the problem of unbalanced training data because the new dataset is usually much smaller than the old dataset. At the same time, due to considerations of data security and privacy, the past datasets may not be available for training. Another method is to use transfer learning with new scenario data. Although this method can reduce the time cost, it will cause the problem of overfitting of the speech recognition model. Using hotword decoding is also a feasible way, but hotword decoding can only operate on the decoding path when the keyword appears in the decoding path to achieve keyword recall. When the keyword does not exist in the decoding path or the keyword probability is low, the hotword decoding method cannot achieve keyword recall. Summary of the Invention
[0003] Embodiments of the present disclosure provide methods and apparatuses for speech recognition.
[0004] In a first aspect, embodiments of the present disclosure provide a method for speech recognition, including: obtaining a keyword to be recognized; searching for audio containing the keyword in an original scenario training set to form a first training set, and obtaining audio containing the keyword and not in the original scenario training set to form a first test set; performing a first round of incremental training on a first speech recognition model in the original scenario based on the first training set to obtain a second speech recognition model; using the second speech recognition model to recognize the first test set, adding the audio in which the keyword in the first test set is correctly recognized to the first training set to obtain a second training set and a second test set; performing a second round of incremental training on the second speech recognition model based on the second training set to obtain a third speech recognition model; and inputting the second test set into the third speech recognition model to obtain an initial recognition result.
[0005] In some embodiments, the method further includes: adjusting the initial recognition result through a hotword decoder and calculating the recall rate of the keyword.
[0006] In some embodiments, adding the audio in which the keyword in the first test set is correctly recognized to the first training set to obtain a second training set and a second test set includes: adjusting the recognition result obtained by the second speech recognition model for recognizing the first test set through a hotword decoder; determining the audio in which the keyword is correctly recognized according to the adjusted recognition result; adding the determined audio to the first training set to obtain a second training set; and deleting the determined audio from the first training set to obtain a second test set.
[0007] In some embodiments, the number of audio in the first training set is greater than a first threshold.
[0008] In some embodiments, the method further includes: if the number of audio containing the keyword in the original scenario training set is not greater than the first threshold, recording the audio containing the keyword and adding it to the first training set.
[0009] In some embodiments, the number of audio in the first test set is greater than a second threshold.
[0010] In some embodiments, the method further includes: if all the keywords in the first test set are correctly recognized or the number of audio in the second test set is less than a third threshold, recording the audio containing the keyword and adding it to the second test set.
[0011] In a second aspect, an embodiment of the present disclosure provides a speech recognition device, including: an acquisition unit configured to acquire a keyword to be recognized; a composition unit configured to find the audio containing the keyword from the original scenario training set to form a first training set and acquire the audio containing the keyword and not in the original scenario training set to form a first test set; a first training unit configured to perform a first round of incremental training on a first speech recognition model of the original scenario based on the first training set to obtain a second speech recognition model; a recognition unit configured to use the second speech recognition model to recognize the first test set, add the audio in which the keyword in the first test set is correctly recognized to the first training set to obtain a second training set and a second test set; a second training unit configured to perform a second round of incremental training on the second speech recognition model based on the second training set to obtain a third speech recognition model; and an output unit configured to input the second test set into the third speech recognition model to obtain an initial recognition result.
[0012] In some embodiments, the device further includes a calculation unit configured to: adjust the initial recognition result through a hotword decoder and calculate the recall rate of the keyword.
[0013] In some embodiments, the recognition unit is further configured to: adjust the recognition result obtained by the second speech recognition model for recognizing the first test set through a hotword decoder; determine the audio in which the keyword is correctly recognized according to the adjusted recognition result; add the determined audio to the first training set to obtain a second training set; and delete the determined audio from the first training set to obtain a second test set.
[0014] In some embodiments, the number of audios in the first training set is greater than a first threshold.
[0015] In some embodiments, the device further includes a first recording unit configured to: if the number of audios containing the keyword in the original scenario training set is not greater than the first threshold, record the audio containing the keyword and add it to the first training set.
[0016] In some embodiments, the number of audios in the first test set is greater than a second threshold.
[0017] In some embodiments, the device further includes a second recording unit configured to: if all the keywords in the first test set are correctly recognized or the number of audios in the second test set is less than a third threshold, record the audio containing the keyword and add it to the second test set.
[0018] In a third aspect, an embodiment of the present disclosure provides an electronic device for speech recognition, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of the first aspect.
[0019] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable medium storing a computer program thereon, wherein the program, when executed by a processor, implements the method according to any one of the first aspect.
[0020] The method and device for speech recognition provided by the embodiments of the present disclosure perform the first-round incremental training by using a small amount of training data with keywords. After completing the first-round training, the model obtained from the first-round training is used, and the data correctly recalled in the test set is added to the training set for the second-round incremental training. This can not only greatly reduce the training time and the required data volume, but also solve the overfitting problem of incremental learning and the limitations of hotword decoding. Description of the Drawings
[0021] Other features, objects, and advantages of the present disclosure will become more apparent from the following detailed description of non-limiting embodiments read in conjunction with the accompanying drawings:
[0022] Figure 1 is an exemplary system architecture diagram to which an embodiment of the present disclosure can be applied;
[0023] Figure 2 is a flowchart of an embodiment of a method for speech recognition according to the present disclosure;
[0024] Figure 3 is a schematic diagram of an application scenario of a method for speech recognition according to the present disclosure;
[0025] Figure 4 is a flowchart of another embodiment of a method for speech recognition according to the present disclosure;
[0026] Figure 5 is a hot word decoding flowchart of a method for speech recognition according to the present disclosure;
[0027] Figure 6 is a schematic structural diagram of an embodiment of a device for speech recognition according to the present disclosure;
[0028] Figure 7 is a schematic structural diagram of a computer system of an electronic device suitable for implementing an embodiment of the present disclosure. Detailed Embodiments
[0029] The present disclosure will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the relevant invention and not for limiting the invention. Additionally, it should be noted that for the convenience of description, only parts related to the relevant invention are shown in the drawings.
[0030] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The present disclosure will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0031] Figure 1 Illustrates an exemplary system architecture 100 of a method for speech recognition and a device for speech recognition to which embodiments of the present application can be applied.
[0032] As Figure 1 shown, the system architecture 100 may include terminals 101, 102, a network 103, a database server 104, and a server 105. The network 103 is used to provide a medium for communication links between the terminals 101, 102, the database server 104, and the server 105. The network 103 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0033] User 110 can use terminals 101 and 102 to interact with server 105 via network 103 to receive or send messages, etc. Various client applications can be installed on terminals 101 and 102, such as model training applications, speech recognition applications, shopping applications, payment applications, web browsers, and instant messaging tools, etc.
[0034] Terminals 101 and 102 here can be either hardware or software. When terminals 101 and 102 are hardware, they can be various electronic devices with microphones, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), laptop computers, and desktop computers, etc. When terminals 101 and 102 are software, they can be installed in the above-listed electronic devices. It can be implemented as multiple software or software modules (for example, to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.
[0035] When terminals 101 and 102 are hardware, an audio acquisition device can also be installed on them. The audio acquisition device can be various devices that can implement the function of acquiring audio, such as microphones, etc. User 110 can use the audio acquisition devices on terminals 101 and 102 to acquire the voice of himself or others.
[0036] Database server 104 can be a database server that provides various services. For example, a sample set can be stored in the database server. The sample set contains a large number of samples. Among them, the samples can include audio and the corresponding annotation information. In this way, user 110 can also select samples from the sample set stored in database server 104 through terminals 101 and 102.
[0037] Server 105 can also be a server that provides various services, such as a background server that provides support for various applications displayed on terminals 101 and 102. The background server can use the samples in the sample set sent by terminals 101 and 102 to train the initial model, and can send the training results (such as the generated speech recognition model) to terminals 101 and 102. In this way, users can apply the generated speech recognition model for speech recognition.
[0038] The database server 104 and the server 105 here can also be either hardware or software. When they are hardware, they can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When they are software, they can be implemented as multiple software or software modules (for example, used to provide distributed services), or as a single software or software module. No specific limitation is made here.
[0039] It should be noted that the method for speech recognition provided by the embodiments of the present application is generally executed by the server 105. Correspondingly, the device for speech recognition is generally also set in the server 105.
[0040] It should be pointed out that in the case where the server 105 can implement the related functions of the database server 104, the database server 104 may not be set in the system architecture 100.
[0041] It should be understood that Figure 1 the numbers of the terminals, networks, database servers, and servers in
[0042] Continue to refer to Figure 2 , which shows a process 200 of an embodiment of the method for speech recognition according to the present application. The method for speech recognition may include the following steps:
[0043] Step 201, obtain the keyword to be recognized.
[0044] In this embodiment, the execution subject of the method for speech recognition (such as Figure 1 the server 105 shown) can obtain the keyword to be recognized in various ways. For example, the execution subject can obtain the existing keyword to be recognized stored therein from the database server (such as Figure 1 the database server 104 shown) through a wired connection method or a wireless connection method. For another example, the user can collect the keyword to be recognized through the terminal (such as Figure 1 the terminals 101 and 102 shown). In this way, the execution subject can receive the keyword to be recognized collected by the terminal and store these keywords to be recognized locally.
[0045] When the keyword to be recognized is applied to a new scenario, the recognition effect of the speech recognition model trained in the original scenario on the keyword to be recognized is not good. The keyword is generally a professional term or a product name, such as "ETF", "XY loan", etc.
[0046] Step 202, search the original scenario training set for the audio containing the keyword to form a first training set, and obtain the audio containing the keyword and not in the original scenario training set to form a first test set.
[0047] In this embodiment, the present application performs incremental learning based on the first speech recognition model in the original scenario. Using the original speech recognition model, it is trained on a small batch of new scenario training sets to maximize the recognition performance in the original scenario while fitting the new scenario. Incremental learning means that the model learns new task or scenario knowledge using new data without accessing the original data, while not forgetting the task or scenario knowledge already learned. It can optimize the performance of the model in the new scenario and retain the accuracy of the model in the original scenario to the greatest extent.
[0048] The first speech recognition model in the original scenario is trained through the original scenario training set. The original scenario training set includes a large number of audio files, each corresponding to annotation information, and the annotation information is used to annotate the content of the audio. Some of the audio files in the original scenario training set include keywords, and these audio files including keywords can be selected to form the first training set. Then, collect audio files that contain keywords and are not in the original scenario training set to form the first test set, which is used to verify the performance of the trained model.
[0049] In some optional implementation manners of this embodiment, the number of audio files in the first training set is greater than a first threshold (for example, 30). If the training data is insufficient, the training effect is not ideal, and it is necessary to record audio files containing the keywords and add them to the first training set so that the total number of audio files containing the keywords is greater than the first threshold.
[0050] In some optional implementation manners of this embodiment, the number of audio files in the first test set is greater than a second threshold (for example, 15). The second threshold can be less than the first threshold. It is necessary to have a sufficient test set to verify the performance of the model, and also to ensure that after some data in the test set is added to the training set, there are still enough audio files in the test set to verify the performance of the model.
[0051] Step 203: Perform the first round of incremental training on the first speech recognition model in the original scenario based on the first training set to obtain a second speech recognition model.
[0052] In this embodiment, incremental training does not require a large amount of training data to be provided before the start of the training process, but continuously uses new training data for training over time. The audio files and annotation information in the first training set are used as the input and expected output of the first speech recognition model respectively, and the first speech recognition model is trained in a supervised manner to obtain a second speech recognition model. The specific training process is prior art and will not be elaborated here.
[0053] Step 204: Use the second speech recognition model to recognize the first test set, add the audio in which the keywords in the first test set are correctly recognized to the first training set, and obtain a second training set and a second test set.
[0054] In this embodiment, use the trained second speech recognition model to recognize the first test set. For the audio in which the keywords are correctly recognized, add it to the first training set to obtain a second training set, and delete the audio from the first test set to obtain a second test set. That is to say, after the first incremental training, the number of audio in the training set is increased, and the number of audio in the test set is decreased.
[0055] In some optional implementation manners of this embodiment, if all the keywords in the first test set are correctly recognized or the number of audio in the second test set is less than a third threshold, record the audio containing the keywords and add it to the second test set. The third threshold (for example, 5) can be less than the second threshold (for example, 15) and also less than the first threshold (for example, 30). If the keywords of all the audio in the test set are correctly recognized by the model, or the number of audio in the new test set is less than 5, record new audio containing the keywords to make the number of audio in the new test set not less than 5.
[0056] In some optional implementation manners of this embodiment, adjust the recognition result obtained by the second speech recognition model for the first test set through a hot word decoder; determine the audio in which the keywords are correctly recognized according to the adjusted recognition result; add the determined audio to the first training set to obtain a second training set; and delete the determined audio from the first training set to obtain a second test set. During the model recognition process, this method uses a hot word decoder as shown in Figure 5 The hot word decoder matches the end of the decoding path at each time step during the speech recognition decoding process. If the end of the decoding path is the specified hot word, a corresponding score is added to this path. As shown in Figure 5 , the specified decoding hot word is "ETF". At a certain time step, "ETF" appears at the end of a certain decoding path. After being matched by the hot word decoder, the score is improved. When the decoding search is completed, the decoder outputs the decoding path with the highest score as the recognition result, which is the path with the keyword that obtains the hot word reward. In this way, through the adjustment of the acoustic model and the adjustment of the hot word, this method effectively improves the recall rate of the keywords.
[0057] Step 205: Perform a second round of incremental training on the second speech recognition model based on the second training set to obtain a third speech recognition model.
[0058] In this embodiment, the audio and annotation information in the second training set are respectively used as the input and expected output of the second speech recognition model, and the second speech recognition model is supervised to train to obtain the third speech recognition model. The recognition effect of the third speech recognition model on keywords has been significantly improved.
[0059] In the method of speech recognition in this embodiment, in the case of fewer training samples, an incremental learning method is used to adjust the speech recognition model, and a method of combining hot words is used to improve the recall rate of keywords; at the same time, compared with other methods, this method can effectively retain the recognition performance of the new model in the old scenario and solve the problem that hot word decoding is limited by the acoustic probability distribution.
[0060] Step 206, input the second test set into the third speech recognition model to obtain an initial recognition result.
[0061] In this embodiment, the second test set is used to verify the performance of the third speech recognition model to obtain an initial recognition result, that is, the accuracy of keyword recognition. For further reference Figure 3 , Figure 3 is a schematic diagram of an application scenario of the speech recognition method according to this embodiment. The specific process is as follows:
[0062] 1) The user can input the keyword "ETF" to be recognized to the server through the terminal device.
[0063] 2) For the specified keyword, find the training data containing the keyword in the original scenario training set to form the training set for the first incremental training. If the number of audio containing the keyword in the training set is less than 30, it is necessary to record the audio containing the keyword so that the number of audio in the training set is not less than 30. At the same time, collect no less than 15 pieces of data containing the keyword and not in the original scenario training set as the test set.
[0064] 3) Use the speech recognition model in the original scenario for incremental training, and use the training set in 1) for the first round of incremental training.
[0065] 4) After the training is completed, use the new model obtained in 3) to recognize the test set. For the audio whose keyword is correctly recognized, add it to the training set in 2); for the audio that is not correctly recognized, add it to the new test set. If the keywords of all the audio in the test set are correctly recognized by the model, or the number of audio in the new test set is less than 5, record the new audio containing the keyword so that the number of audio in the new test set is not less than 5.
[0066] In the process of model recognition, this method uses such as Figure 4The hotword decoder shown. During the speech recognition decoding process, the hotword decoder matches the end of the decoding path at each time step. If the end of the decoding path is the specified hotword, a corresponding score boost is given to that path. As Figure 4 shown in Figure 4 , the specified decoding hotword is "ETF". At a certain time step, "ETF" appears at the end of a certain decoding path. After being matched by the hotword decoder, the score is improved. When the decoding search is completed, the decoder outputs the decoding path with the highest score as the recognition result, which is the path with the keyword that obtains the hotword reward. In this way, through the adjustment of the speech recognition model (especially the acoustic model) and the adjustment of the hotword, this method effectively improves the recall rate of the keyword.
[0067] 5) Use the training set in 4) to perform a second round of incremental training on the model obtained in 3).
[0068] 6) After the training is completed, the model in 5) is obtained, which is the speech recognition model.
[0069] Continue to refer to Figure 4 , as an implementation of the methods shown in the above figures, the present application provides a flowchart of another embodiment of a method for speech recognition.
[0070] As Figure 4 shown, the method 400 for speech recognition in this embodiment may include:
[0071] Step 401, obtain the keyword to be recognized.
[0072] Step 402, search in the original scenario training set for the audio containing the keyword to form a first training set, and obtain the audio containing the keyword and not in the original scenario training set to form a first test set.
[0073] Step 403, perform a first round of incremental training on the first speech recognition model of the original scenario based on the first training set to obtain a second speech recognition model.
[0074] Step 404, use the second speech recognition model to recognize the first test set, and add the audio in the first test set where the keyword is correctly recognized to the first training set to obtain a second training set and a second test set.
[0075] Step 405, perform a second round of incremental training on the second speech recognition model based on the second training set to obtain a third speech recognition model.
[0076] Steps 401 - 405 are basically the same as steps 201 - 205, so they will not be elaborated here.
[0077] Step 406, input the second test set into the third speech recognition model to obtain an initial recognition result.
[0078] In this embodiment, the performance of the third speech recognition model is verified using the second test set to obtain an initial recognition result, that is, the accuracy of keyword recognition.
[0079] Step 407, adjust the initial recognition result through a hotword decoder and calculate the recall rate of the keyword.
[0080] In this embodiment, a hotword decoder as Figure 5 shown is used. During the speech recognition decoding process, the hotword decoder matches the end of the decoding path at each time step. If the end of the decoding path is the specified hotword, a corresponding score increase is given to this path. As Figure 5 shown, the specified decoding hotword is "ETF". At a certain time step, "ETF" appears at the end of a certain decoding path. After being matched by the hotword decoder, the score is improved. When the decoding search is completed, the decoder outputs the decoding path with the highest score as the recognition result, which is the path with the keyword that obtains the hotword reward. In this way, through the adjustment of the acoustic model and the hotword, the recall rate of the keyword is effectively improved.
[0081] Optionally, if the recall rate does not reach the expected threshold, steps 404 - 407 can be repeatedly executed. Continue to select the audio that accurately recognizes the keyword from the test set and add it to the training set, and retrain the model with the updated training set until the recall rate of the trained model reaches the expected threshold.
[0082] As can be seen from Figure 4 compared with the corresponding embodiment, the process 400 of the speech recognition method in this embodiment reflects the step of adjusting the recognition result of the speech recognition model through a hotword decoder. Thus, the solution described in this embodiment can improve the recall rate of the keyword. Figure 2
[0083] Continuing to refer to Figure 6 , as an implementation of the methods shown in the above figures, an embodiment of a speech recognition device is provided in this application. This device embodiment corresponds to the method embodiment shown in Figure 2 and this device can be specifically applied to various electronic devices.
[0084] As Figure 6 As shown in the figure, the voice recognition device 600 in this embodiment may include: an acquisition unit 601, a composition unit 602, a first training unit 603, an identification unit 604, a second training unit 605, and an output unit 606. Among them, the acquisition unit 601 is configured to acquire keywords to be recognized; the composition unit 602 is configured to search the original scenario training set for audio containing the keywords to form a first training set, and acquire audio containing the keywords and not in the original scenario training set to form a first test set; the first training unit 603 is configured to perform the first round of incremental training on the first voice recognition model of the original scenario based on the first training set to obtain a second voice recognition model; the identification unit 604 is configured to use the second voice recognition model to identify the first test set, and add the audio in the first test set where the keywords are correctly recognized to the first training set to obtain a second training set and a second test set; the second training unit 605 is configured to perform the second round of incremental training on the second voice recognition model based on the second training set to obtain a third voice recognition model; the output unit 606 is configured to input the second test set into the third voice recognition model to obtain an initial recognition result.
[0085] In this embodiment, the specific processing of the acquisition unit 601, the composition unit 602, the first training unit 603, the identification unit 604, the second training unit 605, and the output unit 606 of the voice recognition device 600 may refer to Figure 2 Steps 201, 202, 203, 204, 205, and 206 in the corresponding embodiment.
[0086] In some optional implementation manners of this embodiment, the device 600 further includes a calculation unit (not shown in the figure), which is configured to: adjust the initial recognition result through a hotword decoder and calculate the recall rate of the keywords.
[0087] In some optional implementation manners of this embodiment, the identification unit 604 is further configured to: adjust the recognition result obtained by the second voice recognition model for identifying the first test set through a hotword decoder; determine the audio in which the keywords are correctly recognized according to the adjusted recognition result; add the determined audio to the first training set to obtain a second training set; and delete the determined audio from the first training set to obtain a second test set.
[0088] In some optional implementation manners of this embodiment, the number of audio in the first training set is greater than a first threshold.
[0089] In some alternative implementation manners of this embodiment, the apparatus 600 further includes a first recording unit (not shown in the drawings), which is configured to: if the number of audios containing the keyword in the original scenario training set is not greater than a first threshold, record the audio containing the keyword and add it to the first training set.
[0090] In some alternative implementation manners of this embodiment, the number of audios in the first test set is greater than a second threshold.
[0091] In some alternative implementation manners of this embodiment, the apparatus 600 further includes a second recording unit (not shown in the drawings), which is configured to: if all the keywords in the first test set are correctly recognized or the number of audios in the second test set is less than a third threshold, record the audio containing the keyword and add it to the second test set.
[0092] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device and a readable storage medium.
[0093] Figure 7 FIG. shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processing, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0094] As Figure 7 shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0095] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as a keyboard, mouse, etc.; output unit 707, such as various types of displays, speakers, etc.; storage unit 708, such as a disk, optical disc, etc.; and communication unit 709, such as a network card, modem, wireless communication transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0096] Computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 701 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 701 executes the various methods and processes described above, such as the method of speech recognition. For example, in some embodiments, the method of speech recognition can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by computing unit 701, one or more steps of the method of speech recognition described above can be executed. Alternatively, in other embodiments, computing unit 701 can be configured to execute the method of speech recognition in any other suitable way (e.g., by means of firmware).
[0097] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0098] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0099] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0100] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, speech input, or tactile input).
[0101] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0102] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a server of a distributed system, or a server incorporating a blockchain. The server can also be a cloud server, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology. The server can be a server of a distributed system, or a server incorporating a blockchain. The server can also be a cloud server, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology.
[0103] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0104] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for speech recognition, comprising: obtaining a keyword to be recognized; searching in the original scenario training set for audios containing the keyword to form a first training set, and obtaining audios containing the keyword and not in the original scenario training set to form a first test set; performing a first round of incremental training on the first speech recognition model of the original scenario based on the first training set to obtain a second speech recognition model; using the second speech recognition model to recognize the first test set, and adding the audios in the first test set whose keywords are correctly recognized to the first training set to obtain a second training set and a second test set; performing a second round of incremental training on the second speech recognition model based on the second training set to obtain a third speech recognition model; inputting the second test set into the third speech recognition model to obtain an initial recognition result; adjusting the initial recognition result through a hotword decoder, and calculating the recall rate of the keyword, wherein, during the speech recognition decoding process, the hotword decoder matches the end of the decoding path at each time step. If the end of the decoding path is the specified hotword, a corresponding score is given to this path, and the decoding path with the highest score is output as the recognition result.
2. The method according to claim 1, wherein, the step of adding the audios in the first test set whose keywords are correctly recognized to the first training set to obtain a second training set and a second test set includes: adjusting the recognition result obtained by the second speech recognition model for the first test set through a hotword decoder; determining the audios whose keywords are correctly recognized according to the adjusted recognition result; adding the determined audios to the first training set to obtain a second training set; deleting the determined audios from the first training set to obtain a second test set.
3. The method according to claim 1, wherein, the number of audios in the first training set is greater than a first threshold.
4. The method according to claim 3, wherein, the method further includes: if the number of audios containing the keyword in the original scenario training set is not greater than the first threshold, recording audios containing the keyword and adding them to the first training set.
5. The method according to claim 1, wherein, the number of audios in the first test set is greater than a second threshold.
6. The method according to claim 1, wherein, the method further includes: if all the keywords in the first test set are correctly recognized or the number of audios in the second test set is less than a third threshold, recording audios containing the keyword and adding them to the second test set.
7. A speech recognition device, comprising: an obtaining unit configured to obtain a keyword to be recognized; a composing unit configured to search in the original scenario training set for audios containing the keyword to form a first training set, and obtain audios containing the keyword and not in the original scenario training set to form a first test set; a first training unit configured to perform a first round of incremental training on the first speech recognition model of the original scenario based on the first training set to obtain a second speech recognition model; An identification unit, configured to use the second speech recognition model to identify the first test set, add the audio in which the keywords in the first test set are correctly recognized to the first training set, and obtain a second training set and a second test set; A second training unit, configured to perform a second round of incremental training on the second speech recognition model based on the second training set, and obtain a third speech recognition model; An output unit, configured to input the second test set into the third speech recognition model to obtain an initial recognition result; A calculation unit, configured to adjust the initial recognition result through a hotword decoder and calculate the recall rate of the keywords. In the speech recognition decoding process, the hotword decoder matches the end of the decoding path at each time step. If the end of the decoding path is the specified hotword, a corresponding score is given to this path, and the decoding path with the highest score is output as the recognition result.
8. An electronic device for speech recognition, comprising: One or more processors; A storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.
9. A computer-readable medium, on which a computer program is stored, wherein, When the program is executed by a processor, it implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Keyword sample determining method, speech recognition method, devices, equipment and medium
CN109979440A
Method for constructing speech recognition model in specific field
CN111627427A