Training Method of Speech Recognition Model, Speech Recognition Method and Device

By inserting noise information in speech recognition model training and determining noise identification, the noise word issue problem of end-to-end speech recognition system in a noisy environment is solved, and the robustness and user experience of the speech recognition model are improved.

CN115831103BActive Publication Date: 2025-07-11ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211678808.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2025-07-11
Estimated Expiration
2042-12-26

AI Technical Summary

Technical Problem

Existing end-to-end speech recognition systems are prone to noise-producing problems in noisy environments, affecting the user experience.

Method used

By obtaining reference speech and noise information, the noise information is inserted into the reference speech, the processed speech including speech segments and noise segments is generated, the preset noise identification of the noise segment is determined, and the model is trained based on the reference speech, the processed speech and the preset noise identification are obtained to obtain the speech recognition model.

Benefits of technology

It significantly improves the robustness of the speech recognition model to noise, effectively reduces and even avoids the situation of noise recognition words, and improves the user experience and practicality of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115831103B_ABST
    Figure CN115831103B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a method for training a speech recognition model, a speech recognition method, and a device. The method for training a speech recognition model includes: obtaining a reference speech and noise information, where the reference speech includes speech information that can be recognized as a reference text; inserting the noise information into the reference speech to generate a processed speech including speech segments and noise segments; determining a preset noise identifier corresponding to the noise segments in the processed speech; and performing model training based on the reference speech, the processed speech, and the preset noise identifier to obtain a speech recognition model, where the speech recognition model is used to recognize an input speech as text. The technical solution provided in this embodiment obtains a speech recognition model through noise display modeling operations, thereby improving the robustness of the speech recognition model to noise or background noise. During the speech recognition operation, the situation of recognizing words from noise is effectively reduced or even avoided, ensuring a good experience for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and in particular, to a method for training a speech recognition model, a speech recognition method, and a device therefor. Background Art

[0002] With the rapid development of science and technology, the application scenarios of speech recognition technology are increasing, such as: multi-person far and near field recognition in offline meetings, recognition in noisy office environments, online meetings with different pick-up devices, etc. These scenarios pose relatively high requirements on the robustness of the speech recognition system.

[0003] Currently, since the end-to-end speech recognition system is sensitive to noise, during the speech recognition process, even if the speech has been filtered by voice activity detection (VAD), data simulation, acoustic model optimization, etc., some words will still be decoded from pure noise segments from time to time. Although it has little impact on the overall character error rate (CER) of the speech recognition system, the above situation seriously affects the user experience. Therefore, during the end-to-end speech recognition operation, the problem of words being decoded from noise needs to be solved urgently. Summary of the Invention

[0004] Embodiments of the present invention provide a method for training a speech recognition model, a speech recognition method, and a device therefor, which effectively reduce or even avoid the situation of words being recognized from noise during speech recognition, thereby ensuring a good user experience.

[0005] In a first aspect, an embodiment of the present invention provides a method for training a speech recognition model, including:

[0006] Obtaining a reference speech and noise information, where the reference speech includes speech information that can be recognized as a reference text;

[0007] Inserting the noise information into the reference speech to generate a processed speech including speech segments and noise segments;

[0008] Determining a preset noise identifier corresponding to the noise segment in the processed speech;

[0009] Training a model based on the reference speech, the processed speech, and the preset noise identifier to obtain a speech recognition model, where the speech recognition model is used to recognize an input speech as text.

[0010] In a second aspect, an embodiment of the present invention provides a training device for a speech recognition model, including:

[0011] A first acquisition module, configured to acquire a reference voice and noise information, where the reference voice includes voice information that can be recognized as a reference text;

[0012] A first generation module, configured to insert the noise information into the reference voice to generate a processed voice including voice segments and noise segments;

[0013] A first processing module, configured to determine a preset noise identifier corresponding to the noise segment in the processed voice;

[0014] A first training module, configured to perform model training based on the reference voice, the processed voice, and the preset noise identifier to obtain a speech recognition model, where the speech recognition model is used to recognize an input voice as text.

[0015] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor; wherein, the memory is used to store one or more computer instructions, and when the one or more computer instructions are executed by the processor, the training method of the speech recognition model in the first aspect above is implemented.

[0016] In a fourth aspect, an embodiment of the present invention provides a computer storage medium, configured to store a computer program, and when the computer program is executed by a computer, the training method of the speech recognition model in the first aspect above is implemented.

[0017] In a fifth aspect, an embodiment of the present invention provides a computer program product, including: a computer program, when the computer program is executed by a processor of an electronic device, enabling the processor to execute the steps in the training method of the speech recognition model in the first aspect above.

[0018] In a sixth aspect, an embodiment of the present invention provides a speech recognition method, including:

[0019] Acquiring an audio to be recognized;

[0020] Inputting the audio to be recognized into the speech recognition model to recognize voice segments and noise segments included in the audio to be recognized, and transcribing the voice segments to obtain text information corresponding to the audio to be recognized;

[0021] Wherein, the speech recognition model is obtained through learning and training with a reference voice, a processed voice, and a preset noise identifier, the processed voice is obtained by inserting noise information into the reference voice, the processed voice includes voice segments and noise segments, and the preset noise identifier corresponds to the noise segment.

[0022] In a seventh aspect, an embodiment of the present invention provides a speech recognition device, including:

[0023] A second acquisition module, configured to acquire the audio to be recognized;

[0024] A second processing module, configured to input the audio to be recognized into the speech recognition model to recognize the speech segment and noise segment included in the audio to be recognized, and transcribe the speech segment to obtain text information corresponding to the audio to be recognized;

[0025] Wherein, the speech recognition model is obtained through learning and training with reference to speech, processed speech, and a preset noise identifier. The processed speech is obtained by inserting noise information into the reference speech. The processed speech includes a speech segment and a noise segment, and the preset noise identifier corresponds to the noise segment.

[0026] In a tenth aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor; wherein, the memory is configured to store one or more computer instructions, and when the one or more computer instructions are executed by the processor, the speech recognition method in the above sixth aspect is implemented.

[0027] In an eleventh aspect, an embodiment of the present invention provides a computer storage medium for storing a computer program, and when the computer program is executed by a computer, the speech recognition method in the above sixth aspect is implemented.

[0028] In a twelfth aspect, an embodiment of the present invention provides a computer program product, including: a computer program, when the computer program is executed by a processor of an electronic device, the processor is caused to execute the steps in the speech recognition method in the above sixth aspect.

[0029] The training method, speech recognition method and device for the speech recognition model provided in this embodiment, by acquiring reference speech and noise information, inserting the noise information into the reference speech to generate processed speech including a speech segment and a noise segment; then determining a preset noise identifier corresponding to the noise segment in the processed speech, and performing model training based on the reference speech, processed speech, and the preset noise identifier, realizes obtaining a speech recognition model through explicit modeling of noise, which can significantly improve the robustness of the speech recognition model to noise or background noise. The above speech recognition model can be implemented as an end-to-end speech recognition model, and then speech recognition operations can be performed based on the speech recognition model, effectively reducing or even avoiding the situation of recognizing words from noise, thereby ensuring a good experience for users, effectively improving the practicability of this method, and being conducive to market promotion and application. Description of the Drawings

[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0031] Figure 1 Scenario schematic diagram of a method for training a speech recognition model provided by an embodiment of the present invention;

[0032] Figure 2 Flow schematic diagram of a method for training a speech recognition model provided by an embodiment of the present invention;

[0033] Figure 3 Schematic of the processed speech provided by an embodiment of the present invention Figure 1 ;

[0034] Figure 4 Schematic of the processed speech provided by an embodiment of the present invention Figure 2 ;

[0035] Figure 5 Flow schematic diagram of another method for training a speech recognition model provided by an embodiment of the present invention;

[0036] Figure 6 Flow schematic diagram of a speech recognition method provided by an embodiment of the present invention;

[0037] Figure 7 Structural schematic diagram of a device for training a speech recognition model provided by an embodiment of the present invention;

[0038] Figure 8 For Figure 7 Structural schematic diagram of an electronic device corresponding to the device for training a speech recognition model shown in the embodiment;

[0039] Figure 9 Structural schematic diagram of a speech recognition device provided by an embodiment of the present invention;

[0040] Figure 10 For Figure 9 Structural schematic diagram of an electronic device corresponding to the speech recognition device shown in the embodiment. Detailed implementation manners

[0041] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0042] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. "Plural" generally includes at least two, but does not exclude the case of including at least one.

[0043] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0044] Depending on the context, the words "if" and "when" as used herein may be interpreted as "when" or "when...", or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".

[0045] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a commodity or system including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such commodity or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the commodity or system including the said element.

[0046] In addition, the sequence of steps in the following method embodiments is only an example and not strictly limited.

[0047] To be able to understand the specific implementation process of the technical solution in this embodiment, the following first explains the related technologies:

[0048] With the rapid development of science and technology, the application scenarios of speech recognition technology are becoming more and more numerous. For example: multi-person near and far-field recognition in offline meetings, recognition in noisy office environments, online meetings with different pick-up devices, etc. These scenarios pose higher requirements for the robustness of the speech recognition system.

[0049] Currently, due to the fact that end-to-end speech recognition systems are relatively sensitive to noise, during the speech recognition process, even if the speech has been filtered by the Voice Activity Detection (VAD) principle, data simulation, acoustic model optimization, etc., there are problems such as autoregressive missing words, noise-induced words, and low-quality sensitivity in end-to-end speech recognition technology. For example, in real-time meeting scenarios, for pure noise segments, some words such as "um", "yes" will sometimes pop up when being displayed on the screen. Although it has little impact on the overall Character Error Rate (CER) of the speech recognition system, such problems have affected the user experience to a certain extent in actual applications.

[0050] To solve the above technical problems, for the speech recognition model, robust speech recognition construction operations can be carried out. Among them, noise-robust speech recognition is a popular research topic in speech recognition and one of the core challenges faced in the actual implementation of speech recognition. Existing industrial-grade speech recognition systems usually enhance the robustness of the model through a data-driven approach. Currently, the ways to enhance the robustness of the model can include two types: (1) annotate the data in many real scenarios and then perform modeling operations through the annotated data; (2) generate noisy training data through signal simulation and then perform modeling operations based on the noisy training data.

[0051] However, the above implementation method (1) has the following problems: not only is the data collection difficult, but the manual annotation cost is also high. Related technologies often use implementation method (2) for modeling operations. However, there is often a mismatch between the simulated data and the real scenario. Therefore, although the noise robustness of the system can be improved to a certain extent, it is still relatively limited.

[0052] To solve the above technical problems, this embodiment provides a method for training a speech recognition model, a speech recognition method, and a device. The execution subject of the method for training a speech recognition model can be a training device for a speech recognition model. The training device for a speech recognition model can be implemented as a local server or a server in the cloud. At this time, the method for training a speech recognition model can be executed in the cloud. In the cloud, several computing nodes (cloud servers) can be deployed, and each computing node has processing resources such as computing and storage. In the cloud, a service can be organized by multiple computing nodes. Of course, a single computing node can also provide one or more services. The way the cloud provides the service can be to provide a service interface externally, and users can call the service interface to use the corresponding service. The service interface includes forms such as a Software Development Kit (SDK) and an Application Programming Interface (API).

[0053] Specifically, as shown in the attached Figure 1 figure, the training device for the speech recognition model can be communicatively connected to a client or a requester. For the solution provided in this embodiment of the present invention, the cloud can provide a service interface for the training service of the speech recognition model. Users can trigger a request to call the training service interface of the speech recognition model from the cloud by calling the training service interface of the speech recognition model through the client / requester. The cloud determines the computing node that responds to the request and uses the processing resources in the computing node to execute the specific processing operations for training the speech recognition model.

[0054] The client / requester can be any computing device with a certain data transmission ability. Specifically, in implementation, the client / requester can be a mobile phone, a personal computer (PC), a tablet computer, a set application program, etc. In addition, the basic structure of the client / requester can include: at least one processor. The number of processors depends on the configuration and type of the client / requester. The client / requester can also include a memory, which can be volatile, such as RAM, or non-volatile, such as Read-Only Memory (ROM), flash memory, etc., or can also include both types at the same time. The memory usually stores an operating system (OS), one or more application programs, and can also store program data, etc. In addition to the processing unit and the memory, the client / requester also includes some basic configurations, such as a network card chip, an IO bus, a display component, and some peripheral devices. Optionally, some peripheral devices can include, for example, a keyboard, a mouse, a stylus, a printer, etc. Other peripheral devices are well known in the art and will not be elaborated here.

[0055] A training device for a speech recognition model refers to a device that can provide training services for a speech recognition model in a network virtual environment, usually a device that uses the network for information planning and training operations of the speech recognition model. Physically, the training device for the speech recognition model can be any device that can provide computing services, respond to service requests, and perform processing. For example, it can be a cluster server, a conventional server, a cloud server, a cloud host, a virtual center, etc. The composition of the training device for the speech recognition model mainly includes a processor, a hard disk, a memory, a system bus, etc., which is similar to a general computer architecture.

[0056] In the above-mentioned embodiment of the present application, the client can be network-connected to the training device for the speech recognition model, and this network connection can be a wireless or wired network connection. If the client is communicatively connected to the training device for the speech recognition model, the network mode of the mobile network can be any one of 2G (GSM), 2.5G (GPRS), 3G (WCDMA, TD-SCDMA, CDMA2000, UTMS), 4G (LTE), 4G+ (LTE+), WiMax, 5G, 6G, etc.

[0057] In the embodiment of the present application, the client can obtain reference speech and noise information. Among them, the reference speech includes speech information that can be recognized as a reference text. The above-mentioned "can be recognized" means that the speech information can be accurately speech-recognized and transcribed into a reference text. Specifically, the speech information can be speech from different countries and regions, and the text written can also be text information corresponding to the speech information. For example: Chinese text, English text, German text, etc. In addition, the reference speech and noise information can be obtained through a speech acquisition device. For the noise information, it can be an interference signal, specifically including environmental interference sounds and non-environmental interference sounds. The above-mentioned environmental interference sounds can include event-related interference sounds, and the event-related interference sounds can include at least one of the following: ringing, coughing, knocking, telephone beeping, etc. The above-mentioned non-environmental interference sounds can refer to human voices that cannot be normally recognized, which can include simulated human voice noise obtained by performing preset processing operations on actual human voice data, etc. After obtaining the reference speech and noise information, in order to enable the training operation of the speech recognition model, the reference speech and noise information can be sent to the training device for the speech recognition model as training data for the model.

[0058] A training device for a speech recognition model is configured to receive and obtain a reference speech and noise information sent by a client, and then analyze and process the noise information and the reference speech. Specifically, the noise information can be inserted into the reference speech to generate a processed speech including speech segments and noise segments. To enable explicit modeling of the noise, a preset noise identifier corresponding to the noise segments in the processed speech can be determined, and then a display model can be established based on the reference speech, the processed speech, and the preset noise identifier, thereby obtaining a speech recognition model for speech recognition operations.

[0059] The technical solution provided in this embodiment inserts the noise information into the reference speech by obtaining the reference speech and the noise information to generate a processed speech including speech segments and noise segments. Then, a preset noise identifier corresponding to the noise segments in the processed speech is determined, and model training is performed based on the reference speech, the processed speech, and the preset noise identifier, achieving the acquisition of a speech recognition model through explicit modeling of the noise. This can significantly improve the robustness of the speech recognition model to noise or background noise. The above speech recognition model can be implemented as an end-to-end speech recognition model, and then speech recognition operations can be performed based on the speech recognition model, effectively reducing or even avoiding the situation of incorrect character recognition due to noise, thereby ensuring a good user experience, effectively improving the practicability of this method, and facilitating market promotion and application.

[0060] The following describes in detail some embodiments of the present invention with reference to the accompanying drawings. Without conflict between the embodiments, the features in the following embodiments and the embodiments can be combined with each other. Additionally, the step sequence in the following method embodiments is only an example and is not strictly limited.

[0061] Figure 2 It is a schematic flowchart of a method for training a speech recognition model provided by an embodiment of the present invention; refer to the attached Figure 2 As shown, this embodiment provides a method for training a speech recognition model. The execution subject of this method can be a training device for a speech recognition model. It can be understood that the training device for a speech recognition model can be implemented as software, or a combination of software and hardware. Specifically, when the training device for a speech recognition model is implemented as hardware, it can be various electronic devices with speech recognition model training operations, including but not limited to tablet computers, personal computers (PCs), servers, etc. When the training device for a speech recognition model is implemented as software, it can be installed in the above-mentioned exemplified electronic devices. Based on the above training device for a speech recognition model, the method for training a speech recognition model in this embodiment can include the following steps:

[0062] Step S201: Obtain a reference speech and noise information, where the reference speech includes speech information that can be recognized as a reference text.

[0063] Step S202: Insert the noise information into the reference speech to generate a processed speech including speech segments and noise segments.

[0064] Step S203: Determine a preset noise identifier corresponding to the noise segment in the processed speech.

[0065] Step S204: Perform model training based on the reference speech, the processed speech, and the preset noise identifier to obtain a speech recognition model, which is used to recognize the input speech as text.

[0066] The specific implementation principles and implementation effects of the above steps will be described in detail below:

[0067] Step S201: Obtain a reference speech and noise information, where the reference speech includes speech information that can be recognized as a reference text.

[0068] Since the speech recognition model is used to recognize the input speech as text, therefore, in order to implement the training operation of the speech recognition model, the reference speech and noise information can be obtained, where the reference speech can include speech information that can be recognized as a reference text. The above speech information can refer to speech information in any country and any region, for example: Chinese speech information, English speech information, Korean speech information, Russian speech information, French speech information, Cantonese speech information, etc. Correspondingly, the reference text can also refer to text in any country and any region, for example: Chinese text, English text, Korean text, Russian text, French text, Cantonese text, etc.

[0069] In addition, for the reference speech, it can not only include speech information that can be recognized as a reference text, but also include other interfering sounds, such as: environmental noise, signal noise, etc. Specifically, the acquisition method of the reference speech in this embodiment is not limited. In some instances, the reference speech can be obtained by collecting through a speech collector. At this time, the training device of the speech recognition model is communicatively connected to the speech collector, and the reference speech in a preset scenario (office scenario, meeting scenario, speech scenario, etc.) and a preset environment (office environment, meeting room environment, concert hall scenario) can be obtained through the speech collector. In other instances, the reference speech can not only be obtained by collecting through a speech collector, but also be an automatically synthesized speech. At this time, a preset application program for synthesizing the reference speech can be configured in the training device of the speech recognition model, and the reference speech can be obtained through the preset application program.

[0070] In an actual scenario, different types of noise information often accompany a user's speech. At this time, in order to improve the training quality and effect of a speech recognition model, noise information can be obtained simultaneously with or after obtaining the reference speech. The noise information can include at least one of the following: environmental interference sounds and non-environmental interference sounds. The above-mentioned environmental interference sounds can include event-related interference sounds. Specifically, the event-related interference sounds can include at least one of the following: ringtones, coughing sounds, knocking sounds, door-opening sounds, door-closing sounds, telephone beeping sounds, etc.; environmental noise can not only include event-related interference sounds, but also include specific noises in a preset scenario, such as wind noise in an outdoor scenario, subway running sounds when a user takes the subway, etc.; non-environmental interference sounds refer to speech information that cannot be normally recognized as text, such as the noisy sounds of a user in a cafeteria, the noisy sounds in a large conference, etc. In some instances, non-environmental interference sounds can include simulated human voice noise.

[0071] In addition, the specific method for obtaining the noise information in this embodiment is not limited. In some instances, the noise information can be obtained by collecting through a voice collector. At this time, the training device of the speech recognition model is communicatively connected to the voice collector, and the noise information in a preset scenario (such as an office scenario, a meeting scenario, a speech scenario, etc.) and a preset environment (such as an office environment, a meeting room environment, a concert hall scenario) can be obtained through the voice collector. In other instances, the noise information can not only be obtained by collecting through a voice collector, but also be automatically synthesized noise information. At this time, a preset application program for synthesizing noise information can be configured in the training device of the speech recognition model, and the noise information can be obtained through the preset application program. In still other instances, the noise information can also be obtained by analyzing and processing the reference speech. Specifically, when the non-environmental interference sounds include simulated human voice noise, obtaining the noise information can include: obtaining the human voice data included in the reference speech; performing at least one of the processing operations of low-volume amplitude modulation, adding noise, and distortion on the human voice data to obtain the simulated human voice noise corresponding to the reference speech.

[0072] Since the simulated human voice noise is obtained by analyzing and processing normal human voices, in order to accurately obtain the simulated human voice noise, the human voice data included in the reference voice can be obtained first. Specifically, an information extraction operation can be performed on the reference voice to obtain the human voice data, and then at least one of the processing operations of low-volume amplitude modulation, noise addition, and distortion can be performed on the human voice data, so as to obtain the simulated human voice noise. Specifically, one implementation method: the human voice data can be subjected to low-volume amplitude modulation or noise addition or distortion operation to obtain the simulated human voice noise; another implementation method: the human voice data can be subjected to low-volume amplitude modulation and noise addition processing, or low-volume amplitude modulation and distortion processing, or noise addition and distortion processing operations to obtain the simulated human voice noise; yet another implementation method: the human voice data can be subjected to low-volume amplitude modulation, noise addition, and distortion processing to obtain the simulated human voice noise. It can be understood that the simulated human voice noise obtained by different processing methods is different.

[0073] It should be noted that the above-mentioned reference voice can not only refer to the voice information sent by humans, but also refer to the voice information emitted by animals. For example, cat meows, dog barks, etc. are all applicable to the voice recognition model training method provided in this application, except that the voice recognition objects targeted by the trained voice recognition models are different.

[0074] Step S202: Insert the noise information into the reference voice to generate a processed voice including voice segments and noise segments.

[0075] After obtaining the noise information and the reference voice, in order to accurately imitate the real dialogue scenario, the noise information can be inserted into the reference voice, so as to generate and obtain a processed voice including voice segments and noise segments. In some instances, the noise information can be randomly inserted into the reference voice. At this time, inserting the noise information into the reference voice to generate a processed voice including voice segments and noise segments can include: randomly or evenly inserting the noise information into the reference voice to obtain a processed voice including voice segments and noise segments.

[0076] In still some other instances, since there are many types of noise information and the duration of the voice segments of the reference voice is limited, in order to ensure the generation quality and effect of the processed voice, inserting the noise information into the reference voice to generate a processed voice including human voice segments and noise segments can include: obtaining the noise insertion position of the reference voice; determining the noise insertion segment based on the noise information; inserting the noise insertion segment into the reference voice based on the noise insertion position to generate the processed voice.

[0077] Among them, after obtaining the reference speech, the reference speech can be analyzed and processed, so that the noise insertion position of the reference speech can be obtained. The above noise insertion position can include at least one of the following: before or after the first sentence speech of the reference speech, before or after the last sentence speech of the reference speech, and between sentences of the reference speech. Specifically, the specific acquisition method of the noise insertion position in this embodiment is not limited. In some examples, the above noise insertion position can be a pre-configured or default position. For example, a noise insertion position list is pre-stored, and the mapping relationship between multiple reference speeches with different durations and the noise insertion position is stored in the above noise insertion position list. After obtaining the reference speech, the noise insertion position corresponding to the reference speech can be obtained by looking up the table. In some other examples, the noise insertion position can be obtained through the interaction operation of the user. At this time, obtaining the noise insertion position of the reference speech can include: displaying a human-computer interaction interface, obtaining the execution operation input by the user in the human-computer interaction interface, and obtaining the noise insertion position of the reference speech based on the execution operation.

[0078] Since there are many types of the obtained noise information, in order to ensure the generation quality and effect of the processed speech, after obtaining the noise information, the noise information can be analyzed and processed, so that the noise insertion segment can be determined. The noise insertion segment is a part of the noise information. In some examples, one or more noise insertion segments are obtained by randomly extracting operations in the noise information, and the durations corresponding to any two noise insertion segments can be the same or different. In some other examples, in order to avoid the noise segment ratio being too large and reducing the signal-to-noise ratio, thereby affecting the accuracy of speech recognition, determining the noise insertion segment based on the noise information in this embodiment can include: obtaining a duration parameter for limiting the duration of the noise insertion segment; randomly generating at least one noise insertion segment that meets the duration parameter based on the noise information.

[0079] Among them, for each noise insertion segment, there is corresponding duration information. To avoid the situation of affecting the accuracy of speech recognition due to reducing the signal-to-noise ratio, a duration parameter for limiting the duration of the noise insertion segment can be obtained. This duration parameter can be a pre-configured parameter, or it can also be a parameter configured or adjusted in real time through a human-computer interaction interface. In some instances, the duration parameter is often between 0s and 5s. Commonly, the duration parameter can be set to 2s. After obtaining the duration parameter, at least one noise insertion segment that meets the duration parameter is randomly generated based on the noise information. For example, when the duration parameter is 2s, a noise insertion segment with a duration of 1s, a noise insertion segment with a duration of 2s, or a noise insertion segment with a duration of 0.5s can be generated, etc. When the number of noise insertion segments is multiple, the multiple noise insertion segments can be the same or different, that is, the noise types and noise information included in any two noise insertion segments can be the same or different, or the duration parameters corresponding to any two noise insertion segments can be the same or different.

[0080] After obtaining the noise insertion position and the noise insertion segment, the noise insertion segment can be inserted into the reference speech based on the noise insertion position, so that the processed speech can be generated. Example 1, as shown in the attached Figure 3 figure, the noise insertion position can be before the first sentence of the reference speech. After obtaining the noise insertion segment and the reference speech, the noise insertion segment can be inserted before the first sentence of the reference speech, so that the processed speech can be obtained. Example 2, as shown in the attached Figure 4 figure, the noise insertion position can be before the first sentence of the reference speech and after the last sentence of the reference speech. After obtaining the noise insertion segment 1, the noise insertion segment 2 and the reference speech, the noise insertion segment 1 can be inserted before the first sentence of the reference speech, and the noise insertion segment 2 can be inserted after the last sentence of the reference speech, so that the processed speech can be obtained.

[0081] It should be noted that the number of noise insertion positions corresponds to the number of noise insertion segments, and the specific number of noise insertion positions can be determined by the user's selection or randomly determined, so as to ensure the flexible diversity of the processed speech and cover most application scenarios.

[0082] Step S203: Determine a preset noise identifier corresponding to the noise segment in the processed speech.

[0083] In order to obtain a speech recognition model through the display modeling operation of noise, after obtaining the processed speech, the noise segments in the processed speech can be analyzed and processed, so as to determine a preset noise identifier corresponding to the noise segments in the processed speech. In some instances, the preset noise identifier can be obtained through manual standard operations. At this time, determining the preset noise identifier corresponding to the noise segments in the processed speech can include: obtaining the annotation operation input by the user for the noise segments in the processed speech, and obtaining the preset noise identifier based on the annotation operation. In still other instances, the preset noise identifier can not only be obtained through manual standard operations, but also be obtained by analyzing and processing the noise segments through a pre-trained machine learning model or neural network model. At this time, determining the preset noise identifier corresponding to the noise segments in the processed speech can include: obtaining a pre-trained machine learning model or neural network model, and inputting the noise segments of the processed speech into the machine learning model or neural network model, so as to obtain the preset noise identifier corresponding to the noise segments in the processed speech.

[0084] Step S204: Based on the reference speech, the processed speech, and the preset noise identifier, perform model training to obtain a speech recognition model, which is used to recognize the input speech as text.

[0085] After obtaining the reference speech, the processed speech, and the preset noise identifier, a model training operation can be performed based on the reference speech, the processed speech, and the preset noise identifier. Since the preset noise identifier corresponds to the noise segments in the processed speech, the display modeling operation for noise is realized, so that a speech recognition model that can stably and accurately recognize the input speech as text can be obtained.

[0086] In still other instances, after obtaining the speech recognition model, the trained speech recognition model can be used for speech recognition operations. At this time, the method in this embodiment can further include: obtaining the audio to be recognized; inputting the audio to be recognized into the speech recognition model to identify the speech segments and noise segments included in the audio to be recognized, and transcribing the speech segments to obtain the text information corresponding to the audio to be recognized.

[0087] Specifically, the audio to be recognized refers to the voice information that needs to be normally recognized and written into text. In some instances, the audio to be recognized can be the audio that is pre-collected and stored in a third device (such as a vehicle device, a conference room device, etc.). The third device is communicatively connected to the training device of the speech recognition model. At this time, the audio to be recognized can be actively or passively obtained through the third device. In some other instances, the audio to be recognized can be collected in real time by a voice collector. At this time, the voice collector can be configured inside the training device of the speech recognition model, or the voice collector can be communicatively connected to the training device of the speech recognition model. After the voice collector collects the audio to be recognized in real time, the obtained audio to be recognized can be directly transmitted to the training device of the speech recognition model to perform speech recognition operations through the obtained speech recognition model.

[0088] After obtaining the audio to be recognized, the audio to be recognized can be input into the speech recognition model. After the speech recognition model obtains the audio to be recognized, it can analyze and process the audio to be recognized, so as to determine the speech segments and noise segments included in the audio to be recognized. Among them, the number of speech segments and noise segments can be one or more. After recognizing the speech segments and noise segments, the speech recognition model can transcribe the speech segments normally, so as to obtain the text information corresponding to the audio to be recognized. For the noise segments, the speech recognition model will not transcribe the noise segments, thus effectively avoiding the situation of noise generating characters.

[0089] It should be noted that when performing the training operation of the speech recognition model, not only can the preset noise identifier corresponding to the noise segment in the processed speech be obtained, but also the preset noise identifier corresponding to the pure noise segment can be obtained. Combining the pure noise segment and the preset noise identifier corresponding to the pure noise segment to perform display modeling operations. At this time, the training method of the speech recognition model can include: obtaining reference speech and noise information, where the reference speech includes the speech information that can be recognized as reference text; inserting the noise information into the reference speech to generate processed speech including speech segments and noise segments; determining the first noise identifier corresponding to the noise information and the second noise identifier corresponding to the noise segments in the processed speech; performing model training based on the reference speech, the processed speech, the noise information, the first noise identifier, and the second noise identifier to obtain a speech recognition model, and the speech recognition model is used to recognize the input speech as text.

[0090] Among them, the specific implementation methods, implementation effects, and implementation principles of the steps in this implementation method are similar to those of the steps in the above implementation method. For details, please refer to the above description and will not be elaborated here.

[0091] The training method of the speech recognition model provided in this embodiment obtains reference speech and noise information, inserts the noise information into the reference speech to generate processed speech including speech segments and noise segments; then determines a preset noise identifier corresponding to the noise segments in the processed speech; and performs model training based on the reference speech, the processed speech, and the preset noise identifier to obtain a speech recognition model, effectively realizing the acquisition of a speech recognition model with high robustness through noise display modeling operations. This significantly improves the robustness of the speech recognition model to noise or background noise. During the process of performing speech recognition operations based on the speech recognition model, it can reduce or even avoid the situation of recognizing words from noise, thereby ensuring a good user experience and effectively improving the practicality of this method, which is conducive to market promotion and application.

[0092] Figure 5 It is a schematic flowchart of another training method of the speech recognition model provided in an embodiment of the present invention; on the basis of the above embodiment, refer to the attached Figure 5 As shown, after obtaining the speech recognition model, in order to improve and ensure the speech recognition effect and quality of the speech recognition model, this embodiment may further include a solution for optimizing the speech recognition model. Specifically, the method in this embodiment may include:

[0093] Step S501: During the process of performing speech recognition on a speech segment using the speech recognition model, obtain a predicted value for identifying the number of characters corresponding to the speech segment through the speech recognition model.

[0094] Step S502: Determine an offset parameter and a regularization parameter for optimizing the predicted value.

[0095] Step S503: Optimize the predicted value based on the offset parameter and the regularization parameter to obtain an optimized predicted value.

[0096] Step S504: Obtain an optimized speech recognition model based on the optimized predicted value.

[0097] Among them, the speech recognition model includes an identification module and a transcription module. The above identification model is used to implement speech recognition operations, and the above transcription module is used to transcribe the recognized speech into text information. During the process of performing speech recognition on a speech segment using the speech recognition model, a predicted value for identifying the number of characters corresponding to the speech segment can be obtained through the transcription module, and this predicted value can be called the cif weight.

[0098] During the normal transcription of speech segments, multiple prediction values are often obtained. When some smaller prediction values are accumulated, it may trigger the appearance of characters, which can easily lead to the situation of noise-induced character appearance. To avoid the above situation, it is necessary to adjust the prediction values. At this time, the offset parameter and the regularization parameter for optimizing the prediction values can be determined. The above offset parameter and regularization parameter can be pre-configured parameters, and the above offset parameter and regularization parameter can be stored in a preset area. At this time, by accessing the preset area, the offset parameter and the regularization parameter for optimizing the prediction values can be obtained.

[0099] In some other instances, the offset parameter and the regularization parameter can be obtained based on the user's parameter configuration operation. At this time, determining the offset parameter and the regularization parameter for optimizing the prediction values can include: obtaining the parameter configuration page for implementing the parameter configuration operation, obtaining the parameter configuration operation input by the user on the parameter configuration page, and obtaining the offset parameter and the regularization parameter for optimizing the prediction values based on the parameter configuration operation.

[0100] After obtaining the offset parameter and the regularization parameter, the prediction values can be optimized based on the offset parameter and the regularization parameter, so that the optimized prediction values can be obtained. In some instances, optimizing the prediction values based on the offset parameter and the regularization parameter to obtain the optimized prediction values can include: obtaining the difference between the prediction value and the offset parameter; performing a rectified linear unit (ReLU) operation on the difference to obtain the processed difference; and determining the product value of the processed difference and the regularization parameter as the so-called optimized prediction value.

[0101] For example, taking the offset parameter as p noise and the regularization parameter as p smooth and the prediction value as f`(x) as an example, after obtaining the prediction value f`(x) and the offset parameter p noise , the difference between the prediction value f`(x) and the offset parameter p noise can be obtained, that is, f`(x) - p noise . Then, the ReLU operation can be performed on the difference to obtain the processed difference. Specifically, the processed difference is relu(f`(x) - p noise ); then, the product value between the processed difference and the regularization parameter p smooth is obtained, that is, relu(f`(x) - p noise ) * p smooth . Then, the product value of the processed difference and the regularization parameter p smooth can be determined as the optimized prediction value, that is, f(x) = relu(f`(x) - p noise ) * p smooth, thus effectively ensuring the accuracy and reliability of obtaining the optimized predicted value. Among them, the above-mentioned p noise = 0.05, p smooth = 0.9. Of course, those skilled in the art can also adjust the values of p noise and p smooth according to specific application scenarios or application requirements. For example, p noise can be 0.02, 0.03 or 0.01, etc. Similarly, p smooth can be 0.8, 0.7, 0.6 or 0.9, etc., to meet the application requirements of different application scenarios, which will not be elaborated here.

[0102] In some other examples, the optimization operation of the predicted value can be implemented not only through the above formula, but also through a pre-trained machine learning model or neural network model. At this time, optimizing the predicted value based on the offset parameter and the regularization parameter to obtain the optimized predicted value may include: obtaining a machine learning model or neural network model for implementing the optimization operation of the predicted value. After obtaining the offset parameter and the regularization parameter, the offset parameter, the regularization parameter, and the predicted value can be input into the machine learning model or neural network model, so that the optimized predicted value output by the machine learning model or neural network model can be obtained.

[0103] After obtaining the optimized predicted value, the optimized predicted value can be analyzed and processed, so that an optimized speech recognition model can be obtained. When using the optimized speech recognition model for speech recognition operations, it can effectively reduce or even avoid the situation of recognizing words from noise during the speech recognition process.

[0104] In this embodiment, during the process of using the speech recognition model to perform speech recognition on a speech segment, a predicted value for identifying the number of characters corresponding to the speech segment is obtained through the speech recognition model, and then an offset parameter and a regularization parameter for optimizing the predicted value are determined. The predicted value is optimized based on the offset parameter and the regularization parameter to obtain the optimized predicted value, and an optimized speech recognition model is obtained based on the optimized predicted value, thus effectively realizing the optimization operation of the speech recognition model and further improving the noise robustness of the speech recognition model.

[0105] In specific applications, taking the offset parameter p noise as 0.05 and the regularization parameter p smoothTaking 0.9 as an example, the embodiment of this application provides a method for training a speech recognition model. This method is applicable to the conference scenario. By combining data simulation and model modeling solutions, a speech recognition model with high robustness can be obtained through noise display modeling operations. When using the speech recognition model for speech recognition operations, the problem of noise-induced word generation can be effectively reduced or even avoided, thereby improving the noise robustness of the speech recognition model. Specifically, the method for training the speech recognition model may include the following steps:

[0106] Step 1: Obtain reference speech and noise information. The reference speech may include speech information that can be recognized as reference text.

[0107] Among them, for the noise information in the model training data, the noise information may include environmental noise and non-environmental noise. Environmental noise may include event-related interference sounds. For event-related interference sounds, they can be collected through a speech acquisition device, such as: ringtones, coughs, knocks, phone beeps, etc. Non-environmental noise refers to speech information that cannot be normally recognized as text. Non-environmental noise may include simulated human voice noise and actual human voice noise. Among them, actual human voice noise may include pure human voice noise collected through a semantic acquisition device, and simulated human voice noise may be simulated partial human voice noise obtained by performing low-volume amplitude modulation / noise addition / distortion processing on human voices.

[0108] In addition, the reference speech may be speech training data used to implement model training operations. The reference speech can be obtained by extracting from the preset model training data. Specifically, a reference speech with a preset duration can be extracted from the model training data. The preset duration can be 5k hours, 1w hours, etc. The specific value of the preset duration may vary based on different application scenarios.

[0109] Step 2: In the noise information in the model training data, extract one or more noise segments of 0 - 2s.

[0110] Among them, any one noise segment may include one or more different types of noise information. It can be understood that the more types of noise included in a noise segment, the better the effect of model training.

[0111] Step 3: Randomly splice the one or more extracted noise segments with a preset duration at the beginning / end / between sentences of the reference speech to obtain processed data. The obtained processed data may include speech segments and noise segments.

[0112] Among them, the preset duration can be any value within the range of 0 - 5s. Since the pause time during people's speaking or thinking process is usually within 1s or 2s, when the number of noise segments is multiple, the preset duration corresponding to most of the noise segments can be 2s, while the preset duration corresponding to a small number of noise segments can be configured as 3s, 4s or 5s. To avoid reducing the signal-to-noise ratio, for noise segments with a longer preset duration, the number is smaller; for noise segments with a shorter preset duration, the number can be larger.

[0113] Step 4: Determine the first noise identifier corresponding to the noise segment in the processed data and the second noise identifier corresponding to the noise information.

[0114] Among them, before the modeling operation, add a first noise identifier of "noise" to the noise segments in the processed data. Specifically, to ensure the quality and effect of model training and implement the display modeling operation for noise, during model training, some noise data can be directly used as training data, for example: some pure noise segments, etc. For the above noise information, a second noise identifier of "noise" can also be added. The above first noise identifier and second noise identifier can be the same or different.

[0115] Step 5: Perform a display modeling operation based on the processed data, reference speech, noise information, first noise identifier, and second noise identifier, so as to obtain an end-to-end speech recognition model. This speech recognition model can absorb noise through the "noise" noise identifier.

[0116] After obtaining the speech recognition model, the end-to-end speech recognition model can be used for speech recognition operations. Specifically, for an audio containing both human voices and noise, through the speech recognition model, the human voice frequency band can be used as a positive sample for normal transcription, while the noise segment is not transcribed as a negative sample. For example, <gbg>Taking the noise identification as an example, the speech recognition model at this time can recognize all noise information as <gbg>, at this time, when using the speech recognition model to recognize three speech to be processed, the intermediate recognition results corresponding to each of the three speech to be processed can be obtained. For example, the intermediate recognition results can include intermediate recognition result 1, intermediate recognition result 2, and intermediate recognition result 3. Among them, intermediate recognition result 1 can specifically be "document content <gbg>"as follows"; the intermediate recognition result 2 can be "turn on the light"; the intermediate recognition result 3 can be "to <gbg>”; After decoding the above intermediate recognition results, filtering can be processed to remove <gbg>Symbols, so that the recognition and transcription results of the three voices to be processed can be obtained respectively, and the transcription results are "The content of the document is as follows", "Turn on the light", and "Yes"; generally speaking, for a complete voice containing noise segments (for example: noise segment + voice segment + noise segment + voice segment), when the intermediate result is obtained through the voice recognition model, the gbg symbol can be used to represent the noise segment. Then, when decoding the intermediate result, the gbg symbol can be filtered out, thus effectively realizing the rejection operation of noise through the added noise identification and avoiding the situation of noise generating characters.

[0117] In addition, to ensure the accuracy of voice recognition, the method in this embodiment may further include an operation of optimizing the voice recognition model. At this time, the method in this embodiment may include:

[0118] Step 11: Obtain the voice to be recognized.

[0119] Step 12: When using the voice recognition model to analyze and process the voice to be recognized, obtain a predicted value for identifying the number of characters corresponding to the voice to be recognized.

[0120] Among them, the predicted value can be implemented as the weight of (Continuous Integrate-and-Fire, abbreviated as cif). The above cif weight draws on the principle of the spiking neural network. For voice tasks, the input of each frame of the cif network can be understood as a membrane potential. By accumulating the membrane potential of each frame, when the accumulated membrane potential exceeds a preset threshold, a neuron is activated (reaching the boundary of a character).

[0121] Step 13: Determine the offset parameter and regularization parameter for optimizing the predicted value.

[0122] Among them, after obtaining the cif weight f`(x), the cif weight f`(x) can be analyzed and processed to implement the voice transcription operation. However, when some small burr values accumulate, it may also trigger character output, which is likely to cause the situation of noise generating extra characters. To avoid the above situation, the predicted value can be subtracted by an offset, that is, the predicted value - offset parameter, so as to obtain the difference f`(x) - p noise , and then the above difference can be subjected to a rectified linear unit operation. Specifically, the rectified linear unit function relu function can be used to process the difference, so as to obtain relu(f`*x) - p noise ), in order to keep the magnitude of the data consistent, after obtaining the processed difference above, the processed difference can be multiplied by a preset regularization parameter, so as to obtain the optimized predicted value. Specifically, the optimized predicted value f(x) = relu(f`(x)p noise )*p smotch , where p noise = 0.05; p smotch = 0.9.

[0123] Step 14: Since the predicted value is obtained through the speech recognition model, after obtaining the optimized predicted value, the optimized predicted value can be analyzed and processed, so that an optimized speech recognition model can be obtained.

[0124] The technical solution provided by the embodiment of this application realizes the explicit modeling operation of noise by combining data simulation technology and model modeling technology, so that an end-to-end speech recognition model can be obtained, and then the end-to-end speech recognition model can be used to process the problem of noise character output. Compared with the implementation methods in the prior art, it can better improve the recognition and utilization of the speech recognition model for noise data, thereby enhancing the noise robustness of the model. It can also be seen from the experimental indicators of the system that the speech recognition model obtained by training in the above manner can further obtain an obvious optimization effect on noise character output. Specifically, when using the speech recognition model obtained in this application for speech recognition operations, the rejection rate relative to the reference noise on the pure noise test set is increased from the original 8.91% to 52.6%, thereby significantly avoiding or reducing the situation of noise character output, further improving the practicability of this method, and being conducive to market promotion and application.

[0125] Figure 6 It is a schematic flowchart of a speech recognition method provided by an embodiment of the present invention; referring to the attached Figure 6 As shown, this embodiment provides a speech recognition method. The execution subject of this method can be a speech recognition device. It can be understood that the speech recognition device can be implemented as software, or a combination of software and hardware. Specifically, when the speech recognition device is implemented as hardware, it can specifically be various electronic devices with speech recognition operations, including but not limited to tablet computers, personal computers PC, servers, etc. When the speech recognition device is implemented as software, it can be installed in the above-mentioned exemplified electronic devices. Based on the above speech recognition device, the speech recognition method in this embodiment can include the following steps:

[0126] Step S601: Obtain the audio to be recognized.

[0127] Among them, the audio to be recognized refers to the voice information that needs to be normally recognized and written into text. In this embodiment, the acquisition method of the audio to be recognized is not limited. In some examples, the audio to be recognized can be the audio pre-collected and stored in the third device. The third device is communicatively connected to the training device of the speech recognition model. At this time, the audio to be recognized can be actively or passively obtained through the third device. In some other examples, the audio to be recognized can be collected in real time by a voice collector. At this time, the voice collector can be configured inside the training device of the speech recognition model, or the voice collector can be communicatively connected to the training device of the speech recognition model. After the voice collector collects the audio to be recognized in real time, the obtained audio to be recognized can be directly transmitted to the training device of the speech recognition model to perform a speech recognition operation through the obtained speech recognition model.

[0128] Step S602: Input the audio to be recognized into the speech recognition model to identify the speech segments and noise segments included in the audio to be recognized, and transcribe the speech segments to obtain the text information corresponding to the audio to be recognized;

[0129] Among them, the speech recognition model is obtained by learning and training with reference speech, processed speech, and a preset noise identifier. The processed speech is obtained by inserting noise information into the reference speech. The processed speech includes speech segments and noise segments, and the preset noise identifier corresponds to the noise segments. It should be noted that the specific training method of the speech recognition model is similar to the training method of the speech recognition model shown above. For details, reference can be made to the above description and will not be elaborated here. Figures 1 - 5 After obtaining the audio to be recognized, the audio to be recognized can be input into the speech recognition model. After the speech recognition model obtains the audio to be recognized, it can analyze and process the audio to be recognized, so as to determine the speech segments and noise segments included in the audio to be recognized. Among them, the number of speech segments and noise segments can be one or more. After identifying the speech segments and noise segments, the speech recognition model can normally transcribe the speech segments, so as to obtain the text information corresponding to the audio to be recognized. For the noise segments, the speech recognition model will not transcribe the noise segments, thus effectively avoiding the situation of noise generating characters.

[0130]

[0131] ​In the speech recognition method of this embodiment, by obtaining the audio to be recognized, the audio to be recognized can then be input into a speech recognition model to recognize the speech segments and noise segments included in the audio to be recognized, and transcribe the speech segments to obtain the text information corresponding to the audio to be recognized. Since the speech recognition model has high robustness against noise, when using the speech recognition model for speech recognition operations, the situation of noise generating characters is effectively avoided, and the quality and efficiency of the speech recognition operations are further improved.

[0132] Figure 7 It is a schematic structural diagram of a training device for a speech recognition model provided by an embodiment of the present invention; refer to the attached Figure 7 As shown, this embodiment provides a training device for a speech recognition model. The training device for the speech recognition model can execute the above Figure 2 shown speech recognition model training method. The training device for the speech recognition model can include: a first acquisition module 11, a first generation module 12, a first processing module 13, and a first training module 14. Specifically,

[0133] The first acquisition module 11 is used to acquire reference speech and noise information, and the reference speech includes speech information that can be recognized as reference text;

[0134] The first generation module 12 is used to insert the noise information into the reference speech to generate processed speech including speech segments and noise segments;

[0135] The first processing module 13 is used to determine a preset noise identifier corresponding to the noise segment in the processed speech;

[0136] The first training module 14 is used to perform model training based on the reference speech, the processed speech, and the preset noise identifier to obtain a speech recognition model, and the speech recognition model is used to recognize the input speech as text.

[0137] In some instances, the noise information includes at least one of the following: environmental interference sounds, non-environmental interference sounds.

[0138] In some instances, the non-environmental interference sound includes simulated human voice noise. When the first acquisition module 11 acquires the noise information, the first acquisition module 11 is used to execute: acquire the human voice data included in the reference speech; perform at least one of the processing operations of low-volume amplitude modulation, adding noise, and distortion on the human voice data to obtain the simulated human voice noise corresponding to the reference speech.

[0139] In some instances, when the first generation module 12 inserts noise information into the reference speech to generate the processed speech including human voice segments and noise segments, the first generation module 12 is configured to perform: obtaining the noise insertion position of the reference speech; determining the noise insertion segment based on the noise information; inserting the noise insertion segment into the reference speech based on the noise insertion position to generate the processed speech.

[0140] In some instances, the noise insertion position includes at least one of the following: before or after the first sentence speech of the reference speech, before or after the last sentence speech of the reference speech, and between sentences of the reference speech.

[0141] In some instances, when the first generation module 12 determines the noise insertion segment based on the noise information, the first generation module 12 is configured to perform: obtaining the duration parameter for defining the duration of the noise insertion segment; randomly generating at least one noise insertion segment that meets the duration parameter based on the noise information.

[0142] In some instances, when the number of noise insertion segments is multiple, the multiple noise insertion segments are the same or different.

[0143] In some instances, after obtaining the speech recognition model, the first acquisition module 11 and the first processing module 13 in this embodiment are configured to perform the following steps:

[0144] The first acquisition module 11 is configured to acquire the audio to be recognized.

[0145] The first processing module 13 is configured to input the audio to be recognized into the speech recognition model to recognize the speech segments and noise segments included in the audio to be recognized, and transcribe the speech segments to obtain the text information corresponding to the audio to be recognized.

[0146] In some instances, after obtaining the speech recognition model, the first processing module 13 in this embodiment is configured to perform the following steps: during the process of performing speech recognition on the speech segment by using the speech recognition model, obtaining a predicted value for identifying the number of characters corresponding to the speech segment through the speech recognition model; determining the offset parameter and the regularization parameter for optimizing the predicted value; optimizing the predicted value based on the offset parameter and the regularization parameter to obtain the optimized predicted value; obtaining the optimized speech recognition model based on the optimized predicted value.

[0147] In some instances, when the first processing module 13 optimizes the predicted value based on the offset parameter and the regularization parameter to obtain the optimized predicted value, the first processing module 13 is configured to perform: obtaining the difference between the predicted value and the offset parameter; performing a rectified linear unit operation on the difference to obtain the processed difference; determining the product value of the processed difference and the regularization parameter as the optimized predicted value.

[0148] Figure 7 The device shown can execute Figures 1 - 5 the method of the embodiment shown. For parts not described in detail in this embodiment, reference can be made to the relevant descriptions of the Figures 1 - 5 embodiment shown. For the execution process and technical effects of this technical solution, refer to the description in the Figures 1 - 5 embodiment shown, which will not be elaborated here.

[0149] In a possible design, Figure 7 the structure of the training device of the voice recognition model shown can be implemented as an electronic device, and this electronic device can be various devices such as a controller, a personal computer, a server, etc. As Figure 8 shown, this electronic device may include: a first processor 21 and a first memory 22. Among them, the first memory 22 is used to store a program for the corresponding electronic device to execute the voice recognition model training method provided in the Figures 1 - 5 embodiment shown. The first processor 21 is configured to execute the program stored in the first memory 22.

[0150] The program includes one or more computer instructions. Among them, when one or more computer instructions are executed by the first processor 21, the following steps can be realized: obtaining reference speech and noise information, where the reference speech includes speech information that can be recognized as reference text; inserting the noise information into the reference speech to generate processed speech including speech segments and noise segments; determining a preset noise identifier corresponding to the noise segment in the processed speech; performing model training based on the reference speech, the processed speech, and the preset noise identifier to obtain a voice recognition model, and the voice recognition model is used to recognize the input speech as text.

[0151] Further, the first processor 21 is also used to execute all or part of the steps in the foregoing Figures 1 - 5 embodiment shown.

[0152] Among them, the structure of the electronic device may further include a first communication interface 23, which is used for the electronic device to communicate with other devices or communication networks.

[0153] In addition, an embodiment of the present invention provides a computer storage medium for storing computer software instructions used by an electronic device, which includes a program involved in executing the voice recognition model training method in the Figures 1 - 5 embodiment shown.

[0154] In addition, an embodiment of the present invention provides a computer program product, including: a computer-readable storage medium storing computer instructions, and when the computer instructions are executed by one or more processors, one or more processors are caused to execute the steps in the voice recognition model training method in the Figures 1 - 5 method embodiment shown.

[0155] Figure 9 Schematic structural diagram of a voice recognition device provided by an embodiment of the present invention; refer to the appendix Figure 9 As shown, this embodiment provides a voice recognition device, and this voice recognition device can execute the above Figure 6 As shown, the voice recognition device provided in this embodiment may include: a second acquisition module 31 and a second processing module 32. Specifically,

[0156] The second acquisition module 31 is used to acquire the audio to be recognized;

[0157] The second processing module 32 is configured to input the audio to be recognized into a voice recognition model to recognize the voice segment and noise segment included in the audio to be recognized, and transcribe the voice segment to obtain text information corresponding to the audio to be recognized;

[0158] Among them, the voice recognition model is obtained through learning and training with reference voice, processed voice, and a preset noise identifier. The processed voice is obtained by inserting noise information into the reference voice. The processed voice includes a voice segment and a noise segment, and the preset noise identifier corresponds to the noise segment.

[0159] Figure 9 The device shown can execute Figure 6 The method of the embodiment shown. For parts not described in detail in this embodiment, reference may be made to the relevant descriptions of the corresponding Figure 6 The execution process and technical effects of this technical solution are referred to the descriptions in the embodiment shown in Figure 6 and will not be elaborated here.

[0160] In a possible design, Figure 9 The structure of the voice recognition device shown can be implemented as an electronic device, and this electronic device can be various devices such as a mobile phone, a tablet computer, a server, etc. As shown in 10, this electronic device may include: a second processor 41 and a second memory 42. Among them, the second memory 42 is used to store a program for the corresponding electronic device to execute the voice recognition method provided in the above Figure 6 shown embodiment, and the second processor 41 is configured to execute the program stored in the second memory 42.

[0161] The program includes one or more computer instructions. When the one or more computer instructions are executed by the second processor 41, the following steps can be implemented: obtaining the audio to be recognized; inputting the audio to be recognized into a speech recognition model to recognize the speech segments and noise segments included in the audio to be recognized, and transcribing the speech segments to obtain text information corresponding to the audio to be recognized; wherein, the speech recognition model is obtained through learning and training with reference speech, processed speech, and a preset noise identifier. The processed speech is obtained by inserting noise information into the reference speech. The processed speech includes speech segments and noise segments, and the preset noise identifier corresponds to the noise segments.

[0162] Further, the second processor 41 is further configured to execute all or part of the steps in the foregoing Figure 6 illustrated embodiments.

[0163] Wherein, the structure of the electronic device may further include a second communication interface 43 for the electronic device to communicate with other devices or communication networks.

[0164] In addition, an embodiment of the present invention provides a computer storage medium for storing computer software instructions used by the electronic device, which includes a program involved in the speech recognition method in the foregoing Figure 6 illustrated method embodiments.

[0165] In addition, an embodiment of the present invention provides a computer program product, including: a computer-readable storage medium storing computer instructions. When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the steps in the speech recognition method in the foregoing Figure 6 illustrated method embodiments.

[0166] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0167] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, can also be implemented by a combination of hardware and software. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program codes.

[0168] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable devices generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0169] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0170] These computer program instructions can also be loaded onto a computer or other programmable device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0171] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0172] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.

[0173] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.< / gbg> < / gbg> < / gbg> < / gbg> < / gbg>

Claims

1. A training method for a speech recognition model, characterized in that, Including: Obtain reference speech and noise information, where the reference speech includes speech information that can be recognized as reference text; Insert the noise information into the reference speech to generate processed speech including speech segments and noise segments; Determine a preset noise identifier corresponding to the noise segment in the processed speech; Perform model training based on the reference speech, the processed speech, and the preset noise identifier to obtain a speech recognition model, which is used to recognize input speech as text; During the process of using the speech recognition model to perform speech recognition on a speech segment, obtain a predicted value for identifying the number of characters corresponding to the speech segment through the speech recognition model; Determine an offset parameter and a regularization parameter for optimizing the predicted value; Optimize the predicted value based on the offset parameter and the regularization parameter to obtain an optimized predicted value; Based on the optimized predicted value, obtain an optimized speech recognition model.

2. The method according to claim 1, wherein The noise information includes at least one of the following: environmental interference sounds, non-environmental interference sounds.

3. The method according to claim 2, wherein The non-environmental interference sounds include simulated human voice noise. Obtaining the noise information includes: Obtain the human voice data included in the reference speech; Perform at least one of the processing operations of low-volume amplitude modulation, adding noise, and distortion on the human voice data to obtain simulated human voice noise corresponding to the reference speech.

4. The method according to claim 1, wherein Insert the noise information into the reference speech to generate processed speech including human voice segments and noise segments, including: Obtain the noise insertion position of the reference speech; Based on the noise information, determine a noise insertion segment; Insert the noise insertion segment into the reference speech based on the noise insertion position to generate the processed speech.

5. The method according to claim 4, wherein The noise insertion position includes at least one of the following: before or after the first sentence speech of the reference speech, before or after the last sentence speech of the reference speech, between sentences of the reference speech.

6. The method according to claim 4, wherein Based on the noise information, determine a noise insertion segment, including: Obtain a duration parameter for limiting the duration of the noise insertion segment; Randomly generate at least one noise insertion segment that satisfies the duration parameter based on the noise information.

7. The method according to claim 4, wherein When the number of noise insertion segments is multiple, the multiple noise insertion segments are the same or different.

8. The method according to claim 1, characterized in that, After obtaining the speech recognition model, the method further includes: Obtain an audio to be recognized; Input the audio to be recognized into the speech recognition model to recognize the speech segments and noise segments included in the audio to be recognized, and transcribe the speech segments to obtain text information corresponding to the audio to be recognized.

9. The method according to claim 1, wherein Optimizing the predicted value based on the offset parameter and the regularization parameter to obtain an optimized predicted value, including: Obtain the difference between the predicted value and the offset parameter; Perform a rectified linear unit operation on the difference to obtain a processed difference; Determine the product value of the processed difference and the regularization parameter as the optimized predicted value.

10. A voice recognition method, characterized in that, Including: Obtain an audio to be recognized; Input the audio to be recognized into a speech recognition model to recognize the speech segments and noise segments included in the audio to be recognized, and transcribe the speech segments to obtain text information corresponding to the audio to be recognized; Among them, the speech recognition model is obtained through learning and training with reference speech, processed speech, and a preset noise identifier. The processed speech is obtained by inserting noise information into the reference speech. The processed speech includes speech segments and noise segments, and the preset noise identifier corresponds to the noise segments. The speech recognition model is optimized through the following steps: In the process of performing speech recognition on speech segments using the speech recognition model, obtain a predicted value for identifying the number of characters corresponding to the speech segments through the speech recognition model; determine an offset parameter and a regularization parameter for optimizing the predicted value; optimize the predicted value based on the offset parameter and the regularization parameter to obtain an optimized predicted value; Based on the optimized predicted value, obtain an optimized speech recognition model.

11. An electronic device, characterized in that, Including: A memory and a processor; among them, the memory is used to store one or more computer instructions, and when the one or more computer instructions are executed by the processor, the method according to any one of claims 1-10 above is implemented.

Citation Information

Patent Citations

  • Processing method and device for voice recognition in vehicle and electronic equipment

    CN108022591A

  • Voice conversion method and device, computer equipment and storage medium

    CN114283811A

  • Speech processing model training method, and speech data noise reduction method and device

    CN114464168A