Speech recognition model training method, server and computer readable storage medium
The processing of training data through the time stamp-driven collaborative masking technology solves the problem of the speech recognition model being allergic to high-frequency voice requests in the on-board dialogue system, reduces the false touch rate, and improves user experience and model generalization capabilities.
Patent Information
- Application Number
- CN202510253088.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-17
AI Technical Summary
In the existing in-vehicle dialogue system, the voice recognition model is too sensitive to high-frequency voice requests, resulting in an increase in the false touch rate and poor user experience.
Through timestamp-driven collaborative masking technology, the target training data is processed to determine the target mask training data, which is used to train the speech recognition model. The method includes obtaining the target training data and related timestamp information, determining the target mask training data, and using it for training the speech recognition model.
It reduces the error touch rate of high-frequency voice requests by the speech recognition model, improves the user experience, and retains data diversity while reducing error touch, and improves the generalization ability of the model.
Smart Images

Figure CN120164458A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle control, and particularly relates to a training method for a speech recognition model, a server, and a computer-readable storage medium. Background Art
[0002] To improve the driving experience, an in-vehicle dialogue system is provided for users to control the vehicle through voice interaction. In related technologies, desensitized data obtained online is usually used to train the speech recognition model in the in-vehicle dialogue system to improve the in-vehicle dialogue system's ability to understand users' voice requests. However, the directly obtained online desensitized data often focuses on voice requests such as opening the window. Thus, the speech recognition model will be too sensitive to such voice requests, resulting in an increase in the accidental touch rate and poor user experience. Summary of the Invention
[0003] The present application provides a training method for a speech recognition model, a server, and a computer-readable storage medium.
[0004] An embodiment of the present application provides a training method for a speech recognition model, the method comprising:
[0005] Determine target timestamp information associated with the target training data according to the obtained target training data, wherein the target training data includes target audio data and target text data corresponding to the target audio data;
[0006] Determine target masked training data according to the target training data and the target timestamp information;
[0007] Train the speech recognition model according to the target masked training data.
[0008] In this way, the server determines target timestamp information associated with the target training data according to the obtained target training data, wherein the target training data includes target audio data and target text data corresponding to the target audio data. Then, the server determines target masked training data according to the target training data and the target timestamp information. Finally, the server trains the speech recognition model according to the target masked training data. In this way, through the timestamp-driven collaborative masking technology, the target training data is actively processed during the speech recognition model training stage to determine the target masked training data, which can enable the speech recognition model to learn deep context semantics and anti-interference features, reduce the accidental touch rate, and improve the user experience. Moreover, while reducing accidental touches, data diversity is retained and the model generalization ability is improved.
[0009] In some embodiments, the target training data is obtained from a pre-constructed training database, which is constructed based on original training data. The original training data includes original audio data and original text data corresponding to the original audio data. The training database is constructed through the following steps:
[0010] Based on a preset model, perform time alignment processing on the original audio data and the original text data in units of words to determine timestamp information;
[0011] Construct the training database according to the original training data and the timestamp information.
[0012] In this way, the server performs time alignment processing on the original audio data and the original text data in units of words based on a preset model to determine timestamp information. Then, the server constructs a training database according to the original training data and the timestamp information. In this way, by pre-constructing a training database, it is possible to avoid performing time alignment operations every time training is carried out, so as to improve the training efficiency.
[0013] In some embodiments, the constructing the training database according to the original training data and the timestamp information includes:
[0014] Perform association processing on the current original training data and the timestamp information corresponding to the current original training data to determine associated training data, where the original training data includes the current original training data;
[0015] Construct the training database according to the associated training data.
[0016] In this way, the server performs association processing on the current original training data and the timestamp information corresponding to the current original training data to determine associated training data, where the original training data includes the current original training data. Then, the server constructs a training database according to the associated training data. In this way, by associating the current original training data and the corresponding timestamp information, an operable and accurate mapping is achieved, laying a foundation for subsequent masking processing.
[0017] In some embodiments, the determining the target masked training data according to the target training data and the target timestamp information includes:
[0018] Based on a predefined keyword table, determine the target masked training data according to the target training data and the target timestamp information, where the keyword table includes predefined keywords and keyword masking probabilities corresponding to each keyword, and the keyword masking probabilities are determined based on the occurrence frequencies of the keywords in the pre-constructed training database.
[0019] Thus, based on a predefined keyword table, the server determines target masked training data according to the target training data and the target timestamp information, where the keyword table includes predefined keywords and keyword masking probabilities corresponding to each keyword, and the keyword masking probabilities are determined based on the occurrence frequencies of the keywords in a pre-constructed training database. In this way, by using the predefined keywords, the keyword masking probabilities corresponding to each keyword, and the target timestamp information to perform masking processing on the target training data, the mis-triggering rate of the speech recognition model can be reduced, so that various speech requests can be accurately recognized and processed, improving the robustness of the speech recognition model.
[0020] In some embodiments, the determining the target masked training data according to the target training data and the target timestamp information based on the predefined keyword table includes:
[0021] When the current target text data includes the keyword, based on the keyword masking probability, the target masked training data is determined according to the target training data and the target timestamp information.
[0022] Thus, when the current target text data includes a keyword, based on the keyword masking probability, the server determines the target masked training data according to the target training data and the target timestamp information. In this way, when the current target text data includes a keyword, based on the keyword masking probability, randomly performing masking processing on the target training data can reduce the mis-triggering rate of the speech recognition model, so that various speech requests can be accurately recognized and processed, improving the robustness of the speech recognition model.
[0023] In some embodiments, the determining the target masked training data according to the target training data and the target timestamp information when the current target text data includes the keyword includes:
[0024] Based on the keyword masking probability, determine the target keyword;
[0025] Perform a first masking process on the target keyword in the target text data to determine temporary target text data;
[0026] According to the target timestamp information, perform a second masking process on the audio segment corresponding to the keyword in the target audio data to determine temporary target audio data;
[0027] According to the temporary target text data and the temporary target audio data, determine the target masked training data.
[0028] In this way, based on the keyword masking probability, the server determines the target keyword. Then, the server performs a first masking process on the target keyword in the target text data to determine the temporary target text data. Next, the server performs a second masking process on the audio segment corresponding to the keyword in the target audio data according to the target timestamp information to determine the temporary target audio data. Finally, the server determines the target masked training data based on the temporary target text data and the temporary target audio data. In this way, by masking the keyword in the target training data and the audio segment corresponding to the keyword in the target audio data, the false touch rate of the speech recognition model can be reduced, so that various speech requests can be accurately recognized and processed, and the robustness of the speech recognition model can be improved.
[0029] In some embodiments, the performing a first masking process on the target keyword in the target text data to determine the temporary target text data includes:
[0030] Performing a discard process on the target keyword in the target text data to determine the temporary target text data;
[0031] The performing a second masking process on the audio segment corresponding to the keyword in the target audio data according to the target timestamp information to determine the temporary target audio data includes:
[0032] Performing a discard or noise replacement process on the audio segment corresponding to the keyword in the target audio data according to the target timestamp information to determine the temporary target audio data.
[0033] In this way, the server performs a discard process on the target keyword in the target text data to determine the temporary target text data. Then, the server performs a discard or noise replacement process on the audio segment corresponding to the keyword in the target audio data according to the target timestamp information to determine the temporary target audio data. In this way, by discarding the target keyword in the target text data and discarding or replacing the audio segment corresponding to the keyword in the target audio data with noise, the false touch rate of the keyword can be effectively reduced, while other information is retained, avoiding too much impact on the recognition rate of the speech recognition model.
[0034] In some embodiments, the method further includes:
[0035] When the current target text data does not include the keyword, or the target keyword is not determined based on the keyword masking probability, training the speech recognition model according to the target training data.
[0036] Thus, when the current target text data does not include the keyword, or the target keyword is not determined based on the keyword masking probability, the server trains the speech recognition model according to the target training data. In this way, when the current target text data does not include the keyword, or the target keyword is not determined based on the keyword masking probability, by directly using the target training data to train the speech recognition model, the robustness of the model can be improved, enabling it to handle various scenarios and speech variations.
[0037] An embodiment of the present application provides a server, which includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, the above-mentioned method for training the speech recognition model is implemented.
[0038] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for training the speech recognition model as described above are implemented.
[0039] Additional aspects and advantages of the embodiments of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, wherein:
[0041] Figure 1 is one of the flow diagrams of the method for training the speech recognition model according to some embodiments of the present application;
[0042] Figure 2 is another flow diagram of the method for training the speech recognition model according to some embodiments of the present application;
[0043] Figure 3 is the flow diagram of the forced processing of the Kaldi model according to some embodiments of the present application;
[0044] Figure 4 is a third flow diagram of the method for training the speech recognition model according to some embodiments of the present application;
[0045] Figure 5 is a fourth flow diagram of the method for training the speech recognition model according to some embodiments of the present application;
[0046] Figure 6 is a fifth flow diagram of the method for training the speech recognition model according to some embodiments of the present application;
[0047] Figure 7It is the sixth flowchart of the training method of the speech recognition model according to some embodiments of the present application;
[0048] Figure 8 It is the seventh flowchart of the training method of the speech recognition model according to some embodiments of the present application;
[0049] Figure 9 It is the eighth flowchart of the training method of the speech recognition model according to some embodiments of the present application. Specific embodiments
[0050] The following describes in detail the embodiments of the present application. The examples of the embodiments are shown in the accompanying drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application, and cannot be understood as a limitation to the embodiments of the present application.
[0051] To improve the driving experience, an in-vehicle dialogue system is provided for users to control the vehicle through voice interaction. In the related art, desensitized data obtained online is usually used to train the speech recognition model in the in-vehicle dialogue system to improve the in-vehicle dialogue system's ability to understand users' voice requests. However, there is often a problem with the directly obtained online desensitized data: the data set often focuses on some high-frequency voice requests, such as "open the window" and "close the door". This can cause the speech recognition model to be too sensitive to these voice requests and easily cause accidental touches. For example, when the user is chatting and says "the doors and windows at home seem to be unlocked", only the words "window" and "unlock" appear, but the speech recognition model wrongly recognizes it as opening the window, resulting in an accidental touch and a poor user experience. In addition, since the training data set lacks other voice requests with relatively low frequencies, the speech recognition model's ability to understand these low-frequency voice requests will also be affected, further reducing the user experience.
[0052] Based on the above problems, please refer to Figure 1 , the embodiments of the present application provide a training method for a speech recognition model, and the method includes:
[0053] 01: According to the obtained target training data, determine the target timestamp information associated with the target training data;
[0054] 02: According to the target training data and the target timestamp information, determine the target masked training data;
[0055] 03: Train the speech recognition model according to the target masked training data.
[0056] The embodiments of the present application also provide a server, including a memory and a processor. The training method of the speech recognition model according to the embodiments of the present application can be implemented by the server according to the embodiments of the present application. Specifically, a computer program is stored in the memory, and the processor is configured to determine target timestamp information associated with the target training data according to the obtained target training data. And determine target masked training data according to the target training data and the target timestamp information. The processor is further configured to train the speech recognition model according to the target masked training data.
[0057] The embodiments of the present application also provide a model training device. The training method of the speech recognition model according to the embodiments of the present application can be implemented by the model training device according to the embodiments of the present application. Specifically, the model training device includes a determination module and a training module. The determination module is configured to determine target timestamp information associated with the target training data according to the obtained target training data. The determination module is further configured to determine target masked training data according to the target training data and the target timestamp information. The training module is configured to train the speech recognition model according to the target masked training data.
[0058] Specifically, the speech recognition model refers to a model used to recognize user speech requests in an in-vehicle dialogue system, which can convert the audio signal of the user speech request into corresponding audio features, analyze the semantic information of the user speech request, and finally convert the user speech request into corresponding text information based on the determined audio features and semantic information.
[0059] The target training data refers to the training data selected from a large amount of training data in the training database at present for training the speech recognition model, including target audio data, target text data corresponding to the target audio data, and associated target timestamp information.
[0060] The target timestamp information refers to the audio segmentation markers generated by the Force Alignment technology, which accurately mark the start and end time points (in milliseconds) of each word or phrase in the target training data in the target audio data, and can be used to locate specific parts in the audio, such as {"navigation": [100, 200], "to": [230, 280], "area A": [300, 400], "address a": [410, 500]}. In some embodiments, the target training data may include the target timestamp information, that is, the target timestamp information is a part of the target training data. The target timestamp information may also exist independently in the training database and be associated with the target training data through a pre-constructed "key".
[0061] Target masked training data refers to the target training data after masking processing, which can reduce the overfitting of the speech recognition model to high-frequency speech requests, while retaining data diversity, thereby reducing the mis-touch rate of the speech recognition model and enhancing the generalization ability for low-frequency instructions.
[0062] High-frequency speech requests refer to speech instructions that users often use and have a very high frequency, such as "Open the window", "Close the door", and "Navigate to XX address", etc.
[0063] The server obtains target training data from a large amount of training data in a pre-constructed training database, including target audio data and corresponding target text data, and determines target timestamp information associated with the target training data. For example, the server obtains target audio data "Navigate to Area A Address a" and corresponding text annotation "Navigate to Area A Address a" from a large amount of training data in a pre-constructed training database, and obtains associated target timestamp information {"Navigate": [100, 200], "to": [230, 280], "Area A": [300, 400], "Address a": [410, 500]}.
[0064] Next, the server determines the target masked training data according to the target training data and the target timestamp information. Continuing the above example, the server determines the target masked training data according to the target training data and the target timestamp information, including text data "to Area A Address a" and corresponding audio data "to Area A Address a".
[0065] Finally, the server trains the speech recognition model according to the target masked training data. The speech recognition model will infer the complete text from the target masked training data after masking processing, avoiding over-reliance on high-frequency words. Continuing the above example, the server trains the speech recognition model according to the text data "to Area A Address a" and corresponding audio data "to Area A Address a".
[0066] In summary, in the method and server for training a speech recognition model provided by the embodiments of the present application, the server determines target timestamp information associated with the target training data according to the obtained target training data, where the target training data includes target audio data and target text data corresponding to the target audio data. Then, the server determines target masked training data according to the target training data and the target timestamp information. Finally, the server trains the speech recognition model according to the target masked training data. In this way, through the timestamp-driven collaborative masking technology, the target training data is actively processed during the speech recognition model training stage to determine the target masked training data, which enables the speech recognition model to learn deep context semantics and anti-interference features, reduces the false touch rate, and improves the user experience. Moreover, while reducing false touches, data diversity is retained and the model generalization ability is improved.
[0067] Please refer to Figure 2 , in some embodiments, the target training data is obtained from a pre-constructed training database, and the training database is constructed based on the original training data, where the original training data includes original audio data and original text data corresponding to the original audio data. The training database is constructed through the following steps:
[0068] 06: Based on a preset model, perform time alignment processing on the original audio data and the original text data in units of words to determine timestamp information;
[0069] 07: Construct a training database according to the original training data and the timestamp information.
[0070] In some embodiments, the model training device further includes a database construction module, and the database construction module is further configured to perform time alignment processing on the original audio data and the original text data in units of words based on a preset model to determine timestamp information. And construct a training database according to the original training data and the timestamp information.
[0071] In some embodiments, the processor is further configured to perform time alignment processing on the original audio data and the original text data in units of words based on a preset model to determine timestamp information. And construct a training database according to the original training data and the timestamp information.
[0072] Specifically, the original training data refers to the actual interaction data generated by the user when using the in-vehicle dialogue system, such as voice commands, voice replies, etc., including original audio data and original text data corresponding to the original audio data. In some embodiments, the original text data is obtained by manual annotation to obtain text labels. For example, if the user says "Navigate to Area A, Address a", then the corresponding original audio data is the audio waveform data of this sentence, and the original text data is the text "Navigate to Area A, Address a".
[0073] The preset model refers to the pre-trained Kaldi model before time alignment processing, which can perform time alignment processing on the original audio data and the original text data to obtain the start and end timestamps of each word.
[0074] The Kaldi model is a speech recognition toolkit that focuses on speech-to-text conversion based on hybrid Gaussian-Hidden Markov models and deep neural network technologies, supporting the full process from audio feature extraction to text decoding. The modular architecture of the Kaldi model makes it perform excellently in tasks such as forced alignment. Please refer to Figure 3 Figure shows the process of the Kaldi model performing forced alignment processing on the original audio data and the original text data.
[0075] Forced alignment processing is a speech recognition technology that aligns the original audio data and the corresponding original text data to determine the start and end times of each word in the audio. That is, it "matches" the original audio data and the original text data to find their corresponding relationships.
[0076] Word-based means splitting the original training data into individual words and performing time alignment processing for each word. Since keywords usually correspond to specific functions or commands, masking on a word-by-word basis can more precisely remove high-frequency words that have a greater impact on accidental touches, such as "navigation" and "play", while retaining other words and speech information, thus reducing the accidental touch rate while minimizing the impact on the recognition rate as much as possible.
[0077] Time alignment processing refers to the forced alignment processing of the original audio data and the original text data using the Kaldi model.
[0078] Timestamp information refers to the start time and end time of each word in the audio data output after forced alignment processing. For example, for the user's voice request "Navigate to address B", timestamp information such as "Navigation: 100 - 200ms, To: 230 - 280ms, Address B: 300 - 400ms" can be generated.
[0079] Obtain the original training data, including the original audio data and the corresponding original text data. Continuing the above example, obtain the user's voice request "Navigate to area A address a" and the corresponding annotated text "Navigate to area A address a".
[0080] Next, use a preset model to perform forced alignment on the original audio data and text data to determine the timestamp information for each word at the word level. Continuing with the above example, the timestamp information corresponding to the user's voice request "Navigate to Area A, Address a" is "Navigate: 100 - 200ms; to: 230 - 280ms; Area A: 300 - 400ms; Address a: 410 - 500ms".
[0081] Subsequently, based on the original training data and the timestamp information, the server constructs a training database. Continuing with the above example, based on the user's voice request "Navigate to Area A, Address a", the corresponding labeled text "Navigate to Area A, Address a", and the timestamp information "Navigate: 100 - 200ms; to: 230 - 280ms; Area A: 300 - 400ms; Address a: 410 - 500ms", a training database is constructed.
[0082] In this way, the server performs time alignment processing on the original audio data and the original text data at the word level based on the preset model to determine the timestamp information. Then, the server constructs a training database according to the original training data and the timestamp information. In this way, by pre - constructing the training database, it is possible to avoid performing time alignment operations every time training is carried out, so as to improve the training efficiency.
[0083] Please refer to Figure 4 , in some embodiments, step 07 (constructing a training database according to the original training data and the timestamp information) includes:
[0084] 071: Perform an association process on the current original training data and the timestamp information corresponding to the current original training data to determine the associated training data;
[0085] 072: Construct a training database according to the associated training data.
[0086] In some embodiments, the database construction module is also used to perform an association process on the current original training data and the timestamp information corresponding to the current original training data to determine the associated training data. And construct a training database according to the associated training data.
[0087] In some embodiments, the processor is also used to perform an association process on the current original training data and the timestamp information corresponding to the current original training data to determine the associated training data. And construct a training database according to the associated training data.
[0088] Specifically, the current original training data refers to the original training data that is currently being processed among a large amount of original training data.
[0089] Association processing refers to associating the original training data with its corresponding timestamp information. In some embodiments, the association processing may directly merge the original training data and its corresponding timestamp information to generate a new training data, and use this new training data as the associated training data. In some embodiments, the association processing may also associate the original training data and its corresponding timestamp information by setting a special "key" to form structured data, where the original training data and the timestamp information are stored separately in the training database.
[0090] The server performs association processing on the current original training data and the timestamp information corresponding to the current original training data to determine the associated training data, where the original training data includes the current original training data. Then, the server constructs a training database based on the associated training data.
[0091] In this way, the server performs association processing on the current original training data and the timestamp information corresponding to the current original training data to determine the associated training data, where the original training data includes the current original training data. Then, the server constructs a training database based on the associated training data. In this way, by associating the current original training data and the corresponding timestamp information, an operable and accurate mapping is achieved, laying a foundation for subsequent masking processing.
[0092] Please refer to Figure 5 , in some embodiments, step 02 (determining the target masked training data according to the target training data and the target timestamp information) includes:
[0093] 021: Based on a predefined keyword table, determine the target masked training data according to the target training data and the target timestamp information.
[0094] In some embodiments, the determination module is further configured to determine the target masked training data based on a predefined keyword table according to the target training data and the target timestamp information.
[0095] In some embodiments, the processor is further configured to determine the target masked training data based on a predefined keyword table according to the target training data and the target timestamp information.
[0096] Specifically, the keyword table refers to a predefined list that contains some specific keywords and the keyword masking probabilities corresponding to each keyword. The keyword table can guide the speech recognition model which keywords need to be masked during the training process and the probability of masking.
[0097] Among them, the keyword masking probability is determined based on the occurrence frequency of the keyword in a pre-constructed training database. The higher the occurrence frequency of the keyword, the higher the masking probability. For example, the pre-determined keywords are "navigation, open, window, and backrest". Among them, "navigation" appears 100 times, "open" appears 200 times, "window" appears 110 times, and "backrest" appears 50 times. Then, according to the high and low occurrence frequencies of these keywords, the keyword masking probability of "open" is determined to be 60%, the keyword masking probability of "window" is 35%, the keyword masking probability of "navigation" is 30%, and the keyword masking probability of "backrest" is 10%.
[0098] According to the pre-constructed training database, count the occurrence frequency of each keyword, and determine the keyword table based on the keyword and the occurrence frequency.
[0099] Next, based on the keyword table, the target training data, and the target timestamp information, determine which keywords need to be masked and the size of the masking probability. That is, check whether the current text data contains keywords and whether masking is required.
[0100] In this way, based on the predefined keyword table, the server determines the target masked training data according to the target training data and the target timestamp information. Among them, the keyword table includes predefined keywords and the keyword masking probability corresponding to each keyword. The keyword masking probability is determined based on the occurrence frequency of the keyword in the pre-constructed training database. In this way, by using the predefined keywords, the keyword masking probability corresponding to each keyword, and the target timestamp information to mask the target training data, the mis-touch rate of the speech recognition model can be reduced, so that various speech requests can be accurately recognized and processed, and the robustness of the speech recognition model can be improved.
[0101] Please refer to Figure 6 , in some embodiments, step 021 (based on the predefined keyword table, determine the target masked training data according to the target training data and the target timestamp information) includes:
[0102] 0211: When the current target text data includes keywords, based on the keyword masking probability, determine the target masked training data according to the target training data and the target timestamp information.
[0103] In some embodiments, the determination module is further configured to, when the current target text data includes keywords, based on the keyword masking probability, determine the target masked training data according to the target training data and the target timestamp information.
[0104] In some embodiments, the processor is further configured to, when the current target text data includes a keyword, determine target masked training data based on the keyword masking probability, the target training data, and the target timestamp information.
[0105] Specifically, when it is confirmed that the current target text data includes a keyword, the target training data is masked according to the masking probability of the keyword to determine the target masked training data. Continuing with the above example, if the current text data is "Navigate to area A, address a", after keyword recognition processing, it is determined that the keyword "Navigate" is included. Then, according to the keyword masking probability of 30% corresponding to "Navigate", the target training data is masked to determine the target masked data.
[0106] In this way, when the current target text data includes a keyword, based on the keyword masking probability, the server determines the target masked training data according to the target training data and the target timestamp information. In this way, when the current target text data includes a keyword, based on the keyword masking probability, randomly masking the target training data can reduce the false trigger rate of the speech recognition model, so as to accurately recognize and process various speech requests and improve the robustness of the speech recognition model.
[0107] Please refer to Figure 7 , in some embodiments, step 0211 (when the current target text data includes a keyword, determine target masked training data based on the keyword masking probability, the target training data, and the target timestamp information) includes:
[0108] 02111: Determine the target keyword based on the keyword masking probability;
[0109] 02112: Perform a first masking process on the target keyword in the target text data to determine the temporary target text data;
[0110] 02113: According to the target timestamp information, perform a second masking process on the audio segment corresponding to the keyword in the target audio data to determine the temporary target audio data;
[0111] 02114: Determine the target masked training data according to the temporary target text data and the temporary target audio data.
[0112] In some embodiments, the determining module is further configured to determine a target keyword based on a keyword masking probability, and perform a first masking process on the target keyword in the target text data to determine temporary target text data. The determining module is further configured to perform a second masking process on the audio segment corresponding to the keyword in the target audio data according to the target timestamp information to determine temporary target audio data, and determine target masking training data according to the temporary target text data and the temporary target audio data.
[0113] In some embodiments, the processor is further configured to determine a target keyword based on a keyword masking probability, and perform a first masking process on the target keyword in the target text data to determine temporary target text data. The processor is further configured to perform a second masking process on the audio segment corresponding to the keyword in the target audio data according to the target timestamp information to determine temporary target audio data, and determine target masking training data according to the temporary target text data and the temporary target audio data.
[0114] Specifically, the target keyword refers to the keyword currently being processed. If the target text data includes multiple predefined keywords, then there may be multiple target keywords or none at all, which is determined based on the keyword masking probability.
[0115] Based on the keyword masking probability, the server determines the target keyword, which can be understood as determining whether to process the keyword in the current text data based on the keyword masking probability. If some keywords need to be masked, then these keywords that need to be masked are used as the target keywords. For example, if the target text data is "Close the door and turn on the navigation system, navigate to address C", then the target keyword may be "Close, door, and navigation", or it may be "navigation".
[0116] After determining the target keyword, the server performs a first masking process on the target keyword in the target text data to determine temporary target text data. Continuing with the above example, if the target text data is "Navigate to area A, address a" and the determined target keyword is "navigation", then the server will perform a first masking process on "navigation" in the target text data "Navigate to area A, address a" to determine the temporary target text data "to area A, address a".
[0117] Then, the server performs a second masking process on the audio segment corresponding to the keyword in the target audio data according to the target timestamp information to determine temporary target audio data. Continuing with the above example, if the target text data is "Navigate to area A, address a" and the determined target keyword is "navigation", then the server will perform a second masking process on "navigation" in the target audio data "Navigate to area A, address a" to determine the temporary target audio data "to area A, address a".
[0118] Finally, the server determines the target mask training data based on the temporary target text data and the temporary target audio data. Continuing with the above example, based on the temporary target text data "to address a in area A" and the temporary target audio data "to address a in area A", the target mask training data is determined. The target mask training data may be in the form of "audio file; annotation information: to address a in area A".
[0119] In this way, based on the keyword mask probability, the server determines the target keywords. Then, the server performs a first masking process on the target keywords in the target text data to determine the temporary target text data. Next, the server performs a second masking process on the audio segments corresponding to the keywords in the target audio data according to the target timestamp information to determine the temporary target audio data. Finally, the server determines the target mask training data based on the temporary target text data and the temporary target audio data. In this way, by masking the keywords in the target training data and the audio segments corresponding to the keywords in the target audio data, the mis-touch rate of the speech recognition model can be reduced, so that various speech requests can be accurately recognized and processed, and the robustness of the speech recognition model can be improved.
[0120] Please refer to Figure 8 , in some embodiments, step 02112 (performing a first masking process on the target keywords in the target text data to determine the temporary target text data) includes:
[0121] 021121: Discard the target keywords in the target text data to determine the temporary target text data;
[0122] Step 02113 (performing a second masking process on the audio segments corresponding to the keywords in the target audio data according to the target timestamp information to determine the temporary target audio data) includes:
[0123] 021131: According to the target timestamp information, perform a discard or noise replacement process on the audio segments corresponding to the keywords in the target audio data to determine the temporary target audio data.
[0124] In some embodiments, the determination module is further configured to discard the target keywords in the target text data to determine the temporary target text data. And according to the target timestamp information, perform a discard or noise replacement process on the audio segments corresponding to the keywords in the target audio data to determine the temporary target audio data.
[0125] In some embodiments, the processor is further configured to discard the target keywords in the target text data to determine the temporary target text data. And according to the target timestamp information, perform a discard or noise replacement process on the audio segments corresponding to the keywords in the target audio data to determine the temporary target audio data.
[0126] Specifically, the discard process refers to directly deleting the selected keywords in the target text data to generate corresponding data. For example, directly deleting "navigation" from the target text data "Navigate to area A, address a" to generate temporary target text data "To area A, address a".
[0127] The audio segment refers to finding the audio segment corresponding to the masked keyword according to the timestamp information, such as "navigation" corresponding to 100 - 200ms.
[0128] The audio segment discard process refers to setting the audio segment to zero. For example, if "navigation" corresponds to 100 - 200ms and the corresponding audio segment length is 100ms, then mask the audio segment corresponding to "navigation" as 0.
[0129] The noise replacement process refers to covering the original audio segment with white noise or random noise. For example, adding white noise or random noise to the audio segment of 100 - 200ms corresponding to "navigation".
[0130] Continuing the above example, if the target text data is "Navigate to area A, address a" and "navigation" is the target keyword, the model will mask the word "navigation" to obtain temporary target text data "To area A, address a". At the same time, according to the target timestamp information, mask the audio segment of 100 - 200ms corresponding to "navigation" to obtain temporary target audio data.
[0131] In this way, the server performs a discard process on the target keywords in the target text data to determine the temporary target text data. Then, the server performs a discard or noise replacement process on the audio segments corresponding to the keywords in the target audio data according to the target timestamp information to determine the temporary target audio data. In this way, by discarding the target keywords in the target text data and discarding or replacing the audio segments corresponding to the keywords in the target audio data with noise, the false touch rate of the keywords can be effectively reduced while retaining other information and avoiding too much impact on the recognition rate of the speech recognition model.
[0132] Please refer to Figure 9 , in some embodiments, the method further includes:
[0133] 0212: When the current target text data does not include keywords, or the target keyword is not determined based on the keyword masking probability, train the speech recognition model according to the target training data.
[0134] In some embodiments, the training module is used to train the speech recognition model according to the target training data when the current target text data does not include keywords, or the target keyword is not determined based on the keyword masking probability.
[0135] In some embodiments, the processor is further configured to train the speech recognition model according to the target training data when the current target text data does not include a keyword or when the target keyword cannot be determined based on the keyword masking probability.
[0136] Specifically, if the target text data does not contain a keyword or if the target keyword cannot be determined based on the keyword masking probability, the original target training data is directly used to train the speech recognition model.
[0137] Continuing with the above example, if the target text data in the target training data is "wipe the glass with the windshield wiper", then this data does not contain any keyword in the predefined keyword list, so this target training data is directly used to train the speech recognition model.
[0138] If the target text data is "navigate to area A, address a", and "navigate" is a keyword in the keyword list, but since the masking probability is 30%, based on this masking probability, it is determined that "navigate" is not masked, that is, the target keyword is not determined. In this case, this data is also directly used to train the speech recognition model.
[0139] In this way, when the current target text data does not include a keyword or when the target keyword cannot be determined based on the keyword masking probability, the server trains the speech recognition model according to the target training data. In this way, when the current target text data does not include a keyword or when the target keyword cannot be determined based on the keyword masking probability, by directly using the target training data to train the speech recognition model, the robustness of the model can be improved so that it can handle various scenarios and speech variations.
[0140] This application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for training the speech recognition model as described above are implemented.
[0141] It can be understood that the computer program includes computer program code. The computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable storage medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution media, etc.
[0142] In the description of this specification, the descriptions referring to terms such as "specifically", "further", "specially", "understandably", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms are not necessarily intended to refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0143] Any process or method description shown in the flowchart or described in other ways herein can be understood to represent a module, segment, or portion of executable instructions including one or more steps for implementing a specific logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, not in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0144] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for training a speech recognition model, characterized in that: The method comprises: Determining target timestamp information associated with the target training data according to the acquired target training data, wherein the target training data includes target audio data and target text data corresponding to the target audio data; Determining target mask training data according to the target training data and the target timestamp information; The speech recognition model is trained according to the target mask training data.
2. The method according to claim 1, characterized in that The target training data is obtained from a pre-constructed training database, the training database is constructed based on original training data, the original training data includes original audio data and original text data corresponding to the original audio data, and the training database is constructed by the following steps: Based on a preset model, time alignment processing is performed on the original audio data and the original text data in units of words to determine timestamp information; The training database is constructed according to the original training data and the timestamp information.
3. The method according to claim 2, characterized in that The step of constructing the training database according to the original training data and the timestamp information includes: Performing association processing on current original training data and timestamp information corresponding to the current original training data to determine associated training data, wherein the original training data includes the current original training data; The training database is constructed according to the associated training data.
4. The method according to claim 1, characterized in that: The step of determining target mask training data according to the target training data and the target timestamp information includes: Based on a predefined keyword table, the target mask training data is determined according to the target training data and the target timestamp information, wherein the keyword table includes predefined keywords and a keyword mask probability corresponding to each of the keywords, and the keyword mask probability is determined based on the frequency of occurrence of the keyword in a pre-constructed training database.
5. The method according to claim 4, characterized in that The determining the target mask training data based on the predefined keyword table and according to the target training data and the target timestamp information includes: In the case that the current target text data includes the keyword, the target mask training data is determined based on the keyword mask probability, according to the target training data and the target timestamp information.
6. The method according to claim 5, characterized in that In the case where the current target text data includes the keyword, determining the target mask training data based on the keyword mask probability, according to the target training data and the target timestamp information, includes: Determining a target keyword based on the keyword mask probability; Performing a first masking process on the target keyword in the target text data to determine temporary target text data; According to the target timestamp information, performing a second masking process on the audio segment corresponding to the keyword in the target audio data to determine temporary target audio data; The target mask training data is determined according to the temporary target text data and the temporary target audio data.
7. The method according to claim 6, characterized in that The step of performing a first masking process on the target keyword in the target text data to determine temporary target text data includes: discarding the target keyword in the target text data to determine temporary target text data; The step of performing a second mask process on the audio segment corresponding to the keyword in the target audio data according to the target timestamp information to determine the temporary target audio data includes: According to the target timestamp information, the audio segment corresponding to the keyword in the target audio data is discarded or replaced with noise to determine temporary target audio data.
8. The method according to claim 5, characterized in that The method further comprises: When the current target text data does not include the keyword, or the target keyword is not determined based on the keyword mask probability, the speech recognition model is trained according to the target training data.
9. A server, characterized in that: The server includes a processor and a memory, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.