Low-resource scene hot word enhancement method and device based on synthetic data
By extracting the hot word list of the target scene corpus and using a large language model to generate sample text and audio, the problems of low accuracy of hot word recognition and high cost of training data generation in the prior art are solved, and efficient and accurate hot word recognition is achieved.
Patent Information
- Application Number
- CN202510227505.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-27
AI Technical Summary
The existing speech recognition system has low accuracy in hot word recognition in dynamic scenarios, and generating high-quality hot word corpus and audio training data requires a lot of manual annotation and recording, which is costly and inefficient.
By obtaining the target scene corpus, extracting the scene hot word list, generating a sample text set based on the first large language model, and synthesizing the corresponding sample audio set, using this as training data to train the initial speech recognition model to obtain the final hot word recognition model.
It realizes efficient and high-quality generation of hot word audio in low-resource scenarios, significantly improving the accuracy of hot word recognition in specific scenarios of hot word recognition models.
Smart Images

Figure CN120220657A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and particularly to a method and device for hot word enhancement in low-resource scenarios based on synthetic data. Background Art
[0002] With the rapid development of speech recognition technology, automatic speech recognition systems have been widely applied in fields such as intelligent voice assistants, speech input methods, intelligent customer service, meeting records, and speech search. Existing speech recognition systems still face many challenges in hot word recognition in dynamic scenarios, such as specific personal names, place names, brand names, or domain-specific terms.
[0003] However, due to the low frequency or scene specificity of hot words, existing ASR (Automatic Speech Recognition System) models cannot effectively recognize them, and generating high-quality hot word corpora and audio training data for specific scenarios usually requires a large amount of manual annotation and recording work, with high training costs and low efficiency, resulting in a significant reduction in the accuracy of hot word recognition in specific scenarios. Summary of the Invention
[0004] The present invention provides a method and device for hot word enhancement in low-resource scenarios based on synthetic data to solve the defect of a significant reduction in the accuracy of hot word recognition in specific scenarios in the prior art.
[0005] The present invention provides a method for hot word enhancement in low-resource scenarios based on synthetic data, including: Obtaining a target scenario corpus; Extracting a list of scene hot words from the target scenario corpus, and defining a generation task based on the knowledge classification and / or target scenario of the list of scene hot words; Based on a first large language model, applying the list of scene hot words and the generation task to generate a sample text set corresponding to the list of scene hot words; Synthesizing a sample audio set corresponding to the sample text set, and using the sample audio set as training data and the sample text set as labels to train an initial speech recognition model to obtain a final hot word recognition model.
[0006] According to the method for hot word enhancement in low-resource scenarios based on synthetic data provided by the present invention, the step of, based on a first large language model, applying the list of scene hot words and the generation task to generate a sample text set corresponding to the list of scene hot words includes: Based on the generation task and the style of the prompt sentence pattern, constructing a set of generation prompt texts that conform to a preset structure template; Based on the first large language model, apply the scenario hot word list and the generated prompt text set to generate the sample text set; The first large language model is trained based on a general first large language model and the target scenario corpus.
[0007] According to a low-resource scenario hot word enhancement method based on synthetic data provided by the present invention, the step of generating the sample text set by applying the scenario hot word list and the generated prompt text set based on the first large language model includes: Based on the first large language model, apply the scenario hot word list and the generated prompt text set to generate an initial sample text set; Calculate the text similarity between any two initial sample texts in the initial sample text set; Eliminate any one of the two initial sample texts with a text similarity higher than the preset text similarity threshold to obtain a filtered sample text set; Match each filtered sample text in the filtered sample text set with the scenario hot word list, and select the filtered sample text with a successful match as the sample text set.
[0008] According to a low-resource scenario hot word enhancement method based on synthetic data provided by the present invention, the step of training the initial speech recognition model with the sample audio set as the training data and the sample text set as the label to obtain the final hot word recognition model includes: Obtain an initial speech recognition model, where the initial speech recognition model includes an audio encoder and a second large language model; Based on the audio encoder, encode any sample audio in the sample audio set to generate an audio feature sequence; Based on the audio feature sequence, the second large language model, and the recognition prompt text, generate the output recognition text of the any sample audio; the recognition prompt text is constructed based on the scenario hot word list of the any sample audio; Based on the output recognition text and the sample text corresponding to the any sample audio, calculate the recognition loss, and adjust the model parameters of the initial speech recognition model based on the recognition loss to obtain the hot word recognition model.
[0009] According to a low-resource scenario hot word enhancement method based on synthetic data provided by the present invention, the step of generating the output recognition text of the any sample audio based on the audio feature sequence, the second large language model, and the recognition prompt text includes: Based on the second large language model, decode the audio feature sequence to obtain a decoded candidate sequence; Enhance the transcription priority of the decoding candidate sequences within the range of the scene hot word list in the recognition prompt text to obtain an adjusted decoding sequence; Based on the second large language model and the adjusted decoding sequence, perform text transcription to generate the output recognition text of any of the sample audios.
[0010] According to a low-resource scene hot word enhancement method based on synthetic data provided by the present invention, the synthesizing the sample audio set corresponding to the sample text set includes: Based on the sample text set, the sound attribute settings corresponding to the target scene, and the emotional style, generate a personalized sample audio set; Inject scene noise into the personalized sample audio set to obtain the sample audio set.
[0011] According to a low-resource scene hot word enhancement method based on synthetic data provided by the present invention, the injecting scene noise into the personalized sample audio set to obtain the sample audio set includes: Inject scene noise into the personalized sample audio set to obtain a noise sample audio set; Extract the hot word audio segments in each noise sample audio in the noise sample audio set; Perform pronunciation adjustment and / or speech rate adjustment on the hot word audio segments to obtain enhanced audio segments; Based on the enhanced audio segments, respectively replace the hot word audio segments in each noise sample audio corresponding to the enhanced audio segments to obtain the sample audio set.
[0012] According to a low-resource scene hot word enhancement method based on synthetic data provided by the present invention, the extracting the scene hot word list of the target scene corpus includes: Construct an extraction prompt text based on the target scene to which the target scene corpus belongs; Based on the first large language model, apply the extraction prompt text and the target scene corpus to extract the scene hot word list.
[0013] The present invention also provides a low-resource scene hot word enhancement device based on synthetic data, including: An acquisition unit that acquires a target scene corpus; An extraction unit that extracts the scene hot word list of the target scene corpus and defines a generation task based on the knowledge classification and / or target scene of the scene hot word list; A sample generation unit that, based on the first large language model, applies the scene hot word list and the generation task to generate a sample text set corresponding to the scene hot word list; A training unit synthesizes a sample audio set corresponding to the sample text set, and uses the sample audio set as training data and the sample text set as labels to train an initial speech recognition model to obtain a final hotword recognition model.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for enhancing hotwords in a low-resource scenario based on synthetic data as described in any one of the above.
[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for enhancing hotwords in a low-resource scenario based on synthetic data as described in any one of the above.
[0016] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method for enhancing hotwords in a low-resource scenario based on synthetic data as described in any one of the above.
[0017] The method and device for enhancing hotwords in a low-resource scenario based on synthetic data provided by the present invention extract a list of scenario hotwords from the target scenario corpus, define a generation task based on the knowledge classification and / or the target scenario of the scenario hotword list; generate a sample text set based on the scenario hotword list and the generation task of the first large language model; synthesize a sample audio set corresponding to the sample text set, and use the sample audio set as training data and the sample text set as labels to train an initial speech recognition model to obtain a final hotword recognition model, realizing efficient, high-quality, and diverse sample data generation in the target scenario, and greatly improving the hotword recognition accuracy of the hotword recognition model in the target scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 is a flowchart of the method for enhancing hotwords in a low-resource scenario based on synthetic data provided by the present invention; Figure 2 is a structural diagram of the device for enhancing hotwords in a low-resource scenario based on synthetic data provided by the present invention; Figure 3 is a structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the protection scope of the present invention.
[0021] It should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features.
[0022] In view of the above problems, the present invention provides a method for enhancing hot words in a low-resource scenario based on synthetic data to achieve high-quality and efficient hot word audio generation in a specific scenario, thereby improving the accuracy of hot word recognition of the hot word recognition model in a specific scenario. Figure 1 is a schematic flowchart of the method for enhancing hot words in a low-resource scenario based on synthetic data provided by the present invention, as Figure 1 shown, the method includes: Step 110, obtaining a target scenario corpus; Specifically, first, text data related to the target scenario can be collected from relevant websites or databases through web crawling technology, or initial scenario corpus can be obtained through manual collation. Among them, the text data related to the target scenario can be technical documents, industry reports, etc. Then, the initial scenario corpus can be subjected to text preprocessing, such as word segmentation, stop word removal, and stemming, to clean text noise and obtain a target scenario corpus with higher quality. It should be noted that the target scenario corpus should cover common terms, professional terms, etc. in the target scenario. The target scenario here can be related conference scenarios in the computer field, medical report meeting scenarios, etc.
[0023] Step 120, extracting a list of scenario hot words from the target scenario corpus, and defining a generation task based on the knowledge classification and / or target scenario of the list of scenario hot words; Specifically, first, the obtained target scenario corpus and extraction prompt text can be input into a first large language model, and the first large language model extracts a list of scenario hot words from the target scenario corpus according to the instructions of the extraction prompt text. Or, text mining techniques, such as algorithms like TF-IDF and TextRank, can be used to extract keywords from the target scenario corpus to obtain a list of scenario hot words.
[0024] Then, tasks can be generated through the knowledge classification of the extracted list of scene hot words and / or the definition of the target scene. For example, by inputting the list of scene hot words into the first large language model, the knowledge classification of the list of scene hot words and / or the target scene can be output by the first large language model. In this embodiment, preferably, the generation task is jointly defined through the knowledge classification of the list of scene hot words and the target scene.
[0025] Among them, the classification of the list of scene hot words refers to the knowledge domain attribute to which the list of scene hot words belongs. For example, it belongs to the medical field, the agricultural field, the economic field, etc. The knowledge classification can be used to reflect the relevant background knowledge of the list of scene hot words and clarify the theme of the sample text set to be generated. In addition, the target scene refers to the scene where the list of scene hot words is actually applied, such as an economic meeting, etc. Therefore, the generation task here is used to clarify the specific requirements of the sample text set corresponding to the list of scene hot words at the levels of knowledge classification and target scene.
[0026] Step 130, based on the first large language model, apply the list of scene hot words and the generation task to generate the sample text set corresponding to the list of scene hot words; Here, the first large language model can be any mature model with powerful language generation capabilities. For example, it can be large language models such as Deep Seek and Wenxin Yiyan.
[0027] It can be understood that the extracted list of scene hot words only contains the hot words in the target scene. The hot words can be word segments, or phrases or sentences composed of word segments. In the actual application of hot word recognition, the hot words are usually embedded in the conversations of the speaker in the target scene. Therefore, it is necessary to construct a sample text set that contains hot words and is in the target scene.
[0028] Specifically, by inputting the extracted list of scene words and the generation task into the first large language model, the first large language model generates the sample text corresponding to the list of scene hot words according to the instructions of the generation task, and constructs the generated sample text into a sample text set, then the text containing the target scene hot words in the target scene is obtained.
[0029] It should be noted that compared with the sample text constructed manually, the sample text set generated by the first large language model according to the generation task defined has higher text quality and more diverse texts. Therefore, the performance of model training based on this sample text set is better, and the accuracy of hot word recognition is higher. In addition, a corresponding sample audio set can be generated from the sample text set, and the sample text set can be used as the label of the sample audio set, eliminating the need for manual annotation of the sample audio set, which greatly improves the efficiency of constructing sample data.
[0030] Step 140: Synthesize the sample audio set corresponding to the sample text set. Use the sample audio set as training data and the sample text set as labels to train the initial speech recognition model to obtain the final hotword recognition model.
[0031] Specifically, through text-to-speech technology, the sample text set can be converted into the corresponding sample audio set. For example, each sample text in the sample text set can be separately converted into audio to obtain the sample audio corresponding to each sample text, and all the synthesized sample audios are constructed to form the sample audio set. In the model training stage, the sample audio set can be used as training data, and the sample text corresponding to each sample audio in the sample audio set is used as a label to train the initial speech recognition model to achieve iterative optimization of the model parameters and obtain the final hotword recognition model.
[0032] The method provided by the embodiments of the present invention extracts the list of scene hotwords in the target scene corpus, defines the generation task based on the knowledge classification and / or target scene of the scene hotword list; generates the sample text set based on the first large language model, the application scene hotword list and the generation task; synthesizes the sample audio set corresponding to the sample text set, uses the sample audio set as training data and the sample text set as labels to train the initial speech recognition model to obtain the final hotword recognition model, realizes efficient, high-quality and diverse sample data generation in the target scene, and greatly improves the hotword recognition accuracy of the hotword recognition model in the target scene.
[0033] Based on any of the above embodiments, step 130 includes: Construct a set of generation prompt texts that conform to the preset structure template based on the generation task and the prompt sentence style; Based on the first large language model, apply the scene hotword list and the set of generation prompt texts to generate the sample text set; The first large language model is trained based on a general first large language model and the target scene corpus.
[0034] Here, the generation task can be used to reflect the specific requirements for generating each prompt text, including information such as the theme, style, application scene, keywords of the text. The prompt sentence style is used to reflect the expression way of generating the prompt text, such as constructing diverse prompt texts like "question - answer" dialogues, scenario descriptions, command requests, etc. The preset structure template here refers to the preset framework and format for constructing the set of generation prompt texts, including the structure and sentence pattern of the text, and the preset structure template can ensure the logic and domain relevance of the generated sample text, such as "[Hotword] applications include...".
[0035] In addition, to cover the requirements of specific scenarios or fields, the first large language model here can be obtained by fine-tuning and training the general first large language model based on the target scenario corpus, so that the trained first large language model can generate higher-quality sample texts. For example, the general first large language model can be fine-tuned or a domain-specific model can be used, such as the general first large language model in the medical, legal and other fields, to generate more accurate term contexts and text corpora, that is, to generate more accurate and target-field-specific prompt texts.
[0036] Specifically, by combining the generation task and the prompt sentence style, a set of generation prompt texts that conform to the preset structure template can be constructed to guide the first large language model to generate sample texts with different grammatical structures, expressions and contexts. For example, by combining the generation task and the prompt sentence style, a series of preset structure templates can be designed. These templates can include parts such as the beginning, the middle content, and the end, and each part has corresponding sentence patterns and vocabulary requirements. According to the preset structure template, specific prompt content is filled in to form a set of generation prompt texts, and the prompt texts in the set of generation prompt texts should be able to guide the first large language model to generate sample texts that meet the content requirements and text format requirements.
[0037] Then, the scene hot word list and any one of the generation prompt texts in the set of generation prompt texts can be input into the first large language model. The first large language model generates sample texts according to the instructions of any one of the generation prompt texts, and constructs all the generated prompt texts to obtain a set of prompt texts.
[0038] It should be noted that by designing a generation strategy that combines structured templates and open-ended prompts, a set of generation prompt texts that conform to the preset structure template is constructed. Among them, the preset structure template can ensure the logic of the generated prompt texts, and the open-ended prompts realized through the generation task and the prompt sentence style can stimulate the first large language model to generate freely, enriching the meaning and expression forms of the texts.
[0039] To further improve the data quality of the sample text set, based on any of the above embodiments, generating the sample text set based on the first large language model, applying the scene hot word list and the set of generation prompt texts includes: Based on the first large language model, applying the scene hot word list and the set of generation prompt texts to generate an initial sample text set; Calculate the text similarity between any two initial sample texts in the initial sample text set; Remove any one of the two initial sample texts with a text similarity higher than the preset text similarity threshold to obtain a filtered sample text set; For each screened sample text in the screened sample text set, match it with the scene hot word list respectively, and select the screened sample text with successful matching as the sample text set.
[0040] Specifically, first, generate an initial sample text set through the first large language model according to the instructions of the scene hot word list and the generated prompt text set. Then, semantic filtering and hot word verification of the initial sample text set are implemented to further improve the text quality of the sample text set.
[0041] In detail, the text similarity between any two initial sample texts in the initial sample text set can be calculated. For example, it can be calculated through the cosine similarity algorithm or other semantic similarity algorithms. Then, any one of the two initial sample texts with text similarity higher than the preset text similarity threshold is removed to obtain the screened sample text set. Thus, the screened sample text set does not contain duplicate text content.
[0042] Furthermore, for each screened sample text in the screened sample text set, match it with the scene hot word list respectively, and select the screened sample text with successful matching as the sample text set. Thus, each sample text in the sample text set contains at least one hot word in the scene hot word list.
[0043] The method provided by the embodiment of the present invention performs semantic filtering and hot word verification on the initial sample text set, ensuring that the generated corpus not only conforms to the target scene, but also can cover various grammatical structures, context scenarios and expression styles, providing high-quality text input for subsequent synthetic audio training.
[0044] Based on any of the above embodiments, in step 140, using the sample audio set as training data and the sample text set as labels, training the initial speech recognition model to obtain the final hot word recognition model, including: Obtain an initial speech recognition model, where the initial speech recognition model includes an audio encoder and a second large language model; Based on the audio encoder, encode any sample audio in the sample audio set to generate an audio feature sequence; Based on the audio feature sequence, the second large language model and the recognition prompt text, generate the output recognition text of the any sample audio; the recognition prompt text is constructed based on the scene hot word list of the any sample audio; Based on the output recognition text and the sample text corresponding to the any sample audio, calculate the recognition loss, and adjust the model parameters of the initial speech recognition model based on the recognition loss to obtain the hot word recognition model.
[0045] Specifically, first, obtain an initial speech recognition model. The initial speech recognition model includes an audio encoder and a second large language model. Here, the initial speech recognition model can be the Qwen-2-Audio model. The audio encoder here can be Whisper-large-v3.
[0046] Then, through the audio encoder, any sample audio in the sample audio set can be encoded to generate an audio feature sequence. It can be understood that taking the qwen2-audio-instruct model as the initial speech recognition model as an example, by using Whisper-large-v3 as the audio feature extractor, the sample audio can be converted into a feature representation that can be processed by the second large language model, that is, an audio feature sequence. The cross-modal adaptation layer performs feature alignment and conversion between the audio encoder and the second large language model to ensure that the audio feature sequence can be used for text generation tasks. Qwen-7B can be used as the core of the second large language model. Using the audio feature sequence of the audio encoder and combining the input of the recognition prompt text, the audio-to-text task is realized, that is, the hot word recognition text is obtained.
[0047] However, it should be noted that the existing speech recognition systems have poor adaptability to specific scenarios and are difficult to dynamically adjust the hot word recognition performance according to the changes in the scenarios. Therefore, the hot word recognition performance in the target scenario can be improved through the following steps to improve the accuracy of the hot word recognition text in the output recognition text.
[0048] Specifically, the audio feature sequence and the recognition prompt text corresponding to the sample audio can be input into the second large language model. Through the second large language model, according to the prompt of the recognition prompt text, the audio feature sequence is decoded based on the tokenizer and word embedding layer of the second large language model to generate the output recognition text of the sample audio. For example, it can be to enhance the priority of the audio feature sequence within the range of the scene hot word list in the decoding candidate sequence, so that this audio feature sequence is more preferentially decoded into any hot word in the scene hot word list, thereby improving the accuracy of hot word recognition.
[0049] It should be noted that the recognition prompt text here is constructed based on the scene hot word list corresponding to the sample audio and the recognition task prompt word, so as to indicate that the second large language model should give priority to considering the decoding candidate sequence in the scene hot word list during decoding. The recognition task prompt word here can be, for example, "Convert the input audio feature sequence into text".
[0050] Finally, the recognition loss can be calculated through the output recognition text and the sample text corresponding to the sample audio. Based on the recognition loss, the model parameters of the initial speech recognition model are adjusted to obtain the final hot word recognition model.
[0051] It should be noted that, in view of the scene characteristics of background noise and pronunciation variants, mixed scene data is used for training during the fine-tuning process to improve the robustness of the model in complex environments, such as noisy backgrounds and variable speech rates. During training, the CTC (Connectionist Temporal Classification) loss function can be used to optimize the character-by-character alignment of speech-to-text mapping, and at the same time, combined with the learning rate adjustment and annealing strategy, to balance the convergence speed and stability of the model.
[0052] In addition, after fine-tuning, scene-based metrics, such as WER (Word Error Rate) and SER (Service) evaluation metrics, can be used to evaluate the model to ensure a significant improvement in its recognition performance in the target scene. Finally, through the in-depth optimization of the generated hot word text and synthesized audio, the fine-tuned initial speech recognition model can accurately adapt to specific scenes, improve the recognition ability for key terms, diverse pronunciations, and complex backgrounds, and provide high-quality support for practical applications.
[0053] Based on any of the above embodiments, generating the output recognition text of any of the sample audios based on the audio feature sequence, the second large language model, and the recognition prompt text includes: Decoding the audio feature sequence based on the second large language model to obtain a decoded candidate sequence; Enhancing the transcription priority of the decoded candidate sequence within the range of the scene hot word list in the recognition prompt text to obtain an adjusted decoded sequence; Performing text transcription based on the second large language model and the adjusted decoded sequence to generate the output recognition text of any of the sample audios.
[0054] Specifically, for decoding any sample audio to obtain the output recognition text, the audio feature sequence and the recognition prompt text can be input into the second large language model. Then, the second large language model decodes the audio feature sequence to obtain a decoded candidate sequence. It can be understood that there are multiple decoded candidate sequences corresponding to a single audio feature sequence, and the probabilities of each decoded candidate sequence, that is, the transcription priorities, are different.
[0055] Furthermore, the transcription priority of the decoded candidate sequence within the range of the scene hot word list in the recognition prompt text is enhanced to obtain an adjusted decoded sequence. Finally, the second large language model performs text transcription on the adjusted decoded sequence to generate the output recognition text of this sample audio.
[0056] For example, the audio feature sequence and the recognition prompt text can be input into the second large language model. In the beam search stage of the decoder, dynamic hot word enhancement will give priority to the candidate paths containing hot words, that is, the decoded candidate sequences. For example, if "electrocardiogram" is a hot word, when decoding and generating candidate sequences, the candidate sequences containing "electrocardiogram" will be given a higher priority, such as increasing the probability distribution of this candidate sequence, so that it is more likely to appear in the final transcription, ensuring its accuracy in the recognition result, and at the same time enhancing the robustness and generalization performance of the model in complex environments. Finally, the adjusted decoded sequence can be text-transcribed by the second large language model to generate the output recognition text of the sample audio.
[0057] The method provided by the embodiment of the present invention enhances the transcription priority of the audio feature sequence within the range of the scene hot word list of any sample audio to obtain an adjusted decoded sequence, realizes the accurate recognition of hot word voices in a specific scene, and further improves the hot word recognition performance.
[0058] Based on any of the above embodiments, in step 140, synthesizing the sample audio set corresponding to the sample text set includes: Generating a personalized sample audio set based on the sample text set, the voice attribute settings corresponding to the target scene, and the emotional style; Injecting scene noise into the personalized sample audio set to obtain the sample audio set.
[0059] Here, the voice attribute settings include attributes such as the timbre and intonation of the speaker, such as the voice attributes of male, female, old person, and child. The emotional style here can be styles such as formal, relaxed, pleasant, and tense. In addition, the scene noise here can be noise such as conference background noise, e-commerce customer service ambient sound, and medical device sound.
[0060] Specifically, first, the voice attribute settings and the emotional style can be adjusted according to the target user group or application scenario. Then, each sample text in the sample text set is converted into audio according to the selected voice attribute settings and emotional style, and the audio after all sample texts are converted is used as the personalized sample audio set. For example, through the multi-speaker feature of the VITS model, diverse voice samples can be simulated with different timbres and intonations (such as male, female, old person, child), and the emotional expression can be adjusted (such as formal, relaxed, pleasant, tense) to achieve highly natural and personalized speech synthesis. Or, the sample text can be combined into the corresponding sample audio through the CosyVoice model.
[0061] It can be understood that the personalized sample audio set contains diverse text contents, voice attributes, and emotional styles that conform to the target scene, providing rich materials for subsequent steps.
[0062] Next, according to the target scenario to which the sample text set belongs, the possible scenario noises in the target scenario can be selected, such as conference background noise, e-commerce customer service ambient sound, medical device sound, etc. Inject the scenario noise into the personalized sample audio set to obtain the sample audio set. For example, the scenario noise can be dynamically added to the personalized sample audio in the personalized sample audio set to enhance the authenticity and applicability of the audio. Thus, each audio in the obtained sample audio set has the specified text content, sound attributes, emotional styles, and scenario noises.
[0063] It should be noted that through the text-to-speech conversion technology, the sample text is converted into high-quality synthetic sample audio, and the audio includes multiple timbres, multiple intonations, scenario-based noises, and pronunciation variants, constructing a diverse hotword training data set and reducing the manual recording and annotation costs.
[0064] The method provided by the embodiments of the present invention generates a personalized sample audio set through the sample text set, the sound attribute settings corresponding to the target scenario, and the emotional styles, injects the scenario noise into the personalized sample audio set, and obtains the sample audio close to the actual application scenario, improving the authenticity and immersion of the audio. Moreover, by injecting the scenario noise, the audio conditions in the actual application environment can be simulated, which helps to evaluate the performance of the model in a complex environment, and further improves the hotword recognition ability of the speech recognition model under complex conditions such as a noisy environment, multiple timbres, or intonation changes.
[0065] Based on any of the above embodiments, the injecting the scenario noise into the personalized sample audio set to obtain the sample audio set includes: Injecting the scenario noise into the personalized sample audio set to obtain a noise sample audio set; Extracting the hotword audio segments in each noise sample audio in the noise sample audio set; Performing pronunciation adjustment and / or speech rate adjustment on the hotword audio segments to obtain enhanced audio segments; Based on the enhanced audio segments, respectively replacing the hotword audio segments in each noise sample audio corresponding to the enhanced audio segments to obtain the sample audio set.
[0066] Specifically, first, inject the scenario noise into the personalized sample audio set to obtain a noise sample audio set. Then, extract the hotword audio segments in each noise sample audio in the noise sample audio set, that is, extract the audio segments containing hotwords.
[0067] Further, perform pronunciation adjustment and / or speech rate adjustment on the hot word audio segment to obtain an enhanced audio segment. Specifically, for the pronunciation variants of hot words, such as dialects, accents, and speech rate changes, generate diverse audio samples to improve the robustness of the speech recognition model to different pronunciation scenarios. For example, by leveraging the efficient inference ability of VITS, a large number of synthetic audio samples can be quickly generated in batches through an automated process, greatly reducing manual intervention and time costs, and finally generating a sample audio set with rich diversity and realism.
[0068] Finally, based on the enhanced audio segment, replace the hot word audio segment in each noise sample audio corresponding to the enhanced audio segment to obtain a sample audio set, further improving the recognition accuracy of the speech recognition model for hot word recognition and enhancing the robustness and accuracy of the speech recognition model in actual speech recognition applications.
[0069] Based on any of the above embodiments, in step 120, extracting the scene hot word list of the target scene corpus includes: Construct an extraction prompt text based on the target scene to which the target scene corpus belongs; Based on the first large language model, apply the extraction prompt text and the target scene corpus to extract the scene hot word list.
[0070] Specifically, an extraction prompt text can be constructed according to the target scene to which the target scene corpus belongs. Then, by inputting the extraction prompt text and the target scene corpus into the first large language model, the first large language model extracts hot words from the target scene corpus according to the prompt of the extraction prompt text, and uses the extracted hot words as the scene hot word list.
[0071] It should be noted that through the comprehensive analysis of the scene corpus, user input, and domain knowledge, a hot word list highly relevant to the target scene is dynamically extracted to ensure the adaptability and practicality of the hot words in the scene hot word list to the specific scene.
[0072] Based on any of the above embodiments, Figure 2 is a schematic structural diagram of a low-resource scene hot word enhancement device based on synthetic data provided by the present invention, as Figure 2 shown, the device includes: An acquisition unit 210, which acquires a target scene corpus; An extraction unit 220, which extracts the scene hot word list of the target scene corpus and defines a generation task based on the knowledge classification and / or target scene of the scene hot word list; A sample generation unit 230, which generates a sample text set corresponding to the scene hot word list based on the first large language model, applying the scene hot word list and the generation task; The training unit 240 synthesizes a sample audio set corresponding to the sample text set, and uses the sample audio set as training data and the sample text set as labels to train an initial speech recognition model to obtain a final hotword recognition model.
[0073] The device provided by the embodiment of the present invention extracts the scene hotword list of the target scene corpus, defines a generation task based on the knowledge classification and / or target scene of the scene hotword list; generates a sample text set based on the first large language model, the scene hotword list and the generation task; synthesizes a sample audio set corresponding to the sample text set, and uses the sample audio set as training data and the sample text set as labels to train an initial speech recognition model to obtain a final hotword recognition model, realizing efficient, high-quality and diverse sample data generation in the target scene, and greatly improving the hotword recognition accuracy of the hotword recognition model in the target scene.
[0074] Based on any of the above embodiments, the sample generation unit is specifically configured to: Construct a generation prompt text set that conforms to a preset structure template based on the generation task and the prompt sentence style; Based on the first large language model, apply the scene hotword list and the generation prompt text set to generate the sample text set; The first large language model is trained based on a general first large language model and the target scene corpus.
[0075] Based on any of the above embodiments, the sample generation unit is further specifically configured to: Based on the first large language model, apply the scene hotword list and the generation prompt text set to generate an initial sample text set; Calculate the text similarity between any two initial sample texts in the initial sample text set; Remove any one of the two initial sample texts with a text similarity higher than a preset text similarity threshold to obtain a filtered sample text set; Match each filtered sample text in the filtered sample text set with the scene hotword list, and select the filtered sample text that matches successfully as the sample text set.
[0076] Based on any of the above embodiments, the training unit is specifically configured to: Obtain an initial speech recognition model, where the initial speech recognition model includes an audio encoder and a second large language model; Based on the audio encoder, encode any sample audio in the sample audio set to generate an audio feature sequence; Generate the output recognition text for any of the sample audios based on the audio feature sequence, the second large language model, and the recognition prompt text; the recognition prompt text is constructed based on the scene hot word list of any of the sample audios; Calculate a recognition loss based on the output recognition text and the sample text corresponding to any of the sample audios, and adjust the model parameters of the initial speech recognition model based on the recognition loss to obtain the hot word recognition model.
[0077] Based on any of the above embodiments, the training unit is further specifically configured to: Decode the audio feature sequence based on the second large language model to obtain a decoded candidate sequence; Enhance the transcription priority of the decoded candidate sequence within the range of the scene hot word list in the recognition prompt text to obtain an adjusted decoded sequence; Perform text transcription based on the second large language model and the adjusted decoded sequence to generate the output recognition text for any of the sample audios.
[0078] Based on any of the above embodiments, the training unit is further specifically configured to: Generate a personalized sample audio set based on a sample text set, the sound attribute settings corresponding to the target scene, and the emotional style; Inject scene noise into the personalized sample audio set to obtain the sample audio set.
[0079] Based on any of the above embodiments, the training unit is further specifically configured to: Inject scene noise into the personalized sample audio set to obtain a noise sample audio set; Extract the hot word audio segments from each noise sample audio in the noise sample audio set; Perform pronunciation adjustment and / or speech rate adjustment on the hot word audio segments to obtain enhanced audio segments; Based on the enhanced audio segments, replace the hot word audio segments in each noise sample audio corresponding to the enhanced audio segments to obtain the sample audio set.
[0080] Based on any of the above embodiments, the extraction unit is specifically configured to: Construct an extraction prompt text based on the target scene to which the target scene corpus belongs; Based on the first large language model, apply the extraction prompt text and the target scene corpus to extract the scene hot word list.
[0081] Figure 3 Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 3As shown in the figure, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communications interface 320, and the memory 330 complete their mutual communication through the communication bus 340. The processor 310 may call the logical instructions in the memory 330 to execute a method for enhancing hot words in a low-resource scenario based on synthetic data. The method includes: obtaining a target scenario corpus; extracting a list of scenario hot words from the target scenario corpus, and defining a generation task based on the knowledge classification and / or the target scenario of the list of scenario hot words; based on a first large language model, applying the list of scenario hot words and the generation task to generate a sample text set corresponding to the list of scenario hot words; synthesizing a sample audio set corresponding to the sample text set, and using the sample audio set as training data and the sample text set as labels to train an initial speech recognition model to obtain a final hot word recognition model.
[0082] In addition, when the logical instructions in the above-mentioned memory 330 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0083] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for enhancing hot words in a low-resource scenario based on synthetic data provided by the above-mentioned various methods. The method includes: obtaining a target scenario corpus; extracting a list of scenario hot words from the target scenario corpus, and defining a generation task based on the knowledge classification and / or target scenario of the list of scenario hot words; based on a first large language model, applying the list of scenario hot words and the generation task to generate a sample text set corresponding to the list of scenario hot words; synthesizing a sample audio set corresponding to the sample text set, and using the sample audio set as training data and the sample text set as labels to train an initial speech recognition model to obtain a final hot word recognition model.
[0084] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the method for enhancing hot words in a low-resource scenario based on synthetic data provided by the above-mentioned various methods. The method includes: obtaining a target scenario corpus; extracting a list of scenario hot words from the target scenario corpus, and defining a generation task based on the knowledge classification and / or target scenario of the list of scenario hot words; based on a first large language model, applying the list of scenario hot words and the generation task to generate a sample text set corresponding to the list of scenario hot words; synthesizing a sample audio set corresponding to the sample text set, and using the sample audio set as training data and the sample text set as labels to train an initial speech recognition model to obtain a final hot word recognition model.
[0085] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0086] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A low-resource scenario hot word enhancement method based on synthetic data, characterized in that: include: Obtain the target scene corpus; Extracting a scene hot word list of the target scene corpus, and defining a generation task based on the knowledge classification and / or target scene of the scene hot word list; Based on the first large language model, applying the scene hot word list and the generation task, generating a sample text set corresponding to the scene hot word list; A sample audio set corresponding to the sample text set is synthesized, and the sample audio set is used as training data and the sample text set is used as a label to train an initial speech recognition model to obtain a final hot word recognition model.
2. The low-resource scenario hot word enhancement method based on synthetic data according to claim 1 is characterized in that: The step of applying the scene hot word list and the generation task based on the first large language model to generate a sample text set corresponding to the scene hot word list includes: Based on the generation task and prompt sentence style, construct a generation prompt text set that conforms to a preset structure template; Based on the first large language model, the scene hot word list and the generation prompt text set are applied to generate the sample text set; The first large language model is trained based on a general large language model and the target scenario corpus.
3. The low-resource scenario hot word enhancement method based on synthetic data according to claim 2 is characterized in that: The step of generating the sample text set based on the first large language model by applying the scene hot word list and the generation prompt text set includes: Based on the first large language model, the scene hot word list and the generation prompt text set are applied to generate an initial sample text set; Calculate the text similarity between any two of the initial sample texts in the initial sample text set; Eliminate any initial sample text from the pairwise initial sample texts whose text similarity is higher than a preset text similarity threshold to obtain a screening sample text set; Each screening sample text in the screening sample text set is matched with the scene hot word list, and the screening sample text with successful matching is selected as the sample text set.
4. The low-resource scenario hot word enhancement method based on synthetic data according to any one of claims 1 to 3, characterized in that: The initial speech recognition model is trained using the sample audio set as training data and the sample text set as labels to obtain a final hot word recognition model, including: Acquire an initial speech recognition model, wherein the initial speech recognition model includes an audio encoder and a second large language model; Based on the audio encoder, encode any sample audio in the sample audio set to generate an audio feature sequence; Based on the audio feature sequence, the second large language model and the recognition prompt text, generating an output recognition text of any sample audio; the recognition prompt text is constructed based on the scene hot word list and task instructions of any sample audio; Based on the output recognition text and the sample text corresponding to any one of the sample audios, the recognition loss is calculated, and based on the recognition loss, the model parameters of the initial speech recognition model are adjusted to obtain the hot word recognition model.
5. The low-resource scenario hot word enhancement method based on synthetic data according to claim 4 is characterized in that: The step of generating the output recognition text of any sample audio based on the audio feature sequence, the second large language model and the recognition prompt text comprises: Decoding the audio feature sequence based on the second large language model to obtain a decoding candidate sequence; enhancing the transcription priority of the decoding candidate sequence within the scope of the scene hot word list in the recognition prompt text to obtain an adjusted decoding sequence; Text transcription is performed based on the second large language model and the adjusted decoding sequence to generate an output recognition text of any one of the sample audios.
6. The low-resource scenario hot word enhancement method based on synthetic data according to any one of claims 1 to 3, characterized in that: The synthesizing the sample audio set corresponding to the sample text set includes: Generate a personalized sample audio set based on the sample text set, the sound attribute settings corresponding to the target scene, and the emotional style; Scene noise is injected into the personalized sample audio set to obtain the sample audio set.
7. The low-resource scenario hot word enhancement method based on synthetic data according to claim 6 is characterized in that: The injecting scene noise into the personalized sample audio set to obtain the sample audio set includes: injecting scene noise into the personalized sample audio set to obtain a noise sample audio set; Extracting hot word audio segments from each noise sample audio in the noise sample audio set; Performing pronunciation adjustment and / or speech speed adjustment on the hot word audio segment to obtain an enhanced audio segment; Based on the enhanced audio segment, the hot word audio segments in each noise sample audio corresponding to the enhanced audio segment are replaced respectively to obtain the sample audio set.
8. The low-resource scenario hot word enhancement method based on synthetic data according to any one of claims 1 to 3, characterized in that: The step of extracting a scene hot word list of the target scene corpus includes: Constructing an extraction prompt text based on the target scenario to which the target scenario corpus belongs; Based on the first large language model, the extraction prompt text and the target scene corpus are applied to extract the scene hot word list.
9. A low-resource scenario hot word enhancement device based on synthetic data, characterized in that: include: Acquisition unit, acquiring target scene corpus; An extraction unit extracts a scene hot word list of the target scene corpus, and defines a generation task based on the knowledge classification and / or target scene of the scene hot word list; A sample generating unit, based on the first large language model, applies the scene hot word list and the generating task to generate a sample text set corresponding to the scene hot word list; The training unit synthesizes a sample audio set corresponding to the sample text set, uses the sample audio set as training data and the sample text set as a label, trains the initial speech recognition model, and obtains a final hot word recognition model.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the low-resource scenario hot word enhancement method based on synthetic data is implemented as described in any one of claims 1 to 8.
Citation Information
Cited By
Hot word extraction method and device for speech recognition, storage medium and product
CN120636410A