Voice wake-up and model training method and device, related equipment and program product
By synthesizing synthetic audio using the voiceprint features of locally recorded audio, and updating the training model with user reflow audio data, the problem of poor adaptability of the voice wake-up model to local accents and personal speech style in the prior art is solved, and the wake-up effect and robustness are improved.
Patent Information
- Application Number
- CN202510532285.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-06-27
AI Technical Summary
While the existing voice wake-up training program reduces the recording cost, it is difficult to cover local accents or personal speaking style, resulting in poor awakening effect.
By obtaining the voiceprint features of locally recorded audio and synthesize corresponding synthetic audio based on these features, training data for training voice wake-up models is constructed. At the same time, the user's return audio data is collected, their voiceprint features are extracted and additional statements are synthesized to update the training model.
It improves the robustness and wake-up effect of the voice wake-up model, and can better adapt to various usage scenarios and user accents.
Smart Images

Figure CN120220647A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of voice signal processing, and more specifically, to a voice wake-up and model training method, device, related equipment and program product. Background Art
[0002] The increasing maturity of voice wake-up technology has made smart furniture, vehicles, etc. have higher interactivity, greatly facilitating users' daily lives.
[0003] Existing voice wake-up training schemes generally construct training data based on a large number of recorded audio files to train a pre-constructed wake-up model. Undoubtedly, a large amount of recorded audio data will increase the recording time and labor cost. To reduce costs, existing wake-up schemes combine voice synthesis technology to construct training data with synthetic data or a mixture of real and synthetic audio files, effectively saving the recording time and labor cost. Existing voice synthesis technology mainly includes two parts: language analysis and acoustic system. Language analysis performs structural judgment, phoneme conversion, and prosody prediction on the input text, and the acoustic system synthesizes audio based on the text information. The audio synthesized using this voice synthesis technology has a single style and often fails to cover regional accents or strong personal speaking styles, resulting in a poor wake-up effect of the trained voice recognition model for some users. Summary of the Invention
[0004] In view of the above problems, the present application is proposed to provide a voice wake-up and model training method, device, related equipment and program product to improve the wake-up effect of the voice wake-up model. The specific solutions are as follows:
[0005] In a first aspect, a method for training a voice wake-up model is provided, including:
[0006] Obtain first training data, where the first training data includes locally recorded audio and first synthetic audio, and the first synthetic audio is audio synthesized based on the voiceprint feature of the locally recorded audio and a first text, and the first text covers complete statements in the usage scenarios of the voice wake-up model;
[0007] Train a voice wake-up model using the first training data.
[0008] In a possible design, in another implementation manner of the first aspect of the embodiments of the present application, the process of obtaining the first training data includes:
[0009] Obtain locally recorded audio and extract the voiceprint feature of the locally recorded audio;
[0010] Perform voice synthesis based on the voiceprint feature and the first text to obtain first synthetic audio, and the first training data is composed of the first synthetic audio and the locally recorded audio.
[0011] In a possible design, in another implementation of the first aspect of the embodiments of the present application, it further includes:
[0012] Obtain the reflux audio data, where the reflux audio data is the audio data input by the user during the process of using the voice wake-up model;
[0013] Obtain a second text, where the second text at least includes a statement to be supplemented, and the statement to be supplemented is the statement missing from the reflux audio data relative to the first text;
[0014] Extract the voiceprint feature of the reflux audio data, perform voice synthesis based on the voiceprint feature of the reflux audio data and the second text to obtain a second synthesized audio, and form second training data from the second synthesized audio and the reflux audio data;
[0015] Use the second training data to update and train the voice wake-up model.
[0016] In a possible design, in another implementation of the first aspect of the embodiments of the present application, the second text further includes: a statement to be optimized, where the statement to be optimized is a statement with a wake-up rate lower than a set threshold summarized according to the actual use of the user.
[0017] In a possible design, in another implementation of the first aspect of the embodiments of the present application, the process of extracting the voiceprint feature of the reflux audio data includes:
[0018] Group the reflux audio data by speaker, and splice the reflux audio data of the same speaker;
[0019] Extract the voiceprint feature of the spliced audio of each speaker to obtain the voiceprint feature corresponding to each speaker.
[0020] In a possible design, in another implementation of the first aspect of the embodiments of the present application, the process of performing voice synthesis based on the voiceprint feature of the reflux audio data and the second text to obtain a second synthesized audio includes:
[0021] For the statement to be supplemented in the second text, send the statement to be supplemented and the voiceprint feature of each speaker to an audio synthesizer to obtain the synthesized audio of all the statements to be supplemented of each speaker as the second synthesized audio.
[0022] In a possible design, in another implementation of the first aspect of the embodiments of the present application, the second text further includes a statement to be optimized;
[0023] The process of performing voice synthesis based on the voiceprint features of the returned audio data and the second text to obtain the second synthesized audio further includes:
[0024] Determine the relevant speaker of the statement to be optimized, where the relevant speaker is a user whose wake-up rate for the statement to be optimized during the process of using the voice wake-up model is lower than a set threshold;
[0025] For the statement to be optimized in the second text, send the statement to be optimized and the voiceprint features of the relevant speaker into an audio synthesizer to obtain the synthesized audio of the statement to be optimized of the relevant speaker as the second synthesized audio.
[0026] In a second aspect, a voice wake-up method is provided, including:
[0027] Obtain user speech;
[0028] Process the user speech using a voice wake-up model to obtain a wake-up result;
[0029] Wherein, the voice wake-up model is trained using any of the voice wake-up model training methods in the first aspect of the embodiments of the present application.
[0030] In a third aspect, a voice wake-up model training device is provided, including:
[0031] A first training data acquisition unit for acquiring first training data, where the first training data includes locally recorded audio and a first synthesized audio, and the first synthesized audio is an audio synthesized based on the voiceprint features of the locally recorded audio and a first text, and the first text covers complete statements in the usage scenario of the voice wake-up model;
[0032] A model training unit for training a voice wake-up model using the first training data.
[0033] In a fourth aspect, an electronic device is provided, including: a memory and a processor;
[0034] The memory is used to store a program;
[0035] The processor is used to execute the program to implement each step of the voice wake-up model training method described in any item of the first aspect of the present application, or to implement each step of the voice wake-up method described in the second aspect of the present application.
[0036] In a fifth aspect, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, each step of the voice wake-up model training method described in any one of the foregoing first aspects of the present application is implemented, or each step of the voice wake-up method described in the foregoing second aspect of the present application is implemented.
[0037] In a sixth aspect, a computer program product is provided, including a computer program. When the computer program is executed by a processor, each step of the voice wake-up model training method described in any one of the foregoing first aspects of the present application is implemented, or each step of the voice wake-up method described in the foregoing second aspect of the present application is implemented.
[0038] By means of the above technical solution, the first training data used in training the voice wake-up model in the present application simultaneously includes locally recorded audio (such as audio data recorded by real people with a small number of different local accents and different speaking styles) and the first synthesized audio, and the first synthesized audio is the audio synthesized based on the voiceprint characteristics of the locally recorded audio and the first text. The present application only needs to collect a small amount of real user-recorded audio for extracting voiceprint characteristics, and then can synthesize audio of any text (the first text), reducing the cost of manually recording audio. Moreover, the synthesized audio is guided by the voiceprint characteristics of the locally recorded audio, so that the synthesized audio can simulate the accents and speaking styles of the speakers of the locally recorded audio, being closer to the recorded audio of real users. The first training data obtained in this way can cover more local accents and personal speaking styles. In addition, the first text during audio synthesis covers complete statements in the usage scenarios of the voice wake-up model, and examples include various types of wake-up words, command words, free talk, etc., so as to ensure that the synthesized audio can cover complete statements. Training the voice wake-up model with the first training data can improve the robustness of the voice wake-up model and enhance its wake-up effect in various usage scenarios. Description of the Drawings
[0039] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0040] Figure 1 It is a schematic diagram of a two-stage training process of a voice wake-up model provided by an embodiment of the present application;
[0041] Figure 2 It is a schematic diagram of an implementation system architecture of a model training method and a voice wake-up method provided by an embodiment of the present application;
[0042] Figure 3Schematic diagram of a voice wake-up model training method provided by an embodiment of the present application;
[0043] Figure 4 Schematic diagram of the processing flow of an audio synthesizer provided by an embodiment of the present application;
[0044] Figure 5 Schematic diagram of the process of training a voice wake-up model provided by an embodiment of the present application;
[0045] Figure 6 Schematic diagram of another voice wake-up model training method flow provided by an embodiment of the present application;
[0046] Figure 7 Schematic diagram of yet another voice wake-up model training method flow provided by an embodiment of the present application;
[0047] Figure 8 Schematic diagram of the structure of a voice wake-up model training device provided by an embodiment of the present application;
[0048] Figure 9 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0049] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0050] It can be understood that before using the technical solutions disclosed in the embodiments of the present application, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present application should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0051] For example, it is clearly informed to the user in the product instruction manual whether to agree that the system collects the audio data generated during the user's use process for optimizing the voice wake-up model and improving the user experience. Thus, the user can independently choose whether to provide the corresponding information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the technical solutions of the present application.
[0052] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present application. Other ways that meet relevant laws and regulations can also be applied to the implementation manner of the present application.
[0053] It is understandable that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0054] Existing voice wake-up training solutions generally construct training data based on a large number of recorded audio to train a pre-constructed wake-up model. A large amount of training data often leads to high recording costs. To reduce costs, existing wake-up solutions combine speech synthesis technology to construct training data with synthetic data or a mixture of real and synthetic audio, effectively saving recording time and labor costs. Existing speech synthesis technology mainly includes two parts: language analysis and an acoustic system. Language analysis performs structural judgment, phoneme conversion, and prosody prediction on the input text, and the acoustic system synthesizes audio based on the text information. The specific synthesis schemes are mainly divided into waveform concatenation speech synthesis and statistical parameter-based speech synthesis. These synthesis schemes only focus on the content of the synthesized text, and the synthesized audio has a single style. For example, the synthesis is uniformly performed according to fixed audio parameters, and the synthesized audio often fails to cover local accents or strong personal speaking styles, resulting in a poor wake-up effect of the trained speech recognition model for some users.
[0055] Therefore, the present application proposes a voice recognition model wake-up solution based on voice replication. Combining Figure 1 As shown, this solution can include an initial model (initial speech recognition model) training stage and a subsequent model optimization stage (model fine-tuning and updating process). Among them, the model optimization stage is an optional operation.
[0056] In the initial model training stage, training data is constructed based on synthetic audio and locally recorded audio. The synthetic audio can be generated using a voice replication tool, and the voiceprint features used by the voice replication tool are derived from the locally recorded audio. The text input covers complete statements in the usage scenarios of the product (the product to which the voice wake-up model is applied), such as wake-up words, command words, and free-form statements. Here, the free-form statement does not mean an infinite list of statements, but rather refers to expanding multiple statements for the core intention and considering prefixes, suffixes, and insertions. Therefore, the free-form statement is more flexible than the command word and wake-up word, and there will be a predefined total number of statements during the training process. The constructed training data is used to train the voice wake-up model. In the second-stage model optimization process, a small amount of feedback audio data is collected, statements to be supplemented are generated according to requirements, and statements to be optimized with a poor wake-up rate are summarized to jointly form the text input for the second stage. The voiceprint features of the feedback audio data are extracted using the voice replication tool, and audio is synthesized based on the voiceprint features and the text input. The synthesized audio and the feedback audio together form the training data for the second stage, and the voice wake-up model is fine-tuned and updated, ensuring the integrity of the statements in the training data without increasing the recording cost and effectively improving the wake-up effect of the voice wake-up model.
[0057] The present application provides a method for training a voice wake-up model and a voice wake-up method based on the trained voice wake-up model. Among them, both the voice wake-up model training method and the voice wake-up method can be applied to, for example, Figure 2 the system architecture shown. The system may include a terminal 100 and a server 200. The server 200 may include one or more servers ( Figure 1 illustrated by taking one server as an example).
[0058] Either the terminal 100 or the server 200 can be used alone to execute the voice wake-up model training method provided by the embodiments of the present application. In addition, the terminal 100 and the server 200 can also be used in cooperation to execute the voice wake-up model training method provided in the embodiments of the present application.
[0059] The voice wake-up model trained by the voice wake-up model training method can be deployed on the terminal 100 or on the server 200. Then, either the terminal 100 or the server 200 can be used alone to execute the voice wake-up method provided by the embodiments of the present application. Or, the terminal 100 and the server 200 can also be used in cooperation to execute the voice wake-up method provided in the embodiments of the present application.
[0060] In an optional example, the voice wake-up model training method is completed on the server 200, and the trained voice wake-up model is deployed locally on the terminal 100, and the terminal 100 executes the voice wake-up method.
[0061] In another optional example, the voice wake-up model training method is completed on the server 200, and the trained voice wake-up model is deployed to the server 200. Then, the terminal 100 and the server 200 can be used in cooperation to execute the voice wake-up method. For example, the terminal 100 obtains the user's voice and uploads it to the server 200. The server 200 calls the voice wake-up model to perform recognition processing on the user's voice, obtains the wake-up result and sends it to the terminal 100, and the terminal 100 makes a decision on the next operation according to the wake-up result.
[0062] Next, describe Figure 1 the product form of the terminal 100;
[0063] The terminal 100 in the embodiments of the present application can be various products with voice wake-up functions, including but not limited to: mobile phones, tablet computers, household appliances, wearable devices, vehicle-mounted devices, conference terminals, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc. The embodiments of the present application do not impose any restrictions on this.
[0064] Next, the training process of the voice wake-up model will be introduced first. The embodiments of the present application provide a method for training a voice wake-up model. Taking the application of this method to a computer device as an example, the computer device can specifically be Figure 2 the terminal 100 in Figure 3 or a system composed of the terminal 100 and the server 200. Referring to
[0065] Step S100: Obtain first training data, where the first training data includes locally recorded audio and first synthesized audio. The first synthesized audio is audio synthesized based on the voiceprint feature of the locally recorded audio and a first text, and the first text covers complete statements in the usage scenario of the voice wake-up model.
[0066] Among them, the locally recorded audio is locally recorded audio by a small number of recorders. Exemplarily, a small number of recorders can be arranged (such as recorders covering multiple local accents according to the general requirements of the product), and for each entry in the instruction word list of the product (the product to which the voice wake-up model is applied), a certain number of audio are respectively recorded to form the locally recorded audio.
[0067] The first text is the text of the audio to be synthesized, and it can cover complete statements in the usage scenario of the voice wake-up model. Generally, the interaction statements when users use the product include wake-up words, command words, and free talk statements. Then the first text can include wake-up words, command words, and free talk statements.
[0068] The locally recorded audio is used to provide the voiceprint information of the recorder. In this embodiment, audio can be synthesized based on the voiceprint feature extracted from the locally recorded audio and the first text, and is defined as the first synthesized audio. In this way, the first training data can be constructed based on the first synthesized audio and the locally recorded audio for training the voice wake-up model.
[0069] Step S110: Train the voice wake-up model using the first training data.
[0070] The voice wake-up model can adopt various model structures. By training the voice wake-up model with the first training data constructed above, a preliminary voice wake-up model can be obtained for deployment to a terminal or a server and used as a model for performing voice wake-up recognition tasks during the subsequent product usage process.
[0071] In the method provided by the embodiments of the present application, the first training data used for training the voice wake-up model simultaneously includes locally recorded audio (such as audio data recorded by real people with a small number of different local accents and different speaking styles) and the first synthesized audio, and the first synthesized audio is an audio synthesized based on the voiceprint features of the locally recorded audio and the first text. The present application only needs to collect a small amount of real user-recorded audio for extracting voiceprint features, and then can synthesize audio of any text (the first text), reducing the cost of manually recording audio. Moreover, the synthesized audio is guided by the voiceprint features of the locally recorded audio, so that the synthesized audio can simulate the accents and speaking styles of the speakers of the locally recorded audio, being closer to the recorded audio of real users. Thus, the obtained first training data can cover more local accents and personal speaking styles. In addition, the first text during audio synthesis covers complete statements in the usage scenarios of the voice wake-up model, and examples include various types of wake-up words, command words, free talk, etc., so as to ensure that the synthesized audio can cover complete statements. Training the voice wake-up model with the first training data can improve the robustness of the voice wake-up model and enhance its wake-up effect in various usage scenarios.
[0072] In some embodiments of the present application, an optional implementation manner of obtaining the first training data in the above step S100 is introduced, which may specifically include:
[0073] S1. Obtain locally recorded audio and extract the voiceprint features of the locally recorded audio.
[0074] S2. Perform voice synthesis based on the voiceprint features and the first text to obtain the first synthesized audio, and the first training data is composed of the first synthesized audio and the locally recorded audio.
[0075] In the above steps, the process of extracting voiceprint features and synthesizing audio can be implemented by a voice replication tool. The voice replication tool may include a voiceprint extractor and an audio synthesizer, where the voiceprint extractor is used to extract the voiceprint features of the audio.
[0076] Exemplarily, the voiceprint extractor can extract Filter Bank features, HuBERT features, etc. of the locally recorded audio and save these features as voiceprint files.
[0077] It should be noted that the voiceprint features corresponding to different speakers are different. Therefore, for locally recorded audio, it is first grouped according to the speaker, and then the corresponding voiceprint features are extracted from the locally recorded audio of each speaker respectively.
[0078] The audio synthesizer is used to perform speech synthesis based on the voiceprint features and the first text. As shown in Figure 4 shown:
[0079] The processing flow of the audio synthesizer may include a text analysis part and a speech synthesis part.
[0080] The text analysis part performs language processing based on the text input and a predefined dictionary or rules. Examples of processing include text normalization, text to phoneme conversion, etc. Further prosody processing is performed, and finally the text analysis result is obtained, which may include various types of text analysis results such as prosody information, word segmentation, phonemes, etc.
[0081] The speech synthesis part synthesizes the required audio based on the extracted voiceprint features and the text analysis result. The timbre of the synthesized audio is the same as that of the speaker corresponding to the voiceprint features, and the prosody of the synthesized audio is controlled by the input text.
[0082] Corresponding to the above step S2, Figure 4 the input text in may be the first text, and the voiceprint features are the voiceprint features of the locally recorded audio extracted in step S1. The output audio is the first synthesized audio.
[0083] Finally, the first training data is composed of the first synthesized audio and the locally recorded audio.
[0084] By using the first training data acquisition method provided in this embodiment, a first synthesized audio with the same timbre as the locally recorded audio can be synthesized based on the locally recorded audio through a voice replication tool. Thus, on the basis of arranging a small number of recording personnel to record local audio data, a large number of synthesized audios closer to the pronunciation characteristics of real users can be automatically synthesized. As a result, a large amount of first training data can be obtained for training the voice wake-up model and improving the training effect of the voice wake-up model.
[0085] In some embodiments of the present application, an optional implementation manner of training the voice wake-up model using the first training data in the above step S110 is introduced.
[0086] In this embodiment, taking the first text including three statements of a wake-up word, a command word, and free speech as an example, the training process of training the voice wake-up model is introduced. Specifically, it may include the following steps:
[0087] S1. Extract the acoustic features of the audio in the first training data. Examples include fb40 features, Mel-frequency cepstral coefficient (MFCC) features, etc.
[0088] S2. Feed the acoustic features into the constructed wake-up model to obtain the category probabilities of the wake-up word, command word, and free speech respectively.
[0089] S3. Calculate the first loss between the category probability of the wake-up word and the wake-up word category label of the audio, calculate the second loss between the category probability of the command word and the command word category label of the audio, calculate the third loss between the category probability of free speech and the free speech category label of the audio, calculate the total loss based on the first loss, second loss, and third loss, and update the model parameters through the backpropagation algorithm to obtain the trained voice wake-up model.
[0090] Referring to Figure 5 , this embodiment provides an optional structure of the voice wake-up model. The voice wake-up model may include an encoding module and a decoding module.
[0091] The encoding module includes a feature extractor and a speech unit classifier. The feature extractor is used to map the input acoustic features (such as fb40 features) to hidden layer features. The speech unit classifier maps the hidden layer features to the category probabilities of speech units.
[0092] The speech unit classifier may include: a phoneme classifier and / or a syllable classifier. The phoneme classifier is used to map the hidden layer features to the category probabilities of phonemes, and the syllable classifier is used to map the hidden layer features to the category probabilities of syllables.
[0093] The decoder includes a wake-up word decoder, a command word decoder, and a free speech decoder. The wake-up word decoder is used to map the hidden layer features to the category probabilities of wake-up words, the command word decoder is used to map the hidden layer features to the category probabilities of command words, and the free speech decoder is used to map the hidden layer features to the category probabilities of free speech.
[0094] Then, in addition to the above first loss, second loss, and third loss, the total loss in the training stage of the voice wake-up model may further include the speech unit classification loss corresponding to the speech unit classifier. Corresponding to Figure 5 the structure shown, the speech unit classification loss may include: a phoneme classification loss and a syllable classification loss. The phoneme classification loss is calculated based on the phoneme category probabilities output by the phoneme classifier and the phoneme category labels corresponding to the audio, and the syllable classification loss is calculated based on the syllable category probabilities output by the syllable classifier and the syllable category labels corresponding to the audio.
[0095] Of course, the above only provides an optional network structure of the voice wake-up model. In addition, the voice wake-up model may also adopt other network structures, which will not be elaborated one by one in this application.
[0096] Further, in order to control the number of false awakenings of the voice wake-up model, for the above-mentioned trained voice wake-up model, during the test phase, the posterior probability of the voice wake-up model can also be extracted based on a pre-constructed false awakening set. According to the set number of false awakenings, threshold optimization is performed based on this posterior probability to obtain the corresponding threshold. This process may include:
[0097] S1. Construct a false awakening set:
[0098] Collect diverse non-awakening word audio samples (such as noise, words with similar pronunciations to the awakening word, irrelevant speech, etc.), ensuring coverage of scenarios that may cause false triggers.
[0099] S2. Extract the posterior probability:
[0100] Input the false awakening set into the trained voice wake-up model to obtain the posterior probability that each audio sample is determined to be the target awakening word.
[0101] S3. Sort the posterior probabilities:
[0102] Arrange the posterior probabilities that each audio sample in the false awakening set is determined to be the target awakening word in descending order, denoted as P1≥P2≥…≥P M (M is the number of audio samples).
[0103] S4. Determine the threshold according to the sorting result:
[0104] According to the set allowable number of false awakenings N, select the threshold as the probability value of the (N + 1)-th largest probability, that is, Threshold = P N+1 . For example, if 5 false awakenings are allowed, the threshold is the 6th highest probability P6 in the sorting, ensuring that only the first 5 samples exceed the threshold.
[0105] Optionally, after determining the threshold, continuous monitoring can be performed according to the actual application feedback, and the threshold can be re-optimized when necessary to adapt to new scenarios.
[0106] Through the above method, the correct awakening rate can be maximized under the premise of strictly controlling the number of false awakenings, which is applicable to application scenarios sensitive to false triggers.
[0107] In some embodiments of the present application, another training method for the voice wake-up model is introduced. Based on the initial training stage of the voice wake-up model introduced in the foregoing embodiments, in this embodiment, the optimization (update training) process of the voice wake-up model is mainly described, corresponding to Figure 1 the model optimization part in Figure 6 . As shown in
[0108] Step S200. Obtain the reflux audio data.
[0109] Among them, the returned audio data is the audio data input by the user during the process of using the voice wake-up model. Specifically, after the initial voice wake-up model is trained by the model training method introduced in the foregoing embodiment, the voice wake-up model can be applied to specific products (such as smart home appliances, vehicle-mounted terminals, etc.) for the user to actually experience and use. When the user inputs audio data during the use process, such audio data can be collected as the returned audio data and uploaded to the server for model optimization training.
[0110] Step S210, obtain a second text, where the second text includes at least the statement to be supplemented.
[0111] Among them, the statement to be supplemented is the statement missing in the returned audio data relative to the first text.
[0112] Since the returned audio data has strong randomness and the user experiences the product in the initial stage and may not be proficient in using the product, the returned audio data often cannot cover complete statements, so the returned audio data is not convenient to be directly used for training.
[0113] In this step, according to the complete statement of the product usage scenario (that is, the first text mentioned in the foregoing embodiment), the missing statement in the returned audio data is determined and recorded as the statement to be supplemented. The second text includes at least this statement to be supplemented. Using this second text for subsequent speech synthesis, the synthesized audio of the statement to be supplemented can be synthesized, and together with the returned audio data, it forms the second training data, which can ensure that the second training data covers complete statements.
[0114] Step S220, extract the voiceprint feature of the returned audio data, perform speech synthesis based on the voiceprint feature of the returned audio data and the second text to obtain a second synthesized audio, and the second synthesized audio and the returned audio data form the second training data.
[0115] Specifically, in this step, the process of extracting the voiceprint feature of the returned audio data can be implemented by using the voiceprint extractor of the voice replication tool, and the detailed process can refer to the introduction above. The process of performing speech synthesis based on the voiceprint feature of the returned audio data and the second text can refer to Figure 4 the processing flow of the audio synthesizer shown, corresponding to step S220, Figure 4 the input text in can be the second text, and the voiceprint feature is the voiceprint feature of the returned audio data. The output audio is the second synthesized audio.
[0116] Step S230, use the second training data to update and train the voice wake-up model.
[0117] Specifically, the second training data can be used to update and train the voice wake-up model according to the model fine-tuning update algorithm.
[0118] It should be noted that the model update training process of this embodiment can be executed according to the set strategy. For example, according to the product update cycle, the model update training process of steps S200-S230 is executed using the reflow audio data obtained in the previous cycle. In addition, other model update strategies can also be adopted, which are not strictly limited in this application. Each time the model is updated, the voice wake-up model before the update (which can be the voice wake-up model after initial training, or the intermediate version of the voice wake-up model after one or more updates) can be updated.
[0119] The voice wake-up model update training process provided in this embodiment collects a small amount of reflow audio data, and uses a sound replication tool to synthesize a second synthesized audio with the same timbre as the reflow audio data. Based on the second synthesized audio and the reflow audio data, a second training data with complete utterances can be obtained for model update training. Without increasing the recording cost, the integrity of the training data utterances is guaranteed, thereby effectively improving the effect of model update training.
[0120] In a possible implementation, considering that some of the reflow audio data obtained in the above steps may have poor audio quality or be unusable, in order to ensure the completeness of the second training data, the obtained reflow audio data may be processed in this embodiment to remove unusable audio data. The second text is determined based on the processed reflow audio data.
[0121] In another possible implementation, the user may have some statements with poor awakening effect during the actual use of the product, which is affected by various factors such as the user's spoken language, speaking style, and the ability of the voice awakening model after initial training. In order to improve the awakening effect of the voice awakening model on such statements, in this embodiment, such statements can be further added to the second text. Specifically, such statements are defined as statements to be optimized, and the statements to be optimized are statements whose awakening rates are lower than the set threshold according to the actual use of users, and the statements to be optimized are added to the second text.
[0122] In actual use, you can collect statements about poor wake-up effects from different users, summarize all the collected results, and filter out statements with wake-up rates below a set threshold according to a certain strategy as statements to be optimized.
[0123] By adding the statement to be optimized to the second text, the audio corresponding to the statement to be optimized can be synthesized and added to the second training data, thereby achieving data enhancement for statements with poor wake-up rates. After the speech wake-up model is updated and trained using the second training data, the wake-up effect of the speech wake-up model on the statement to be optimized can be improved.
[0124] In some embodiments of the present application, another training method for the voice wake-up model is introduced. Referring to Figure 7 as shown, the method includes the following steps:
[0125] Step S300: Obtain the return audio data and the second text.
[0126] Among them, the definition of the return audio data refers to the relevant introduction above.
[0127] The second text includes at least the sayings to be supplemented. In addition, the second text may also include the sayings to be optimized.
[0128] Step S310: Group the return audio data by speaker, and splice the return audio data of the same speaker.
[0129] Specifically, the return audio data may contain the audio of multiple speakers. In this step, the return audio data can be grouped by speaker, and the return audio data included in the same group comes from the same speaker. In order to improve the accuracy of the voiceprint feature extraction of each speaker, in this step, the return audio data in the same group can be spliced to obtain the spliced audio of the same speaker.
[0130] Step S320: Extract the voiceprint features from the spliced audio of each speaker to obtain the voiceprint features corresponding to each speaker.
[0131] Step S330: For the sayings to be supplemented, referring to the voiceprint features of each speaker, synthesize the synthesized audio of all the sayings to be supplemented of each speaker as the second synthesized audio.
[0132] Specifically, the sayings to be supplemented and the voiceprint features of each speaker can be sent to the audio synthesizer to obtain the synthesized audio of all the sayings to be supplemented of each speaker as the second synthesized audio.
[0133] Step S340: Determine the relevant speaker of the saying to be optimized.
[0134] Among them, the relevant speaker of the saying to be optimized is the user whose wake-up rate for the saying to be optimized is lower than the set threshold during the process of using the voice wake-up model.
[0135] Combined with the determination process of the saying to be optimized above, during the actual product use process, the user can feedback the saying to be optimized with poor wake-up effect, then the user identification and the saying to be optimized he / she feedbacks can be recorded, so as to determine the relevant speaker (user identification) of the saying to be optimized.
[0136] Step S350: For the saying to be optimized, referring to the voiceprint features of the relevant speaker of the saying to be optimized, synthesize the synthesized audio of the saying to be optimized of the relevant speaker as the second synthesized audio.
[0137] Specifically, the statement to be optimized and the voiceprint features of its related speaker can be sent to an audio synthesizer to obtain the synthesized audio of the statement to be optimized by the related speaker, which serves as the second synthesized audio.
[0138] In this step, for the statement to be optimized, the audio of the statement to be optimized is not synthesized for each speaker. Instead, it is targeted. Using the voiceprint features of the related speaker of the statement to be optimized, the synthesized audio of the statement to be optimized by the related speaker is synthesized.
[0139] Exemplarily, the statement to be optimized is "Open the refrigerator". The related speakers of this statement to be optimized (i.e., the users with poor wake-up effects for this statement to be optimized during the actual product usage process) include Speaker A and Speaker B (the number of related speakers can be one or more, taking 2 as an example here). Then, the text: "Open the refrigerator" and the voiceprint features of Speaker A can be sent to the audio synthesizer to obtain the synthesized audio, and, the text: "Open the refrigerator" and the voiceprint features of Speaker B can be sent to the audio synthesizer to obtain the synthesized audio.
[0140] Step S360: The second training data is composed of the second synthesized audio and the reflux audio data.
[0141] Step S370: Use the second training data to update and train the voice wake-up model.
[0142] For the method of this embodiment, for the statement to be supplemented and the statement to be optimized in the second text, audio synthesis is performed according to different strategies respectively. This not only ensures that the second training data can cover the statements to be supplemented of all speakers in the reflux audio data, but also specifically synthesizes the audio of the related speaker for the statement to be optimized, realizing data enhancement for the statement with a poor wake-up rate for a specific speaker. After using this second training data to update and train the voice wake-up model, the overall wake-up effect of the voice wake-up model and the wake-up effects for specific speakers and specific statements can be improved.
[0143] It can be understood that Figure 7 The example is the case where the second text includes both the statement to be supplemented and the statement to be optimized. In some possible implementations, if the second text does not contain the statement to be optimized, steps S340 - S350 can be omitted in the training method.
[0144] In some embodiments of the present application, a voice wake-up method is further provided, which specifically includes:
[0145] Obtain the user's voice.
[0146] Use the voice wake-up model to process the user's voice to obtain a wake-up result.
[0147] Among them, the voice wake-up model is trained by using the training method of any of the foregoing embodiments.
[0148] Since the voice wake-up model can be trained based on the training data of complete statements during the training phase, and its model wake-up effect is better, applying this voice wake-up model can effectively improve the voice wake-up effect.
[0149] Next, a voice wake-up model training device provided in an embodiment of the present application will be described. The voice wake-up model training device described below can be mutually corresponding and referred to the voice wake-up model training method described above.
[0150] See Figure 8 , Figure 8 which is a schematic structural diagram of a voice wake-up model training device disclosed in an embodiment of the present application.
[0151] As Figure 8 shown, the device may include:
[0152] A first training data acquisition unit 11, configured to acquire first training data, where the first training data includes locally recorded audio and first synthesized audio, and the first synthesized audio is audio synthesized based on the voiceprint feature of the locally recorded audio and a first text, and the first text covers complete statements in the usage scenario of the voice wake-up model;
[0153] A model training unit 12, configured to train a voice wake-up model by using the first training data.
[0154] In a possible implementation, the process of the first training data acquisition unit acquiring the first training data includes:
[0155] Acquire locally recorded audio, and extract the voiceprint feature of the locally recorded audio;
[0156] Perform voice synthesis based on the voiceprint feature and the first text to obtain first synthesized audio, and the first training data is composed of the first synthesized audio and the locally recorded audio.
[0157] In a possible implementation, the device of the present application may further include:
[0158] A model update unit, configured to perform update training on the voice wake-up model, and the update training process includes:
[0159] Acquire return audio data, where the return audio data is audio data input by a user during the process of using the voice wake-up model;
[0160] Acquire a second text, where the second text at least includes a statement to be supplemented, and the statement to be supplemented is a statement missing from the return audio data relative to the first text;
[0161] Extract the voiceprint features of the returned audio data, perform speech synthesis based on the voiceprint features of the returned audio data and the second text to obtain a second synthesized audio, and form second training data from the second synthesized audio and the returned audio data;
[0162] Use the second training data to update and train the voice wake-up model.
[0163] In a possible implementation, the second text further includes: a statement to be optimized, which is a statement with a wake-up rate lower than a set threshold summarized according to the actual use of the user.
[0164] In a possible implementation, the process by which the model update unit extracts the voiceprint features of the returned audio data includes:
[0165] Group the returned audio data by speaker, and splice the returned audio data of the same speaker;
[0166] Extract the voiceprint features from the spliced audio of each speaker to obtain the voiceprint features corresponding to each speaker.
[0167] In a possible implementation, the process by which the model update unit performs speech synthesis based on the voiceprint features of the returned audio data and the second text to obtain a second synthesized audio includes:
[0168] For the statement to be supplemented in the second text, send the statement to be supplemented and the voiceprint features of each speaker to an audio synthesizer to obtain the synthesized audio of all statements to be supplemented of each speaker as the second synthesized audio.
[0169] In a possible implementation, if the second text further includes a statement to be optimized, the process by which the model update unit performs speech synthesis based on the voiceprint features of the returned audio data and the second text to obtain a second synthesized audio further includes:
[0170] Determine the relevant speaker of the statement to be optimized, where the relevant speaker is a user with a wake-up rate lower than the set threshold for the statement to be optimized during the use of the voice wake-up model;
[0171] For the statement to be optimized in the second text, send the statement to be optimized and the voiceprint features of the relevant speaker to an audio synthesizer to obtain the synthesized audio of the statement to be optimized of the relevant speaker as the second synthesized audio.
[0172] An electronic device is further provided in an embodiment of the present application. Refer to Figure 9As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, terminals such as mobile phones, tablet computers, household appliances, vehicle-mounted devices, wearable devices, and the like. Figure 9 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0173] As Figure 9 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603, so as to implement the voice wake-up model training method or the voice wake-up method in the foregoing embodiments of the present application. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0174] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 9 the electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.
[0175] In the embodiments of the present application, there is also provided a computer program product including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device is enabled to implement any one of the voice wake-up model training methods or the voice wake-up methods provided in the embodiments of the present application.
[0176] In the embodiments of the present application, there is also provided a computer-readable storage medium. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can be enabled to implement any one of the voice wake-up model training methods or the voice wake-up methods provided in the embodiments of the present application.
[0177] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationships between the modules indicate that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0178] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, software program implementation is a better implementation method in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.
[0179] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0180] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are all or partially generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, a computer, a training device, or a data center to another website, a computer, a training device, or a data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0181] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.
Claims
1. A voice wake-up model training method, characterized in that: include: Acquire first training data, where the first training data includes locally recorded audio and first synthesized audio, where the first synthesized audio is audio synthesized based on a voiceprint feature of the locally recorded audio and first text, where the first text covers complete statements in a usage scenario of the voice wake-up model; The first training data is used to train a voice wake-up model.
2. The method according to claim 1, characterized in that The process of obtaining the first training data includes: Acquire local recorded audio and extract voiceprint features of the local recorded audio; Speech synthesis is performed based on the voiceprint feature and the first text to obtain a first synthesized audio, and the first synthesized audio and the locally recorded audio constitute the first training data.
3. The method according to claim 1, characterized in that Also includes: Acquire reflow audio data, where the reflow audio data is audio data input by a user using the voice wake-up model process; Acquire a second text, wherein the second text at least includes a statement to be supplemented, and the statement to be supplemented is a statement that is missing from the reflow audio data relative to the first text; Extracting voiceprint features of the reflow audio data, performing speech synthesis based on the voiceprint features of the reflow audio data and the second text to obtain second synthesized audio, and forming second training data with the second synthesized audio and the reflow audio data; The voice wake-up model is updated and trained using the second training data.
4. The method according to claim 3, characterized in that The second text also includes: a statement to be optimized, where the statement to be optimized is a statement that the wake-up rate summarized based on actual user use is lower than a set threshold.
5. The method according to claim 3, characterized in that: The process of extracting the voiceprint features of the reflow audio data includes: The returned audio data are grouped according to the speakers, and the returned audio data of the same speaker are spliced; The voiceprint features are extracted from the spliced audio of each speaker to obtain the voiceprint features corresponding to each speaker.
6. The method according to claim 5, characterized in that The process of performing speech synthesis based on the voiceprint feature of the reflow audio data and the second text to obtain the second synthesized audio includes: For the utterances to be added in the second text, the utterances to be added and the voiceprint features of each speaker are sent to an audio synthesizer to obtain synthesized audio of all utterances to be added of each speaker as the second synthesized audio.
7. The method according to claim 6, characterized in that The second text also includes a statement to be optimized; The process of performing speech synthesis based on the voiceprint feature of the reflow audio data and the second text to obtain the second synthesized audio further includes: Determining relevant speakers of the statement to be optimized, wherein the relevant speakers are users whose awakening rate for the statement to be optimized using the speech awakening model process is lower than a set threshold; For the statement to be optimized in the second text, the statement to be optimized and the voiceprint features of the relevant speaker are sent to an audio synthesizer to obtain a synthesized audio of the statement to be optimized by the relevant speaker as the second synthesized audio.
8. A voice wake-up method, characterized in that: include: Get user voice; Processing the user's voice using a voice wake-up model to obtain a wake-up result; Wherein, the voice wake-up model is trained using the voice wake-up model training method described in any one of claims 1-7.
9. A speech wake-up model training device, characterized in that: include: a first training data acquisition unit, configured to acquire first training data, wherein the first training data includes a locally recorded audio and a first synthesized audio, wherein the first synthesized audio is an audio synthesized based on a voiceprint feature of the locally recorded audio and a first text, wherein the first text covers a complete statement in a usage scenario of the voice wake-up model; A model training unit is used to train a voice wake-up model using the first training data.
10. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the various steps of the voice wake-up model training method as described in any one of claims 1 to 7, or to implement the various steps of the voice wake-up method as described in claim 8.
11. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the steps of the voice wake-up model training method according to any one of claims 1 to 7 are implemented, or the steps of the voice wake-up method according to claim 8 are implemented.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the voice wake-up model training method as described in any one of claims 1 to 7 are implemented, or the steps of the voice wake-up method as described in claim 8 are implemented.