Voice wake-up model training method and device, equipment and readable storage medium
By obtaining customized entries and audio corpus uploaded by users and performing noise reverberation simulation, a voice wake-up model that can recognize customized voice commands and perform well in noisy environments is trained, which solves the recognition problem of existing models in customized instructions and noisy environments.
Patent Information
- Application Number
- CN202510222677.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
Existing acoustic models are difficult to recognize customized voice commands during training and perform poorly in noisy environments with low signal-to-noise.
By obtaining customized entries uploaded by the user, determine the pronunciation sequence, and find and obtain audio corpus from the storage system. The audio corpus is subjected to noise reverberation simulation processing to generate simulated corpus carrying noise and reverberation, and finally the speech wake-up model is trained based on this corpus.
The speech wake-up model improves the accuracy of the recognition of customized voice commands, and enhances the robustness of the model in different environments, ensuring that voice commands can be accurately recognized in noisy environments with low signal-to-noise.
Smart Images

Figure CN120071900A_ABST
Abstract
Claims
1. A method for training a voice wake-up model, characterized in that: The method comprises: Obtaining a custom word uploaded by a user, and determining a pronunciation sequence according to the custom word; According to the customized word entry and the pronunciation sequence, searching and acquiring audio corpus from a storage system; the audio corpus at least includes clean corpus, synthetic corpus and self-recorded corpus; Performing noise-reverberation simulation processing on the audio corpus to obtain a simulation corpus carrying noise and reverberation; Based on the simulation corpus, a speech wake-up model is trained.
2. The method for training a voice wake-up model according to claim 1, wherein: Performing noise and reverberation simulation processing on the audio corpus to obtain a simulation corpus carrying noise and reverberation, including: Get environmental reverberation, environmental noise, and specific noise generated by specific devices uploaded by users; The environmental reverberation, environmental noise and specific noise are added to the audio corpus to obtain a simulated corpus carrying noise and reverberation.
3. The method for training a voice wake-up model according to claim 2, wherein: Adding the environmental reverberation, environmental noise and specific noise to the audio corpus to obtain a simulated corpus carrying noise and reverberation, including: Obtain the user's noise requirements; According to the noise requirement, setting the ratio information of the environmental reverberation, environmental noise and specific noise; According to the ratio information, the environmental reverberation, environmental noise and specific noise are added to the audio corpus to obtain a simulated corpus carrying noise and reverberation.
4. The method for training a voice wake-up model according to claim 1, wherein: Obtaining a custom word uploaded by a user, and determining a pronunciation sequence according to the custom word, including: Obtaining a custom entry uploaded by a user, and determining candidate pronunciations of each character in the custom entry; Determine the polyphonetic characters in the customized entry, and delete the non-target pronunciations of the polyphonetic characters from the candidate pronunciations to obtain the target pronunciations of each character in the customized entry; The pronunciation sequence is generated based on the target pronunciation.
5. The method for training a voice wake-up model according to any one of claims 1 to 4, characterized in that: The method further comprises: Converting the voice wake-up model into a code library that can be integrated into a voice product software development kit; Integrate the code library into the voice product software development kit.
6. The method for training a voice wake-up model according to claim 5, characterized in that: The method further comprises: Burn the voice product software development kit into the relevant chip.
7. A training device for a voice wake-up model, characterized in that: The device comprises: A pronunciation determination module, used to obtain a custom word uploaded by a user and determine a pronunciation sequence according to the custom word; An audio search module, used to search and obtain audio corpus from a storage system according to the customized word entry and the pronunciation sequence; the audio corpus at least includes clean corpus, synthetic corpus and self-recorded corpus; A simulation processing module, used for performing noise and reverberation simulation processing on the audio corpus to obtain a simulation corpus carrying noise and reverberation; The model training module is used to train a speech wake-up model based on the simulation corpus.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the training method of the voice wake-up model described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the training method of the voice wake-up model described in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the training method of the voice wake-up model described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Polyphone model training method, and speech synthesis method and device
CN105336322A
Model training method and device, storage medium and electronic device
CN112992170A
Identification model training method, identification method, electronic equipment and storage medium
CN114267342A
Model training method for command word speech enhancement
CN116386613A