Voice wake-up model training method and device, equipment and readable storage medium

By obtaining customized entries and audio corpus uploaded by users and performing noise reverberation simulation, a voice wake-up model that can recognize customized voice commands and perform well in noisy environments is trained, which solves the recognition problem of existing models in customized instructions and noisy environments.

CN120071900APending Publication Date: 2025-05-30AISPEECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510222677.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing acoustic models are difficult to recognize customized voice commands during training and perform poorly in noisy environments with low signal-to-noise.

Method used

By obtaining customized entries uploaded by the user, determine the pronunciation sequence, and find and obtain audio corpus from the storage system. The audio corpus is subjected to noise reverberation simulation processing to generate simulated corpus carrying noise and reverberation, and finally the speech wake-up model is trained based on this corpus.

Benefits of technology

The speech wake-up model improves the accuracy of the recognition of customized voice commands, and enhances the robustness of the model in different environments, ensuring that voice commands can be accurately recognized in noisy environments with low signal-to-noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071900A_ABST
    Figure CN120071900A_ABST
Patent Text Reader

Abstract

The invention discloses a voice wake-up model training method and device, equipment and a readable storage medium, and relates to the technical field of model training. Comprising the following steps: acquiring customized entries uploaded by a user, and determining a pronunciation sequence according to the customized entries; according to the customized vocabulary entry and the pronunciation sequence, searching and obtaining an audio corpus from a storage system; the audio corpus at least comprises a clean corpus, a synthetic corpus and a self-recording corpus; performing noise reverberation simulation processing on the audio corpus to obtain a simulation corpus carrying noise and reverberation; and training to obtain a voice wake-up model based on the simulation corpus. According to the scheme of the invention, the recognition of the specific voice instruction is realized, and the recognition accuracy in a noise environment is improved.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A method for training a voice wake-up model, characterized in that: The method comprises: Obtaining a custom word uploaded by a user, and determining a pronunciation sequence according to the custom word; According to the customized word entry and the pronunciation sequence, searching and acquiring audio corpus from a storage system; the audio corpus at least includes clean corpus, synthetic corpus and self-recorded corpus; Performing noise-reverberation simulation processing on the audio corpus to obtain a simulation corpus carrying noise and reverberation; Based on the simulation corpus, a speech wake-up model is trained.

2. The method for training a voice wake-up model according to claim 1, wherein: Performing noise and reverberation simulation processing on the audio corpus to obtain a simulation corpus carrying noise and reverberation, including: Get environmental reverberation, environmental noise, and specific noise generated by specific devices uploaded by users; The environmental reverberation, environmental noise and specific noise are added to the audio corpus to obtain a simulated corpus carrying noise and reverberation.

3. The method for training a voice wake-up model according to claim 2, wherein: Adding the environmental reverberation, environmental noise and specific noise to the audio corpus to obtain a simulated corpus carrying noise and reverberation, including: Obtain the user's noise requirements; According to the noise requirement, setting the ratio information of the environmental reverberation, environmental noise and specific noise; According to the ratio information, the environmental reverberation, environmental noise and specific noise are added to the audio corpus to obtain a simulated corpus carrying noise and reverberation.

4. The method for training a voice wake-up model according to claim 1, wherein: Obtaining a custom word uploaded by a user, and determining a pronunciation sequence according to the custom word, including: Obtaining a custom entry uploaded by a user, and determining candidate pronunciations of each character in the custom entry; Determine the polyphonetic characters in the customized entry, and delete the non-target pronunciations of the polyphonetic characters from the candidate pronunciations to obtain the target pronunciations of each character in the customized entry; The pronunciation sequence is generated based on the target pronunciation.

5. The method for training a voice wake-up model according to any one of claims 1 to 4, characterized in that: The method further comprises: Converting the voice wake-up model into a code library that can be integrated into a voice product software development kit; Integrate the code library into the voice product software development kit.

6. The method for training a voice wake-up model according to claim 5, characterized in that: The method further comprises: Burn the voice product software development kit into the relevant chip.

7. A training device for a voice wake-up model, characterized in that: The device comprises: A pronunciation determination module, used to obtain a custom word uploaded by a user and determine a pronunciation sequence according to the custom word; An audio search module, used to search and obtain audio corpus from a storage system according to the customized word entry and the pronunciation sequence; the audio corpus at least includes clean corpus, synthetic corpus and self-recorded corpus; A simulation processing module, used for performing noise and reverberation simulation processing on the audio corpus to obtain a simulation corpus carrying noise and reverberation; The model training module is used to train a speech wake-up model based on the simulation corpus.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the training method of the voice wake-up model described in any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the training method of the voice wake-up model described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the training method of the voice wake-up model described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Polyphone model training method, and speech synthesis method and device

    CN105336322A

  • Model training method and device, storage medium and electronic device

    CN112992170A

  • Identification model training method, identification method, electronic equipment and storage medium

    CN114267342A

  • Model training method for command word speech enhancement

    CN116386613A