Method, device, equipment, medium and program product for generating wake-up word audio
By generating wake-up word audio of different voiceprint types in intelligent vehicles, the problem of low recognition accuracy of traditional wake-up word recognition systems in diverse scenarios is solved, improving efficiency and user experience, adapting to multiple voiceprint types, and increasing device wake-up rate.
Patent Information
- Application Number
- CN202410853794.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2044-06-27
AI Technical Summary
Traditional wake word recognition systems have low accuracy in diverse real-world usage scenarios, especially in the presence of background noise, different accents, varying speech rates, or individual voiceprint differences. This leads to a decline in user experience and affects the usability and safety of intelligent driving assistance functions.
By acquiring recorded audio in the target environment, extracting voiceprint features and performing deep learning, different types of wake-up word audio are generated. The voiceprint model is used to replace the recording of real human voiceprints, including noise reduction processing, speech enhancement, deep learning and modeling, to generate wake-up word audio.
It improves the efficiency and accuracy of wake word recognition, saves manpower and time costs, adapts to various voiceprint types, and enhances user experience and device wake-up rate.
Smart Images

Figure CN118711597B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle technology, and in particular to a method, apparatus, device, medium, and program product for generating wake-up word audio. Background Technology
[0002] Currently, voice wake-up is a key component in intelligent vehicle interaction systems, enabling human-machine dialogue. Users activate the in-vehicle voice assistant using specific wake words to issue commands, query information, or control vehicle functions. However, traditional wake word recognition systems face a common problem: in diverse real-world usage scenarios, especially with background noise, different accents, varying speech rates, or individual voiceprint differences, the accuracy of wake word recognition is often unsatisfactory. This instability can lead to a degraded user experience and affect the usability and safety of intelligent driving assistance functions.
[0003] One common strategy in the industry is to train and optimize wake word recognition models by extensively collecting real-person voice samples with different genders, ages, accents, and voiceprint characteristics.
[0004] However, using real human voiceprint recording wastes a lot of manpower and time, and because each person's voiceprint characteristics are different, it is impossible to collect multiple types of voiceprints. In addition, the sample size of real human voiceprint recording is small and the efficiency is low. Summary of the Invention
[0005] This application provides a method, apparatus, device, medium, and program product for generating wake-up word audio, in order to solve the problems of low efficiency in related technologies that use real human voiceprints to record wake-up words.
[0006] The first aspect of this application provides a method for generating wake-up word audio, comprising the following steps: acquiring recorded audio in a target environment; extracting at least one real human voiceprint from the recorded audio, and performing deep learning on the at least one real human voiceprint to obtain voiceprint features corresponding to the at least one real human voiceprint; modeling based on the voiceprint features to obtain at least one voiceprint model, and using the at least one voiceprint model to generate wake-up word audio of at least one voiceprint.
[0007] Optionally, voiceprint features include at least one of the following: spectrum, cepstral, formants, pitch, reflection coefficient, dialect, and prosody.
[0008] Optionally, before extracting at least one real human voiceprint from the recorded audio, the process includes: performing noise reduction on the recorded audio and performing speech enhancement on the noise-reduced recorded audio to obtain the processed recorded audio.
[0009] Optionally, deep learning includes at least one of convolutional neural networks, recurrent neural networks, long short-term memory networks, and deep belief networks; the wake word audio format includes at least one of MP3, WAV, and AAC.
[0010] Optionally, modeling based on voiceprint features to obtain at least one voiceprint model includes: reducing the dimensionality of the voiceprint features; and modeling based on the dimensionality-reduced voiceprint features to obtain at least one voiceprint model, wherein the modeling method includes at least one of Gaussian mixture model, support vector machine, and deep neural network.
[0011] Optionally, after generating wake-up word audio for at least one voiceprint using at least one voiceprint model, the method further includes: waking up the target device using wake-up word audio for at least one voiceprint and detecting the wake-up rate of the target device; if the wake-up rate is lower than a preset value, then regenerating wake-up word audio for different voiceprints.
[0012] A second aspect of this application provides an apparatus for generating wake-up word audio, comprising: an acquisition module for acquiring recorded audio in a target environment; an extraction module for extracting at least one real human voiceprint from the recorded audio and performing deep learning on the at least one real human voiceprint to obtain voiceprint features corresponding to the at least one real human voiceprint; and a generation module for modeling based on the voiceprint features to obtain at least one voiceprint model and using the at least one voiceprint model to generate wake-up word audio of at least one voiceprint.
[0013] Optionally, voiceprint features include at least one of the following: spectrum, cepstral, formants, pitch, reflection coefficient, dialect, and prosody.
[0014] Optionally, it further includes: a processing module, configured to perform noise reduction processing on the recorded audio before extracting at least one real human voiceprint from the recorded audio, and to perform speech enhancement on the noise-reduced recorded audio to obtain the processed recorded audio.
[0015] Optionally, deep learning includes at least one of convolutional neural networks, recurrent neural networks, long short-term memory networks, and deep belief networks; the wake word audio format includes at least one of MP3, WAV, and AAC.
[0016] Optionally, the generation module is further used to: reduce the dimensionality of the voiceprint features; and to model at least one voiceprint model based on the reduced voiceprint features, wherein the modeling method includes at least one of Gaussian mixture model, support vector machine, and deep neural network.
[0017] Optionally, it further includes: a detection module, used to wake up the target device using the wake-up word audio of at least one voiceprint after generating the wake-up word audio of at least one voiceprint using at least one voiceprint model, and to detect the wake-up rate of the target device; if the wake-up rate is lower than a preset value, then regenerate the wake-up word audio of different voiceprints.
[0018] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the program to perform the wake word audio generation method as described in the above embodiments.
[0019] A fourth aspect of this application provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, performs the wake-word audio generation method as described above.
[0020] A fifth aspect of this application provides a computer program product, including a computer program or instructions, which, when executed, implement the wake word audio generation method as described in the above embodiments.
[0021] Therefore, this application has at least the following beneficial effects:
[0022] This application's embodiments can acquire recorded audio from a target environment, extract voiceprint features from the audio, and model these features to obtain voiceprint models of different voiceprint types. Different voiceprint models are then used to generate different wake-up word audio, replacing the process of recording with a large number of real people. This saves time, improves efficiency and cost, and makes it more convenient to acquire different voiceprints. Therefore, it solves the technical problems of low efficiency in related technologies that use real human voiceprints to record wake-up words.
[0023] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0024] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0025] Figure 1 This is a flowchart of a method for generating wake word audio according to an embodiment of this application;
[0026] Figure 2 This is a schematic diagram of a method for generating wake word audio according to an embodiment of this application;
[0027] Figure 3 This is an example diagram of a wake-word audio generation apparatus provided according to an embodiment of this application;
[0028] Figure 4 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Detailed Implementation
[0029] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0030] The following description, with reference to the accompanying drawings, outlines a method, apparatus, device, medium, and program product for generating wake-up word audio according to embodiments of this application. Addressing the issues mentioned in the background art, such as the waste of significant manpower and time in recording with real human voiceprints, the inability to collect multiple types of voiceprints due to individual differences in voiceprint characteristics, and the small sample size and low efficiency of recording with real human voiceprints, this application provides a method for generating wake-up word audio. In this method, recorded audio from a target environment is acquired, voiceprint features are extracted from the recorded audio, and voiceprint models of different voiceprint types are obtained by modeling these features. Different voiceprint models are then used to generate different wake-up word audios. This solves the problems of low efficiency associated with recording wake-up words using real human voiceprints in related technologies.
[0031] Specifically, Figure 1 This is a flowchart illustrating a method for generating wake word audio provided in an embodiment of this application.
[0032] like Figure 1 As shown, the method for generating the wake word audio includes the following steps:
[0033] In step S101, the recorded audio in the target environment is acquired.
[0034] The target environment can be densely populated squares, shopping malls, streets, etc., so that more types of real human voiceprints can be extracted later.
[0035] In step S102, at least one real human voiceprint is extracted from the recorded audio, and deep learning is performed on the at least one real human voiceprint to obtain the voiceprint features corresponding to the at least one real human voiceprint.
[0036] Voiceprint features include at least one of the following: spectrum, cepstral, formants, pitch, reflection coefficient, dialect, and prosody; deep learning includes at least one of convolutional neural networks, recurrent neural networks, long short-term memory networks, and deep belief networks.
[0037] It is understood that the embodiments of this application can extract at least one real human voiceprint from the recorded audio and perform deep learning on the at least one real human voiceprint to obtain voiceprint features corresponding to the at least one real human voiceprint. The extraction of at least one real human voiceprint from the recorded audio can be done using voiceprint recognition software or development packages, such as the librosa library in Python, PythonSpeechFeatures, Kaldi, DeepSpeech, etc.
[0038] In this embodiment of the application, before extracting at least one real human voiceprint from the recorded audio, the process includes: performing noise reduction processing on the recorded audio, and performing speech enhancement on the noise-reduced recorded audio to obtain the processed recorded audio.
[0039] It is understood that, before extracting at least one real human voiceprint from the recorded audio, the embodiments of this application may perform noise reduction processing on the recorded audio and perform speech enhancement on the noise-reduced recorded audio to obtain the processed recorded audio, so as to improve the effect of subsequent real human voiceprint extraction.
[0040] In step S103, at least one voiceprint model is obtained by modeling based on voiceprint features, and at least one voiceprint wake-up word audio is generated using at least one voiceprint model.
[0041] The wake word audio format includes at least one of MP3, WAV, and AAC.
[0042] It is understood that the embodiments of this application can model at least one voiceprint model based on voiceprint features, and use the voiceprint model to generate at least one voiceprint wake-up word audio, so as to improve the wake-up word audio of different voiceprint types, thereby replacing the process of recording a large number of real people, saving time, improving efficiency and cost, and making it more convenient to obtain wake-up words of different voiceprint types, which is convenient for subsequent use of wake-up word audio to improve the wake-up rate of wake-up devices.
[0043] In this embodiment of the application, at least one voiceprint model is obtained by modeling based on voiceprint features, including: reducing the dimensionality of the voiceprint features; and modeling based on the reduced-dimensional voiceprint features to obtain at least one voiceprint model.
[0044] The modeling methods include at least one of Gaussian mixture models, support vector machines, and deep neural networks.
[0045] It is understood that the embodiments of this application can perform dimensionality reduction processing on voiceprint features to reduce the amount of data processing in subsequent modeling and improve the accuracy of modeling, and model at least one voiceprint model by modeling the dimensionality-reduced voiceprint features.
[0046] In this embodiment of the application, after generating wake-up word audio for at least one voiceprint using at least one voiceprint model, the method further includes: waking up the target device using wake-up word audio for at least one voiceprint and detecting the wake-up rate of the target device; if the wake-up rate is lower than a preset value, then regenerating wake-up word audio for different voiceprints.
[0047] The target device can be an electronic device that supports voice wake-up, such as a mobile phone, a speaker, etc.; the preset value can be set according to the specific situation, and there is no specific limitation on it, such as setting it to 80%.
[0048] It is understood that the embodiments of this application can use the generated wake word audio to wake up the target device and detect the wake-up rate of the target device. If the wake-up rate is lower than a preset value, the execution process of this application is reused to generate wake word audio with different voiceprints. The wake-up rate of the target device is improved by using the regenerated wake word audio with different voiceprints.
[0049] It should be noted that the wake word can be a noun, such as "smart speaker", or a sentence, such as "Hi, turn on the smart speaker".
[0050] The following specific embodiment illustrates the method for generating wake-up word audio according to this application, as follows: Figure 2 As shown, it includes:
[0051] 1. Record an audio clip in a crowded place (square, shopping mall, street, etc.);
[0052] 2. Set up a voiceprint extraction module, and use existing voiceprint extraction technology to extract real human voiceprints from the recording one by one by processing the recording through noise reduction, speech enhancement, signal extraction, etc., and output n kinds of real human original voiceprints.
[0053] 3. Use the AI learning module to perform deep learning on the extracted voiceprints, extract features from the preprocessed voiceprint speech signals, and convert them into MFCC coefficients, cepstral coefficients, etc., and use these information features to perform voiceprint modeling to generate n voiceprint models.
[0054] 4. After learning and mastering the skills, use the voiceprint model to output the audio of the voice wake-up word corresponding to n voiceprints through the AI output module.
[0055] 5. Optimize the wake-up rate according to the original training voice wake-up rate procedure;
[0056] 6. If the wake-up rate is still low, repeat step one to record and acquire more voiceprint types.
[0057] According to the wake-up word audio generation method proposed in the embodiments of this application, the voiceprint features in the recorded audio of the target environment can be extracted and the voiceprint features can be modeled to obtain voiceprint models of different voiceprint types. Different wake-up word audios can be generated using different voiceprint models, which replaces the process of recording a large number of real people. This saves time, improves efficiency and cost, and makes it more convenient to obtain different voiceprints.
[0058] Next, with reference to the accompanying drawings, a wake-up word audio generation apparatus according to an embodiment of this application is described.
[0059] Figure 3 This is a block diagram of a wake-word audio generation device according to an embodiment of this application.
[0060] like Figure 3 As shown, the wake-up word audio generation device 10 includes: an acquisition module 100, an extraction module 200, and a generation module 300.
[0061] The acquisition module 100 is used to acquire the recorded audio in the target environment; the extraction module 200 is used to extract at least one real human voiceprint from the recorded audio and perform deep learning on the at least one real human voiceprint to obtain the voiceprint features corresponding to the at least one real human voiceprint; the generation module 300 is used to model based on the voiceprint features to obtain at least one voiceprint model, and use the at least one voiceprint model to generate at least one voiceprint wake word audio.
[0062] In this application embodiment, the voiceprint features include at least one of the following: spectrum, cepstral spectrum, formant, pitch, reflection coefficient, dialect, and prosody.
[0063] In this embodiment of the application, the apparatus 10 further includes a processing module.
[0064] The processing module is used to perform noise reduction on the audio recording before extracting at least one real human voiceprint from the audio recording, and to perform speech enhancement on the noise-reduced audio recording to obtain the processed audio recording.
[0065] In this embodiment, deep learning includes at least one of convolutional neural networks, recurrent neural networks, long short-term memory networks, and deep belief networks; the wake word audio format includes at least one of MP3, WAV, and AAC.
[0066] In this embodiment of the application, the generation module 300 is further configured to: reduce the dimensionality of the voiceprint features; and model at least one voiceprint model based on the reduced voiceprint features, wherein the modeling method includes at least one of Gaussian mixture model, support vector machine, and deep neural network.
[0067] In this embodiment of the application, the apparatus 10 further includes a detection module.
[0068] The detection module is used to wake up the target device using the wake-up word audio of at least one voiceprint after generating the wake-up word audio of at least one voiceprint using at least one voiceprint model, and to detect the wake-up rate of the target device; if the wake-up rate is lower than a preset value, wake-up word audio of different voiceprints is regenerated.
[0069] It should be noted that the foregoing explanation of the method embodiment for generating wake word audio also applies to the wake word audio generation apparatus of this embodiment, and will not be repeated here.
[0070] The wake-up word audio generation device proposed in the embodiments of this application can obtain the recorded audio in the target environment, extract the voiceprint features in the recorded audio, and model the voiceprint features to obtain voiceprint models of different voiceprint types. Different wake-up word audios can be generated using different voiceprint models, which replaces the process of recording a large number of real people, saving time, improving efficiency and cost, and making it more convenient to obtain different voiceprints.
[0071] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include:
[0072] The memory 401, the processor 402, and the computer program stored on the memory 401 and capable of running on the processor 402.
[0073] When the processor 402 executes the program, it implements the wake-word audio generation method provided in the above embodiments.
[0074] Furthermore, electronic devices also include:
[0075] Communication interface 403 is used for communication between memory 401 and processor 402.
[0076] The memory 401 is used to store computer programs that can run on the processor 402.
[0077] The memory 401 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0078] If the memory 401, processor 402, and communication interface 403 are implemented independently, then the communication interface 403, memory 401, and processor 402 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized into address buses, data buses, control buses, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0079] Optionally, in a specific implementation, if the memory 401, processor 402, and communication interface 403 are integrated on a single chip, then the memory 401, processor 402, and communication interface 403 can communicate with each other through an internal interface.
[0080] Processor 402 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0081] This application also provides a computer-readable storage medium storing a computer program or instructions thereon, which, when executed by a processor, implements the above-described method for generating wake word audio.
[0082] This application also provides a computer program product, including a computer program or instructions, which, when executed, implement a method for generating wake word audio, as described above.
[0083] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0084] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0085] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0086] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or more of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0087] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
Claims
1. A method for generating an audio of a wake-up word, characterized in that, The method comprises the following steps: acquiring recorded audio in a target environment, wherein the target environment is a crowded square, a shopping mall or a street; extracting at least one real voiceprint in the recorded audio, and performing deep learning on the at least one real voiceprint to obtain a voiceprint feature corresponding to the at least one real voiceprint; modeling based on the voiceprint feature to obtain at least one voiceprint model, and generating an awakening word audio of the at least one voiceprint by using the at least one voiceprint model.
2. The method of claim 1, wherein, The voiceprint feature comprises at least one of a frequency spectrum, a cepstrum, a formant, a fundamental tone, a reflection coefficient, a dialect and a prosody.
3. The method of claim 1, wherein, Before extracting the at least one real voiceprint in the recorded audio, the method comprises: performing noise reduction processing on the recorded audio, and performing speech enhancement on the recorded audio after the noise reduction processing to obtain processed recorded audio.
4. The method of claim 1, wherein, The deep learning comprises at least one of a convolutional neural network, a recurrent neural network, a long short-term memory network and a deep belief network; and the format of the awakening word audio comprises at least one of MP3, WAV and AAC.
5. The method of claim 1, wherein, The modeling based on the voiceprint feature to obtain at least one voiceprint model comprises: dimensionality reduction on the voiceprint feature; modeling based on the voiceprint feature after the dimensionality reduction to obtain at least one voiceprint model, wherein the modeling method comprises at least one of a Gaussian mixture model, a support vector machine and a deep neural network.
6. The method of claim 1, wherein, After generating the awakening word audio of the at least one voiceprint by using the at least one voiceprint model, the method further comprises: awakening a target device by using the awakening word audio of the at least one voiceprint, and detecting an awakening rate of the target device; if the awakening rate is lower than a preset value, generating an awakening word audio of a different voiceprint.
7. An apparatus for generating an audio of a wake-up word, the apparatus comprising: The method comprises: an acquisition module configured to acquire recorded audio in a target environment, wherein the target environment is a crowded square, a shopping mall or a street; an extraction module configured to extract at least one real voiceprint in the recorded audio, and perform deep learning on the at least one real voiceprint to obtain a voiceprint feature corresponding to the at least one real voiceprint; a generation module configured to model based on the voiceprint feature to obtain at least one voiceprint model, and generate an awakening word audio of the at least one voiceprint by using the at least one voiceprint model.
8. An electronic device, comprising: The method comprises: a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for generating an awakening word audio according to any one of claims 1-6.
9. A computer readable storage medium having stored thereon a computer program or instructions, characterized in that, The computer program or instructions are executed by the processor to implement the method for generating an awakening word audio according to any one of claims 1-6.
10. A computer program product comprising computer programs or instructions, characterized in that, The computer program or instructions are executed to implement the method for generating an awakening word audio according to any one of claims 1-6.
Citation Information
Patent Citations
Vocal print noise reduction method and system based on machine learning and deep learning
CN108831440A
Voiceprint confirmation model training method and device, electronic equipment and storage medium
CN114898757A