Voice wake-up model construction method, device and equipment

By generating a noisy reverberant speech dataset and selecting data that meets the fidelity requirements to train the wake-up model, the problem of low accuracy of voice wake-up devices in complex acoustic environments is solved, and cost-effective improvement in model adaptability is achieved.

CN121306136APending Publication Date: 2026-01-09ZHEJIANG FUTURE ELF ARTIFICIAL INTELLIGENCE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511668872.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing voice wake-up devices have low wake-up accuracy in complex acoustic environments and noisy scenarios, resulting in poor robustness and user experience, and high data collection costs.

Method used

By acquiring a speech dataset, a noise dataset, and a room impact response dataset, a noisy reverberant speech dataset is generated. A second noisy reverberant speech dataset is then generated through the front-end speech signal processing module of the target device. Data that meets the fidelity requirements are selected to form a training set, thereby learning the wake-up model.

Benefits of technology

It effectively reduces data collection costs, enhances the model's adaptability to various real-world application environments, and improves wake-up accuracy and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121306136A_ABST
    Figure CN121306136A_ABST
Patent Text Reader

Abstract

The invention discloses a voice wake-up model construction method, device and equipment. According to the method, under the condition that real environment voice data returned by equipment is limited, noise-containing reverberation voice data is generated through simulation, the real equipment is used for carrying out front-end voice signal processing on the noise-containing reverberation voice data, high-fidelity noise-containing reverberation wake-up voice data processed by the real equipment is screened as a positive training sample, and the positive training sample is used as a negative training sample. And taking noise-containing reverberation non-wake-up voice data processed by real equipment as a negative training sample, and learning from the positive training sample and the negative training sample to obtain a wake-up model. By adopting the processing mode, under the condition that the real environment voice data returned by the equipment is limited, the voice data acquisition cost can be effectively reduced, the adaptability of the model to various practical application environments is enhanced, and the awakening accuracy of the model in various acoustic scenes under the complex noise background is ensured, so that the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, specifically to a method and apparatus for constructing a voice wake-up model, and electronic devices. Background Technology

[0002] Voice wake-up refers to the pre-installation of a wake word in a smart device. When a user utters the corresponding voice command, the smart device can be activated and respond to subsequent voice commands. Devices with voice wake-up functionality use a wake-up model to identify whether the captured voice contains a wake word. A typical approach to building a wake-up model involves generating training data from a large amount of real-world recordings fed back from the device, and then learning the model parameters from this training data.

[0003] However, the inventors of this application discovered at least the following problems with existing solutions during the development of this application: When the device is not deployed for practical application, it is usually impossible to obtain real-world environmental voice data returned by the device. To ensure the voice wake-up model is adaptable to various practical application environments, real-world environmental voice data needs to be obtained through manual recording. Since data acquisition and model training rely on a large amount of real-world recordings, this results in high manual recording costs and makes it difficult to cover the diverse acoustic conditions and noise effects encountered in actual use. This leads to limited versatility and environmental adaptability of the voice wake-up device, especially in complex acoustic environments and noisy scenarios where wake-up accuracy drops significantly. These problems seriously affect the robustness and user experience of the voice wake-up device. Summary of the Invention

[0004] This application provides a method for constructing a voice wake-up model to address the problem in existing technologies that cannot simultaneously achieve low voice data acquisition costs and high model robustness. This application also provides a voice wake-up model construction apparatus and an electronic device.

[0005] This application provides a method for constructing a voice wake-up model, including: Acquire a voice dataset, a noise dataset, and a room impact response dataset for a target acoustic scene. The voice dataset includes wake-up voice data and non-wake-up voice data. Based on the speech dataset, noise dataset, and room impact response dataset, a first noisy band reverberant speech dataset is generated. The front-end speech signal processing module of the target device generates a second noisy reverberant speech dataset based on the first noisy reverberant speech dataset. Based on the fidelity-related data of the second noisy band reverberant speech data corresponding to the wake-up speech data, obtain the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data. A training dataset is formed based on the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data and the second noisy band reverberant speech data that corresponds to the non-wake-up speech data. The wake-up model is learned from the training dataset.

[0006] Optionally, obtain the room shock response dataset for the target acoustic scene, including: Collect at least one first-room impact response dataset for a first acoustic scenario; Based on the first room impact response dataset, generate at least one second room impact response dataset for a second acoustic scenario. The training dataset is formed based on the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data and the second noisy band reverberant speech data that corresponds to the non-wake-up speech data, including: A third training dataset is formed based on the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data and the second noisy band reverberant speech data that corresponds to the non-wake-up speech data. The method further includes: Based on the speech dataset and the room impact response dataset of the first acoustic scene, a speech dataset with reverberation for the first acoustic scene is generated as the first training dataset. Based on the speech dataset and the room impact response dataset of the second acoustic scene, a reverberant speech dataset of the second acoustic scene is generated as the second training dataset. The wake-up model learned from the training dataset includes: The wake-up model is trained based on the first training dataset to obtain the first parameter set; The second parameter set is obtained by training the wake-up model using the first parameter set based on the second training dataset. The second parameter set is obtained by training the wake-up model using the second parameter set based on the third training dataset.

[0007] Optionally, obtain the speech dataset, including: A speech dataset is generated from text using speech synthesis.

[0008] This application provides a method for constructing a voice wake-up model, including: Acquire the speech dataset and collect at least one first room impact response dataset for a first acoustic scenario; Based on the first room impact response dataset, generate at least one second room impact response dataset for a second acoustic scenario. Based on the speech dataset and the first room impact response dataset, a reverberant speech dataset for the first acoustic scene is generated; and based on the speech dataset and the second room impact response dataset, a reverberant speech dataset for the second acoustic scene is generated. A wake-up model is learned from at least one reverberant speech dataset of a first acoustic scenario and at least one reverberant speech dataset of a second acoustic scenario.

[0009] Optionally, learning the wake-up model from at least one reverberant speech dataset of a first acoustic scene and at least one reverberant speech dataset of a second acoustic scene includes: The wake-up model is trained based on a reverberant speech dataset of at least one first acoustic scene to obtain a first parameter set; The second parameter set is obtained by training a wake-up model using the first parameter set based on a reverberant speech dataset of at least one second acoustic scenario.

[0010] Optional, also includes: Obtain the noisy dataset; Based on the reverberant speech dataset and noise dataset of the first acoustic scene, generate a noisy reverberant speech dataset of the first acoustic scene; and / or, based on the reverberant speech dataset and noise dataset of the second acoustic scene, generate a noisy reverberant speech dataset of the second acoustic scene. The wake-up model is learned from at least one reverberant speech dataset of a first acoustic scene and at least one reverberant speech dataset of a second acoustic scene, including: A wake-up model is learned from at least one noisy reverberant speech dataset of a first acoustic scene and at least one noisy reverberant speech dataset of a second acoustic scene, or at least one noisy reverberant speech dataset of a first acoustic scene and at least one reverberant speech dataset of a second acoustic scene, or at least one reverberant speech dataset of a first acoustic scene and at least one noisy reverberant speech dataset of a second acoustic scene.

[0011] Optionally, learning the wake-up model from at least one noisy reverberant speech dataset and a reverberant speech dataset of a first acoustic scene, and at least one noisy reverberant speech dataset and a reverberant speech dataset of a second acoustic scene, includes: The wake-up model is trained based on a reverberant speech dataset of at least one first acoustic scene to obtain a first parameter set; The wake-up model using the first parameter set is trained based on a reverberant speech dataset of at least one second acoustic scene to obtain the second parameter set; A third parameter set is obtained by training a wake-up model using a second parameter set based on at least one noisy reverberant speech dataset of a first acoustic scenario and at least one noisy reverberant speech dataset of a second acoustic scenario.

[0012] Optional, also includes: Noisy reverberant speech data is input into the front-end speech signal processing module of the target device to obtain noisy reverberant speech data after front-end speech signal processing of the target device, which is used as training data for the wake-up model.

[0013] Optional, also includes: Obtain the fidelity-related data of the noisy reverberant speech data; Based on fidelity-related data, the noisy reverberant speech data is filtered to obtain noisy reverberant speech data that meets the fidelity requirements.

[0014] Optionally, the voice dataset includes: wake-up voice data and non-wake-up voice data; The step of filtering the noisy reverberant speech data based on fidelity-related data includes: Based on fidelity-related data, the noisy reverberant speech data corresponding to the wake-up speech data is filtered so that the noisy reverberant speech data corresponding to the non-wake-up speech data can be used as negative samples in model training.

[0015] Optionally, obtaining the voice dataset includes: A speech dataset is generated from text using speech synthesis.

[0016] This application provides a voice wake-up model construction device, including: The data acquisition unit is used to acquire a voice dataset, a noise dataset, and a room impact response dataset of the target acoustic scene. The voice dataset includes wake-up voice data and non-wake-up voice data. The data simulation unit is used to generate a first noisy reverberant speech dataset based on the speech dataset, noise dataset, and room impact response dataset. The device processing simulation unit is used to generate a second noisy reverberant speech dataset based on the first noisy reverberant speech dataset through the front-end speech signal processing module of the target device. The data filtering unit is used to obtain the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data based on the fidelity-related data of the second noisy band reverberant speech data corresponding to the wake-up speech data. The training data acquisition unit is used to form a training dataset based on the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data and the second noisy band reverberant speech data that corresponds to the non-wake-up speech data. The model learning unit is used to learn the wake-up model from the training dataset.

[0017] This application provides a voice wake-up model construction device, including: The data acquisition unit is used to acquire a voice dataset and collect at least one first room impact response dataset for a first acoustic scenario; The data simulation unit is used to generate at least one second room impact response dataset for a second acoustic scene based on the first room impact response dataset; to generate a reverberant speech dataset for the first acoustic scene based on the speech dataset and the first room impact response dataset; and to generate a reverberant speech dataset for the second acoustic scene based on the speech dataset and the second room impact response dataset. A model learning unit is used to learn a wake-up model from at least one reverberant speech dataset of a first acoustic scene and at least one reverberant speech dataset of a second acoustic scene.

[0018] This application provides an electronic device, including: Processor; and A memory for storing a program for implementing the method described in any of the preceding methods, wherein the device is powered on and the program of the method is executed by the processor.

[0019] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the various methods described above.

[0020] This application also provides a computer program product including instructions that, when run on a computer, cause the computer to perform the various methods described above.

[0021] Compared with the prior art, this application has the following advantages: The voice wake-up model construction method provided in this application involves acquiring a voice dataset, a noise dataset, and a room impact response dataset for a target acoustic scene. The voice dataset includes wake-up voice data and non-wake-up voice data. A first noisy reverberant voice dataset is generated based on the voice dataset, noise dataset, and room impact response dataset. A second noisy reverberant voice dataset is generated based on the first noisy reverberant voice dataset using the front-end voice signal processing module of the target device. Second noisy reverberant voice data that meets fidelity requirements corresponding to the wake-up voice data is obtained based on fidelity-related data of the second noisy reverberant voice data corresponding to the wake-up voice data. A training dataset is formed based on the second noisy reverberant voice data that meets fidelity requirements corresponding to the wake-up voice data and the second noisy reverberant voice data corresponding to the non-wake-up voice data. A wake-up model is learned from the training dataset. This approach allows for the generation of noisy reverberant speech data through simulation, even when real-world speech data from the device is limited. Real devices are then used to process this noisy reverberant speech data at the front end. High-fidelity noisy reverberant wake-up speech data processed by the real device is selected as positive training samples, while noisy reverberant non-wake-up speech data processed by the real device is used as negative training samples. The wake-up model is then learned from both positive and negative training samples. Therefore, this effectively reduces data acquisition costs, enhances the model's adaptability to various real-world application environments, and ensures the model's wake-up accuracy in various acoustic scenarios with complex noise backgrounds, thereby improving the user experience.

[0022] The voice wake-up model construction method provided in this application involves acquiring a voice dataset and collecting at least one first room impact response dataset for a first acoustic scenario; generating at least one second room impact response dataset for a second acoustic scenario based on the first room impact response dataset; generating a reverberant voice dataset for the first acoustic scenario based on the voice dataset and the first room impact response dataset; and generating a reverberant voice dataset for the second acoustic scenario based on the voice dataset and the second room impact response dataset; and learning a wake-up model from the reverberant voice datasets of at least one first acoustic scenario and at least one second acoustic scenario. This approach enhances the model's adaptability to unknown reverberation conditions by simulating room impact response data and reverberant voice data, even when the amount of real-world voice data returned by the device is limited. Therefore, it effectively reduces data acquisition costs, enhances the model's adaptability to various practical application environments, and improves user experience. Attached Figure Description

[0023] Figure 1 This is a flowchart illustrating an embodiment of the voice wake-up model construction method provided in this application; Figure 2 This is a schematic diagram illustrating the specific process of an embodiment of the voice wake-up model construction method provided in this application; Figure 3 This is another specific flowchart illustrating an embodiment of the voice wake-up model construction method provided in this application; Figure 4 This is a flowchart illustrating an embodiment of the voice wake-up model construction method provided in this application; Figure 5 This is a schematic diagram illustrating the specific process of an embodiment of the voice wake-up model construction method provided in this application. Detailed Implementation

[0024] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.

[0025] This application provides a method, apparatus, and system for constructing a voice wake-up model, as well as an electronic device. The various solutions are described in detail below in each embodiment.

[0026] First Embodiment Please refer to Figure 1 This is a flowchart of the voice wake-up model construction method of this application. In this embodiment, the method may include the following steps: Step S101: Obtain the voice dataset and collect at least one first room impact response dataset for the first acoustic scene.

[0027] The method provided in this application can be executed by a device with voice wake-up function (hereinafter referred to as a voice wake-up device), or other devices (such as servers) capable of executing the method provided in this application. The voice wake-up device can be a smart speaker, smart air conditioner, smart TV, voice assistant, etc.

[0028] In one example, the voice wake-up device includes a voice wake-up model. The voice wake-up device collects voice data, uses the voice wake-up model to identify whether the voice data contains a device wake-up word, and if the device wake-up word is identified, it responds to the user's command.

[0029] In another example, the voice wake-up device does not include a voice wake-up model. The voice wake-up device collects voice data, uploads it to a server, and uses the server's voice wake-up model to identify whether the voice data contains a device wake-up word. The server then sends the identification result back to the voice wake-up device. If the identification result includes the device wake-up word, the voice wake-up device responds to the user's command.

[0030] In practical applications, voice wake-up devices can be used in a variety of environments. For example, smart speakers can be placed in anechoic chambers, everyday home environments (living room, bedroom), conference rooms, open-plan offices, and semi-outdoor spaces. Different locations typically have different acoustic environments (such as reverberation and noise). Voice wake-up models learned from training data from various acoustic environments are adaptable to a wide range of practical applications; conversely, voice wake-up models learned from training data from a relatively singular acoustic environment are only adaptable to that specific environment and cannot be used in other environments.

[0031] Before voice wake-up devices are put into practical use, it is usually impossible to obtain real-world ambient voice data returned by the device. To ensure that the model training corpus covers diverse acoustic environments and reduce manual recording costs, the method provided in this application generates training corpus for various acoustic environments through data simulation. To this end, the method provided in this application first acquires a voice dataset and collects the room impulse response of the voice wake-up device in one or more typical usage scenarios (first acoustic scenario).

[0032] The voice data obtained by the method provided in this application embodiment can be voice data with reverberation or voice data without reverberation (clean voice data).

[0033] In one example, obtaining a speech dataset can be achieved by generating the speech dataset from the text using speech synthesis. Text-to-speech (TTS) is a technology that automatically converts text information into natural human speech. This approach ensures that the acquired speech data is clean, eliminating the need for reverberation removal; therefore, it effectively improves the efficiency of training data acquisition, thereby enhancing model building efficiency. Furthermore, this approach allows for recording only the RIR data of the first scene, eliminating the need to record the actual speech data; thus, it further reduces manual recording costs.

[0034] In another example, obtaining a speech dataset can be achieved by: acquiring a speech dataset with reverberation; and then removing the reverberation from the reverberated speech dataset. Specifically, this can be achieved through crowdsourcing. Crowdsourced speech is a method of processing or collecting speech-related tasks using a crowdsourcing model. In crowdsourced speech, a large number of speech tasks are distributed to numerous participants; these tasks can include speech data collection, speech transcription, and speech annotation. Crowdsourcing allows for the collection of speech datasets from different languages ​​and accents around the world, enriching the training data for speech recognition technology and providing strong support for the development of speech technology.

[0035] Room Impulse Response (RIR) refers to the output response of a room to a very short "impact signal" (such as clapping). It fully records the entire process of sound reflection, scattering, and attenuation within the room, and is data describing the room's acoustic characteristics. The method provided in this application collects the room impulse response of a voice wake-up device in one or more typical usage scenarios (first acoustic scenario). For example, smart speakers can be applied in various practical environments. If typical usage scenarios include ordinary residences and medium-sized conference rooms, then the room impulse responses of ordinary residences and medium-sized conference rooms are recorded.

[0036] In common home and office environments where voice wake-up devices are prevalent, RIR recordings were performed on scenarios with typical acoustic characteristics. In one example, the acoustic scenario was defined by multiple dimensions, including: location, position within the location, room size, architectural acoustic materials, etc., to ensure that the collected RIR data has sufficient spatial and material diversity. Locations included anechoic chambers, everyday home environments (living room, bedroom), conference rooms, open-plan offices, and semi-outdoor spaces, to cover diverse acoustic conditions ranging from strong to weak reflections; positions within the location included the center of a table, a corner, and variations in distance from the wall; room sizes included small, medium, and large rooms; and architectural acoustic materials included wooden floors, carpets, plasterboard walls, and glass windows. Table 1 below shows the room impulse responses of the voice wake-up devices collected in various first acoustic scenarios according to embodiments of this application.

[0037]

[0038] Table 1. Room impact response under various first acoustic scenarios Step S103: Generate at least one second room impact response dataset for a second acoustic scenario based on the first room impact response dataset.

[0039] Due to the high cost and limited physical space coverage of on-site RIR acquisition, measured RIR data cannot fully cover diverse application scenarios. To address this, the method provided in this application introduces simulated RIR generation technology, which can extend the virtual acoustic environment (one or more second acoustic scenes) through acoustic simulation engines (such as Image Method, Ray Tracing, or physics-based sound field simulation tools), thereby enhancing the model's environmental generalization ability.

[0040] In step S103, based on RIR data (first room impact response dataset) acquired (e.g., recorded) from one or more actual application scenarios (first acoustic scenario), RIR data (second room impact response dataset) for one or more other application scenarios (second acoustic scenario) are simulated and generated. The acquired RIR data is the measured RIR data, and the simulated RIR data is the simulated RIR data. The combined measured and simulated RIR data cover a variety of acoustic scenarios.

[0041] In one example, step S103 can be implemented as follows: setting multi-dimensional scene feature data for the second acoustic scene; generating at least one second room impact response dataset for the second acoustic scene based on the first room impact response dataset and the multi-dimensional scene feature data. The multi-dimensional scene features are used to describe the acoustic scene, including but not limited to one or more of the following dimensions: location, position within the location, room size, building acoustic materials, etc., ensuring that the simulated RIR data has sufficient spatial and material diversity.

[0042] For example, based on the typical acoustic scenario parameters in Table 1 above, acoustic simulation methods can be used to further expand the room impulse response (RIR) data. By reasonably configuring the room geometry (including length, width, and height), the sound absorption coefficients of each surface (corresponding to different materials such as walls, floors, and ceilings), and the relative distance and spatial position combination between the sound source and the microphone, a statistically representative virtual RIR dataset can be generated in batches using a physical acoustic simulation engine (such as the mirror method, ray tracing, or finite element method).

[0043] Step S105: Generate a reverberant speech dataset for the first acoustic scene based on the speech dataset and the first room impact response dataset; and generate a reverberant speech dataset for the second acoustic scene based on the speech dataset and the second room impact response dataset.

[0044] This step generates reverberant speech datasets for multiple acoustic scenarios based on the speech dataset and RIR data from multiple acoustic scenarios. The multiple acoustic scenarios include one or more first acoustic scenarios and one or more second acoustic scenarios. The RIR data for these multiple acoustic scenarios includes real RIR data and virtual RIR data. The virtual RIR data is diverse speech data generated based on extended RIRs, covering a wider range of reverberation times (e.g., T60) and spatial configurations of reverberant speech samples.

[0045] In practice, step S105 can be implemented using relatively mature existing technologies, such as convolving RIR data with speech data to generate reverberant speech data with real reverberation characteristics, so as to approximate the actual acoustic propagation process.

[0046] Step S107: Learn a wake-up model from at least one reverberant speech dataset of a first acoustic scene and at least one reverberant speech dataset of a second acoustic scene.

[0047] This step combines simulated reverberant speech datasets from real-world acoustic scenarios with simulated reverberant speech datasets from simulated acoustic scenarios to train the wake-up model, resulting in a wake-up model that is better adapted to various real-world application environments.

[0048] In practical implementation, mature existing technologies, such as supervised or unsupervised machine learning, can be used to learn the wake-up model from simulated training corpora in multiple scenarios. When processing input reverberant speech data, the wake-up model can first extract acoustic features (such as Mel spectrum, FBANK, or Fbank+Pitch) to train the basic model. Based on the trained wake-up model, syllable-level classification can be performed on audio in real-world application scenarios, and decoding strategies can be used to determine whether a wake-up word is contained within it.

[0049] In one example, the method provided in this application embodiment may further include the following steps: acquiring syllable classification data of speech data; step S107 may be implemented as follows: learning a wake-up model from at least one set of correspondences between reverberant speech data and syllable classification data in a first acoustic scenario and at least one set of correspondences between reverberant speech data and syllable classification data in a second acoustic scenario. This processing method allows for supervised model training; therefore, it can effectively improve syllable classification accuracy.

[0050] Please refer to Figure 2 This is a flowchart illustrating the voice wake-up model construction method of this application. In one example, step S107 may include the following sub-steps: Step S1071: Train the wake-up model based on at least one reverberant speech dataset of the first acoustic scene to obtain the first parameter set.

[0051] The reverberant speech dataset of the first acoustic scene is a training corpus that closely resembles the real application environment. This step S1071 is the training of the basic model. The aim is to introduce real scene (one or more first acoustic scenes) RIR modeling, so that the model can initially learn the acoustic spatial characteristics of the target application field, improve its syllable modeling accuracy in the real environment, and build a basic model that is initially adapted to typical usage environments.

[0052] In one example, step S1071 can be implemented as follows: The wake-up model is trained using a speech dataset and at least one reverberant speech dataset of a first acoustic scene to obtain a first parameter set. This approach allows the basic wake-up model to be trained by combining the original speech with simulated reverberant speech from the first acoustic scene; therefore, it can effectively improve the model's generalization ability.

[0053] Step S1073: Train the wake-up model using the first parameter set based on at least one reverberant speech dataset of the second acoustic scene to obtain the second parameter set.

[0054] Step S1073, based on scene augmentation data (at least one reverberant speech dataset of a second acoustic scene), fine-tunes the base model obtained in step S1071 (if supervised), enhancing the model's adaptability to unknown reverberation conditions. This step effectively compensates for the lack of diversity in experimental data and significantly improves the model's robustness and generalization performance in non-training reverberant environments.

[0055] Please refer to Figure 3 This is another specific flowchart of the voice wake-up model construction method of this application. In another example, the method provided by the embodiments of this application may further include: Step S201: Obtain the noise dataset.

[0056] Noise data refers to irregular sounds in the environment that interfere with normal activities; it is also known as environmental noise. Noise datasets can include a variety of typical background noises, such as traffic noise, machine noise, construction noise, thunder, rain noise, animal calls, etc.

[0057] Step S203: Generate a noisy reverberant speech dataset for the first acoustic scene based on the reverberant speech dataset and noise dataset for the first acoustic scene; and / or, generate a noisy reverberant speech dataset for the second acoustic scene based on the reverberant speech dataset and noise dataset for the second acoustic scene.

[0058] This step superimposes various typical background noises onto the reverberant speech data generated by the aforementioned simulation, covering a signal-to-noise ratio range of 0–15 dB. In specific implementations, only the noisy reverberant speech dataset of the first acoustic scene can be generated, or only the noisy reverberant speech dataset of the second acoustic scene can be generated, or both the noisy reverberant speech datasets of the first and second acoustic scenes can be generated simultaneously.

[0059] Accordingly, step S107 can be implemented as follows: learn a wake-up model from at least one noisy reverberant speech dataset of a first acoustic scene and at least one noisy reverberant speech dataset of a second acoustic scene, or from at least one noisy reverberant speech dataset of a first acoustic scene and at least one reverberant speech dataset of a second acoustic scene, or from at least one reverberant speech dataset of a first acoustic scene and at least one noisy reverberant speech dataset of a second acoustic scene.

[0060] In specific implementation, a wake-up model can be learned from at least one noisy reverberant speech dataset of a first acoustic scene and at least one noisy reverberant speech dataset of a second acoustic scene. Alternatively, a wake-up model can be learned from at least one noisy reverberant speech dataset of a first acoustic scene and at least one reverberant speech dataset of a second acoustic scene. Or, a wake-up model can be learned from at least one reverberant speech dataset of a first acoustic scene and at least one noisy reverberant speech dataset of a second acoustic scene.

[0061] The method provided in this application, through the specific implementation of the above steps S201, S203 and S107, introduces simulated noisy and reverberant speech data modeling, which can effectively improve the wake-up stability and discrimination ability of the model under low signal-to-noise ratio conditions.

[0062] In one example, step S107 may also include the following sub-steps: Step S1075: Train the wake-up model using the second parameter set based on at least one noisy reverberant speech dataset of the first acoustic scene and at least one noisy reverberant speech dataset of the second acoustic scene to obtain the third parameter set.

[0063] Step S1075 is a noise adaptation step based on steps S1071 and S1073 above, which can improve the stability and discrimination ability of the model under low signal-to-noise ratio conditions.

[0064] In one example, the method provided in this application embodiment may further include step S301: inputting noisy reverberant speech data to the front-end speech signal processing module of the target device to obtain noisy reverberant speech data after front-end speech signal processing of the target device. In specific implementation, the wake-up model is learned from at least one noisy reverberant speech dataset of a first acoustic scene and at least one noisy reverberant speech dataset of a second acoustic scene. This can be achieved in the following ways: The wake-up model is learned from the processed reverberant speech dataset of at least one first acoustic scene and at least one processed reverberant speech dataset of a second acoustic scene. The wake-up model is also learned from the same datasets. This processing method allows the simulated noisy and reverberant speech data to be input into a front-end speech signal processing module (such as AEC echo cancellation, NS noise suppression, AGC automatic gain control, etc.) that is identical or similar to the target voice wake-up device. This simulates the complete processing chain of the front-end input signal of the real device, generating speech data that approximates the input of the real device. This makes the positive example posterior output of the model more consistent with the actual application scenario. Therefore, it can further improve the adaptability of the wake-up model to the target device and enhance the stability and discrimination ability of the model when used on the target device under low signal-to-noise ratio conditions.

[0065] In one example, the method provided in this application embodiment may further include the following steps: obtaining fidelity-related data of the noisy reverberant speech data; and filtering the noisy reverberant speech data according to the fidelity-related data to obtain noisy reverberant speech data that meets the fidelity conditions. Fidelity-related data refers to data that reflects the fidelity of the speech signal in the noisy reverberant speech data, such as scale-invariant signal-to-noise ratio (SI-SNR). In specific implementations, the fidelity conditions can be determined according to application requirements. For example, the fidelity condition could be that the scale-invariant signal-to-noise ratio of the noisy reverberant speech data is greater than a scale-invariant signal-to-noise ratio threshold. This processing method enables data quality assessment and screening based on fidelity. Fidelity is used to evaluate the fidelity of speech signals with reverberation and noise bands, selecting high-fidelity samples with complete speech information for model training, thus avoiding the negative impact of low-quality data. Therefore, the quality of training data can be further improved, thereby improving the quality of the wake-up model, namely the accuracy of syllable classification and the stability of wake-up in complex noise environments.

[0066] In one example, the method provided in this application embodiment includes step S401: obtaining the scale-invariant signal-to-noise ratio (SNR) of the noisy reverberant speech data (noisy reverberant speech data processed by the front-end speech signal processing module or noisy reverberant speech data generated by simulation); and filtering the noisy reverberant speech data according to the SNR to obtain noisy reverberant speech data that meets the fidelity requirements. This processing method enables data quality assessment and filtering based on SI-SNR. SI-SNR is used to evaluate the fidelity of the noisy reverberant speech signal, selecting high-fidelity samples with complete speech information for model training, thus avoiding the negative impact of low-quality data. Therefore, the quality of training data can be further improved, thereby improving the quality of the wake-up model, i.e., the accuracy and stability of syllable classification in complex noise environments.

[0067] In practice, the noisy reverberant speech data generated in step S105 can be filtered based on the scale-invariant signal-to-noise ratio, instead of filtering the processed noisy reverberant speech data obtained in step S301.

[0068] In one example, the speech dataset includes not only speech data containing device wake words (wake-up speech data) but also speech data without device wake words (non-wake-up speech data). In this case, step S401 can be implemented as follows: Noisy reverberant speech data corresponding to the wake-up speech data is filtered so that noisy reverberant speech data corresponding to the non-wake-up speech data is used as negative samples in model training. Specifically, the non-wake-up speech data and pure noise data can undergo the same RIR superposition (simulation to generate reverberant speech data) and noise superposition (simulation to generate noisy reverberant speech data) processes described above, and even the same front-end processing procedures described above. When filtering the simulated noisy reverberant speech data based on the scale-invariant signal-to-noise ratio, noisy reverberant speech data from both the non-wake-up speech data and pure noise data can be retained and used as negative samples in model training. This approach enhances the wake-up model's discriminative ability by introducing negative samples, improving the model's ability to distinguish between wake words and non-wake-up content under noise interference. This noise adaptation strategy not only simulates the input characteristics of real devices, but also significantly improves the accuracy of syllable classification and wake-up stability of the model in complex noise environments through high-quality sample screening and negative example modeling.

[0069] As can be seen from the above embodiments, the voice wake-up model construction method provided in this application acquires a voice dataset and collects at least one first room impact response dataset for a first acoustic scenario; generates at least one second room impact response dataset for a second acoustic scenario based on the first room impact response dataset; generates a reverberant voice dataset for the first acoustic scenario based on the voice dataset and the first room impact response dataset; and generates a reverberant voice dataset for the second acoustic scenario based on the voice dataset and the second room impact response dataset; and learns a wake-up model from at least one reverberant voice dataset for the first acoustic scenario and at least one reverberant voice dataset for the second acoustic scenario. This processing method enhances the model's adaptability to unknown reverberation conditions by simulating room impact response data and reverberant voice data, even when the device has limited real-world voice data feedback. Therefore, it effectively reduces data acquisition costs, enhances the model's adaptability to various practical application environments, and thus improves user experience.

[0070] Second Embodiment In the above embodiments, a method for constructing a voice wake-up model is provided. Correspondingly, this application also provides a device for constructing a voice wake-up model. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0071] This application also provides a voice wake-up model construction device, including: The data acquisition unit is used to acquire the voice dataset and collect the first room impact response dataset of the first acoustic scenario; The data simulation unit is used to generate at least one second room impact response dataset for a second acoustic scene based on the first room impact response dataset; to generate a reverberant speech dataset for the first acoustic scene based on the speech dataset and the first room impact response dataset; and to generate a reverberant speech dataset for the second acoustic scene based on the speech dataset and the second room impact response dataset. The model learning unit is used to learn a wake-up model from a reverberant speech dataset of a first acoustic scene and a reverberant speech dataset of at least one second acoustic scene.

[0072] In one example, the model learning unit is specifically used to train a wake-up model based on a reverberant speech dataset of at least one first acoustic scenario to obtain a first parameter set; and to train a wake-up model using the first parameter set based on a reverberant speech dataset of at least one second acoustic scenario to obtain a second parameter set.

[0073] In one example, the apparatus further includes: a noise data acquisition unit for acquiring a noise dataset; a noise simulation unit for generating a noisy reverberant speech dataset for a first acoustic scene based on the reverberant speech dataset and the noise dataset for a first acoustic scene; and / or, generating a noisy reverberant speech dataset for a second acoustic scene based on the reverberant speech dataset and the noise dataset for a second acoustic scene; the model learning unit is specifically used to learn a wake-up model from at least one noisy reverberant speech dataset for a first acoustic scene and at least one noisy reverberant speech dataset for a second acoustic scene, or from at least one noisy reverberant speech dataset for a first acoustic scene and at least one reverberant speech dataset for a second acoustic scene, or from at least one reverberant speech dataset for a first acoustic scene and at least one noisy reverberant speech dataset for a second acoustic scene.

[0074] In one example, the model learning unit is specifically used to train a wake-up model based on at least one reverberant speech dataset of a first acoustic scene to obtain a first parameter set; to train a wake-up model using the first parameter set based on at least one reverberant speech dataset of a second acoustic scene to obtain a second parameter set; and to train a wake-up model using the second parameter set based on at least one noisy reverberant speech dataset of a first acoustic scene and at least one noisy reverberant speech dataset of a second acoustic scene to obtain a third parameter set.

[0075] In one example, the apparatus further includes: a device processing simulation unit, used to input noisy reverberant speech data to the front-end speech signal processing module of the target device to obtain noisy reverberant speech data after front-end speech signal processing of the target device, as training data for the wake-up model.

[0076] In one example, the apparatus further includes: a data filtering unit, configured to acquire fidelity-related data of the noisy reverberant speech data; and to perform data filtering on the noisy reverberant speech data based on the fidelity-related data to obtain noisy reverberant speech data that meets the fidelity conditions.

[0077] In one example, the speech dataset includes: wake-up speech data and non-wake-up speech data; the data filtering unit is specifically used to filter the noisy reverberant speech data corresponding to the wake-up speech data according to fidelity-related data, so that the noisy reverberant speech data corresponding to the non-wake-up speech data can be used as negative samples to participate in model training.

[0078] In one example, obtaining the speech dataset includes: generating a speech dataset from text using speech synthesis.

[0079] Third Embodiment In the above embodiments, a method for constructing a voice wake-up model is provided. Correspondingly, this application also provides a method for constructing a voice wake-up model. This device corresponds to Embodiment 1 of the above method. Since Embodiment 3 is basically similar to Embodiment 1, it is described simply, and relevant parts can be referred to in the description of Embodiment 1. The method embodiments described below are merely illustrative.

[0080] Please refer to Figure 4 This is a flowchart of the voice wake-up model construction method of this application. This application also provides a voice wake-up model construction method, including: Step S401: Obtain a voice dataset, a noise dataset, and a room impact response dataset for the target acoustic scene. The voice dataset includes wake-up voice data and non-wake-up voice data.

[0081] Step S403: Generate the first noisy reverberant speech dataset based on the speech dataset, noise dataset, and room impact response dataset.

[0082] In specific implementation, step S403 can be implemented using the method described in Embodiment 1. For example, first, a reverberant speech dataset is generated based on the speech dataset and the room impact response dataset; then, a noisy reverberant speech dataset is generated based on the noise dataset and the reverberant speech dataset. As another example, a noisy speech dataset is generated first based on the speech dataset and the noise dataset; then, a noisy reverberant speech dataset is generated based on the room impact response dataset and the noisy speech dataset.

[0083] Step S405: The front-end speech signal processing module of the target device generates a second noisy reverberant speech dataset based on the first noisy reverberant speech dataset.

[0084] Step S405 is consistent with the signal processing performed by the front-end voice signal processing module of the target device in the first embodiment of the method, and will not be described again here.

[0085] Step S407: Based on the fidelity-related data of the second noisy band reverberant speech data corresponding to the wake-up speech data, obtain the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data.

[0086] Step S407 is consistent with the high-fidelity data screening in Method Example 1, and will not be repeated here.

[0087] Step S409: Form a training dataset based on the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data and the second noisy band reverberant speech data that corresponds to the non-wake-up speech data.

[0088] Step S409 is consistent with the process of forming positive and negative training samples in Method Example 1, and will not be repeated here.

[0089] Step S411: Learn the wake-up model from the training dataset.

[0090] Step S411 can learn the wake-up model from the training dataset obtained in step S409. Using this approach, only one model training step is needed, effectively improving model building efficiency.

[0091] Please refer to Figure 5This is a flowchart illustrating the specific process of constructing the voice wake-up model according to this application. In one example, step S401, obtaining the room impact response dataset for the target acoustic scene, includes: acquiring at least one first room impact response dataset for a first acoustic scene; and generating at least one second room impact response dataset for a second acoustic scene based on the first room impact response dataset. This approach allows for the acquisition of RIR values ​​from a portion of real-world scenarios and the simulation of RIR values ​​from another portion, thereby extending the RIR simulation to multiple scenarios and enhancing the model's adaptability to unknown reverberation conditions.

[0092] Accordingly, step S409 can be implemented as follows: A third training dataset is formed based on the second noisy reverberant speech data corresponding to the wake-up speech data and satisfying the fidelity condition, and the second noisy reverberant speech data corresponding to the non-wake-up speech data. Correspondingly, the method may further include step S501: generating a reverberant speech dataset for the first acoustic scene based on the speech dataset and the room impact response dataset for the first acoustic scene, as the first training dataset; step S503: generating a reverberant speech dataset for the second acoustic scene based on the speech dataset and the room impact response dataset for the second acoustic scene, as the second training dataset; step S411 can be implemented as follows: training the wake-up model using the first training dataset to obtain a first parameter set; training the wake-up model using the first parameter set using the second training dataset to obtain a second parameter set; training the wake-up model using the second parameter set using the third training dataset to obtain a second parameter set. This three-iteration training approach allows for several improvements. First, training on a simulated reverberant speech dataset of real-world scenarios improves syllable classification accuracy. Second, training on the same dataset enhances the model's adaptability to various acoustic environments. Finally, training on high-fidelity noisy, reverberant wake-up speech (positive examples) and noisy, reverberant non-wake-up speech (negative examples) from all scenarios improves wake-up stability in complex noisy environments. In summary, with limited real-world speech data from devices, collecting only a small number of scene RIRs ensures accurate wake-up in various acoustic scenarios with complex noise backgrounds.

[0093] In one example, obtaining the speech dataset can be achieved by generating the speech dataset from the text using speech synthesis. This approach automates the generation of speech data, further reducing the cost of manual recording.

[0094] As can be seen from the above embodiments, the voice wake-up model construction method provided in this application acquires a voice dataset, a noise dataset, and a room impact response dataset of a target acoustic scene. The voice dataset includes wake-up voice data and non-wake-up voice data. Based on the voice dataset, noise dataset, and room impact response dataset, a first noisy reverberant voice dataset is generated. Through the front-end voice signal processing module of the target device, a second noisy reverberant voice dataset is generated based on the first noisy reverberant voice dataset. Based on the fidelity-related data of the second noisy reverberant voice data corresponding to the wake-up voice data, second noisy reverberant voice data that meets the fidelity condition corresponding to the wake-up voice data is obtained. Based on the second noisy reverberant voice data that meets the fidelity condition corresponding to the wake-up voice data and the second noisy reverberant voice data corresponding to the non-wake-up voice data, a training dataset is formed. A wake-up model is learned from the training dataset. This approach allows for the generation of noisy reverberant speech data through simulation, even when real-world speech data from the device is limited. Real devices are then used to process this noisy reverberant speech data at the front end. High-fidelity noisy reverberant wake-up speech data processed by the real device is selected as positive training samples, while noisy reverberant non-wake-up speech data processed by the real device is used as negative training samples. The wake-up model is then learned from both positive and negative training samples. Therefore, this effectively reduces data acquisition costs, enhances the model's adaptability to various real-world application environments, and ensures the model's wake-up accuracy in various acoustic scenarios with complex noise backgrounds, thereby improving the user experience.

[0095] Fourth embodiment In the above embodiments, a method for constructing a voice wake-up model is provided. Correspondingly, this application also provides a device for constructing a voice wake-up model. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0096] This application also provides a voice wake-up model construction device, including: The data acquisition unit is used to acquire a voice dataset, a noise dataset, and a room impact response dataset of the target acoustic scene. The voice dataset includes wake-up voice data and non-wake-up voice data. The data simulation unit is used to generate a first noisy reverberant speech dataset based on the speech dataset, noise dataset, and room impact response dataset. The device processing simulation unit is used to generate a second noisy reverberant speech dataset based on the first noisy reverberant speech dataset through the front-end speech signal processing module of the target device. The data filtering unit is used to obtain the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data based on the fidelity-related data of the second noisy band reverberant speech data corresponding to the wake-up speech data. The training data acquisition unit is used to form a training dataset based on the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data and the second noisy band reverberant speech data that corresponds to the non-wake-up speech data. The model learning unit is used to learn the wake-up model from the training dataset.

[0097] In one example, the data acquisition unit is specifically used to collect at least one first room impact response dataset for a first acoustic scene; and to generate at least one second room impact response dataset for a second acoustic scene based on the first room impact response dataset. Correspondingly, the training data acquisition unit is specifically used to form a third training dataset based on second noisy reverberant speech data corresponding to wake-up speech data and second noisy reverberant speech data corresponding to non-wake-up speech data that meets fidelity conditions. Accordingly, the device may further include: a first training dataset generation unit, used to generate a reverberant speech dataset for the first acoustic scene based on the speech dataset and the room impact response dataset for the first acoustic scene, as the first training dataset; a second training dataset generation unit, used to generate a reverberant speech dataset for the second acoustic scene based on the speech dataset and the room impact response dataset for the second acoustic scene, as the second training dataset; the model learning unit is specifically used to train a wake-up model based on the first training dataset to obtain a first parameter set; train a wake-up model using the first parameter set based on the second training dataset to obtain a second parameter set; and train a wake-up model using the second parameter set based on the third training dataset to obtain a second parameter set.

[0098] In one example, the data acquisition unit is specifically used to generate a speech dataset from text through speech synthesis.

[0099] Fifth embodiment In the above embodiments, a method for constructing a voice wake-up model is provided. Correspondingly, this application also provides an electronic device. This device corresponds to the embodiments of the above method. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The device embodiments described below are merely illustrative.

[0100] The electronic device of this embodiment includes: a memory and a processor; the memory is used to store a program for implementing the voice wake-up model construction method, and the device is powered on and runs the program of the above-mentioned voice wake-up model construction method through the processor.

[0101] Memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0102] In specific implementations, the electronic device may further include one or more of the following components: a power supply component, an input / output (I / O) interface, and a communication component. The power supply component provides power to various components of the electronic device. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device. The I / O interface provides an interface between the processor 503 and peripheral interface modules, which may be a keyboard, click wheel, buttons, etc. The communication component is configured to facilitate wired or wireless communication between the electronic device and user devices (such as smartphones, tablets, etc.).

[0103] Sixth Embodiment This application also provides a computer-readable storage medium. Since the embodiments of the computer-readable storage medium are substantially similar to the method embodiments, the description is relatively simple; relevant details can be found in the description of the method embodiments. The computer-readable storage medium embodiments described below are merely illustrative.

[0104] In this embodiment, a non-transitory computer-readable storage medium including instructions is provided, such as a memory including instructions, which can be executed by a processor of an electronic device to complete the voice wake-up model construction method provided by this disclosure. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0105] It should be noted that the embodiments of this application may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., with the user's explicit consent, with the user being properly notified, etc.).

[0106] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.

[0107] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0108] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0109] 1. Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.

[0110] 2. Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

Claims

1. A method for constructing a voice wake-up model, characterized in that, include: Acquire a voice dataset, a noise dataset, and a room impact response dataset for a target acoustic scene. The voice dataset includes wake-up voice data and non-wake-up voice data. Based on the speech dataset, noise dataset, and room impact response dataset, a first noisy band reverberant speech dataset is generated. The front-end speech signal processing module of the target device generates a second noisy reverberant speech dataset based on the first noisy reverberant speech dataset. Based on the fidelity-related data of the second noisy band reverberant speech data corresponding to the wake-up speech data, obtain the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data. A training dataset is formed based on the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data and the second noisy band reverberant speech data that corresponds to the non-wake-up speech data. The wake-up model is learned from the training dataset.

2. The method according to claim 1, characterized in that, Obtain the room shock response dataset for the target acoustic scene, including: Collect at least one first-room impact response dataset for a first acoustic scenario; Based on the first room impact response dataset, generate at least one second room impact response dataset for a second acoustic scenario. The method further includes: Based on the speech dataset and the room impact response dataset of the first acoustic scene, a speech dataset with reverberation for the first acoustic scene is generated as the first training dataset. Based on the speech dataset and the room impact response dataset of the second acoustic scene, a reverberant speech dataset of the second acoustic scene is generated as the second training dataset. The training dataset is formed based on the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data and the second noisy band reverberant speech data that corresponds to the non-wake-up speech data, including: A third training dataset is formed based on the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data and the second noisy band reverberant speech data that corresponds to the non-wake-up speech data. The wake-up model learned from the training dataset includes: The wake-up model is trained based on the first training dataset to obtain the first parameter set; The second parameter set is obtained by training the wake-up model using the first parameter set based on the second training dataset. The second parameter set is obtained by training the wake-up model using the second parameter set based on the third training dataset.

3. The method according to claim 2, characterized in that, Obtain the speech dataset, including: A speech dataset is generated from text using speech synthesis.

4. A method for constructing a voice wake-up model, characterized in that, include: Acquire the speech dataset and collect at least one first room impact response dataset for a first acoustic scenario; Based on the first room impact response dataset, generate at least one second room impact response dataset for a second acoustic scenario. Based on the speech dataset and the first room impact response dataset, a reverberant speech dataset for the first acoustic scene is generated; and based on the speech dataset and the second room impact response dataset, a reverberant speech dataset for the second acoustic scene is generated. A wake-up model is learned from at least one reverberant speech dataset of a first acoustic scenario and at least one reverberant speech dataset of a second acoustic scenario.

5. The method according to claim 4, characterized in that, The wake-up model is learned from at least one reverberant speech dataset of a first acoustic scene and at least one reverberant speech dataset of a second acoustic scene, including: The wake-up model is trained based on a reverberant speech dataset of at least one first acoustic scene to obtain a first parameter set; The second parameter set is obtained by training a wake-up model using the first parameter set based on a reverberant speech dataset of at least one second acoustic scenario.

6. The method according to claim 4, characterized in that, Also includes: Obtain the noisy dataset; Based on the reverberant speech dataset and noise dataset of the first acoustic scene, generate a noisy reverberant speech dataset of the first acoustic scene. And / or, generate a noisy reverberant speech dataset for the second acoustic scene based on the reverberant speech dataset and the noise dataset for the second acoustic scene; The wake-up model is learned from at least one reverberant speech dataset of a first acoustic scene and at least one reverberant speech dataset of a second acoustic scene, including: A wake-up model is learned from at least one noisy reverberant speech dataset of a first acoustic scene and at least one noisy reverberant speech dataset of a second acoustic scene, or at least one noisy reverberant speech dataset of a first acoustic scene and at least one reverberant speech dataset of a second acoustic scene, or at least one reverberant speech dataset of a first acoustic scene and at least one noisy reverberant speech dataset of a second acoustic scene.

7. The method according to claim 6, characterized in that, The wake-up model is learned from at least one noisy band reverberant speech dataset of a first acoustic scene and at least one noisy band reverberant speech dataset of a second acoustic scene, including: A wake-up model is learned from at least one noisy band reverberant speech dataset and one reverberant speech dataset of a first acoustic scenario, and at least one noisy band reverberant speech dataset and one reverberant speech dataset of a second acoustic scenario, including: The wake-up model is trained based on a reverberant speech dataset of at least one first acoustic scene to obtain a first parameter set; The wake-up model using the first parameter set is trained based on a reverberant speech dataset of at least one second acoustic scenario to obtain the second parameter set. A third parameter set is obtained by training a wake-up model using a second parameter set based on at least one noisy reverberant speech dataset of a first acoustic scenario and at least one noisy reverberant speech dataset of a second acoustic scenario.

8. The method according to claim 6, characterized in that, Also includes: Noisy reverberant speech data is input into the front-end speech signal processing module of the target device to obtain noisy reverberant speech data after front-end speech signal processing of the target device, which is used as training data for the wake-up model.

9. The method according to any one of claims 6 to 8, characterized in that, Also includes: Obtain the fidelity-related data of the noisy reverberant speech data; Based on fidelity-related data, the noisy reverberant speech data is filtered to obtain noisy reverberant speech data that meets the fidelity requirements.

10. The method according to claim 9, characterized in that, The voice dataset includes: wake-up voice data and non-wake-up voice data; The step of filtering the noisy reverberant speech data based on fidelity-related data includes: Based on fidelity-related data, the noisy reverberant speech data corresponding to the wake-up speech data is filtered so that the noisy reverberant speech data corresponding to the non-wake-up speech data can be used as negative samples in model training.

11. The method according to claim 4, characterized in that, The step of generating at least one second room impact response dataset for a second acoustic scenario based on the first room impact response dataset includes: Set up multi-dimensional scene feature data for the second acoustic scene; Based on the first room impact response dataset and multi-dimensional scene feature data, generate at least one second room impact response dataset for a second acoustic scene.

12. A voice wake-up model construction device, characterized in that, include: The data acquisition unit is used to acquire a voice dataset, a noise dataset, and a room impact response dataset of the target acoustic scene. The voice dataset includes wake-up voice data and non-wake-up voice data. The data simulation unit is used to generate a first noisy reverberant speech dataset based on the speech dataset, noise dataset, and room impact response dataset. The device processing simulation unit is used to generate a second noisy reverberant speech dataset based on the first noisy reverberant speech dataset through the front-end speech signal processing module of the target device. The data filtering unit is used to obtain the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data based on the fidelity-related data of the second noisy band reverberant speech data corresponding to the wake-up speech data. The training data acquisition unit is used to form a training dataset based on the second noisy band reverberant speech data that meets the fidelity condition corresponding to the wake-up speech data and the second noisy band reverberant speech data that corresponds to the non-wake-up speech data. The model learning unit is used to learn the wake-up model from the training dataset.

13. A voice wake-up model construction device, characterized in that, include: The data acquisition unit is used to acquire a voice dataset and collect at least one first room impact response dataset for a first acoustic scenario; The data simulation unit is used to generate at least one second room impact response dataset for a second acoustic scene based on the first room impact response dataset; to generate a reverberant speech dataset for the first acoustic scene based on the speech dataset and the first room impact response dataset; and to generate a reverberant speech dataset for the second acoustic scene based on the speech dataset and the second room impact response dataset. A model learning unit is used to learn a wake-up model from at least one reverberant speech dataset of a first acoustic scene and at least one reverberant speech dataset of a second acoustic scene.

14. An electronic device, characterized in that, include: processor; as well as A memory for storing a program for implementing the method according to any one of claims 1 to 11, wherein the device is powered on and the program for running the method is executed by the processor.