A device control method, a recognition model training method, and related devices

By collecting and processing the target user's voice, and utilizing recognition models and multi-channel training technology, the problem of noise interference in voice wake-up control was solved, enabling accurate wake-up and control of the device, and improving the accuracy of voice recognition and the quality of user voice.

CN122090830APending Publication Date: 2026-05-26ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG GEELY HLDG GRP CO LTD
Filing Date
2026-03-02
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, voice wake-up control functions are susceptible to environmental noise and interference, which prevents devices from accurately recognizing control information and responding to user commands in a timely manner. Furthermore, existing recognition models cannot adapt to multi-channel environments, lack the ability to utilize fine information, and have insufficient noise resistance.

Method used

The system collects the target user's voice and inputs it into the recognition model. It outputs control voice and target voice channels or time-frequency masking codes. Clear voice is transmitted through channels with high signal-to-noise ratio. The system combines multi-channel voice training models to construct multi-channel voice samples, trains the recognition model to improve recognition accuracy, and suppresses interference noise through sound source localization.

Benefits of technology

It improves the accuracy of recognizing control information in voice, suppresses interference noise, enables accurate device wake-up and control, and enhances the quality of user voice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090830A_ABST
    Figure CN122090830A_ABST
Patent Text Reader

Abstract

This application provides a device control method, a recognition model training method, and related devices. The device control method includes: acquiring speech to be processed emitted by a target user; inputting the speech to be processed into a recognition model to obtain recognition information output by the recognition model, wherein the recognition information includes control speech and a target speech channel, or the recognition information includes control speech, a target speech channel, and a time-frequency masking code, wherein the target speech channel is a speech transmission channel with a signal-to-noise ratio greater than a signal-to-noise ratio threshold, and the time-frequency masking code is used for sound source localization of the target user; and waking up or controlling the device's functions based on the recognition information. This application improves the accuracy of recognizing control information in speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and in particular to a device control method, a recognition model training method, and related equipment. Background Technology

[0002] With the acceleration of intelligentization, voice wake-up control has become a basic function of robots, in-vehicle systems, smart homes, and various portable terminals.

[0003] In the exemplary technology, the voice wake-up control function provides users with a seamless and natural interactive experience by allowing the device to listen to voice in the environment and then activating subsequent interaction processes in the device after recognizing control information in the voice.

[0004] However, the presence of echoes, noise, and interference in the environment causes the voice being monitored by the device to contain echoes, noise, and interference, making it impossible to accurately identify control information from the voice and causing the device to fail to respond to the user's voice commands in a timely manner. Summary of the Invention

[0005] Based on the above-mentioned technological status, this application provides a device control method, a recognition model training method, and related equipment to solve the problem of low recognition accuracy of control information in speech.

[0006] To achieve the above-mentioned technical objectives, this application proposes the following technical solution: In a first aspect, this application provides a device control method, comprising: Collect the voice messages to be processed from the target user; The speech to be processed is input into the recognition model to obtain the recognition information output by the recognition model. The recognition information includes control speech and target speech channel, or the recognition information includes control speech, target speech channel and time-frequency masking code. The target speech channel is a speech transmission channel with a signal-to-noise ratio greater than the signal-to-noise ratio threshold. The time-frequency masking code is used to locate the sound source of the target user. Based on the identification information, the device's functions can be activated or controlled.

[0007] In some implementations, the identification information includes control voice and target voice channel, and the step of waking up or controlling the device's functions based on the identification information includes: The target voice channel is determined from the identification information; Based on the speech processing method associated with the target speech channel, the control speech is processed to obtain the target speech; The target voice is sent to the server based on the target voice channel, so that the server can wake up or control the functions of the device based on the target voice.

[0008] In some implementations, the identification information includes control voice, target voice channel, and time-frequency masking code. The step of waking up or controlling the device's functions based on the identification information includes: The device's functions are activated or controlled based on the control voice and the target voice channel. The sound source of the target user is located based on the time-frequency masking code.

[0009] Secondly, this application provides a method for training a recognition model, including: Multiple prompt voices containing user voice elements are acquired, and a target control voice corresponding to the prompt voice is generated based on the user voice elements of the prompt voices and a specified control phrase. The user voice elements include at least one of the first user's timbre, pitch, and speech rate. A multi-channel speech is constructed based on the target control speech and the reference speech, and training samples corresponding to the multi-channel speech are constructed. The reference speech includes at least one of interference sound, echo, and noise. The preset model is trained based on each of the training samples to obtain a recognition model. The recognition model is used to recognize information based on the voice output of the second user. The recognition information is used to wake up or control the functions of the device. The recognition information includes control voice and target voice channel, or the recognition information includes control voice, target voice channel and time-frequency masking code. The target voice channel is a voice transmission channel with a signal-to-noise ratio greater than the signal-to-noise ratio threshold. The time-frequency masking code is used to locate the sound source of the second user.

[0010] In some implementations, generating the target control voice corresponding to the prompt voice based on the user voice elements of the prompt voice and the specified control phrases includes: Based on the user voice elements of the prompt voice and the specified control phrases, an initial control voice is generated; Different types of text conversion models are used to perform text recognition on the initial control speech to obtain multiple text contents; Determine the intersection text among multiple text contents, and determine the matching degree between the intersection text and the control phrase; In response to the matching degree being less than a preset threshold, the initial control voice is determined as the target control voice.

[0011] In some implementations, constructing the training samples corresponding to the multi-channel speech includes: The multi-channel voice is input into a pre-trained preset model to obtain the wake-up rate of the device by the multi-channel voice. In response to the wake-up rate being lower than the wake-up threshold, training samples are constructed based on the multi-channel speech.

[0012] In some implementations, constructing multi-channel speech based on the target control speech and reference speech includes: The target control voice is subjected to preset processing to obtain processed voice. The preset processing includes changing the sound elements in the target control voice and / or splicing any prompt voice with the target control voice. The multi-channel speech is constructed based on the processed speech and the reference speech.

[0013] In some implementations, acquiring multiple prompt voices containing user voice elements includes: Obtain the original voice recordings of first users in different geographical regions; Each of the original voice recordings is processed to obtain multiple prompt voice recordings. The duration of each prompt voice recording is within a preset duration range, and the number of statements from the first user contained in each prompt voice recording is within a preset number range.

[0014] Thirdly, this application provides an electronic device, including a memory and a processor, wherein, The memory is connected to the processor and is used to store programs; The processor is configured to implement the device control method as described in the first aspect or any implementation thereof, or the recognition model training method as described in the second aspect or any implementation thereof, by running a program in the memory.

[0015] Fourthly, this application provides a computer program product, which, when executed by a processor, implements the device control method as described in the first aspect or any implementation thereof, or implements the recognition model training method as described in the second aspect or any implementation thereof.

[0016] Fifthly, this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the device control method as described in the first aspect or any implementation thereof, or implements the recognition model training method as described in the second aspect or any implementation thereof.

[0017] This application provides a device control method, a recognition model training method, and related equipment. The method involves acquiring voice input from a target user, inputting the voice into a recognition model to obtain recognition information, and then using control information within the recognition information to wake up or control the device's functions. In this application, the recognition information includes control voice and a target voice channel. The target voice channel is a voice transmission channel with a signal-to-noise ratio (SNR) greater than a threshold. Therefore, even if the voice to be processed contains noise, clear control voice can be transmitted through the target voice channel, enabling the device to be woken up or controlled based on the clear control voice, thus improving the accuracy of recognizing control information in the voice. Furthermore, the recognition information may also include a time-frequency masking code for sound source localization of the user. Therefore, directional voice acquisition of the user through sound source localization can suppress interference noise and improve the user's voice quality, further enhancing the accuracy of recognizing control information in the voice. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0019] Figure 1 A flowchart of a device control method provided in this application embodiment Figure 1 .

[0020] Figure 2 A flowchart of a device control method provided in this application embodiment Figure 2 .

[0021] Figure 3 This is a diagram of the speech system architecture used in conjunction with the recognition model and sound source localization involved in this application.

[0022] Figure 4 The flowchart of a model training method provided in the embodiments of this application Figure 1 .

[0023] Figure 5 The flowchart of a model training method provided in the embodiments of this application Figure 2 .

[0024] Figure 6 The flowchart of a model training method provided in the embodiments of this application Figure 3 .

[0025] Figure 7 The flowchart of a model training method provided in the embodiments of this application Figure 4 .

[0026] Figure 8 This is a schematic diagram illustrating the generation of target control voice in an embodiment of this application.

[0027] Figure 9 This is a schematic diagram of the structure of the recognition model involved in the embodiments of this application.

[0028] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0030] It should be noted that the user information (including but not limited to electrical equipment information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0031] With the acceleration of intelligentization, voice wake-up control has become a basic function of robots, in-vehicle systems, smart homes, and various portable terminals.

[0032] In the exemplary technology, the voice wake-up control function provides users with a seamless and natural interactive experience by allowing the device to listen to voice in the environment and then activating subsequent interaction processes in the device after recognizing control information in the voice.

[0033] However, the presence of echoes, noise, and interference in the environment causes the voice being monitored by the device to contain echoes, noise, and interference, making it impossible to accurately identify control information from the voice and causing the device to fail to respond to the user's voice commands in a timely manner.

[0034] Furthermore, building a system with high accuracy in recognizing control information faces multiple challenges in practice. First, at the data level, training control information recognition models heavily relies on large-scale, high-quality real-person speech data. This data collection needs to cover different accents, speech rates, and pronunciation habits, resulting in high costs and lengthy cycles, becoming a major bottleneck restricting rapid model iteration and application expansion. Second, at the functional level, existing models take single-channel data as input and typically output a simple binary judgment result: "wake-up" or "not wake-up." Single-channel input cannot fully utilize multi-channel speech information, and the trained model cannot well adapt to wake-up in different multi-channel environments.

[0035] Models trained on single-channel data simplify the output method. While they achieve basic trigger judgment, they lose a large amount of valuable and detailed information in the speech signal, such as the precise start and end positions of the wake word (a type of control information) on the time axis, or its information in high-dimensional feature space. This lack of information prevents downstream tasks (such as sound source localization and directional noise reduction) from fully utilizing the speech features of the wake word itself, treating it only as an isolated switch signal, thus limiting further improvement in the performance of the entire interactive system.

[0036] At the algorithm model level, many existing recognition models are single-channel wake-up, which cannot be well adapted to wake-up in multi-channel scenarios. At the same time, existing technologies directly output wake-up results and sound source localization information through a single model, and their output structure is closed, making it difficult to flexibly integrate and coordinate with other algorithm modules. In addition, in the decoupled structure, when detecting wake-up and performing sound source localization, only the start and end information of the wake word is simply used, without using the mask information of each time-frequency point, resulting in insufficient noise resistance.

[0037] To address the above problems, this application proposes a recognition model training method and a device control method. The device control method proposed in this application will be described in detail below with reference to various embodiments.

[0038] Reference Figure 1 , Figure 1 A flowchart of a device control method provided in this application embodiment Figure 1 Equipment control methods include: Step S101: Collect the voice to be processed sent by the target user.

[0039] In this embodiment, the executing entity is a device control unit, which can be any server or terminal device deployed with a recognition model. The recognition model is a model used to recognize the user's control voice. For ease of description, the term "control unit" will be used to refer to the device control unit below.

[0040] The control device is equipped with a voice acquisition module, which collects the voice emitted by the target user in real time. The collected voice is defined as the voice to be processed.

[0041] Step S102: Input the speech to be processed into the recognition model to obtain the recognition information output by the recognition model. The recognition information includes control speech and target speech channel, or the recognition information includes control speech, target speech channel and time-frequency masking code. The target speech channel is a speech transmission channel with a signal-to-noise ratio greater than the signal-to-noise ratio threshold. The time-frequency masking code is used to locate the sound source of the target user.

[0042] After obtaining the speech to be processed, the speech is input into the recognition model. The recognition model recognizes the speech and outputs the recognition information. The recognition information includes control speech and target speech channel, or the recognition information includes control speech, target speech channel and time-frequency masking code. The target speech channel is the speech transmission channel with a signal-to-noise ratio greater than the signal-to-noise ratio threshold. The time-frequency masking code is used to locate the sound source of the target user.

[0043] Step S103: Based on the identification information, wake up or control the functions of the device.

[0044] After receiving the identification information, the control device wakes up or controls the functions of the device based on the identification information.

[0045] In one example, control voice is extracted from the recognition information. This control voice can be a set of control words used to control the device. The control device transmits the control voice to the device via a target voice channel, enabling the device to wake up or control functions based on the control words in the control voice. For example, the control words might be "Xiao a Xiao a, play music A." If the control device finds that "Xiao a" is the wake-up word for device C, it sends the control words to device C via the target channel, causing device C to respond to the control words and play music A.

[0046] In another example, control voice is transmitted to the device through a target voice channel, and the target user is located based on a time-frequency masking code, thereby enabling the control device to locate and track the target user and then collect the voice emitted by the target user in a targeted manner.

[0047] In this embodiment, the voice to be processed emitted by the target user is collected, and the voice is input into a recognition model to obtain recognition information output by the model. Then, based on the control information in the recognition information, the device's functions are activated or controlled. In this embodiment, the recognition information includes control voice and a target voice channel. The target voice channel is a voice transmission channel with a signal-to-noise ratio greater than a threshold. Therefore, even if the voice to be processed contains noise, clear control voice can be transmitted through the target voice channel, thereby activating or controlling the device based on clear control voice, improving the accuracy of recognizing control information in the voice. Furthermore, the recognition information may also include a time-frequency masking code for sound source localization of the user. Therefore, directional voice acquisition of the user through sound source localization can suppress interference noise and improve the user's voice quality, further enhancing the accuracy of recognizing control information in the voice.

[0048] Reference Figure 2 , Figure 2 A flowchart of a device control method provided in this application embodiment Figure 2 .based on Figure 1 In the embodiment shown, step S103 includes: Step S201: Determine the target speech channel from the recognition information.

[0049] In this embodiment, the recognition information includes control speech and a target speech channel. The recognition model identifies the speech to be processed and also outputs the speech channel n' with the best signal-to-noise ratio. For example, the control device determines the target speech channel from the recognition information. The target speech channel is the speech channel with a signal-to-noise ratio greater than a signal-to-noise ratio threshold. If there are multiple speech channels with a signal-to-noise ratio greater than the threshold, the speech channel with the highest signal-to-noise ratio is selected as the target speech channel.

[0050] Step S202: Based on the speech processing method associated with the target speech channel, the control speech is processed to obtain the target speech.

[0051] After determining the target speech channel, the control speech is processed based on the speech processing method associated with the target speech channel to obtain the target speech.

[0052] For example, the target speech can be obtained by processing the control speech through sub-band synthesis and gain control of the target speech channel.

[0053] Step S203: Send the target voice to the server based on the target voice channel so that the server can wake up or control the device's functions based on the target voice.

[0054] After obtaining the target voice, the target voice is sent to the server so that the server can wake up or control the device's functions based on the target voice.

[0055] In this embodiment, the voice processing method associated with the target voice channel is used to process the control voice process, thereby enabling the server to obtain clearer voice and accurately control the device.

[0056] In one embodiment, the identification information includes control voice, target voice channel, and time-frequency masking code. After the control device wakes up or controls the device's functions based on the control voice and target voice channel, it can also locate the sound source of the target user using the time-frequency masking code. The specific steps of waking up or controlling the device's functions using control voice and target voice channel are described above and will not be repeated here. It should be noted that the step of waking up or controlling the device's functions is not sequential with sound source localization.

[0057] Reference Figure 3 , Figure 3 This is a diagram of a speech system architecture that uses a recognition model in conjunction with sound source localization. The speech system performs zone configuration, which refers to the parameter setting and optimization of the system's effective pickup and playback frequency range, frequency band sensitivity allocation, or frequency band division of multi-channel audio. The core purpose is to match the speech task requirements of specific scenarios and improve signal quality or recognition accuracy. After zone configuration, zone localization can be performed. Zone localization refers to determining the frequency distribution range of a sound signal or locating the characteristic frequency band corresponding to a sound source through algorithms or hardware design, thereby achieving targeted signal processing, noise suppression, or sound source identification. In the zone localization scenario, a control device is set up, and the sequence number of the located zone is sent to a cloud server.

[0058] The voice system also includes multiple microphones and multiple speakers. Speech is acquired through multiple microphones, decomposed into sub-bands from 48kHz to 15kHz, and then subjected to adaptive echo cancellation to obtain the speech to be processed. The speech to be processed is input to the recognition model of the control device in the sound region localization scenario. The recognition model outputs a wake-up mask, which includes a wake-up word, a channel, and a selection. The channel is the optimal voice channel, and the selection refers to the mask.

[0059] The speech to be processed is input into the recognition model and simultaneously input into the blind source separation component. This component separates the speech from four speech channels (for example, four channels, but not limited to four). The separated speech undergoes sub-band synthesis and gain control within each channel to obtain the separated interactive speech. The speech system then sends the interactive speech from the optimal speech channel (selected from the four channels) to the cloud server for subsequent interaction. The mask is used by the speech system for sound source localization.

[0060] In this embodiment, while recognizing the control voice output by the model, a mask is also output, thereby enabling sound source localization based on the mask, which enriches the functionality of the interactive system where the control device is located.

[0061] Reference Figure 4 , Figure 4 The flowchart of a recognition model training method provided in the embodiments of this application Figure 1 .like Figure 4 As shown, the recognition model training method provided in this embodiment includes: Step S301: Obtain multiple prompt voices containing user voice elements, and generate target control voices corresponding to the prompt voices based on the user voice elements of the prompt voices and the specified control phrases. The user voice elements include at least one of the first user's timbre, pitch, and speech rate.

[0062] In this embodiment, the execution entity is a recognition model training device, which can be any server or terminal device with model training capabilities. For ease of description, the term "device" will be used to refer to the recognition model training device below.

[0063] The device acquires multiple prompt voice messages, each containing user voice elements. These user voice elements include at least one of the first user's timbre, pitch, and speech rate. The first user refers to the actual person issuing the prompt voice message.

[0064] The prompt voice can be user voice collected from a database and used to control the device. The device extracts user voice elements from the prompt voice. It should be noted that user voice elements can not only reflect the first user's timbre, pitch, and speech rate, but also represent the first user's gender and age group. In order to cover different voice attribute dimensions such as timbre, speech rate, pitch, gender, and age, the device needs to acquire the voices of first users from different geographical regions, different genders, and different age groups as prompt voices, so that the user voice elements extracted from each prompt voice have diversity and breadth in acoustic features.

[0065] Furthermore, there are certain requirements for the prompt voice. For example, the device acquires the original voice of a first user in different geographical regions; that is, it retrieves device control voice from different storage areas in the voice database as the original voice. Each storage area stores the voice data of the first user in one region; the first user in different geographical regions refers to a user in a different geographical location. After obtaining multiple original voices, each original voice is processed to obtain multiple prompt voices. The duration of each prompt voice is within a preset duration range, and the number of sentences from the first user included in each prompt voice is within a preset number range. For example, the number of sentences from the first user included in each prompt voice is between 10 and 20, and the duration of each prompt voice is between 5 and 10 seconds. This method ensures that the user's voice elements in the prompt voice can be accurately identified and extracted, and that the duration of the prompt voice is not too long to avoid a large workload in constructing training samples.

[0066] After receiving the prompt voice, the user's voice elements are extracted from the prompt voice, and then combined with specified control words to form speech. The synthesized speech is defined as the target control speech. Control words are phrases that can be used to control or wake up device functions. For example, a control word could be "Xiao a, Xiao a" as a wake-up word, or "Play audio A" as a function control phrase. The device uses user voice elements to generate speech containing control words; that is, the generated speech contains user voice elements. It should be noted that there can be multiple prompt voices and multiple control words. Combining user voice elements from different prompt voices with different control words can yield multiple combinations. Each combination can generate one target control speech, meaning the device can obtain multiple target control voices.

[0067] Step S402: Construct multi-channel speech based on the target control speech and reference speech, and construct training samples corresponding to the multi-channel speech. The reference speech includes at least one of interference sound, echo, and noise.

[0068] After obtaining the target control speech, multi-channel speech can be constructed by combining the target control speech with reference speech. The reference speech includes at least one of interference sounds, echoes, and noise. For example, the target control speech and reference speech are synthesized to obtain synthesized speech, which is then simulated as a multi-channel microphone signal to obtain multi-channel speech. It should be noted that multi-channel refers to a multi-channel microphone using multiple spatially distributed pickup units. These pickup units are used to collect the spatial features of the speech, including the direction and distance of the sound source, and the sound field distribution.

[0069] After obtaining the multi-channel speech, training samples corresponding to the multi-channel speech can be constructed. For example, control words in the multi-channel speech can be used as labels, and the multi-channel speech and labels constitute the training samples.

[0070] Step S403: Train the preset model according to each training sample to obtain the recognition model. The recognition model is used to recognize information based on the voice output of the second user. The recognition information is used to wake up or control the function of the device. The recognition information includes control voice and target voice channel, or the recognition information includes control voice, target voice channel and time-frequency masking code. The target voice channel is a voice transmission channel with a signal-to-noise ratio greater than the signal-to-noise ratio threshold. The time-frequency masking code is used to locate the sound source of the second user.

[0071] There are multiple target control voices, and each target control voice can generate a training sample, thus resulting in multiple training samples. The device trains a preset model using these multiple training samples to obtain a recognition model. The preset model can be a deep learning network model or a convolutional neural network model. After obtaining the recognition model, the voice emitted by the second user is input into the recognition model. The recognition model outputs recognition information, which is used to wake up or control the device's functions. The recognition information includes the control voice and the target voice channel, or it includes the control voice, the target voice channel, and a time-frequency masking code. The target voice channel is a voice transmission channel with a signal-to-noise ratio greater than the signal-to-noise ratio threshold. The time-frequency masking code is used to locate the sound source of the second user. The specific applications of the control voice, target voice channel, and time-frequency masking code in the recognition information are explained above and will not be repeated here.

[0072] In this embodiment, multiple prompt voices containing user voice elements are acquired. A target control voice is generated based on these user voice elements, such as timbre, pitch, and speech rate, along with specified control phrases. A multi-channel voice model is constructed using the target control voice, reference voice containing interference, echoes, and noise, and training samples are built for each channel. A preset model is then trained using these training samples for each prompt voice to obtain a recognition model for identifying control information in user speech. In this embodiment, control voice is generated by combining user voice elements and control phrases from the prompt voices. Training samples are constructed based on the reference voice containing interference, echoes, and noise, enabling the recognition model trained on these samples to accurately identify control information from user speech containing interference, echoes, and noise, thus improving the accuracy of control information recognition in speech. Furthermore, the reference voice containing interference, echoes, and noise, along with the control voice, generates multi-channel voice, allowing the recognition model trained on the multi-channel voice samples to accurately identify control information in user speech even in a multi-channel voice environment, further improving the accuracy of control information recognition in speech.

[0073] Reference Figure 5 , Figure 5 The flowchart of a recognition model training method provided in the embodiments of this application Figure 2 ,based on Figure 1 In the embodiment shown, step S401 includes: Step S501: Generate initial control voice based on the user voice elements of the prompt voice and the specified control phrases.

[0074] In this embodiment, the device generates control voice based on the user voice elements of the prompt voice and the specified control phrases. This control voice is defined as the initial control voice. The specific generation of the control voice is described above and will not be repeated here.

[0075] Step S502: Using different types of text conversion models, the initial control speech is processed to perform text recognition, resulting in multiple text contents.

[0076] The device uses various text conversion models to perform text recognition on the initial control voice, thereby obtaining multiple text contents corresponding to the initial control voice.

[0077] Step S503: Determine the intersection text among multiple text contents, and determine the matching degree between the intersection text and the control phrase.

[0078] After obtaining multiple text contents, the intersection of these text contents is obtained, resulting in the intersection text. The phrases in the intersection text refer to phrases present in all the text contents.

[0079] After obtaining the intersection text, the matching degree between the intersection text and the control phrases is calculated. The control phrases refer to the control phrases used to synthesize the initial control speech. For example, the intersection text is converted into a first text feature vector, and the similarity between the first text feature vector and the second text vector converted from the control phrases is calculated. The matching degree is determined based on the similarity, and the greater the similarity, the greater the matching degree.

[0080] Step S504: In response to the matching degree being less than a preset threshold, the initial control voice is determined as the target control voice.

[0081] After determining the matching degree, the matching degree is compared with a preset threshold. If the matching degree is less than the preset threshold, it can be determined that the accuracy of the control words in the initial control speech is low. Therefore, it can be determined that control words with accent, timbre, and speech rate are difficult to recognize. The initial control speech is used as a hard sample for model training, so that the trained model can recognize control words that are not easy to recognize.

[0082] In this embodiment, multiple types of text conversion models are used to perform text recognition on control speech. By utilizing the intersection of different text contents, control speech with difficult-to-recognize control phrases is identified to construct training samples for training the model. This enables the trained model to recognize control phrases that are not easily recognized, thereby improving the recognition accuracy of the model in recognizing control phrases.

[0083] Figure 6 The flowchart of a recognition model training method provided in the embodiments of this application Figure 3 ,based on Figure 4 or Figure 5 In the embodiment shown, step S403 includes: Step S601: Input the multi-channel voice into the pre-trained preset model to obtain the wake-up rate of the device by the multi-channel voice.

[0084] In this embodiment, the preset model needs to undergo multiple rounds of training, with the model obtained from the previous round of training being the pre-trained preset model. To this end, the device inputs multi-channel voice into the preset model obtained from the previous round of training, thereby obtaining the wake-up rate of the device for multi-channel voice.

[0085] For example, multi-channel voice is input to a preset model multiple times. Each time the preset model recognizes the control information, it wakes up or controls the device. If the wake-up or control is successful, the success count is incremented by 1. The ratio between the success count and the number of times multi-channel voice is input to the preset model can be used as the wake-up rate.

[0086] Step S602: In response to the wake-up rate being lower than the wake-up threshold, training samples are constructed based on multi-channel speech.

[0087] After obtaining the wake-up rate, compare the wake-up rate with the wake-up threshold. If the wake-up rate is lower than the wake-up threshold, the control information in the multi-channel speech is difficult to recognize. Therefore, the multi-channel speech can be used as a hard sample, that is, to construct training samples corresponding to the multi-channel speech.

[0088] Existing sample generation methods rely on relatively simple generation models and lack effective data filtering mechanisms, resulting in insufficient diversity and inconsistent quality of training corpora. In this embodiment, corpus data is generated by fusing multiple TTS models, and quality filtering mechanisms such as speech duration and wake-up rate are used to significantly improve the scale and quality of the corpus.

[0089] Figure 7 The flowchart of a recognition model training method provided in the embodiments of this application Figure 4 .based on Figures 4 to 6 In any of the embodiments shown, step S401 includes: Step S701: Perform preset processing on the target control voice to obtain the processed voice. The preset processing includes changing the sound elements in the target control voice and / or splicing any prompt voice with the target control voice.

[0090] In this embodiment, after receiving the target control voice, the device performs preset processing on the target control voice to obtain processed voice. The preset processing includes at least one of the following: Modify the sound elements in the target controlled speech, including timbre, speech rate, and pitch. Changes to sound elements include voice alteration and speech rate alteration. Any prompt voice is spliced ​​with the target control voice. The prompt voice is stored in the target audio library, which contains a large number of TTS voices and a small number of real human voice recordings, all of which are prompt voices.

[0091] It should be noted that when concatenating the prompt speech and the target control speech, the resulting speech carries annotations for both the prompt speech and the target control speech. The annotation for the prompt speech refers to the intersection of multiple text content obtained by performing text recognition on the prompt speech using multiple text conversion models. The annotation for the target control speech is the corresponding control phrase. These annotations in the concatenated speech can be used as labels for training samples.

[0092] Step S702: Construct multi-channel speech based on the processed speech and the reference speech.

[0093] After obtaining the processed speech, a multi-channel speech is constructed based on the processed speech and the reference speech.

[0094] For example, refer to Figure 8A large amount of TTS voice and a small amount of real-person recorded voice are used as prompt voices as the target sound source. Prompt voices are randomly selected from the target sound source. The target control voice is generated by combining the user's voice elements and control phrases in the prompt voices. After the target control voice is changed in voice and speech rate, it is spliced ​​with any prompt voice in the target sound source to obtain the processed voice. The processed voice is then processed by the real or analog transfer function in the convolution module to obtain the target voice. The interference source contains multiple interference sounds. The real or analog transfer function in the convolution module processes the randomly selected interference sounds to obtain the interference speech. In addition, multiple interference sounds can be concatenated. The real or analog transfer function in the convolution module processes the concatenated interference sounds to obtain the interference speech. The reference sound source contains multiple sounds. The sounds are processed by real or analog transfer functions in the convolution module, and the processed sounds are then used to obtain echoes through nonlinear simulation. The noise source contains multiple types of noise, and the scattered noise is obtained by simulating the scattered noise. By randomly adjusting the signal-to-interference ratio, signal-to-return ratio, signal-to-noise ratio, and volume of the target speech, interference speech, echo, and scattered noise, and then synthesizing them in multiple channels, a multi-channel microphone signal can be simulated, which is the multi-channel speech.

[0095] In this embodiment, the target control speech is pre-processed and then synthesized with the reference speech to obtain a variety of multi-channel speech.

[0096] In one embodiment, the recognition model outputs not only control information but also a Mask (time-frequency masking code). Combined with... Figure 9 The network structure of the recognition model shown illustrates the training of the recognition model: 1. After the multi-channel speech in the training samples is processed by signal processing algorithms such as echo cancellation and speech separation, the Filter Bank features (Fbank features, acoustic features) are extracted and input into the recognition model.

[0097] 2. Input Fbank features into the recognition model. The recognition model training process includes the following main steps: 2.1 The input Fbank features are processed through frame stitching and frame skipping operations to obtain the features. Enter The structure maps features onto the input dimension of the FSMN unit, where... ; 2.2. Concatenate L-layer FSMN (Feedforward Sequential Memory Networks) units to model long-term temporal dependencies of keywords: ; 2.3. The last FSMN layer is subjected to max-pooling, converting multi-channel information into single-channel information and outputting the optimal channel. The optimal channel is calculated as follows: ; ; Where y represents the network output before max-pooling, n=1,…,N represents the channel number, and k=1,…,K represents the dimension of y. This represents the frame sequence number. The maximum value is 1 when it originates from this channel, and 0 otherwise.

[0098] 2.4. The single-channel signal features processed by max-pooling are output through two branches. One branch maps the output vector of the FSMN unit to a dimension space that matches the wake word pronunciation unit, and calculates the observation probability of each pronunciation unit. The other branch outputs a Mask (time-frequency masking code). Where k represents the frequency band number and t is the frame sequence number. The loss function of the model is defined as follows: ; in the formula and The calculation methods are as follows: ; ; Where C represents the number of wake word modeling units. To predict probabilities, For real labels, It is the output signal in the time-frequency domain, the wake-up mask output by the neural network. The effect on the input multi-channel data The above information is as follows: ; in, For the target speech signal, This indicates that the average value is calculated for all batches of data.

[0099] In this embodiment, a data processing system integrating TTS generation and intelligent filtering was constructed to address the reliance of model training on real-person data. This system can automatically generate massive amounts of speech covering different accents and speaking speeds using multiple large TTS models as needed. It filters the quality of the generated data based on preset speech duration, multiple ASR (Automatic Speech Recognition) transcription results, wake-up rate, and other metrics, thereby constructing a large-scale, high-quality corpus. This significantly shortens the model development and iteration cycle and substantially reduces data acquisition costs.

[0100] In addition, a multi-channel data simulation system is set up to train the recognition model of this application. This simulation system simulates single-channel, short audio source data into multi-channel, long audio data, and then uses it for model training after passing through signal processing algorithms such as echo cancellation and speech separation.

[0101] Furthermore, the recognition model in this application is a deep neural network with multi-branch output results. The input of the neural network is the Fbank feature of the microphone signal. Through the shared underlying network of the model, it outputs three results in parallel: a wake-up result, a wake-up mask (time-frequency masking code), and channel selection. The wake-up mask (time-frequency masking code) serves as accompanying information for the wake-up event, providing an effective intermediate representation for downstream processing modules.

[0102] It should be noted that the recognition model contains multiple ReLU (Affine) layers, which refer to affine transformation Affine + ReLU activation function. ReLU (Affine) layers are connected to FSMN units, and each FSMN unit is connected to a Max pooling layer. Max pooling connects to one FSMN unit. The Max pooling connection to the FSMN unit consists of a sequentially connected Linear layer, FSMN, and ReLU (Affine). Softmax (Affine) is affine transformation + Softmax activation function, and Sigmoid (Affine) is affine transformation + Sigmoid activation function.

[0103] Compared to exemplary techniques, the model training method provided in this application has advantages in several aspects, as follows: Regarding data generation: Existing methods rely on relatively simple generation models and lack effective data filtering mechanisms, resulting in insufficient diversity and inconsistent quality of training corpora. This application generates data by integrating multiple TTS models and introduces quality filtering mechanisms based on speech duration, multiple ASR models, and wake-up rate, which significantly improves the scale and quality of the corpus. Regarding data augmentation: The multi-channel data framework used in this application employs a large amount of TTS data and a small amount of real-person recording data for the original audio source data, which reduces the data acquisition cost and solves the problem of the scarcity of multi-channel data in practical applications. This method increases the richness of data while realizing the matching training of signals and models, thereby improving the model's performance.

[0104] Regarding the model structure: Existing models use a tightly coupled approach to handle wake-up and sound source localization, resulting in bloated models with poor flexibility. They simply utilize the start and end point information of wake-up, leading to insufficient noise resistance. The network structure of the recognition model in this application is decoupled. While outputting the wake-up discrimination result, it simultaneously generates the wake-up word mask (time-frequency masking code) information and channel selection result. This wake-up mask (time-frequency masking code) can be used as a general intermediate representation, flexibly supporting downstream tasks such as sound source localization, and enhancing the system's scalability and robustness.

[0105] This application provides a schematic diagram of the structure of an electronic device, see [link]. Figure 10 As shown, the electronic device includes a memory 1000 and a processor 1010; wherein the memory 1000 is connected to the processor 1010 and is used to store programs; the processor 1010 is used to implement the model training method or device control method disclosed in any of the above embodiments by running the programs stored in the memory 1000.

[0106] Specifically, the aforementioned electronic device may further include: a bus, a communication interface 1020, an input device 1030, and an output device 1040. The electronic device may also include a data transceiver module, an image monitoring module, and a signal monitoring module.

[0107] The processor 1010, memory 1000, communication interface 1020, input device 1030, and output device 1040 are interconnected via a bus. Among them: A bus can include a pathway for transmitting information between various components in an electronic device.

[0108] The processor 1010 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0109] The processor 1010 may include a main processor, as well as a baseband chip, modem, etc.

[0110] The memory 1000 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 1000 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.

[0111] Input device 1030 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.

[0112] Output device 1040 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.

[0113] The communication interface 1020 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0114] The processor 1010 executes the program stored in the memory 1000 and calls other devices, and can be used to implement each step of any of the model training methods provided in the above embodiments of this application.

[0115] It should be noted that the electronic device can be an in-vehicle terminal, a mobile phone, a wearable device, or a server, etc.; or it can be a vehicle that includes an in-vehicle terminal, etc.

[0116] This application also proposes a chip, which includes a processor and a data interface. The processor reads and runs a program stored in the memory through the data interface to execute the model training method or device control method described in any of the above embodiments. For the specific processing procedures and their beneficial effects, please refer to the above embodiments of the model training method and device control method.

[0117] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the model training method or device control method according to various embodiments of this application as described in any of the above embodiments of this specification.

[0118] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the power device, as a standalone firmware package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0119] Furthermore, embodiments of this application may also be storage media storing computer programs, which are executed by a processor to perform the steps of the model training methods according to various embodiments of this application described in any of the above embodiments of this specification, specifically implementing the steps of the above model training method or device control method.

[0120] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0121] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0122] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.

[0123] The units of the apparatus in the various embodiments of this application can be merged, divided, and deleted according to actual needs.

[0124] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0125] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.

[0126] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or as firmware functional modules or sub-modules.

[0127] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer firmware, or a combination of both. To clearly illustrate the interchangeability of hardware and firmware, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or firmware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0128] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly using hardware, firmware units executed by a processor, or a combination of both. The firmware unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0129] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0130] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A device control method, characterized in that, include: Collect the voice messages to be processed from the target user; The speech to be processed is input into the recognition model to obtain the recognition information output by the recognition model. The recognition information includes control speech and target speech channel, or the recognition information includes control speech, target speech channel and time-frequency masking code. The target speech channel is a speech transmission channel with a signal-to-noise ratio greater than the signal-to-noise ratio threshold. The time-frequency masking code is used to locate the sound source of the target user. Based on the identification information, the device's functions can be activated or controlled.

2. The equipment control method according to claim 1, characterized in that, The identification information includes control voice and target voice channel. The step of waking up or controlling the device's functions based on the identification information includes: The target voice channel is determined from the identification information; Based on the speech processing method associated with the target speech channel, the control speech is processed to obtain the target speech; The target voice is sent to the server based on the target voice channel, so that the server can wake up or control the functions of the device based on the target voice.

3. The equipment control method according to claim 1, characterized in that, The identification information includes control voice, target voice channel, and time-frequency masking code. The step of waking up or controlling the device's functions based on the identification information includes: The device's functions are activated or controlled based on the control voice and the target voice channel. The sound source of the target user is located based on the time-frequency masking code.

4. A method for training a recognition model, characterized in that, include: Multiple prompt voices containing user voice elements are acquired, and a target control voice corresponding to the prompt voice is generated based on the user voice elements of the prompt voices and a specified control phrase. The user voice elements include at least one of the first user's timbre, pitch, and speech rate. A multi-channel speech is constructed based on the target control speech and the reference speech, and training samples corresponding to the multi-channel speech are constructed. The reference speech includes at least one of interference sound, echo, and noise. The preset model is trained based on each of the training samples to obtain a recognition model. The recognition model is used to recognize information based on the voice output of the second user. The recognition information is used to wake up or control the functions of the device. The recognition information includes control voice and target voice channel, or the recognition information includes control voice, target voice channel and time-frequency masking code. The target voice channel is a voice transmission channel with a signal-to-noise ratio greater than the signal-to-noise ratio threshold. The time-frequency masking code is used to locate the sound source of the second user.

5. The recognition model training method according to claim 4, characterized in that, The process of generating the target control voice corresponding to the prompt voice based on the user voice elements of the prompt voice and the specified control phrases includes: Based on the user voice elements of the prompt voice and the specified control phrases, an initial control voice is generated; Different types of text conversion models are used to perform text recognition on the initial control speech to obtain multiple text contents; Determine the intersection text among multiple text contents, and determine the matching degree between the intersection text and the control phrase; In response to the matching degree being less than a preset threshold, the initial control voice is determined as the target control voice.

6. The recognition model training method according to claim 4, characterized in that, The construction of training samples corresponding to the multi-channel speech includes: The multi-channel voice is input into a pre-trained preset model to obtain the wake-up rate of the device by the multi-channel voice. In response to the wake-up rate being lower than the wake-up threshold, training samples are constructed based on the multi-channel speech.

7. The recognition model training method according to claim 4, characterized in that, The step of constructing multi-channel speech based on the target control speech and reference speech includes: The target control voice is subjected to preset processing to obtain processed voice. The preset processing includes changing the sound elements in the target control voice and / or splicing any prompt voice with the target control voice. The multi-channel speech is constructed based on the processed speech and the reference speech.

8. The recognition model training method according to any one of claims 4-7, characterized in that, The acquisition of multiple prompt voices containing user voice elements includes: Obtain the original voice recordings of first users in different geographical regions; Each of the original voice recordings is processed to obtain multiple prompt voice recordings. The duration of each prompt voice recording is within a preset duration range, and the number of statements from the first user contained in each prompt voice recording is within a preset number range.

9. An electronic device, characterized in that, Including memory and processor, among which, The memory is connected to the processor and is used to store programs; The processor is used to implement the device control method as described in any one of claims 1-3 or the recognition model training method as described in any one of claims 4-7 by running the program in the memory.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the device control method as described in any one of claims 1-3 or the recognition model training method as described in any one of claims 4-7.

11. A computer program product, characterized in that, When the computer program is executed by the processor, it implements the device control method as described in any one of claims 1-3 or the recognition model training method as described in any one of claims 4-7.