Model training method, voice wake-up method and electronic equipment
By using two models to identify voice data in the voice assistant, the problem of false wake-up or inability to wake up in the prior art is solved, and higher recognition accuracy and response speed are achieved.
Patent Information
- Application Number
- CN202510387344.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-05-06
AI Technical Summary
Existing voice wake-up methods are prone to error awakening or inability to wake up when identifying user-defined wake-up words.
Two models are used to identify speech data. The first model trains speech through user-defined wake words for personalized training. The second model can show higher accuracy in predicting wake words due to its larger model parameters and stronger learning ability.
It improves the accuracy of speech arousal, reduces the probability of false arousal, and enhances the generalization ability and response speed of the model.
Smart Images

Figure CN119943031A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of terminal technology, and in particular to a model training method, a voice wake-up method and an electronic device. Background Art
[0002] In daily use of mobile phones, users can wake up the voice assistant by pre-setting customized wake-up words to meet their personalized needs.
[0003] In the current implementation, some algorithms can use models to identify the voice stream collected by the microphone and determine whether the voice stream contains a specific wake-up word. The specific wake-up word introduced here can be a user-defined wake-up word. If the voice stream contains a specific wake-up word, the voice assistant can be awakened.
[0004] However, the current voice wake-up method still has the problem of false wake-up or failure to wake up the voice assistant. Summary of the invention
[0005] The present application provides a model training method, a voice wake-up method and an electronic device, which are applied in the field of terminal technology to improve the accuracy of voice wake-up.
[0006] In a first aspect, an embodiment of the present application proposes a model training method. The method comprises:
[0007] When the wake-up word of the voice assistant is adjusted to the first wake-up word, multiple sample training voices are obtained, and the sample training voices include at least the first training voice, and the first training voice contains the first wake-up word, and the first wake-up word is customized by the user.
[0008] For any one of the multiple sample training voices, the sample training voice is input into the first model so that the first model outputs a first prediction result, and the first prediction result is used to indicate whether the sample training voice contains the first wake-up word.
[0009] The sample training speech is input into the second model so that the second model outputs a second prediction result, and the second prediction result is used to indicate whether the sample training speech contains the first wake-up word, wherein the model parameters of the second model are greater than the model parameters of the first model.
[0010] The model parameters of the first model are adjusted according to the first prediction result, the second prediction result and the true label of the sample training speech, where the true label is used to indicate whether the sample training speech contains the first wake-up word.
[0011] In this embodiment, by using sample training speech containing user-defined wake-up words for training, the first model can better capture and understand the voice features of these specific wake-up words. This personalized training method not only meets the diverse needs of users, but also enhances the generalization ability of the model. Regardless of the environment the user is in or the tone of voice used, the model can more accurately identify the wake-up words, thereby improving the response speed and user experience of the voice assistant.
[0012] Secondly, choosing to update the first model locally on the electronic device effectively reduces the pressure on the server. With the popularity and increase in the number of smart devices, the amount of data that needs to be processed on the server side is increasing. Delegating the model update task to the local device can not only alleviate the load on the server, but also reduce the delay in data transmission, thereby improving overall efficiency. This is especially important for voice interaction scenarios that require instant response.
[0013] In addition, given the relatively weak learning ability of the first model, the second model is introduced as a guide. The second model is able to show higher accuracy in predicting the wake-up word due to its larger model parameters and stronger learning ability. By letting the second model guide the parameter adjustment of the first model, the first model can learn the knowledge and experience of the second model.
[0014] In one possible implementation, the moment when the wake-up word of the voice assistant is adjusted to the first wake-up word is the first moment, and the moment when multiple sample training voices are obtained is the second moment.
[0015] The difference between the first moment and the second moment is less than or equal to the first threshold.
[0016] In this implementation, since the difference between the first moment (the moment when the wake-up word is adjusted to the first wake-up word) and the second moment (the moment when multiple sample training voices are obtained) is strictly limited to a range less than or equal to the first threshold, this means that within a very short time after the wake-up word is adjusted, the system can obtain sample training voices containing the new wake-up word. This fast response mechanism ensures that the first model can quickly learn the new wake-up word features, thereby effectively shortening the model's adaptation cycle to the new wake-up word.
[0017] Secondly, this implementation helps maintain the high performance of the model. By timely acquiring and training sample speech containing new wake-up words, the first model can continue to maintain a high recognition rate for the latest wake-up words and avoid performance degradation caused by wake-up word updates. As user needs continue to change, wake-up words may need to be adjusted frequently. By setting the first threshold, the system can flexibly respond to such changes and ensure that model training can be performed quickly after each wake-up word adjustment, thereby maintaining the continuous availability and high performance of the system.
[0018] In a possible implementation, the method further includes: acquiring a first wake-up word from a first storage unit.
[0019] A first wake-up word is provided to the first model and the second model, wherein input data of the first model and the second model also include the first wake-up word.
[0020] In this implementation, the first wake-up word is stored in a fixed area and used by the first model and the second model, which can improve processing efficiency and accuracy. Among them, clear wake-up word information helps the model to more accurately capture relevant voice features, further improving the accuracy and response speed of wake-up word recognition.
[0021] In a possible implementation, the method also includes: obtaining a first wake-up voice within a first time period, the first wake-up voice is voice data collected by the electronic device for the user to wake up the voice assistant, and the first wake-up voice includes a first wake-up word; the first time period is a time period between a first moment and a third moment, the third moment is an end moment of a first collection cycle, and the first moment is within the first collection cycle; or, the first time period is a time period between a fourth moment and a third moment, the fourth moment is a start moment of the first collection cycle, and the first moment is not within the first collection cycle.
[0022] A second sample training voice is generated based on the first wake-up voice, the second sample training data includes at least a third training voice, the third training voice is generated based on the first wake-up voice, and the third training voice includes the first wake-up word.
[0023] The operation of training the first model is repeated according to the second sample training speech.
[0024] In this implementation, the second sample training voice is generated using real user voice data (including the first wake-up voice), and the first model is retrained accordingly to enhance the practicality and adaptability of the model. The model can directly access the user's actual voice pattern, which helps it to more deeply understand and learn the subtle features and changes in the user's voice, thereby improving the recognition accuracy of the first wake-up word. Secondly, using these real data to continuously iterate the training model can gradually adapt to and optimize the response to the user's voice, and effectively utilize the data within the collection cycle to enhance the generalization and robustness of the model.
[0025] In a possible implementation, adjusting the model parameters of the first model according to the first prediction result, the second prediction result, and the true label of the sample training speech includes:
[0026] A first error is determined based on the first prediction result and the second prediction result, where the first error is used to indicate the difference between the prediction result output by the first model and the prediction result output by the second model.
[0027] A second error is determined based on the first prediction result and the true label, where the second error is used to indicate the difference between the prediction result output by the first model and the true label.
[0028] The model parameters of the first model are adjusted according to the first error and the second error.
[0029] In this implementation, by calculating the first error, the difference between the prediction results of the first model and the second model can be intuitively quantified, which helps to discover and correct the deviation of the first model in prediction and promote the improvement of its prediction ability. Secondly, the calculation of the second error can accurately measure the gap between the prediction results of the first model and the true label, providing a direct and accurate basis for adjusting the model parameters. Combining the two, the performance of the first model can be more comprehensively evaluated, and targeted optimization can be performed accordingly.
[0030] In a possible implementation, multiple sample training speech is obtained, including:
[0031] According to the first wake-up word, an initial training speech including the first wake-up word is generated.
[0032] A plurality of preset audio data are respectively added to the initial training speech to obtain a plurality of sample training speech.
[0033] In this implementation, a positive sample training voice is generated based on the first wake-up word, laying a solid foundation for the training of the first model. And by incorporating multiple preset audio data (such as background noise, environmental interference, etc.) into the initial training voice, a complex voice environment that is closer to the actual application scenario is simulated. This not only greatly enriches the diversity of sample training voices, but also exercises the first model's anti-interference ability in complex environments, so as to improve its accuracy and reliability in practical applications.
[0034] In a possible implementation, the plurality of sample training voices also include a second training voice, and the second training voice does not include the first wake-up word.
[0035] In this implementation, voice data that does not contain the first wake-up word is obtained as negative sample training speech to train the first model. The addition of negative samples makes the training set more comprehensive and balanced, which helps the model learn more sophisticated voice feature differentiation capabilities, thereby more accurately identifying the first wake-up word.
[0036] In a second aspect, an embodiment of the present application provides a voice wake-up method. The method includes:
[0037] The collected first voice data is input into the first model so that the first model outputs a first target prediction result for the first voice data. The first target prediction result is used to indicate whether the first voice data contains a first wake-up word. The first wake-up word is user-defined, and the first model is trained according to the above-mentioned model training method.
[0038] When the first target prediction result indicates that the first voice data contains the first wake-up word, the second voice data is input into the second model, so that the second model outputs a second target prediction result for the second voice data, and the second target prediction result is used to indicate whether the second voice data contains the first wake-up word, wherein the second voice data includes the first voice data, the duration of the second voice data is greater than the duration of the first voice data, and the model parameters of the second model are greater than the model parameters of the first model.
[0039] When the second target prediction result indicates that the second voice data contains the first wake-up word, the voice assistant of the electronic device is woken up.
[0040] In this embodiment, the first model is used to perform a preliminary and rapid screening of voice data, which can effectively exclude irrelevant voices that do not contain user-defined wake-up words, which not only reduces the amount of data to be processed later, but also improves the response speed of the entire wake-up process. Secondly, for voice data that is initially judged to contain wake-up words, it is further verified using the second model with higher recognition accuracy and stronger generalization ability. This step greatly improves the accuracy and reliability of wake-up word recognition, effectively reduces the probability of false wake-up, and thus improves the accuracy of voice wake-up.
[0041] In a third aspect, an embodiment of the present application provides a voice wake-up device, which may be an electronic device, or a chip or chip system in an electronic device. The voice wake-up device may include a display unit and a processing unit.
[0042] When the voice wake-up device is an electronic device, the display unit may be a display screen. The display unit is used to perform the display step so that the electronic device implements a voice wake-up method described in the first aspect or any possible implementation of the first aspect.
[0043] When the voice wake-up device is an electronic device, the processing unit may be a processor. The voice wake-up device may also include a storage unit, which may be a memory. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to enable the electronic device to implement a voice wake-up method described in the first aspect or any possible implementation of the first aspect.
[0044] When the voice wake-up device is a chip or chip system in an electronic device, the processing unit may be a processor. The processing unit executes the instructions stored in the storage unit so that the electronic device implements a voice wake-up method described in the first aspect or any possible implementation of the first aspect. The storage unit may be a storage unit in the chip (e.g., a register, a cache, etc.), or a storage unit in the electronic device located outside the chip (e.g., a read-only memory, a random access memory, etc.).
[0045] In a fourth aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, the memory being used to store code instructions, and the processor being used to run the code instructions to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0046] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program or instructions are stored. When the computer program or instructions are run on a computer, the computer executes the method described in the first aspect or any possible implementation of the first aspect.
[0047] In a sixth aspect, an embodiment of the present application provides a computer program product comprising a computer program, which, when the computer program runs on a computer, enables the computer to execute the method described in the first aspect or any possible implementation of the first aspect.
[0048] In a seventh aspect, the present application provides a chip or a chip system, the chip or chip system comprising at least one processor and a communication interface, the communication interface and the at least one processor are interconnected by a line, and the at least one processor is used to run a computer program or instruction to execute the method described in the first aspect or any possible implementation of the first aspect. The communication interface in the chip can be an input / output interface, a pin or a circuit, etc.
[0049] In a possible implementation, the chip or chip system described above in the present application further includes at least one memory, in which instructions are stored. The memory may be a storage unit inside the chip, such as a register, a cache, etc., or a storage unit of the chip (such as a read-only memory, a random access memory, etc.).
[0050] It should be understood that the third to seventh aspects of the present application correspond to the technical solutions of the first and second aspects of the present application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation methods are similar and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 Schematic diagram of application scenarios of the voice wake-up method provided in the embodiment of the present application Figure 1 ;
[0052] Figure 2 Schematic diagram of application scenarios of the voice wake-up method provided in the embodiment of the present application Figure 2 ;
[0053] Figure 3 Schematic diagram of application scenarios of the voice wake-up method provided in the embodiment of the present application Figure 3 ;
[0054] Figure 4 A schematic diagram of the hardware structure of a terminal device provided in an embodiment of the present application;
[0055] Figure 5 A schematic diagram of the software structure of a terminal device provided in an embodiment of the present application;
[0056] Figure 6 A flowchart of a model training method provided in an embodiment of the present application;
[0057] Figure 7 A schematic diagram of a first period provided in an embodiment of the present application;
[0058] Figure 8 A flowchart of a voice wake-up method provided in an embodiment of the present application;
[0059] Fig. 9 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0060] In order to clearly describe the technical solutions of the embodiments of the present application, some terms and technologies involved in the embodiments of the present application are briefly introduced below:
[0061] In the embodiments of the present application, words such as "first" and "second" are used to distinguish the same or similar items with substantially the same functions and effects. For example, the first chip and the second chip are only used to distinguish different chips, and their order is not limited. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.
[0062] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0063] In the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, a--c, bc, or abc, where a, b, and c can be single or multiple.
[0064] The electronic device of the embodiment of the present application may include a handheld device, a vehicle-mounted device, etc. with a data processing function. For example, some electronic devices are: mobile phones, tablet computers, PDAs, laptop computers, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to wireless modems, vehicle-mounted devices, wearable devices, terminal devices in 5G networks or future evolved public land mobile communication networks (PLANs), and wireless terminals in smart cities. The embodiments of the present application do not limit this.
[0065] As an example but not limitation, in the embodiments of the present application, the electronic device may also be a wearable device. Wearable devices may also be referred to as wearable smart devices, which are a general term for wearable devices that are intelligently designed and developed using wearable technology for daily wear, such as glasses, gloves, watches, clothing, and shoes. A wearable device is a portable device that is worn directly on the body or integrated into the user's clothes or accessories. Wearable devices are not only hardware devices, but also powerful functions achieved through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include full-featured, large-sized, and fully or partially independent of smartphones, such as smart watches or smart glasses, as well as devices that only focus on a certain type of application function and need to be used in conjunction with other devices such as smartphones, such as various types of smart bracelets and smart jewelry for vital sign monitoring.
[0066] In addition, in the embodiments of the present application, the electronic device may also be a terminal device in the Internet of Things (IoT) system. IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, thereby realizing an intelligent network that interconnects people and machines and things.
[0067] The electronic device in the embodiments of the present application may also be referred to as: terminal equipment, user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device, etc.
[0068] In an embodiment of the present application, the electronic device or each network device includes a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and a memory (also called main memory). The operating system can be any one or more computer operating systems that implement business processing through processes, such as a Linux operating system, a Unix operating system, an Android operating system, an iOS operating system, or a Windows operating system. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software.
[0069] In order to better understand the technical solution of the present application, the relevant technologies involved in the present application are further introduced in detail below.
[0070] With the popularity of smartphones, users increasingly want to control them through voice commands without manual operation. Voice wake-up technology was born to meet this demand, allowing users to activate devices through preset wake-up words and then perform voice interaction.
[0071] Combine the following Figures 1 to 3 This section introduces the application scenarios of the voice wake-up method. Figure 1 Schematic diagram of application scenarios of the voice wake-up method provided in the embodiment of the present application Figure 1 , Figure 2 Schematic diagram of application scenarios of the voice wake-up method provided in the embodiment of the present application Figure 2 , Figure 3 Schematic diagram of application scenarios of the voice wake-up method provided in the embodiment of the present application Figure 3 .
[0072] like Figure 1 As shown in (a) in the figure, it can be understood as the setting interface of the electronic device. In the setting interface, multiple function controls can be displayed, namely, WLAN, Bluetooth, mobile network, satellite communication, smart interconnection, more connections, smart voice application, desktop and personalization, display and brightness, and sound and vibration. The user can click on any of the multiple function controls to display the content corresponding to the function control. Figure 1 In (a), the user can click on the smart voice application control 101. In response to the user's click operation on the smart voice application control 101, the display can be jumped to Figure 1 The intelligent voice application interface shown in (b) above.
[0073] exist Figure 1 In the diagram (b) in FIG. 1 , text information can be displayed, i.e., “Hello xx” as shown in the figure. Multiple function controls can also be displayed, i.e., voice suggestions, voice assistant, and smart vision as shown in the figure. The user can click on any of the function controls. Figure 1 In (b), the user can click on the voice assistant control 102. In response to the user's click operation on the voice assistant control 102, the display can be jumped to the following Figure 1 The interface shown in (c). Figure 1 (c) in the figure shows the relevant setting page of the voice assistant function. For example, you can click the voice wake-up control to choose to turn on or off the voice wake-up function. You can click the button wake-up control to choose to turn on or off the button wake-up function. And you can also create a desktop shortcut for the voice assistant by clicking the add button.
[0074] Reference Figure 1In (c), the current voice wake-up function is in the off state. The user can click on the voice wake-up function to turn it on. In response to the user's click operation on the voice wake-up control, the interface shown in Figure 1 as in (d) can be displayed. In the schematic diagram of (d) in Figure 1 , a switch control for the voice wake-up function can be displayed, and explanatory text for the function can be shown on the switch control, that is, saying "Hello YOYO" to the phone in the figure to wake up the YOYO assistant. The current switch control for the voice wake-up function is in the off state, and the switch control can be clicked to turn on the voice wake-up function. As shown in Figure 2 in (d), when the user clicks on the switch control, the interface shown in Figure 2 as in (e) can be displayed.
[0075] As shown in Figure 2 in (e), the switch control 201 for the voice wake-up function is in the on state. Also, in the interface of the voice wake-up function, multiple function setting controls are also schematically shown, that is, the wake-up word, re-enter the wake-up word, one-step access, and wake-up sensitivity shown in the figure. The default wake-up word can be selected, or a custom wake-up word can be defined to wake up the voice assistant. Referring to Figure 2 in (e), in the default state, the wake-up word setting control 202 is set to the default wake-up word, that is, "Hello YOYO" shown in the figure. The custom wake-up word can be set by clicking on the custom wake-up word control. Referring to Figure 2 in (e), the user clicks on the custom wake-up word control. In response to the click operation on the custom wake-up word control, the interface shown in Figure 2 as in (f) can be displayed.
[0076] In Figure 2 in the schematic diagram of (f), a window 203 is shown in the voice wake-up interface, and the wake-up word can be set in the window 203. The wake-up word can be entered in the input box 204 in the window. Explanatory text for the function is shown below the input box 204, as shown in the figure "Detecting this wake-up word may affect the wake-up sensitivity. Please use 4 - 6 Chinese characters as the wake-up word, and try to avoid common expressions and reduplicated words. The custom wake-up word may affect the wake-up sensitivity". Also, the custom wake-up word can be saved by clicking the OK button. Or, the editing of the custom wake-up word can be abandoned by clicking the Cancel button. Referring to Figure 2 in (f), the user enters the wake-up word "A Zhuang A Zhuang" in the input box 204 and saves the wake-up word by clicking the OK button. After that, the interface shown in Figure 2 as in (g) can be displayed.
[0077] like Figure 2 As shown in (g), it can be understood as the wake-up word recording interface of the electronic device. Figure 2 (g) in the figure shows a voice input control 205, and text information for explaining the operation of the voice input function is displayed below the voice input control 205, that is, "Press and hold to record the wake-up word. Please say the wake-up word you want to record in a quiet environment" as shown in the figure. In addition, a prompt message is also displayed on the wake-up word recording interface, that is, "Please repeat the wake-up word you want to record 3 times according to the prompt" as shown in the figure. After the wake-up word is successfully recorded, the following message can be displayed: Figure 2 Refer to the interface shown in (h) in the figure. Figure 2 In (h), a text message “wake-up word entry successful” may be displayed in the interface. Afterwards, the recording may be saved by clicking a confirmation control 206 , or the recording may be abandoned by clicking a cancel control 207 .
[0078] The above describes the scenario of setting a custom wake-up word. Figure 3 This article introduces the actual application scenarios of voice wake-up based on custom wake-up words.
[0079] like Figure 3 As shown, it is currently assumed that there are user A and mobile phone A. After user A sets a custom wake-up word for mobile phone A, the custom wake-up word can be used to wake up the voice assistant in mobile phone A. Assume that the current custom wake-up word is set to "A Zhuang A Zhuang". Figure 3 , after user A says "It's a nice day today, A Zhuang A Zhuang", the microphone of mobile phone A can collect the voice data "It's a nice day today, A Zhuang A Zhuang", and can recognize the voice data. When it is detected that the voice data contains a custom wake-up word, the voice assistant can be woken up. For example, after the user says "A Zhuang A Zhuang", the voice assistant can be woken up.
[0080] Based on the application scenarios introduced above, it can be determined that users can wake up the voice assistant in the electronic device through a customized wake-up word to meet the personalized needs of the user. In related implementations, in order to implement the voice wake-up function in the electronic device, a voice recognition model is usually preset. The model can identify the voice data to determine whether it contains the customized wake-up word. Since the voice recognition model needs to process the voice data in real time for recognition, it needs to be running all the time. Considering the battery life of the electronic device, a lightweight voice recognition model can usually be used to implement all-weather voice recognition processing to reduce the power consumption of the model. This means that the scale design (model parameters, computing power, etc.) of this voice recognition model is relatively small, which may lead to a decrease in the recognition accuracy of the model.
[0081] To solve this problem, in one solution, the user's voice data can be collected online as training data for further training of the model. When training the model, it can be done on the cloud or server side based on the collected user data. Afterwards, the trained model can be deployed on the client. However, users can change custom wake-up words at will. If the user updates the wake-up word and the version of the speech recognition model is not updated in time, the model may not be able to accurately recognize the changed custom wake-up word in the voice data. In addition, as the number of users continues to increase, if the custom wake-up word set by each user is trained on the server side, the pressure on the server will increase dramatically, which may lead to problems such as longer response time.
[0082] In response to the above-mentioned problems, this application proposes the following technical concept: two models are used to recognize voice data to accurately determine whether the voice data contains user-defined wake-up words. In the model training process, after the user adjusts the wake-up word, targeted training can be performed locally on the electronic device for the first model to improve the recognition accuracy of the first model for the new wake-up word. In this way, the load pressure on the server can be alleviated at the same time. In addition, in the process of training the first model, the second model can be used as a guide to assist in adjusting the parameters of the first model, thereby further enhancing the recognition accuracy of the first model.
[0083] The model training method and voice wake-up method of the embodiment of the present application can be executed by an electronic device with a microphone, or by a chip, chip system or processor that supports the electronic device to implement the model training method and voice wake-up method, or by a logic module or software that can implement all or part of the functions of the electronic device. This application does not impose specific restrictions on this. The model training method and voice wake-up method of the embodiment of the present application are described in detail below, taking the electronic device as the execution subject.
[0084] The electronic device may be, for example, a terminal device. Figure 4 and Figure 5 Provide a brief introduction to the terminal equipment.
[0085] For example, Figure 4 A schematic diagram of the hardware structure of a terminal device provided in an embodiment of the present application.
[0086] Figure 4A schematic diagram of the structure of a terminal device provided in an embodiment of the present application. The terminal device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0087] It is to be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the terminal device. In other embodiments of the present application, the terminal device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0088] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (application processor, AP), a modem processor, a graphics processor (graphics processing unit, GPU), an image signal processor (image signal processor, ISP), a controller, a video codec, a digital signal processor (digital signal processor, DSP), a baseband processor, and / or a neural network processor (neural-network processing unit, NPU), etc. Among them, different processing units can be independent devices or integrated in one or more processors. In one implementation, for example, the model training method and voice wake-up method provided in the present application can be executed by a processor.
[0089] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel.
[0090] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be arranged in the processor 110, or some functional modules of the audio module 170 can be arranged in the processor 110.
[0091] Microphone 170C, also called "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The terminal device may be provided with at least one microphone 170C. In other embodiments, the terminal device may be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the terminal device may also be provided with three, four or more microphones 170C to realize the collection of sound signals, noise reduction, identification of sound sources, realization of directional recording function, etc.
[0092] The software system of the terminal device may adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture, or a cloud architecture, etc. The embodiment of the present application takes the Android system of the layered architecture as an example to exemplify the software structure of the terminal device.
[0093] For example, Figure 5 A schematic diagram of the software structure of a terminal device provided in an embodiment of the present application.
[0094] like Figure 5 As shown, the layered architecture divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the system may include an application layer, an application framework layer, an Android runtime (Android runtime) and a system library, a hardware abstraction layer (HAL) and a kernel layer. It should be noted that the embodiment of the present application is illustrated by taking the Android system as an example. In other operating systems (such as Hongmeng system, IOS system, etc.), as long as the functions implemented by each functional module are similar to those of the embodiment of the present application, the solution of the present application can also be implemented.
[0095] Among them, the application layer can include a series of application packages.
[0096] like Figure 5 As shown, the application package may include camera, calendar, phone, map, phone, music, settings, mailbox, video, social and other applications. Of course, the application layer may also include other application packages, such as payment applications, shopping applications, banking applications, social applications and other third-party applications, which are not limited in this application. In this embodiment, a voice detection module and a voice assistant may be deployed at the application layer. The voice detection module may, for example, monitor voice signals in real time. The voice assistant may, for example, automatically perform related operations according to the user's instructions.
[0097] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0098] like Figure 5 As shown, the application framework layer may include a window manager, a content provider, a resource manager, a view system, a notification manager, and the like.
[0099] Among them, Android runtime includes core libraries and virtual machines. Android runtime is responsible for the scheduling and management of the Android system.
[0100] The core library consists of two parts: one part is the function that needs to be called by the Java language, and the other part is the Android core library.
[0101] The application layer and the application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object life cycle management, stack management, thread management, security and exception management, and garbage collection.
[0102] The system library may include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0103] Among them, the HAL layer is an encapsulation of the Linux kernel driver, providing an interface to the upper layer and shielding the implementation details of the low-level hardware.
[0104] The HAL layer may include Wi Fi HAL, audio HAL, camera service (Camera HALServer) unit of the HAL layer, and software code library, etc. In this embodiment, a first model and a second model may be deployed at the HAL layer. The first model may, for example, perform real-time processing on the voice data collected by the microphone or voice detection module to identify whether the voice segment contains a user-defined wake-up word. The second model may process the voice segment output by the first model to accurately identify whether the voice segment contains a user-defined wake-up word. When the second model determines that the voice segment contains a custom wake-up word, the display driver may be turned on to display the icon of the voice assistant on the electronic device, and the audio driver may be turned on to reply through the speaker, such as "I am here".
[0105] The kernel layer is the layer between hardware and software. The kernel layer contains at least display driver, camera driver, audio driver, and sensor driver.
[0106] The following, in conjunction with the accompanying drawings, describes in detail the technical solutions of the embodiments of the present application and how the technical solutions of the embodiments of the present application solve the above-mentioned technical problems with specific embodiments. The following specific embodiments can be implemented independently or in combination with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0107] The technical solution of this application involves model training and model application. These two parts are introduced separately below.
[0108] The following is a description of the process of the model training method in conjunction with a specific embodiment. Figure 6 Give a detailed introduction, Figure 6 A schematic diagram of a flow chart of a model training method provided in an embodiment of the present application. The flow chart of the model training method may include the following steps:
[0109] S601. When the wake-up word of the voice assistant is adjusted to the first wake-up word, obtain multiple sample training voices.
[0110] Based on the above introduction, it can be determined that when voice wake-up is implemented, the first model can be used to quickly identify voice data to determine whether it contains the user-defined wake-up word. Since the first model needs to process the continuously collected voice data in real time, in order to reduce the consumption of electronic device resources (including computing power and battery life), the first model is usually relatively small in scale design, which inevitably sacrifices its recognition accuracy to a certain extent.
[0111] In addition, considering that users can set any word as the wake-up word to activate the voice assistant, this further increases the difficulty of the first model to recognize the custom wake-up word. In particular, when the wake-up word selected by the user is very different from the word used when the model was trained, the first model may have difficulty in accurately recognizing it.
[0112] In one implementation, when the wake-up word of the voice assistant is adjusted, the first model can be trained based on the adjusted wake-up word to improve the accuracy of the first model in recognizing the specific wake-up word. The first model can recognize the voice data to output a first prediction result. The first prediction result is used to indicate whether the voice data contains a custom wake-up word.
[0113] In this embodiment, the moment when the wake-up word of the voice assistant is adjusted to the first wake-up word can be defined as the first moment, and the moment when multiple sample training voices are obtained can be defined as the second moment. In order to ensure the timeliness and effectiveness of the training, a condition is set: when the difference between the first moment and the second moment is less than or equal to the first threshold, multiple sample training voices will be obtained and the first model will be trained. The first wake-up word here is a user-defined wake-up word. The first threshold can be a preset value set by an artificial person to control the timing of training.
[0114] Taking the first threshold being preset to 10 milliseconds (ms) as an example, if the wake-up word of the voice assistant is adjusted to the first wake-up word, then within the following 10 milliseconds (or a shorter time, as long as the condition of less than or equal to 10 milliseconds is met), sample voice data can be obtained and the first model can be trained immediately.
[0115] Since the custom wake-up words set by the user are random, it is difficult to obtain a large amount of audio data containing these custom wake-up words by directly collecting them with a microphone. To solve this problem, in one implementation, multiple sample training voices can be generated based on the first wake-up word (i.e., the user-defined wake-up word). These sample training voices are intended to expand the training set containing custom wake-up words in the first model. They simulate the voice samples of users saying custom wake-up words in different scenarios, including but not limited to the user's emotions, pauses and other characteristics.
[0116] For example, in one implementation, text-to-speech technology can be used to generate multiple initial training voices based on the first wake-up word. These initial training voices can describe the first wake-up word in different voice forms, intonations, and voice to increase the diversity and generalization ability of the training.
[0117] Furthermore, in order to simulate the speech conditions in the actual environment and improve the recognition ability of the model, the initial training speech generated above can also be subjected to data enhancement processing. In one implementation, multiple preset audio data can be added to the initial training speech respectively to obtain a sample training speech containing environmental noise. These preset audio data may include but are not limited to white noise, car noise, rain noise, human voice noise, wind noise, music noise, traffic noise, etc. (these noises can be used in any combination). In addition, various types of noise can also include different intensities to more realistically simulate the speech conditions in the actual environment. This embodiment of the present application is not limited to this.
[0118] For example, 200 initial training voices can be generated according to the first wake-up word. Afterwards, noise data can be added to the 200 initial training voices to obtain 200 sample training voices containing environmental noise. Finally, the 200 initial training voices and the 200 sample training voices containing environmental noise can be used as training data to train the first model.
[0119] It is understandable that these data-enhanced sample training voices all contain the first wake-up word, so they can be used as positive samples for training the first model. This processing method not only solves the problem of scarcity of custom wake-up word data, but also helps to improve the recognition accuracy and robustness of the model in actual environments.
[0120] In one implementation, the plurality of sample training voices include not only positive sample voices containing the first wake-up word, but also second training voices as negative samples. The second training voice does not contain the first wake-up word, and it can be any voice data that does not contain the first wake-up word. Exemplarily, 1,000 voice data that do not contain the first wake-up word can be collected and stored in advance as the second training voice, and these data can be pre-saved on an electronic device. When the first model needs to be trained, these second training voices can be used as negative samples for model training together with positive samples containing the first wake-up word.
[0121] It is understandable that every time the user adjusts the wake-up word, the training set can be expanded based on the above method to retrain the first model, which can ensure that the model can accurately recognize the new wake-up word.
[0122] S602: For any one of the multiple sample training speech pieces, input the sample training speech piece into the first model, so that the first model outputs a first prediction result.
[0123] It can be understood that the processing method for each of the multiple sample training speech in step S602 is similar, so any sample training speech is taken as an example for introduction below.
[0124] You can refer to Figure 6 It is understood that after the sample training speech is input into the first model, the first model can recognize the sample training speech and output a first prediction result. The first prediction result is used to indicate whether the sample training speech contains the first wake-up word.
[0125] In this embodiment, the first model may include two main parts, for example, an acoustic model and a decoder. Since the first model needs to be in operation all the time so as to process the voice data collected by the microphone in real time, a lightweight acoustic model is used in the first model to ensure the real-time and efficiency of the processing. The following is an introduction to the process of the first model processing the sample training voice.
[0126] When a piece of sample training data is input into the first model, the sample training data is first input into the acoustic model. The acoustic model can process the sample training data and output a probability distribution. Each probability value in this probability distribution represents the probability that the acoustic model predicts that a certain phoneme may be contained in the sample training data. Among them, the phoneme is the basic unit that constitutes the sound of the language. It is the smallest unit of speech divided according to the natural properties of speech. In the pronunciation process, a syllable often contains one or more phonemes, which together constitute various sounds.
[0127] Next, this probability distribution can be input into the decoder. The decoder will decode the text content in the sample training data based on the prediction results of the acoustic model. Then, the decoder will perform character matching between the decoded text content and the first wake-up word.
[0128] In one implementation, the electronic device may obtain the first wake-up word from the first storage unit and provide it to the first model to perform the above-mentioned character matching process.
[0129] Finally, if the decoded text content matches the first wake-up word, the first model will output a first prediction result, in which the score for the "contains the wake-up word" category will be higher, and the score for the "does not contain the wake-up word" category will be lower. This first prediction result is used to indicate that the sample training speech contains the first wake-up word. On the contrary, if the decoded text content cannot match the first wake-up word, the first model will output a first prediction result, in which the score for the "contains the wake-up word" category will be lower, and the score for the "does not contain the wake-up word" category will be higher. This first prediction result is used to indicate that the sample training speech does not contain the first wake-up word.
[0130] S603: For any one of the multiple sample training speech, input the sample training speech into the second model, so that the second model outputs a second prediction result, wherein the model parameters of the second model are greater than the model parameters of the first model.
[0131] Based on the above introduction, it can be determined that the second model can be used to recognize the speech data, thereby outputting a second prediction result corresponding to the speech data.
[0132] In this embodiment, the number of model parameters of the second model is greater than that of the first model. The greater the number of parameters, the stronger the model's expressiveness and learning ability, which enables the second model to capture more data features and thus handle more complex tasks. In addition, the second model also has a more complex structure, such as a deeper neural network, a more complex connection method, etc., which enables the second model to learn a wider and deeper pattern, and thus generally achieve better performance when processing specific tasks (especially in the fields of natural language processing, computer vision, etc.).
[0133] However, this also requires corresponding hardware and data support, and during operation, the second model consumes higher power. In contrast, the first model has fewer parameters, resulting in limited information learned and relatively low performance in processing tasks.
[0134] It can be understood that the second model can output a relatively accurate prediction result when recognizing the speech data. However, since the first model needs to be run all the time to realize real-time processing of the speech data, the second model with better performance cannot be directly used as the first model.
[0135] In one implementation, the second model can be used to guide the parameter adjustment of the first model. Specifically, by making the prediction result of the first model approach the prediction result of the second model, the first model can learn the knowledge of the second model.
[0136] You can refer to Figure 6 It is understood that the sample training speech can be input into the second model to obtain a second prediction result corresponding to the sample training speech output by the second model.
[0137] In this embodiment, the architecture of the second model is similar to the first model, so in the second prediction result, if a higher score corresponds to the category of "including the wake-up word" and a lower score corresponds to the category of "not including the wake-up word", then the second prediction result can indicate that the sample training speech contains the first wake-up word. On the contrary, if a lower score corresponds to the category of "including the wake-up word" and a higher score corresponds to the category of "not including the wake-up word", then the second prediction result can indicate that the sample training speech does not contain the first wake-up word.
[0138] S604: Adjust the model parameters of the first model according to the first prediction result, the second prediction result, and the true label of the sample training speech.
[0139] In this embodiment, the true label is used to indicate whether the training sample speech contains the first wake-up word. For example, for a training sample speech, its true label is clearly included in the first wake-up word or does not contain the first wake-up word. When the training sample speech is obtained, its corresponding true label can be directly determined.
[0140] In one implementation, the parameters of the first model can be adjusted according to the difference between the first prediction result and the second prediction result, so that the first model can learn the knowledge of the second model, so that the prediction result output by the first model can be closer to the prediction result of the second model.
[0141] In one implementation, for example, the difference between the first prediction result and the second prediction result may be calculated using KL (Kullback–Leibler) divergence to determine the first error. KL divergence may quantify the degree of difference between two probability distributions.
[0142] As an example, in one implementation, the calculation formula of the first error may refer to the following formula 1:
[0143] Formula 1
[0144] In formula 1, represents the KL divergence function, Represents the logarithmic function. is the first type of probability distribution calculated by the second model for the sample training speech, The first type of probability distribution is calculated by the second model for the sample training speech. The second model is calculated for the sample training speech and belongs to the In this embodiment, two categories may be included, one is "including the wake-up word" and the other is "not including the wake-up word". The first model is calculated for the sample training speech and belongs to the The probability of the class. Among them, and The calculation formula of can refer to the following formula 2:
[0145] Formula 2
[0146] In formula 2, Represents the exponential function. is the original output calculated by the first model for the sample training speech, usually a vector, in which each value represents the model's prediction score for a certain category (that is, the unnormalized prediction value of each category). The raw output calculated by the second model for the sample training speech. As parameters, by adjusting the parameters , which can make the first-category probability distribution output by the second model smoother, making it easier for the first model to learn the relative relationship between categories rather than just direct hard labels.
[0147] In one implementation, the parameters of the first model may be adjusted according to the difference between the first prediction result and the true label to ensure that the first model can learn the correct prediction of the true label.
[0148] In one implementation, for example, the second error can be determined by calculating the difference between the first prediction result and the true label using cross entropy. In this case, there is no need to smooth the original output of the first model (i.e., the unnormalized prediction value), so the parameter can be set is 1.
[0149] As an example, in one implementation, the calculation formula of the second error may refer to the following formula 3:
[0150] Formula 3
[0151] In formula 3, Represents the logarithmic function. The true label for the sample training speech, is the second type of probability distribution calculated by the first model for the sample training speech, and its calculation formula can be referred to the following formula 4, for example:
[0152] Formula 4
[0153] In formula 4, Represents the exponential function. is the original output (i.e., unnormalized prediction value) calculated by the first model for the sample training speech. Compared with the second type of probability distribution, the first type of probability distribution contains a parameter , this parameter can make the final first-class probability distribution smoother.
[0154] After obtaining the first error and the second error, a loss function may be determined according to the sum of the first error and the second error to adjust the model parameters of the first model.
[0155] As an example, in one implementation, the calculation formula of the loss function may refer to the following formula 5:
[0156] Formula 5
[0157] In Formula 5, represents the first error, Represents the second error. is a parameter that controls the weights of the first error and the second error. It can be understood that if the first model needs to have a higher accuracy, the weight of the second error can be increased; if the first model is expected to better mimic the prediction results of the second model, the weight of the first error can be increased.
[0158] After calculating the loss function value, you can refer to Figure 6 It can be understood that the parameters of the first model can be updated through the back propagation algorithm to make it as close as possible to the output of the second model and output a more accurate prediction result.
[0159] In this embodiment, the system can obtain sample training speech containing the new wake-up word in a very short time after the wake-up word is adjusted. This ensures that the first model can quickly learn the new wake-up word features, thereby effectively shortening the model's adaptation cycle to the new wake-up word. By training with sample training speech containing user-defined wake-up words, the first model can better capture and understand the voice features of these specific wake-up words to enhance the generalization ability of the model. Regardless of the environment the user is in or the tone of voice used, the model can more accurately identify the wake-up word, thereby improving the response speed and user experience of the voice assistant.
[0160] Secondly, the first model is updated locally on the electronic device, which can effectively reduce the pressure on the server. With the popularity and increase in the number of smart devices, the amount of data that the server needs to process is increasing. Delegating the model update task to the local device can not only relieve the load pressure on the server, but also reduce the delay of data transmission, thereby improving overall efficiency.
[0161] In addition, given the relatively weak learning ability of the first model, the second model is introduced as a guide. The second model is able to show higher accuracy in predicting the wake-up word due to its larger model parameters and stronger learning ability. By letting the second model guide the parameter adjustment of the first model, the first model can learn the knowledge and experience of the second model.
[0162] In the embodiment proposed in this application, the recognition ability of the first model can be further improved through another training scenario. Specifically, the user wake-up voice data collected by the electronic device in actual use can truly reflect the user's pronunciation habits, and using this data as training samples can effectively optimize the recognition effect of the first model.
[0163] In one implementation, a first wake-up voice within a first time period may be obtained, where the first wake-up voice is real voice data collected by an electronic device for a user to wake up a voice assistant, and includes a first wake-up word, which is used to train a first model. Figure 7 The first period is introduced. Figure 7 This is a schematic diagram of the first time period provided in the embodiment of the present application. The specific definition of the first time period is as follows:
[0164] Case 1 (wake-up word adjustment occurs in the first collection cycle):
[0165] If the user has adjusted the wake-up word in the current first collection period, the first period is from the time when the wake-up word is adjusted (the first moment) to the time when the first collection period ends (the third moment). The first collection period can be, for example, one week. Figure 7 , when the first collection period is 1 week (starting at 00:00 on Monday and ending at 23:59 on Sunday) and the wake-up word adjustment occurs within this period, voice data needs to be collected from the adjustment time (first time) to 23:59 on Sunday (third time).
[0166] Case 2 (wake-up word adjustment does not occur in the first collection cycle):
[0167] If the user does not adjust the wake-up word in the current first collection cycle, the first time period covers the entire first collection cycle, that is, from the start time of the cycle (the fourth time) to the end time (the third time). Figure 7 As shown, when the first collection period is 1 week and the wake-up word adjustment occurs in the last week, voice data for the entire week (Monday 00:00 to Sunday 23:59) needs to be collected.
[0168] After the first wake-up voice is collected, a second sample training voice can be generated based on the first wake-up voice. The second sample training data includes at least a third training voice. The third training voice is generated based on the first wake-up voice, and the third training voice includes the first wake-up word.
[0169] In this step, multiple preset audio data can be added to the first wake-up voice respectively. This process is similar to the above-mentioned process of generating sample training voice, and will not be repeated here.
[0170] Finally, the speech can be trained based on the second sample, and the above operation of training the first model can be repeated.
[0171] In this embodiment, the second sample training voice is generated using real user voice data (including the first wake-up voice), and the first model is retrained accordingly to enhance the practicality and adaptability of the model. The model can directly access the user's voice pattern in actual use, which helps it to more deeply understand and learn the subtle features and changes in the user's voice, thereby improving the recognition accuracy of the first wake-up word. Secondly, the continuous iteration of the training model using these real data can gradually adapt to and optimize the response to the user's voice, and effectively utilize the data within the collection cycle to enhance the generalization ability and robustness of the model.
[0172] Based on the model training, the following is combined Figure 8 This section briefly introduces the voice wake-up method. Figure 8 A flowchart of the voice wake-up method provided in an embodiment of the present application.
[0173] S801. Input the collected first speech data into the first model, so that the first model outputs a first target prediction result for the first speech data.
[0174] Based on the above introduction, it can be understood that the first model can process the voice stream input by the microphone in real time. During the processing, the real-time collected voice stream is divided into multiple voice frames, and the first model then processes each voice frame and outputs the first target prediction result corresponding to the voice frame. In this embodiment, the first voice data can be a voice frame.
[0175] For example, please refer to Figure 8 To understand, the continuous voice stream is divided into multiple 10ms voice frames, and the first voice data can be the voice frame to be processed. The first voice data is input into the first model, and the first model can output the first target prediction result corresponding to the first voice data. The first target prediction result is presented as a probability distribution, and based on this probability distribution, it can be determined whether the first voice data contains the user-defined first wake-up word.
[0176] S802: When the first target prediction result indicates that the first voice data contains the first wake-up word, the second voice data is input into the second model so that the second model outputs a second target prediction result for the second voice data. The second target prediction result is used to indicate whether the second voice data contains the first wake-up word. The model parameters of the second model are greater than the model parameters of the first model.
[0177] In this embodiment, the second voice data includes the first voice data, and the duration of the second voice data is longer than the duration of the first voice data. The duration of the second voice data can be set, for example, it can be set to 2 seconds.
[0178] It can be understood that when the second voice data is determined based on the first voice data, the second voice data may be any one of the following three situations: taking the duration of the second voice data as 2 seconds as an example, first, the second voice data may include all voice frames within 2 seconds before the end point of the first voice data; second, the second voice data may include all voice frames within 1 second before the starting point of the first voice data and voice frames within 1 second immediately after the starting point of the first voice data (that is, 1 second before and after, with a total duration of 2 seconds, and this 2-second time period partially overlaps or is adjacent to the first voice data); third, the second voice data may include all voice frames within 2 seconds after the end point of the first voice data.
[0179] Take the second voice data including all voice frames within 2 seconds before the end point of the first voice data as an example, Figure 8 It is understood that when the first target prediction result indicates that the first voice data contains the first wake-up word, all voice frames within 2 seconds before the end point of the current frame can be automatically traced back, and they can be composed into continuous voice segments and input into the second model. The second model can recognize these voice frames again to obtain the second target prediction result.
[0180] If the first target prediction result indicates that the first voice data does not contain the first wake-up word, the newly collected voice frames can be continuously monitored until the first wake-up word is recognized.
[0181] S803: When the second target prediction result indicates that the second voice data contains the first wake-up word, wake up the voice assistant of the electronic device.
[0182] You can refer to Figure 8 It is understood that when the second target prediction result indicates that the first wake-up word is included in the second voice data, the voice assistant can be woken up.
[0183] For example, in the screen-off state or the screen-off AOD (Always on Display) state, after the first wake-up word is recognized, the screen can be lit and the voice assistant icon can be displayed, indicating that the user has been awakened and can reply "I am here" through the speaker output. In the state of displaying the main interface or other application interface, after the custom wake-up word is recognized, the voice assistant icon can be displayed, etc. In this way, the user can continue to send voice commands. The electronic device can perform corresponding operations according to the recognized voice commands.
[0184] When the second target prediction result indicates that the first wake-up word is not included in the voice data, the voice data newly collected by the microphone can continue to be recognized until the first wake-up word is recognized.
[0185] After the first wake-up word is recognized, voiceprint verification is performed to determine whether the first wake-up word included in the collected voice data is spoken by the owner, that is, to verify whether the user is the owner. The voice data recorded by the user in the above introduction can be used to extract the voiceprint information of the owner.
[0186] When it is confirmed that the user is not the owner of the device, the voice assistant will not be woken up. Instead, the current voice data will continue to be collected, and it will be identified whether the newly collected audio includes the first wake-up word and is spoken by the owner of the device, until it is identified that the owner has spoken the first wake-up word.
[0187] In this embodiment, the first model is used to perform a preliminary and rapid screening of voice data, which can effectively exclude irrelevant voices that do not contain user-defined wake-up words, thereby reducing the amount of data to be processed later. Secondly, for voice data that is initially judged to contain wake-up words, the second model with higher recognition accuracy and stronger generalization ability is further used for verification, which can improve the accuracy and reliability of wake-up word recognition, effectively reduce the probability of false wake-up, and thus improve the accuracy of voice wake-up.
[0188] It should be noted that the module names involved in the embodiments of the present application can be defined as other names as long as the functions of each module can be achieved, and there is no specific restriction on the names of the modules.
[0189] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0190] The model training method and voice wake-up method of the embodiment of the present application have been described above. The device for executing the above method provided by the embodiment of the present application is described below. Those skilled in the art can understand that the method and the device can be combined and referenced with each other, and the relevant device provided by the embodiment of the present application can execute the steps in the above model training method and voice wake-up method.
[0191] The model training method and voice wake-up method provided in the embodiments of the present application can be applied to electronic devices with data functions. The electronic devices include terminal devices, and the specific device form of the terminal devices can refer to the above-mentioned related descriptions, which will not be repeated here.
[0192] In one implementation, an embodiment of the present application provides an electronic device, Fig. 9 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application.
[0193] like Fig. 9 As shown, the electronic device 900 includes: a processor 901 and a memory 902; the memory 902 stores computer-executable instructions; the processor 901 executes the computer-executable instructions stored in the memory 902, so that the electronic device 900 executes the above method.
[0194] When the memory 902 is independently provided, the electronic device further includes a bus 903 for connecting the memory 902 and the processor 901 .
[0195] The embodiment of the present application provides a chip. The chip includes a processor, and the processor is used to call a computer program in a memory to execute the technical solution in the above embodiment. Its implementation principle and technical effect are similar to those of the above related embodiments, and will not be repeated here.
[0196] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program. The above method is implemented when the computer program is executed by the processor. The method described in the above embodiment can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. If implemented in software, the function can be stored as one or more instructions or codes on a computer-readable medium or transmitted on a computer-readable medium. Computer-readable media can include computer storage media and communication media, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium that can be accessed by a computer.
[0197] In one possible implementation, a computer-readable medium may include RAM, ROM, compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that is intended to carry or store the required program code in the form of instructions or data structures and can be accessed by a computer. Moreover, any connection is appropriately referred to as a computer-readable medium. For example, if the software is transmitted from a website, server or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technology (such as infrared, radio and microwave), the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technology such as infrared, radio and microwave are included in the definition of medium. Disks and optical disks as used herein include optical disks, laser disks, optical disks, digital versatile disks (DVD), floppy disks and Blu-ray disks, where disks usually reproduce data magnetically, while optical disks reproduce data optically using lasers. Combinations of the above should also be included in the scope of computer-readable media.
[0198] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed, the computer executes the above method.
[0199] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable device to generate a machine, so that the instructions executed by the processing unit of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0200] The above specific implementation methods further illustrate the purpose, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific implementation methods of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solutions of the present invention should be included in the scope of protection of the present invention.
Claims
1. A model training method, characterized in that: Applied to electronic equipment, the method comprises: When the wake-up word of the voice assistant is adjusted to the first wake-up word, a plurality of sample training voices are obtained, wherein the sample training voices include at least the first training voice, the first training voice includes the first wake-up word, and the first wake-up word is user-defined; For any one of the multiple sample training voices, input the sample training voice into a first model so that the first model outputs a first prediction result, where the first prediction result is used to indicate whether the sample training voice contains the first wake-up word; Inputting the sample training speech into a second model so that the second model outputs a second prediction result, wherein the second prediction result is used to indicate whether the sample training speech contains the first wake-up word, wherein a model parameter of the second model is greater than a model parameter of the first model; Adjust the model parameters of the first model according to the first prediction result, the second prediction result and the true label of the sample training speech, where the true label is used to indicate whether the first wake-up word is included in the sample training speech.
2. The method according to claim 1, characterized in that The moment when the wake-up word of the voice assistant is adjusted to the first wake-up word is a first moment, and the moment when the plurality of sample training voices are obtained is a second moment; The difference between the first moment and the second moment is less than or equal to a first threshold.
3. The method according to claim 2, characterized in that The method further comprises: Obtaining a first wake-up word from a first storage unit; The first wake-up word is provided to the first model and the second model, wherein input data of the first model and the second model also include the first wake-up word.
4. The method according to claim 2, characterized in that: The method further comprises: Acquire a first wake-up voice within a first time period, where the first wake-up voice is voice data collected by the electronic device for a user to wake up the voice assistant, and the first wake-up voice includes the first wake-up word; the first time period is a time period between a first moment and a third moment, the third moment is an end moment of a first collection cycle, and the first moment is within the first collection cycle; or, the first time period is a time period between a fourth moment and a third moment, the fourth moment is a start moment of the first collection cycle, and the first moment is not within the first collection cycle; Generate a second sample training voice based on the first wake-up voice, wherein the second sample training data includes at least a third training voice, the third training voice is generated based on the first wake-up voice, and the third training voice includes the first wake-up word; According to the second sample training speech, the operation of training the first model is repeatedly performed.
5. The method according to any one of claims 1 to 4, characterized in that: The adjusting the model parameters of the first model according to the first prediction result, the second prediction result and the true label of the sample training speech includes: Determine a first error according to the first prediction result and the second prediction result, where the first error is used to indicate a difference between a prediction result output by the first model and a prediction result output by the second model; Determine a second error according to the first prediction result and the true label, where the second error is used to indicate a difference between the prediction result output by the first model and the true label; The model parameters of the first model are adjusted according to the first error and the second error.
6. The method according to any one of claims 1 to 4, characterized in that: The step of obtaining a plurality of sample training voices includes: According to the first wake-up word, generating an initial training speech including the first wake-up word; A plurality of preset audio data are respectively added to the initial training speech to obtain the plurality of sample training speech.
7. The method according to any one of claims 1 to 4, characterized in that: The plurality of sample training voices also include a second training voice, and the second training voice does not include the first wake-up word.
8. A voice wake-up method, characterized in that: include: Inputting the collected first voice data into the first model, so that the first model outputs a first target prediction result for the first voice data, the first target prediction result is used to indicate whether the first voice data contains a first wake-up word, the first wake-up word is user-defined, and the first model is trained according to the method described in any one of claims 1 to 7; In a case where the first target prediction result indicates that the first wake-up word is included in the first voice data, the second voice data is input into the second model, so that the second model outputs a second target prediction result for the second voice data, and the second target prediction result is used to indicate whether the second voice data contains the first wake-up word, wherein the second voice data includes the first voice data, the duration of the second voice data is greater than the duration of the first voice data, and the model parameters of the second model are greater than the model parameters of the first model; When the second target prediction result indicates that the first wake-up word is included in the second voice data, the voice assistant of the electronic device is woken up.
9. An electronic device, characterized in that: The electronic device comprises: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code comprises computer instructions, and the one or more processors call the computer instructions so that the electronic device executes the method as described in any one of claims 1 to 8.
10. A chip system, characterized in that: The chip system is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions so that the electronic device executes the method as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium comprises computer instructions, and when the computer instructions are executed on an electronic device, the electronic device is caused to perform the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Keyword detection method capable of supporting self-defined wake-up words
CN111933124A
Speech recognition model optimization method, training method, equipment and medium
CN114120979A
Speech recognition method and electronic equipment
CN117012189A
Voice data processing method and device, equipment and medium
CN117524228A
Display device and voice wake-up method
CN118366434A