Voice interaction device and control method and control apparatus thereof
By working together with a low-power wake-up module and a high-performance main chip, accurate wake-up of voice interaction devices is achieved, solving the problems of false wake-up and wake-up failure caused by low-power chips, improving user experience and reducing power consumption.
Patent Information
- Application Number
- CN202210148501.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-17
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2042-02-17
AI Technical Summary
In existing voice interaction devices, the limited computing power of low-power chips leads to problems such as false wake-up during standby and inability to wake up normally, affecting the user experience.
The system employs a low-power wake-up module and a high-performance main chip working together. The wake-up module initially filters audio signals, while the main chip further determines whether to wake up the device and detects audio signals in the U-boot process to start the host.
It improves the wake-up efficiency of voice interaction devices, avoids false wake-ups and failures to wake up, enhances user experience, and reduces device power consumption.
Smart Images

Figure CN114373462B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of voice interaction. More particularly, it relates to a voice interaction device and a control method and control apparatus thereof. BACKGROUND
[0002] With the continuous development of voice interaction technology, various types of electronic devices are configured with voice interaction functions to achieve voice control of the electronic devices, thereby better meeting user needs.
[0003] However, since high-performance chips have high power consumption, in current voice interaction devices, a low-power chip is usually configured to wake up the voice interaction device. However, the computing capability of the low-power chip is limited, and when the voice interaction device is woken up, standby false wake-up and failure to wake up normally may occur, which seriously affects user experience. SUMMARY
[0004] The exemplary embodiments of the present application provide a voice interaction device and a control method and control apparatus thereof, which can improve wake-up efficiency while ensuring low power consumption.
[0005] In a first aspect, the embodiments of the present application provide a voice interaction device, comprising: a wake-up module, a main chip, and a host;
[0006] The wake-up module is configured to control the main chip to enter a U-boot process in response to collecting a first audio signal; in the U-boot process, the main chip is configured to: initialize the voice interaction device and obtain a second audio signal; in response to the second audio signal reaching an energy threshold, extract an audio feature of the second audio signal; determine whether the audio feature meets a preset condition, and in response to the audio feature meeting the preset condition, control the host to start according to the second audio signal, the second audio signal containing the first audio signal and an audio signal currently collected by the wake-up module.
[0007] In some possible implementation manners, the main chip comprises: an audio signal detection module and a far-field voice module;
[0008] The main chip is configured to, in response to the audio signal detection module extracting the audio feature of the second audio signal in response to the second audio signal reaching the energy threshold, determine whether the audio feature meets the preset condition, and in response to the audio feature meeting the preset condition, control the host to start according to the second audio signal, specifically comprising:
[0009] The audio signal detection module determines whether the audio feature meets the preset condition; in response to the audio feature meeting the preset condition, the far-field voice module is started; after being started, the far-field voice module controls the host to start according to the second audio signal.
[0010] In some possible implementation ways, the audio signal detection module, in response to the second audio signal reaching the energy threshold, and in extracting the audio feature of the second audio signal, specifically comprises: performing frame processing on the second audio signal to obtain a speech frame corresponding to the second audio signal; calculating a short-time energy and / or a short-time zero-crossing rate of the second audio signal according to the speech frame; in response to the short-time zero-crossing rate being greater than or equal to a preset zero-crossing rate and / or the short-time energy being greater than or equal to a preset short-time energy, determining that the second audio signal reaches the energy threshold, and extracting the audio feature of the second audio signal.
[0011] In some possible implementation ways, the audio signal detection module is further configured to: in response to the short-time zero-crossing rate being less than the preset zero-crossing rate and / or the short-time energy being less than the preset short-time energy, determine that the second audio signal does not reach the energy threshold, and control the main chip to interrupt the U-boot process.
[0012] In some possible implementation ways, the audio signal detection module, in determining whether the audio feature meets the preset condition, specifically comprises: determining a similarity between the audio feature and a preset feature sequence according to the audio feature and the preset feature sequence; in response to the similarity being greater than or equal to a preset similarity, determining that the audio feature meets the preset condition; and in response to the similarity being less than the preset similarity, determining that the audio feature does not meet the preset condition.
[0013] In some possible implementation ways, the audio signal detection module is further configured to: in response to the audio feature not meeting the preset condition, control the main chip to interrupt the U-boot process.
[0014] In some possible implementation ways, the wake-up module comprises: an audio acquisition circuit and an activation circuit; and the wake-up module is configured to, in response to receiving the first audio signal and in controlling the main chip to enter the U-boot process, specifically comprise: the audio acquisition circuit, in response to receiving the first audio signal, determining whether the first audio signal contains human voice and / or a target wake-up word; and the activation circuit, in response to the audio signal containing the human voice and / or the target wake-up word, sending an activation instruction to the main chip, the activation instruction being used to instruct the main chip to enter the U-boot process.
[0015] In some possible implementation ways, the audio signal detection module, in response to the audio feature meeting the preset condition and in starting the far-field voice module, specifically comprises: the audio signal detection module, in response to the audio feature meeting the preset condition, obtaining a third audio signal meeting the preset condition in the second audio signal; starting the far-field voice module and sending the third audio signal to the far-field voice module.
[0016] The far-field voice module controls the host to start after being started, and specifically includes: the far-field voice module receives a third audio signal, and controls the host to start according to a wake-up word in the third audio signal.
[0017] In a second aspect, the present application provides a control method of a voice interaction device, the voice interaction device comprising a wake-up module, a main chip and a host, the wake-up module being configured to collect an audio signal; the control method comprising: in response to receiving an activation instruction of the wake-up module, controlling the main chip to enter a U-boot process, the activation instruction being sent by the wake-up module after receiving a first audio signal; in the U-boot process of the main chip, implementing initialization of the voice interaction device and obtaining a second audio signal; in response to the second audio signal reaching an energy threshold, extracting an audio feature of the second audio signal; determining whether the audio feature meets a preset condition, the second audio signal comprising the first audio signal and an audio signal currently collected by the wake-up module; and in response to the audio feature meeting the preset condition, controlling the host to start according to the second audio signal.
[0018] In a third aspect, the present application provides a control device of a voice interaction device, the voice interaction device comprising a wake-up module, a main chip and a host, the wake-up module being configured to collect an audio signal;
[0019] The control device comprises: a receiving unit configured to, in response to receiving an activation instruction of the wake-up module, control the main chip to enter a U-boot process, the activation instruction being sent by the wake-up module after receiving a first audio signal;
[0020] An initialization unit is configured to, in the U-boot process of the main chip, implement initialization of the voice interaction device.
[0021] A first processing unit is configured to obtain a second audio signal, and in response to the second audio signal reaching an energy threshold, extract an audio feature of the second audio signal, determine whether the audio feature meets a preset condition, and the second audio signal comprising the first audio signal and an audio signal currently collected by the wake-up module.
[0022] A second processing unit is configured to, in response to the audio feature meeting the preset condition, control the host to start.
[0023] In a fourth aspect, the present application provides a computer-readable storage medium, the storage medium storing a computer program, the computer program being executed by the main chip to implement the control method of the voice interaction device according to the second aspect.
[0024] In a fifth aspect, the present application provides a computer program product comprising a computer program, the computer program being executed by the processor to implement the control method of the voice interaction device according to the second aspect.
[0025] The voice interaction device and the control method and the control device thereof provided by the application, the voice interaction device comprises a wake-up module, a main chip and a host, the wake-up module controls the main chip to enter a U-boot process in response to collecting a first audio signal; in the U-boot process, the main chip first initializes the voice interaction device and acquires a second audio signal; in response to the second audio signal reaching an energy threshold, an audio feature of the second audio signal is extracted, and it is determined whether the audio feature meets a preset condition; in response to the audio feature meeting the preset condition, the host is started according to the second audio signal. In the scheme, the wake-up module and the main chip cooperate to wake up the voice interaction module, which can guarantee the wake-up performance, thereby avoiding the situation that the voice interaction device is falsely woken up or cannot be normally woken up, and improving the user's interaction experience. In addition, the detection of the audio signal is realized in the U-boot process, thereby starting the host, which can improve the wake-up speed of the voice interaction device, and effectively prevent invalid noise from interfering with the starting process of the host.
[0026] These and other aspects of the application will become more fully understood from the following (multiple) embodiment descriptions. BRIEF DESCRIPTION OF DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the application or the implementation manner in the related art, the following will briefly introduce the drawings needed to be used in the embodiment or related art description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can also be obtained by those skilled in the art according to these drawings.
[0028] Figure 1 The application provides an application scenario schematic diagram of the voice interaction device for the embodiments of the application;
[0029] Figure 2 The application provides a structure schematic diagram of the voice interaction device for the embodiments of the application Figure 1 ;
[0030] Figure 3 The application provides a flow schematic diagram of the control method of the voice interaction device for the embodiments of the application Figure 1 ;
[0031] Figure 4 The application provides a structure schematic diagram of the voice interaction device for the embodiments of the application Figure 2 ;
[0032] Figure 5 The application provides a flow schematic diagram of the control method of the voice interaction device for the embodiments of the application Figure 2 ;
[0033] Figure 6A structural schematic diagram of a control device of a voice interaction device is provided for an embodiment of the present application. DETAILED DESCRIPTION
[0034] For the purpose of making the objects, embodiments and advantages of the present application more clear, the following will combine the drawings in the exemplary embodiments of the present application to make a clear and complete description of the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part of the embodiments of the present application, but not all the embodiments.
[0035] Based on the exemplary embodiments described in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative labor fall within the scope of the claims of the present application. In addition, although the disclosure in the present application is introduced according to one or more examples, it should be understood that each aspect of the disclosure can also constitute a complete embodiment independently.
[0036] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the following described embodiments, and is not intended to limit the embodiments of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and general meanings.
[0037] The terms "first", "second", "third" and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit the specific order or sequence, unless otherwise indicated. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, for example, those other than the order given in the embodiment illustration or description of the present application can be implemented.
[0038] In addition, the terms "include" and "have" and any variations thereof are intended to cover but not exclusive inclusion, for example, the product or device including a series of components does not have to be limited to those components clearly listed, but can include other components not clearly listed or inherent to these products or devices.
[0039] The term "module" used in the present application refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or combination of hardware or / and software code capable of performing functions related to the element.
[0040] The following detailed embodiments are used to illustrate the voice interaction device control method and implementation principle in the embodiments of the present application.
[0041] First, the application scenarios involved in the present application are explained:
[0042] Figure 1An application scenario diagram of a voice interaction device is provided for embodiments of the present application. As shown in Figure 1 the scenario includes a voice interaction device 100 and a user.
[0043] It should be noted that the voice interaction device 100 shown in the figure can be any electronic device with voice interaction function, for example, a mobile phone, a tablet computer, a television set, etc. display device, or a refrigerator, a washing machine, an air conditioner, etc. household appliance, and the embodiments of the present application do not make specific limitation.
[0044] It should be understood that according to different types of voice interaction device 100, there are multiple ways to interact with the voice interaction device 100. On the one hand, when the voice interaction device 100 has the functions of sound collection and voice output, the user can directly interact with the voice interaction device 100 by voice. Specifically, the voice control command of the user is collected by the sound collector on the voice interaction device 100, and the corresponding operation is performed according to the voice control command, for example, the operation of turning off or turning on according to the voice control command, or outputting interactive voice through the voice playing unit.
[0045] On the other hand, the scenario can also include a control terminal, for example, a remote controller, a mobile terminal such as a mobile phone or a tablet computer, which is provided with a voice collection module. The user can collect the voice control command of the user through the voice collection module on the control terminal, and then send the corresponding control instruction to the voice interaction device 100 through the control terminal, so as to realize the voice control of the voice interaction device 100.
[0046] In the related art, when the voice interaction device 100 is in the off or standby state, it is usually necessary to wake up the voice interaction device first. Specifically, the voice wake-up word can be directly input to the voice interaction device, and the wake-up module on the voice interaction device receives the wake-up word to wake up the voice interaction device, or the voice wake-up word can be input on the control terminal, and the wake-up module of the terminal device controls the wake-up of the voice interaction device according to the voice wake-up word. However, since the high-performance main chip has high power consumption, in order to reduce power consumption, whether the wake-up module on the control terminal or the wake-up module on the voice interaction device is usually configured with a low-power wake-up unit.
[0047] However, the operation ability of the low-power wake-up unit is limited and cannot process multi-channel microphone data and back sampling signal data. In a non-quiet environment and device playing condition, the wake-up rate decreases sharply, and even the situation of unable to wake up or false wake-up occurs, which seriously affects the user experience.
[0048] Therefore, the embodiment of the present application provides a voice interaction device and a control method and a control device thereof. The voice interaction device is woken up by a low-power wake-up module and a high-performance main chip. The low-power wake-up module determines whether the wake-up condition is met. When the wake-up condition is met, the high-performance main chip is started, and the main chip further determines whether the host is started. Thus, the situation that the voice interaction device is falsely woken up or cannot be normally woken up is avoided, and the user experience is improved. Meanwhile, a U-boot process is customized to guide the booting process, and the detection of the audio signal is implemented in the U-boot process, so that the host is started. Compared with the prior art in which the detection of the audio signal is implemented after the host is started, the wake-up speed of the voice interaction device is improved, and the interference of invalid noise on the booting process of the host is effectively prevented.
[0049] In addition, when the wake-up module does not control the main chip to enter the U-boot process, the high-performance main chip is in the closed state, and the power consumption of the voice interaction device is reduced.
[0050] It should be noted that the wake-up module can be a module in the voice interaction device or a module in the control terminal, and is not limited in actual application. Next, taking the wake-up module as a module in the voice interaction device as an example, the technical solution of the present application and how the technical solution of the present application solves the above technical problems are described in detail with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described again in some embodiments. The embodiments of the present application will be described below with reference to the drawings.
[0051] Figure 2 The structure of the voice interaction device provided by the embodiment of the present application is shown in the figure Figure 1 . As Figure 2 shown, the voice interaction device includes a wake-up module 201, a main chip 202, and a host 203.
[0052] In some embodiments, the wake-up module 201 and the main chip 202 are in communication connection, and the main chip 202 and the host 203 are in communication connection.
[0053] The wake-up module 201 is configured to collect an audio signal.
[0054] In the embodiments of the present application, the specific type of the main chip 202 is not limited. For example, in some embodiments, the main chip 202 can be one or more ASICs (Application Specific Integrated Circuits), or one or more DSPs (Digital Signal Processors), or one or more FPGAs (Field Programmable Gate Arrays), etc. For another example, when the above module is implemented in the form of a processing element scheduling code, the main chip 202 can be a general-purpose processor, such as a CPU or other processor that can call program code. For another example, the modules can be integrated together to implement a SOC (System-on-a-Chip) or the like.
[0055] It should be noted that the host 203 can be a general term for one or more elements, and the type of the elements is not limited in the embodiments of the present application. For example, the host 203 can include at least one of a display screen, a tuner, a communicator, a detector, and a memory.
[0056] For example, the detector is a temperature sensor, a humidity sensor, etc.; the communicator is at least one of a Wifi chip, a Bluetooth communication protocol chip, a wired Ethernet communication protocol chip, other network communication protocol chips or a near field communication protocol chip, and an infrared receiver; the tuner can receive signals through wired or wireless receiving mode, and perform modulation and demodulation processing such as amplification, frequency mixing and resonance, which will not be listed one by one here.
[0057] Next, the control method of the voice interaction device is described in detail in combination with specific embodiments:
[0058] Figure 3 The flowchart of the control method of the voice interaction device provided in the embodiments of the present application Figure 1 As shown in Figure 3 , the control method of the voice interaction device 100 provided in the embodiments of the present application specifically includes the following steps:
[0059] S301, the wake-up module controls the main chip to enter the U-boot process in response to the collected first audio signal.
[0060] In actual application, the wake-up module in the embodiments of the present application can be set as a low-power wake-up unit. Since its power consumption is low, it can detect the surrounding audio signal in real time and will not consume too much power.
[0061] Optionally, the wake-up module can include a microphone or a detection circuit, which is low in cost and can reduce the cost of the voice interaction device.
[0062] In some embodiments, the first audio signal can be any type of audio signal.
[0063] That is, any audio signal collected by the wake-up module around the voice interaction device can be used to wake up the voice interaction device, and the main chip needs to further determine whether the voice interaction device needs to be woken up. Therefore, after any audio signal is collected, it is regarded as the first audio signal, and the main chip is controlled to enter the U-boot process. In this way, the omission of the audio signal used to wake up the voice interaction device can be avoided.
[0064] In other embodiments, the first audio signal can be a human voice signal and / or an audio signal containing a wake-up word. That is, when the wake-up module collects an audio signal around the voice interaction device, the audio signal needs to be preliminarily screened to determine the human voice signal and / or the audio signal containing the wake-up word as the first audio signal.
[0065] Specifically, in the first aspect, the wake-up module detects the surrounding audio signal in real time, and determines whether the audio signal contains human voice. If the audio signal contains human voice, it means that the current audio signal can be used to wake up the voice interaction device. At this time, the main chip needs to further determine whether the voice interaction device needs to be woken up.
[0066] In this scheme, the wake-up module preliminarily judges the collected audio signal, which can accurately detect human voice and exclude noise interference in the environment, thereby avoiding the situation that the main chip is started multiple times due to noise, and can maximize the reduction of power consumption of the voice interaction device.
[0067] In the second aspect, the wake-up module detects the surrounding audio signal in real time, and determines whether the audio signal contains a target wake-up word or a wake-up word similar to the target wake-up word. If it contains, it is determined that the current audio signal is very likely to be used to wake up the voice interaction device. At this time, the main chip needs to further determine whether the voice interaction device needs to be woken up.
[0068] The target wake-up word is a wake-up word preset for the voice interaction device. The specific content of the wake-up word is not limited in the embodiments of the present application. For example, it can be "start", "shutdown", "standby", etc. It can also be a wake-up word for controlling parameters of the voice interaction device, such as "volume", "temperature", "resolution", etc. Details are not repeated here.
[0069] Compared with human voice detection, the main chip is started by detecting the wake-up word in the audio signal, which is more accurate and can reduce the probability of false start of the main chip, and reduce the power consumption of the voice interaction device.
[0070] Specifically, in the embodiment of the present application, the wake-up module can send an activation instruction to the main chip to instruct the main chip to enter the U-boot process.
[0071] For example, if the audio signal is "please start at 10 o'clock", the wake-up module detects the wake-up word "start", and determines that the current audio signal may be used to wake up the voice interaction device, and sends an activation instruction to the main chip.
[0072] S302, the main chip enters the U-boot process, initializes the voice interaction device, and obtains a second audio signal.
[0073] Correspondingly, after receiving the activation instruction sent by the wake-up module, the main chip enters the U-boot process, wherein the U-Boot process is mainly used to start the main bootloader.
[0074] The hardware to be initialized can be executed according to the hardware configuration of the voice interaction device, and the hardware to be initialized is different for different types of voice interaction devices, which usually includes display screen, audio module, communication module, etc., and can also include camera, radio frequency module, etc. The embodiment of the present application is not limited.
[0075] It should be noted that the second audio signal includes the first audio signal and the audio signal currently collected by the wake-up module.
[0076] Specifically, the wake-up module will buffer the detected audio signal in real time while detecting the surrounding audio signal, and send the collected audio signal to the main chip at the same time as sending the activation instruction to the main chip.
[0077] In some scenarios, when a user wakes up a voice interaction device through a voice command, the control instruction may be included in the wake-up voice, for example, when the wake-up voice is "×× device, please start and adjust the temperature to ×× degree", when the wake-up device collects the voice signal, the voice signal is taken as the first audio signal, and the main chip is controlled to enter the U-boot process, thereby controlling the voice interaction device to start. Only by sending the audio signal to the main chip can the wake-up voice device be controlled accurately at the same time to adjust the temperature to ×× degree, otherwise the instruction cannot be executed. Therefore, in the embodiment of the present application, the first audio signal collected by the wake-up device also needs to be sent to the main chip to prevent the control failure caused by missing control instructions, improve the interaction efficiency, and ensure the control effect.
[0078] In some scenarios, the control instruction is not included in the user's wake-up voice, for example, when the wake-up voice is "X X device, please turn on", when the wake-up device collects the voice signal, the voice signal is taken as the first audio signal, and the main chip is controlled to enter the U-boot process, so as to control the voice interaction device to turn on. However, no other control instructions are included in this voice, therefore, the wake-up module needs to continue collecting the subsequent audio signals and sending the audio signals to the main chip in real time, so that the main chip can realize more accurate control according to the audio signals.
[0079] Optionally, the second audio signal can be sent to the main chip in the activation instruction, or can be sent to the main chip separately, and the embodiments of the present application are not limited specifically.
[0080] S303, the main chip extracts the audio features of the second audio signal in response to the second audio signal reaching the energy threshold.
[0081] The inventor finds that since the wake-up module only performs preliminary screening on the first audio signal, when the wake-up module only realizes human voice detection, as long as the audio data contains human voice, it will send an activation instruction to the main chip to start the main chip; or, when the wake-up module realizes wake-up word detection, due to its low performance, it may also make a wrong judgment, at this time, it will also send an activation instruction to start the main chip. In the above two examples, the second audio signal received by the main chip is not used to wake up the voice interaction device.
[0082] Therefore, in the embodiments of the present application, the main chip needs to further judge whether the second audio signal is used to wake up the voice interaction device according to the energy value of the second audio signal. It should be noted that the type of energy threshold is not limited specifically in the embodiments of the present application, for example, it can be the short-time zero-crossing rate and / or the short-time energy corresponding to the second audio signal.
[0083] Specifically, when determining whether the second audio signal reaches the energy threshold, the following steps are specifically included:
[0084] (1) frame processing is performed on the second audio signal to obtain the voice frame corresponding to the second audio signal;
[0085] (2) the short-time energy and / or the short-time zero-crossing rate of the second audio signal is calculated according to the voice frame;
[0086] (3) in response to the short-time zero-crossing rate being greater than or equal to a preset zero-crossing rate, and / or the short-time energy being greater than or equal to a preset short-time energy, it is determined that the second audio signal reaches the energy threshold, and the audio features of the second audio signal are extracted.
[0087] It should be noted that the specific manner of obtaining the short-time energy and the short-time zero-crossing rate corresponding to the speech frame, and the manner of obtaining the audio feature of the second audio signal, are not described in detail in the embodiments of the present application.
[0088] It should be understood that when the short-time zero-crossing rate is greater than or equal to the preset zero-crossing rate, and / or, the short-time energy is greater than or equal to the preset short-time energy, it indicates that the second audio signal is a valid human voice signal, otherwise, it indicates that the second audio signal is an invalid human voice signal. In the embodiments of the present application, the second audio signal is further judged by the energy threshold, and only when the second audio signal is a valid human voice signal, the subsequent starting process is performed, which can further prevent the host from being mistakenly started, and improve the user's interactive experience.
[0089] S304, the main chip determines whether the audio feature meets the preset condition, and controls the host to start according to the second audio signal in response to the audio feature meeting the preset condition.
[0090] It should be noted that the type of the preset condition is not limited in the embodiments of the present application. For example, the preset condition can be set as: comparing the similarity of the audio feature and the preset feature sequence, and when the similarity meets the preset similarity, it indicates that the second audio feature meets the preset condition.
[0091] It should be noted that the preset feature sequence has multiple obtaining manners. In some embodiments, for the same voice interactive device, its control instruction usually has certain similarity. For example, for the voice interactive air conditioner, its control instruction is usually used to adjust the temperature, such as "adjust the temperature to × × degrees", "lower (increase) the temperature", etc., and for the voice interactive television, its control instruction is usually used to change the program type or adjust the television parameters, such as "adjust to × × channel", "adjust to × × program", "lower (increase) brightness, resolution, sound", etc.
[0092] In the embodiments of the present application, the audio feature corresponding to these control instructions can be taken as the preset feature sequence, thereby as the reference data, when the second audio signal is obtained, the audio feature corresponding to the second audio signal is compared with the preset feature sequence, and when the similarity is greater than or equal to the preset similarity, it indicates that the current second audio signal is used to control the voice interactive device, thereby realizing accurate control.
[0093] In other embodiments, the historical voice control instructions of the user can be taken as the reference data, and the preset feature sequence is obtained according to these historical voice control instructions, when the second audio signal is obtained, the audio feature corresponding to the second audio signal is compared with the preset feature sequence, and when the similarity is greater than or equal to the preset similarity, it indicates that the current second audio signal is used to control the voice interactive device.
[0094] In the embodiments of the present application, since the user of the same voice interactive device is usually a fixed user, the control instructions of the same user when controlling the voice interactive device at different times are similar. The historical voice control instructions of the users are used as the reference data, so that the accuracy of the judgment result can be ensured.
[0095] The control method of the voice interactive device provided in the embodiments of the present application can realize the wake-up of the voice interactive device by the low-power wake-up module and the high-performance main chip together. The wake-up module is used to determine whether the wake-up condition is met, and when the wake-up condition is met, the main chip is further used to determine whether to start the host, so that the situation of false wake-up or the voice interactive device cannot be normally woken up is avoided, and the user experience is improved. Meanwhile, the U-boot process is customized to guide the booting process, and the detection of the audio signal is realized in the U-boot process, so that the host is started. Compared with the prior art in which the detection of the audio signal is performed after the host is started, the wake-up speed of the voice interactive device can be improved, and the interference of invalid noise on the booting process of the host can be effectively prevented.
[0096] In addition, since the main chip is in the off state when the wake-up module does not control the main chip to enter the U-boot process, the power consumption of the voice interactive device can be reduced.
[0097] As an alternative to steps S303 and S304, whether the second audio signal contains the target voice data can also be determined, and then whether the second audio signal is used to wake up the voice interactive device is determined.
[0098] Specifically, the semantic analysis of the second audio signal is performed, when the second audio signal contains the target voice data used to wake up the voice interactive device, it is determined that the audio feature meets the preset condition, and the host is started according to the second audio signal in response to the audio feature meeting the preset condition. In the embodiments of the present application, the semantic analysis of the second audio data can be performed, so that the user's intention can be more accurate, and the accurate control of the voice interactive device can be realized.
[0099] In some optional embodiments, when the wake-up module determines that the received audio signal is not used to wake up the voice interactive device, the currently cached audio signal can be deleted, so as to reduce the storage pressure of the wake-up module.
[0100] In some optional embodiments, when the energy value of the second audio signal does not satisfy the energy threshold, and / or the main chip determines that the audio feature does not satisfy the preset condition, an indication information can be sent to the wake-up module. The indication information is used to instruct the wake-up module to stop sending the audio signal. Correspondingly, after the wake-up module receives the indication information, it stops sending the currently collected audio signal to the main chip. Through this setting, the wake-up module can be controlled to stop sending the audio signal in time, and the power consumption of the wake-up module and the main chip can be reduced to a certain extent.
[0101] In some optional embodiments, when the main chip determines that the audio feature does not satisfy the preset condition, the main chip can also be turned off, so as to reduce the power consumption of the main chip, until the main chip receives the next activation instruction sent by the wake-up module, and then the same processing is performed according to the above steps.
[0102] Figure 4 Structure diagram of the voice interaction device provided by the embodiments of the present application Figure 2 As shown in Figure 4 , in the voice interaction device 200 provided by the embodiments of the present application, the main chip 202 includes an audio signal detection module 2021 and a far-field voice module 2022.
[0103] The audio signal detection module 2021 is at least one computing core in the main chip 202 for processing the audio signal.
[0104] In an optional embodiment, the wake-up module 201 includes an audio collection circuit 2011 and an activation circuit 2012.
[0105] The audio collection circuit 2011 is applied to collect the audio signal, and in response to receiving the first audio signal, it determines whether the first audio signal contains human voice and / or the target wake-up word.
[0106] The activation circuit 2012 is used to send an activation instruction to the main chip in response to the audio signal containing human voice and / or the target wake-up word, and the activation instruction is used to instruct the main chip to enter the U-boot process.
[0107] It should be noted that the schemes performed by the audio collection circuit 2011 and the activation circuit 2012 in the embodiments of the present application are similar to the schemes performed by the wake-up module 201 in the embodiments shown in Figure 3 , and the specific implementation can be referred to the above embodiments, which will not be described here.
[0108] Next, the control method of the voice interaction device in the embodiments shown in Figure 5 will be described in more detail. Figure 4 The control method of the voice interaction device provided by the embodiments of the present application Figure 5 Figure 2 As shown in Figure 5 The control method provided by the embodiment of the application comprises the following steps:
[0109] S501, the wake-up module controls the main chip to enter a U-boot process in response to the first audio signal being collected.
[0110] S502, in the U-boot process, the main chip initializes the voice interaction device and acquires a second audio signal.
[0111] Specifically, when the wake-up module 201 responds to the first audio signal being collected, an activation instruction is sent to the audio signal detection module 2021 in the main chip 202, so that the main chip enters the U-boot process, and at the same time, in the U-boot process, the audio signal detection module 2021 is first woken up.
[0112] It should be noted that the scheme of waking up the audio signal detection module 2021 in steps S501-S502 is similar to the scheme of waking up the main chip 202 in steps S301-S302 in the embodiment shown in Figure 3 and the principle, and specific reference can be made to the above embodiment, which will not be described here.
[0113] S503, the audio signal detection module determines whether the second audio signal reaches an energy threshold.
[0114] S504, the audio signal detection module controls the main chip to interrupt the U-boot process in response to the second audio signal not reaching the energy threshold.
[0115] It should be noted that when the second audio signal does not reach the energy threshold, it means that the second audio signal is not used to control the voice interaction device, and the start-up process of the voice interaction device can be stopped by interrupting the U-boot process. In the embodiment of the application, since the energy value of the second audio signal can be used to accurately determine whether the second audio signal is used to control the voice interaction device, the determination can be made in the early stage of the start-up process of the voice interaction device, so that the subsequent start-up process is not needed, which can prevent the voice interaction device from being mistakenly started, and since only the audio signal detection module needs to be woken up in this process, the energy consumption of the main chip can be reduced.
[0116] S505, the audio signal detection module extracts the audio features of the second audio signal in response to the second audio signal reaching the energy threshold.
[0117] The energy threshold can be the short-term zero-crossing rate corresponding to the second audio signal, and / or the short-term energy, etc.
[0118] Specifically, in the embodiments of the present application, when the short-time zero-crossing rate is greater than or equal to the preset zero-crossing rate, and / or, the short-time energy is greater than or equal to the preset short-time energy, it is determined that the second audio signal reaches the energy threshold; correspondingly, when the short-time zero-crossing rate is less than the preset zero-crossing rate, and / or, the short-time energy is less than the preset short-time energy, it is determined that the second audio signal does not reach the energy threshold.
[0119] It should be noted that the specific scheme and beneficial effects of the audio signal detection module obtaining the short-time energy and the short-time zero-crossing rate can be referred to the scheme and beneficial effects of the step S303 in the embodiment shown in Figure 3 The steps S303-S304 in the embodiment shown in
[0120] S506, the audio signal detection module determines whether the audio feature meets the preset condition.
[0121] It should be noted that the steps S503-S505 are similar to the scheme and principle performed by the main chip 202 in the steps S303-S304 in the embodiment shown in Figure 3 The steps S303-S304 in the embodiment shown in
[0122] S507, the audio signal detection module controls the main chip to interrupt the U-boot process in response to the audio feature not meeting the preset condition.
[0123] S508, the audio signal detection module starts the far-field voice module in response to the audio feature meeting the preset condition.
[0124] It should be noted that when the audio feature meets the preset condition, it means that the second audio signal is used to control the voice interactive device, at this time, the far-field voice module 2022 in the main chip 202 is further woken up, so as to provide more accurate voice service through the far-field voice module 2022.
[0125] When the audio feature does not meet the preset condition, it means that the second audio signal is not used to control the voice interactive device, at this time, the U-boot process can also be interrupted, so as to stop the starting process of the voice interactive device.
[0126] In the embodiments of the present application, the second audio signal can be further judged, which can prevent the voice interactive device from being mistakenly started due to inaccurate energy value judgment of the second audio signal, and at the same time, the U-boot process can be interrupted in time in this process, so as to prevent the far-field voice module of the main chip from being woken up, which can reduce the energy consumption of the main chip.
[0127] S509, the far-field voice module controls the host to start according to the second audio signal after starting.
[0128] Specifically, after the far-field voice module is started, the control instruction in the second audio signal is acquired, and the starting of the host is controlled based on the control instruction.
[0129] The inventor finds that, since the audio signal is collected in real time by the wake-up module in the second audio signal, in the process of starting the voice interaction device, the wake-up module can receive multiple audio signals, which leads to the fact that the second audio signal can include part of the audio signals that do not meet the preset condition, and these audio signals can interfere with the process of starting the host by the far-field voice module. In view of this, in some embodiments, the step S507 specifically includes the following steps to solve the above problem:
[0130] (1) The audio signal detection module acquires the third audio signal in the second audio signal that meets the preset condition.
[0131] (2) The audio signal detection module starts the far-field voice module, and sends the third audio signal to the far-field voice module.
[0132] Correspondingly, the step S509 is specifically: the far-field voice module receives the third audio signal, and controls the starting of the host according to the wake-up word in the third audio signal.
[0133] In the embodiments of the present application, since the main chip is a high-performance engine based on wake-up word detection, it can perform noise reduction processing and wake-up word detection on the cached full-path microphone audio signal, which can guarantee the wake-up performance, thereby avoiding the situation that the voice interaction device is falsely woken up or cannot be normally woken up. In addition, the main chip is started by the wake-up module according to the audio signal, so that the main chip is in the off state when no starting instruction is received, which will not produce too much power consumption, thereby reducing the power consumption while guaranteeing the wake-up performance of the voice interaction device, and improving the user experience.
[0134] Secondly, in the process of starting the main chip, the audio signal detection module in the main chip is started first, and the second audio signal is preliminarily judged through this module, and when it is passed, the far-field voice module of the main chip is started, in this way, the power consumption of the main chip can be reduced to the greatest extent while guaranteeing the wake-up performance of the voice interaction device, and more accurate control can be realized through the far-field voice module.
[0135] Figure 6 The structural schematic diagram of the control device of the voice interaction device provided by the embodiments of the present application is provided. The voice interaction device includes a wake-up module, a main chip, and a host, wherein the wake-up module is used to collect an audio signal.
[0136] As shown in Figure 6 The control device 600 provided by the embodiments of the present application can include:
[0137] The receiving unit 601 is configured to, in response to receiving an activation instruction of the wake-up module, control the main chip to enter a U-boot process, the activation instruction being sent by the wake-up module after receiving the first audio signal; the initializing unit 602 is configured to, in the U-boot process of the main chip, implement initialization of the voice interaction device; the first processing unit 603 is configured to acquire a second audio signal, and in response to the second audio signal reaching an energy threshold, extract an audio feature of the second audio signal, and determine whether the audio feature meets a preset condition, the second audio signal including the first audio signal and an audio signal currently collected by the wake-up module; and the second processing unit 604 is configured to, in response to the audio feature meeting the preset condition, control the host to start.
[0138] In some possible implementation manners, the first processing unit 603 includes an audio signal detection module, and the second processing unit 604 includes a far-field voice module; the first processing unit 603 is specifically configured to: the audio signal detection module, in response to the second audio signal reaching the energy threshold, extracts the audio feature of the second audio signal, and determines whether the audio feature meets the preset condition; and the second processing unit 604 is specifically configured to: the audio signal detection module determines whether the audio feature meets the preset condition; in response to the audio feature meeting the preset condition, starts the far-field voice module; and the far-field voice module, after being started, controls the host to start according to the second audio signal.
[0139] In some possible implementation manners, the first processing unit 603 is specifically configured to: perform frame processing on the second audio signal to obtain a voice frame corresponding to the second audio signal; calculate a short-time energy and / or a short-time zero-crossing rate of the second audio signal according to the voice frame; and in response to the short-time zero-crossing rate being greater than or equal to a preset zero-crossing rate and / or the short-time energy being greater than or equal to a preset short-time energy, determine that the second audio signal reaches the energy threshold, and extract the audio feature of the second audio signal.
[0140] In some possible implementation manners, the first processing unit 603 is further configured to: in response to the short-time zero-crossing rate being less than the preset zero-crossing rate and / or the short-time energy being less than the preset short-time energy, determine that the second audio signal does not reach the energy threshold, and control the main chip to interrupt the U-boot process.
[0141] In some possible implementation manners, the first processing unit 603 is specifically configured to: determine a similarity between the audio feature and a preset feature sequence according to the audio feature and the preset feature sequence; in response to the similarity being greater than or equal to a preset similarity, determine that the audio feature meets the preset condition; and in response to the similarity being less than the preset similarity, determine that the audio feature does not meet the preset condition.
[0142] In some possible implementation manners, the first processing unit 603 is further configured to: in response to the audio feature not meeting the preset condition, control the main chip to interrupt the U-boot process.
[0143] In some possible implementation manners, the wake-up module comprises: an audio acquisition circuit and an activation circuit; the audio acquisition circuit is configured to determine whether the human voice and / or the target wake-up word are contained in the first audio signal in response to receiving the first audio signal; and the activation circuit is configured to send an activation instruction to the main chip in response to the human voice and / or the target wake-up word being contained in the audio signal, the activation instruction being used to instruct the main chip to enter the U-boot process.
[0144] In some possible implementation manners, the first processing unit 603 is specifically configured to: in response to the audio feature satisfying the preset condition, acquire a third audio signal satisfying the preset condition from the second audio signal; start the second processing unit 603, and send the third audio signal to the second processing unit 605; and the second processing unit 604, after being started, receives the third audio signal, and controls the host to start according to the wake-up word in the third audio signal.
[0145] It should be noted that the control device of the voice interaction device provided in the embodiment can be used to execute the voice interaction device control method described above, and the implementation manners and technical effects are similar, which will not be described here again.
[0146] It should be understood that the division of each module of the above device is only a logical function division, and all or part of the modules can be integrated into one physical entity, or can be physically separated. And these modules can all be implemented in the form of software called by a processing element; all can be implemented in the form of hardware; some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the processing module can be a separately set processing element, or can be integrated in a chip of the above device, and in addition, the processing module can be stored in the memory of the above device in the form of program code, and the function of the processing module is called and executed by a processing element of the above device. The implementation of other modules is similar. In addition, all or part of the modules can be integrated together, or can be independently implemented. The processing element here can be an integrated circuit with signal processing capability. In the implementation process, each step of the above method or each module can be completed by the integrated logic circuit of hardware or the instruction of software in the processing element.
[0147] For example, the above modules can be one or more integrated circuits configured to implement the above methods, such as one or more ASICs (Application Specific Integrated Circuits), or one or more DSPs (Digital Signal Processors), or one or more FPGAs (Field Programmable Gate Arrays), or the like. For another example, when a certain module above is implemented in the form of a processing element scheduling code, the processing element can be a general purpose host processor, such as a CPU, or other host processor that can invoke code. For another example, the modules can be integrated together to be implemented in the form of an SOC (System on a Chip).
[0148] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable devices. The computer program can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer program can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (such as floppy disk, hard disk, magnetic tape), optical media (such as DVD), or semiconductor media (such as solid state disk (SSD)) and the like.
[0149] The embodiments of the present application also provide a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is executed by a host processor, the control method of the voice interaction device is implemented according to any one of the method embodiments.
[0150] The embodiments of the present application also provide a chip running an instruction, and the chip is used to execute the control method of the voice interaction device according to any one of the method embodiments.
[0151] The embodiment of the present application further provides a computer program product, which comprises a computer program stored in a computer readable storage medium, and at least one host chip can read the computer program from the computer readable storage medium, and the at least one host chip executes the computer program to realize the control method of the voice interaction device provided by any one of the method embodiments.
[0152] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0153] For the convenience of explanation, the above description has been made in combination with specific embodiments. However, the above exemplary discussion is not intended to exhaust or limit the embodiments to the specific forms disclosed above. Various modifications and variations can be derived according to the above teachings. The selection and description of the above embodiments are to better explain the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.
Claims
1. A voice interactive device, characterized by The application relates to a voice interaction device, comprising: a wake-up module, a main chip and a host; the wake-up module is configured to control the main chip to enter a U-boot process in response to collecting a first audio signal; in the U-boot process, the main chip is configured to: initialize hardware of the voice interaction device, and acquire a second audio signal; perform frame processing on the second audio signal to obtain a voice frame corresponding to the second audio signal; the second audio signal contains the first audio signal and an audio signal currently collected by the wake-up module; according to the voice frame, the short-time energy of the second audio signal is calculated, and / or the short-time zero-crossing rate is calculated; in response to the short-time zero-crossing rate being greater than or equal to a preset zero-crossing rate, and / or the short-time energy being greater than or equal to a preset short-time energy, it is determined that the second audio signal reaches an energy threshold, and an audio feature of the second audio signal is extracted; it is determined whether the audio feature meets a preset condition, and the host is started according to the second audio signal in response to the audio feature meeting the preset condition; in response to the short-time zero-crossing rate being less than the preset zero-crossing rate, and / or the short-time energy being less than the preset short-time energy, it is determined that the second audio signal does not reach the energy threshold, and the main chip is interrupted from the U-boot process; in response to the audio feature not meeting the preset condition, the main chip is interrupted from the U-boot process.
2. The voice interaction device of claim 1, wherein, The main chip comprises an audio signal detection module and a far-field voice module; the main chip is configured to extract the audio feature of the second audio signal in response to the second audio signal reaching the energy threshold, and specifically comprises: the audio signal detection module extracts the audio feature of the second audio signal in response to the second audio signal reaching the energy threshold; the main chip is configured to determine whether the audio feature meets a preset condition, and start the host according to the second audio signal in response to the audio feature meeting the preset condition, and specifically comprises: the audio signal detection module determines whether the audio feature meets the preset condition; the far-field voice module is started in response to the audio feature meeting the preset condition; the far-field voice module controls the host to start according to the second audio signal after being started.
3. The voice interactive device of claim 1, wherein, When the audio signal detection module determines whether the audio feature meets the preset condition, specifically comprises: according to the audio feature and a preset feature sequence, a similarity between the audio feature and the preset feature sequence is determined; in response to the similarity being greater than or equal to a preset similarity, it is determined that the audio feature meets the preset condition; in response to the similarity being less than the preset similarity, it is determined that the audio feature does not meet the preset condition.
4. The voice interaction device of any one of claims 1 to 3, wherein, The wake-up module comprises an audio acquisition circuit and an activation circuit; when the wake-up module controls the main chip to enter the U-boot process in response to receiving the first audio signal, specifically comprises: the audio acquisition circuit determines whether the first audio signal contains human voice and / or a target wake-up word in response to receiving the first audio signal; The activation circuit sends an activation instruction to the main chip in response to the audio signal containing human voice and / or a target wake-up word, and the activation instruction is used to instruct the main chip to enter a U-boot process.
5. The voice interaction device of claim 2 or 3, wherein, The audio signal detection module specifically includes the following steps when starting the far-field voice module in response to the audio feature meeting a preset condition: The audio signal detection module acquires a third audio signal meeting the preset condition from the second audio signal in response to the audio feature meeting the preset condition; The far-field voice module controls the host to start according to the second audio signal after being started. The far-field voice module receives the third audio signal and controls the host to start according to a wake-up word in the third audio signal. The voice interaction device includes a wake-up module, a main chip, and a host, and the wake-up module is used to collect an audio signal.
6. A control method of a voice interaction device, characterized by, The control method includes the following steps: in response to receiving an activation instruction of the wake-up module, controlling the main chip to enter a U-boot process, and the activation instruction is sent by the wake-up module after receiving a first audio signal; In the U-boot process of the main chip, the initialization of the hardware of the voice interaction device is implemented, a second audio signal is acquired, and the second audio signal is frame-processed to obtain a voice frame corresponding to the second audio signal; the short-time energy and / or the short-time zero-crossing rate of the second audio signal are calculated according to the voice frame; in response to the short-time zero-crossing rate being greater than or equal to a preset zero-crossing rate and / or the short-time energy being greater than or equal to a preset short-time energy, it is determined that the second audio signal reaches an energy threshold, and an audio feature of the second audio signal is extracted; It is determined whether the audio feature meets a preset condition, and the second audio signal contains the first audio signal and an audio signal currently collected by the wake-up module; In response to the audio feature meeting the preset condition, the host is controlled to start according to the second audio signal; In response to the short-time zero-crossing rate being less than the preset zero-crossing rate and / or the short-time energy being less than the preset short-time energy, it is determined that the second audio signal does not reach the energy threshold, and the main chip is controlled to interrupt the U-boot process; In response to the audio feature not meeting the preset condition, the main chip is controlled to interrupt the U-boot process. The voice interaction device includes a wake-up module, a main chip, and a host, and the wake-up module is used to collect an audio signal.
7. A control device of a voice interaction device, characterized in that The control device includes: A receiving unit is configured to control the main chip to enter a U-boot process in response to receiving an activation instruction of the wake-up module, and the activation instruction is sent by the wake-up module after receiving a first audio signal; An initialization unit is configured to implement the initialization of the hardware of the voice interaction device in the U-boot process of the main chip. The first processing unit is configured to acquire a second audio signal, and in response to the second audio signal reaching an energy threshold, extract an audio feature of the second audio signal, and determine whether the audio feature meets a preset condition, the second audio signal containing the first audio signal and an audio signal currently collected by the wake-up module; The second processing unit is configured to control the host to start in response to the audio feature meeting the preset condition; The first processing unit is specifically configured to perform frame processing on the second audio signal to obtain a speech frame corresponding to the second audio signal, calculate a short-time energy and / or a short-time zero-crossing rate of the second audio signal according to the speech frame, and in response to the short-time zero-crossing rate being greater than or equal to a preset zero-crossing rate and / or the short-time energy being greater than or equal to a preset short-time energy, determine that the second audio signal reaches the energy threshold, and extract the audio feature of the second audio signal; The first processing unit is further configured to, in response to the short-time zero-crossing rate being less than the preset zero-crossing rate and / or the short-time energy being less than the preset short-time energy, determine that the second audio signal does not reach the energy threshold, and control the main chip to interrupt a U-boot process; The first processing unit is further configured to, in response to the audio feature not meeting the preset condition, control the main chip to interrupt the U-boot process.
8. A computer storage medium, characterized in that The computer storage medium includes instructions, which, when executed on a computer, cause the computer to perform the control method of the voice interaction device according to claim 6.
9. A computer program product, characterised in that, The computer program, when executed by a processor, implements the control method of the voice interaction device according to claim 6.
Citation Information
Patent Citations
Speech recognition power management
CN105009204A
Voice recognition method and device for smart home
CN110853631A