Voice interaction device, system, method, cloud server and medium
By adding a low-power voice processor and a wireless communication module to the voice interaction device, the voice signal is sent to the cloud server for processing, which solves the problem of high power consumption of the voice interaction device and achieves the effect of extending the battery life.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2024-12-06
- Publication Date
- 2026-06-09
AI Technical Summary
Existing voice interaction devices consume a lot of power when interacting with users, which affects battery life.
By adding a low-power voice processor and a low-power wireless communication module to the voice interaction device, a connection is established with the cloud server, and the voice signals collected by the microphone are sent to the cloud server for processing, reducing the power consumption of the local device.
While ensuring real-time voice interaction, it reduces the power consumption of electronic devices and extends battery life.
Smart Images

Figure CN122179247A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a voice interaction device, system, method, cloud server, and medium. Background Technology
[0002] With the continuous advancement of voice interaction technology, more and more electronic devices will be equipped with voice interaction functions, enabling users to control electronic devices to perform corresponding operations through voice commands, such as playing music or checking the weather, which can improve the convenience and efficiency of users using electronic devices.
[0003] Currently, when electronic devices equipped with voice interaction capabilities interact with users via voice, the processor in the device needs to receive the user's voice input in real time, perform voice recognition to obtain voice commands, and then perform corresponding operations based on the voice commands. However, this voice interaction method consumes a lot of power, thus affecting the battery life of electronic devices. Summary of the Invention
[0004] This application provides a voice interaction device, system, method, cloud server, and medium that can reduce the power consumption of electronic devices while meeting the real-time requirements of voice interaction, thereby extending the battery life of electronic devices.
[0005] In a first aspect, embodiments of this application provide a voice interaction device, including: a voice processor, a first communication module, at least one microphone, and a voice interaction system, wherein the voice interaction system includes a second communication module;
[0006] The first input terminal of the voice processor is connected to the microphone, and the first output terminal of the voice processor is connected to the first communication module. The voice processor is used to determine that the microphone has acquired a first voice signal, and to send the first voice signal acquired by the microphone to the cloud server through the connection established between the first communication module and the cloud server.
[0007] The voice interaction system is connected to the cloud server through the second communication module, and the voice interaction system is used to receive the first response signal sent by the cloud server to perform voice interaction.
[0008] Secondly, embodiments of this application provide a cloud server, including: a second speech recognition module and a second speech response module;
[0009] The second speech recognition module is used to convert the speech signal sent by the speech interaction device into text information;
[0010] The second voice response module is used to generate a first response signal based on the text information output by the second voice recognition module, and send the first response signal to the voice interaction device so that the voice interaction device can perform voice interaction based on the first response signal.
[0011] Thirdly, embodiments of this application provide a voice interaction system, including the voice interaction device as described in the first aspect embodiment above, and the cloud server as described in the second aspect embodiment above.
[0012] Fourthly, embodiments of this application provide a voice interaction method applied to the voice interaction device described in the first aspect of the embodiment, the method comprising:
[0013] When the microphone acquires the first voice signal, it sends the first voice signal acquired by the microphone to the cloud server;
[0014] Voice interaction is performed based on the first response signal sent by the cloud server.
[0015] Fifthly, embodiments of this application provide a voice interaction method applied to a cloud server as described in the second aspect of the embodiment above, the method comprising:
[0016] Receives voice signals sent by a voice interaction device, wherein the voice signals include a first voice signal;
[0017] A first response signal is generated based on the speech signal;
[0018] The first response signal is sent to the voice interaction device so that the voice interaction device can perform voice interaction based on the first response signal.
[0019] Sixthly, embodiments of this application provide a computer-readable storage medium for storing a computer program that causes a computer to execute the voice interaction method as described in the fourth aspect embodiments and their implementations, or the voice interaction method as described in the fifth aspect embodiments and their implementations.
[0020] In a seventh aspect, embodiments of this application provide a computer program product containing program instructions, which, when executed on an electronic device, cause the electronic device to perform the voice interaction method as described in the foregoing fourth aspect embodiments and their respective implementations, or the voice interaction method as described in the foregoing fifth aspect embodiments and their respective implementations.
[0021] The technical solution disclosed in this application adds a voice processor and a first communication module to the voice interaction device based on at least one microphone and a voice interaction system. The voice processor determines that the microphone has acquired a first voice signal, and the first communication module establishes a connection with the cloud server to send the first voice signal acquired by the microphone to the cloud server. Then, the second communication module connects to the voice interaction system of the cloud server to receive the first response signal sent by the cloud server for voice interaction. In this way, while meeting the real-time requirements of voice interaction, the power consumption of electronic devices can be reduced, thereby extending the battery life of electronic devices. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a schematic diagram of the structure of a first voice interaction device provided in an embodiment of this application;
[0024] Figure 2 This is a schematic diagram of the structure of a second voice interaction device provided in an embodiment of this application;
[0025] Figure 3 This is a schematic diagram of the structure of a third voice interaction device provided in an embodiment of this application;
[0026] Figure 4 This is a schematic diagram of the structure of a fourth voice interaction device provided in an embodiment of this application;
[0027] Figure 5 This is a schematic diagram of the structure of the fifth voice interaction device provided in the embodiments of this application;
[0028] Figure 6 This is a schematic diagram of the structure of the sixth voice interaction device provided in the embodiments of this application;
[0029] Figure 7 This is a structural schematic diagram of the seventh voice interaction device provided in the embodiments of this application;
[0030] Figure 8 This is a schematic diagram of the structure of a cloud server provided in an embodiment of this application;
[0031] Figure 9 This is a schematic diagram of the structure of a voice interaction system provided in an embodiment of this application;
[0032] Figure 10A flowchart illustrating a voice interaction method provided in an embodiment of this application;
[0033] Figure 11 A flowchart of another voice interaction method provided in an embodiment of this application. Detailed Implementation
[0034] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0035] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0036] In this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or solution described as "exemplary" or "for example" in this application should not be construed as being better or more advantageous than other embodiments or solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0037] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more, that is, at least two. "At least one" means one or more. "Any" means any one or any several.
[0038] Currently, when electronic devices equipped with voice interaction capabilities interact with users via voice, the processor within the device needs to receive user-authorized voice input in real time, perform voice recognition and other processing on the authorized voice to obtain voice commands, and then perform corresponding operations based on the voice commands. However, this voice interaction method consumes a significant amount of power, thus affecting the battery life of the electronic device.
[0039] To address the aforementioned technical problems, embodiments of this application provide a voice interaction device, system, method, cloud server, and medium to solve the problem that voice interaction requires a large amount of electrical energy, thereby affecting the battery life of electronic devices.
[0040] It should be understood that before using the technical solutions disclosed in the embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained. For example, in response to receiving a user's active request, a prompt message may be sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media that perform the operations of the technical solutions of this application, based on the prompt message.
[0041] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device. It is understood that the above notification and user authorization process is merely illustrative and does not limit the implementation of this application; other methods that comply with relevant laws and regulations may also be applied to the implementation of this application. It is understood that the user personal information data involved in this technical solution (including but not limited to the data itself, its acquisition, storage, and use) shall comply with the requirements of relevant laws and regulations and shall not violate public order and good morals.
[0042] The technical solutions provided by the embodiments of this application will be described in detail below through some examples. The embodiments described below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0043] Figure 1 This is a schematic diagram of the structure of a first voice interaction device provided in an embodiment of this application. Figure 1 As shown, the voice interaction device 1000 provided in this application embodiment may include: a voice processor 110, a first communication module 120, at least one microphone 130 and a voice interaction system 140, wherein the voice interaction system 140 includes a second communication module 141.
[0044] The first input terminal of the voice processor 110 is connected to the microphone 130, and the first output terminal of the voice processor 110 is connected to the first communication module 120. The voice processor 110 is used to determine that the microphone 130 has acquired the first voice signal, and to send the first voice signal acquired by the microphone 130 to the cloud server through the connection established with the cloud server by the first communication module 120.
[0045] The voice interaction system 140 is connected to the cloud server through the second communication module 141. The voice interaction system 140 is used to receive the first response signal sent by the cloud server for voice interaction.
[0046] In this embodiment, the voice processor 110 is a low-power voice processor; the first communication module 120 is a low-power wireless communication module, and the power consumption of the first communication module 120 is less than the power consumption of the second communication module 141. It should be understood that the power consumption of the first communication module 120 being less than the power consumption of the second communication module 141 could mean that the first communication module 120 and the second communication module 141 are two identical modules with different startup power consumption, or that the first communication module 120 and the second communication module 141 are two different modules. This application does not impose any limitations on this.
[0047] The low-power wireless communication module can be, but is not limited to, a low-power Wi-Fi communication module and a low-power LoRa module.
[0048] In this application, the second communication module 141 may be, but is not limited to, a Wi-Fi communication module, a 4G wireless communication module, or a 5G wireless communication module.
[0049] The aforementioned first response signal can be understood as a first voice response signal, that is, a voice reply signal generated by the cloud server based on any voice signal sent by the voice interaction device 1000. Optionally, the first response signal can respond to voice data, respond to text data, or respond to both voice data and text data. That is, the first response signal can be at least one of responding to voice data and responding to text data.
[0050] In this application, any voice signal sent by the voice interaction device 1000 can be a first voice signal, or it can be the first voice signal plus at least one other voice signal. For example, when the cloud server sends any voice signal based on the voice interaction device 1000, which is the first voice signal plus at least one other voice signal, the at least one other voice signal can be any number of voice signals following the first voice signal.
[0051] In some optional embodiments, when the voice interaction device 1000 does not receive interactive voice input from the user, the voice interaction device 1000 may be in a standby state. Furthermore, while the voice interaction device 1000 is in standby mode, each microphone 130 in the voice interaction device 1000 can continuously collect ambient voice signals. When any microphone 130 collects a voice signal, it can send that voice signal as a first voice signal to the voice processor 110. Then, the voice processor 110 determines that the user may need to perform voice interaction based on the first voice signal. At this time, the voice interaction device 1000 can be controlled to switch from standby mode to working mode, and the first voice signal collected by any microphone 130 can be sent to the cloud server through the communication connection established between the first communication module 120 and the cloud server.
[0052] Considering that each microphone 130 in the voice interaction device 1000 can perform real-time voice signal acquisition when in working condition, after the voice processor 110 sends the first voice signal acquired by any microphone 130 to the cloud server through the communication connection established between the first communication module 120 and the cloud server, any microphone 130 may continuously acquire at least one other voice signal. Therefore, the voice processor 110 will correspondingly send at least one other voice signal acquired by any microphone 130 to the cloud server through the communication connection established between the first communication module 120 and the cloud server.
[0053] When the cloud server receives a first voice signal, or the first voice signal and at least one other voice signal, sent by the voice interaction device 1000, it can generate a first response signal based on the first voice signal, or based on the first voice signal and at least one other voice signal. Then, the cloud server can send the generated first response to the voice interaction system 140 in the voice interaction device 1000 through the communication connection established between itself and the second communication module 141, so that the voice interaction system 140 receives the first response signal sent by the cloud server and performs voice interaction with the user.
[0054] In other words, the voice interaction system 140 in the voice interaction device 1000 of this application can be used to perform voice interaction based on a first response signal sent by a cloud server, wherein the first response signal sent by the cloud server is generated based on any voice signal acquired by at least one microphone 130. In this application, any voice signal includes the first voice signal and other voice signals besides the first voice signal. Thus, by adding a low-power voice processor (i.e., voice processor 110) and a low-power wireless communication module (first communication module 120) to the voice interaction system 140 and at least one microphone 130 of the voice interaction device 1000, when the low-power voice processor acquires the interactive voice sent by any microphone, it can upload the interactive voice to the cloud server in real time through the communication connection established with the cloud server by the low-power wireless communication module, so that the cloud server can generate corresponding interactive response data, and the voice interaction device can perform voice interaction with the user based on the interactive response data generated by the cloud server. This eliminates the need to use the high-power second communication module 141 in the voice interaction device 1000, thereby reducing the power consumption of the electronic device while meeting the real-time requirements of voice interaction, and thus extending the battery life of the electronic device.
[0055] In some alternative embodiments, considering that maintaining a communication connection between the first communication module 120 and the cloud server in the voice interaction device 1000 at all times may increase the power consumption of the voice interaction device 1000 and affect its battery life, the voice processor 110 in this application is further configured to: establish a connection with the cloud server through the first communication module 120 when it is determined that the second voice signal acquired by the microphone 130 meets the first preset condition, and send the first voice signal acquired by the microphone 130 to the cloud server, wherein the second voice signal precedes the first voice signal.
[0056] Optionally, in this application, each microphone 130 continuously collects ambient voice signals from the moment it begins operation. When any microphone 130 collects a second voice signal preceding the first voice signal (i.e., any microphone 130 collects the second voice signal first), it can send the second voice signal to the voice processor 110. The voice processor 110 then performs keyword detection and other operations on the second voice signal to determine if a preset keyword exists. When at least one preset keyword is determined to exist in the second voice signal, it indicates that the user needs to perform a voice interaction operation with the voice interaction device 1000. At this time, the voice processor 110 can establish a communication connection with the cloud server through the first communication module 120. Furthermore, when the voice processor 110 receives the first voice signal sent by the microphone 130, it can send the first voice signal to the cloud server through the first communication module 120.
[0057] When the cloud server receives a first voice signal or a combination of the first voice signal and at least one other voice signal from the voice interaction device 1000, it can generate a first response signal based on the first voice signal, or based on the first voice signal and at least one other voice signal. Then, the cloud server can send the generated first response signal to the voice interaction system 140 in the voice interaction device 1000 through the communication connection established between itself and the second communication module 141, so that the voice interaction system 140 can receive the first response signal sent by the cloud server and perform voice interaction with the user. Thus, by detecting whether the second voice signal includes preset keywords through the voice processor 110, the voice interaction device 1000 can be controlled to quickly enter the working state from the standby state when the user needs voice interaction, thereby further reducing the power consumption of the voice interaction device 1000 and extending its battery life. Furthermore, it can also reduce interference from irrelevant voice data and improve the accuracy and clarity of voice interaction.
[0058] In some alternative embodiments, such as Figure 2 As shown, the voice interaction system 140 provided in this application embodiment may further include: an interaction module 142.
[0059] The interaction module 142 is connected to the cloud server through the second communication module 141, and is used to receive the first response signal sent by the cloud server and perform voice interaction based on the first response signal.
[0060] Optionally, since the first response signal can be at least one of response voice data and response text data, the interaction module 142 can determine the type of the first response signal after receiving it from the cloud server. If the first response signal is determined to be response voice data, the response voice data is played to the user through the speaker in the voice interaction system 140. If the first response signal is determined to be response text data, the response text data is converted into response voice data and played to the user through the speaker in the voice interaction system 140. If the first response signal is determined to be both response voice data and response text data, the response text data is displayed to the user through the interactive screen provided by the interaction module 142 while the response voice data is played to the user through the speaker, thus achieving the purpose of voice interaction with the user.
[0061] In some alternative embodiments, such as Figure 3 As shown, the voice processor 110 provided in this application embodiment may include a voice detection component 111 and a keyword detection component 112.
[0062] The input end of the voice detection component 111 is connected to the microphone 130. The voice detection component 111 is used to perform voice detection on the second voice signal acquired by the microphone 130 to obtain the second voice segment.
[0063] The input end of the keyword detection component 112 is connected to the output end of the speech detection component 111. The keyword detection component 112 is used to perform keyword detection on the second speech segment output by the speech detection component 111. When it is determined that the second speech segment includes at least one preset keyword, it is determined that the second speech signal meets the first preset condition, and a connection is established with the cloud server through the first communication module 120.
[0064] The aforementioned speech detection component 111 can be understood as a speech activity detection (VAD) component, used to identify the speech and non-speech components in the audio signal (i.e., the speech signal in this application), and to determine the identified speech component as the second speech segment. The non-speech component can be background noise or silence.
[0065] The aforementioned first preset condition may refer to the detection of a trigger keyword in the second voice signal that establishes a communication connection between the voice processor 110 and the cloud server. This trigger keyword can be understood as a preset keyword. Establishing the communication connection between the voice processor 110 and the cloud server may be achieved through the first communication module 120.
[0066] In some optional embodiments, after the voice detection component 111 receives a second voice signal sent by any microphone 130, it can identify the second voice signal to obtain the voice portion including human voice and the non-voice portion excluding human voice from the second voice signal, and determine the voice portion including human voice as the second voice segment. Then, the voice detection component 111 outputs the second voice segment to the keyword detection component 112, so that the keyword detection component 112 detects whether the second voice segment includes at least one preset keyword, and then performs different operations according to the detection result.
[0067] Optionally, the keyword detection component 112 detects whether the second speech segment includes at least one preset keyword, which may include one of the following methods:
[0068] The first method involves obtaining the feature vector of the second speech segment and the feature vector of each preset keyword, and calculating the similarity between the feature vector of the second speech segment and the feature vector of each preset keyword. When any similarity is greater than a similarity threshold, the preset keyword corresponding to that similarity is determined to be included in the second speech segment. When each similarity is less than or equal to the similarity threshold, the second speech segment is determined to not include the preset keywords.
[0069] Secondly, when the keyword detection component 112 is any keyword detection model, the second speech segment can be input into the keyword detection model to process the second speech segment and output the keyword detection result for the second speech segment.
[0070] The keyword detection results include: a first detection result and a second detection result. The first detection result indicates that the second speech segment includes preset keywords, along with specific preset keyword information. The second detection result indicates that the second speech segment does not include preset keywords.
[0071] After obtaining the detection results, if the keyword detection component 112 determines that the second speech segment includes at least one preset keyword, then the second speech signal is determined to meet the first preset condition. At this time, a connection is established with the cloud server through the first communication module 120, so that the first speech signal subsequently obtained from the microphone 130 can be sent to the cloud server based on the established communication connection. This allows the cloud server to perform speech understanding based on any speech signal obtained from the voice interaction device 1000 and generate a corresponding first response signal. Conversely, if the keyword detection component 112 determines that the second speech segment does not include the preset keyword, then the second speech signal is determined not to meet the first preset condition. At this time, a connection is not established with the cloud server through the first communication module 120, and the second speech signal collected by each microphone 130 continues to be received, and it is determined whether the second speech signal meets the first preset condition. In this way, the communication connection between the voice processor 110 and the cloud server is established only through the first communication module 120 when the voice processor 110 determines that the second voice signal includes preset keywords. This avoids the problem of excessive power consumption caused by constantly establishing a communication connection between the voice processor 110 and the cloud server through the first communication module 120, thereby effectively extending the standby time of the voice interaction device 1000.
[0072] In some embodiments, after the voice processor 110 establishes a communication connection between itself and the cloud server via the first communication module 120 based on a second voice signal including preset keywords, the voice processor 110 may receive a first voice signal transmitted by any microphone 130. That is, the voice processor 110 may receive a first voice signal input by the user and authorized by the user, which is actually used for voice interaction. In this regard, the output of the voice detection component 111 in this embodiment can be connected to the first communication module 120, see [link to relevant documentation]. Figure 3 As shown.
[0073] The voice detection component 111 is also used to perform voice detection on the first voice signal acquired by the microphone 130, obtain a first voice segment, and send the first voice segment to the cloud server through the first communication module 120, so that the cloud server can generate a first response signal based on the first voice segment.
[0074] In other words, after establishing a communication connection with the cloud server through the first communication module 120, when the voice detection component 111 in the voice processor 110 receives the first voice signal sent by any microphone 130, it can first perform voice detection on the first voice signal to obtain a first voice segment. Then, the first voice segment is sent to the cloud server through the first communication module 120, so that the cloud server can quickly generate a first response signal based on the first voice segment. In this way, the processing complexity of the interactive voice data sent by the voice interaction device 1000 by the cloud server can be reduced, and the transmission efficiency and response efficiency of the interactive voice data can be improved.
[0075] Considering that the speech segment acquired by the speech detection component 111 from the speech signal transmitted by the microphone 130 may contain noise, and that the speech signal may include a first speech signal, a second speech signal, and other speech signals besides the first and second speech signals, the speech processor 110 in the optional embodiments of this application may further include a noise reduction component 113, such as... Figure 4 As shown.
[0076] The input of the noise reduction component 113 is connected to the output of the speech detection component 111, and the noise reduction component 113 is used to perform noise reduction processing on the second speech segment output by the speech detection component.
[0077] The input of the keyword detection component 112 is connected to the output of the noise reduction component 113. The keyword detection component 112 is used to detect keywords in the second speech segment after noise reduction output by the noise reduction component 113. When the second speech segment after noise reduction includes at least one preset keyword, it is determined that the second speech signal meets the first preset condition, and a connection is established with the cloud server through the first communication module 120.
[0078] In other words, this application sets a noise reduction component 113 between the speech detection component 111 and the keyword detection component 112 in the speech processor 110. The noise reduction component 113 performs noise reduction processing on the second speech segment output by the speech detection component 111 to filter out background noise, such as environmental noise, in the second speech segment. This makes the second speech segment after noise reduction clearer and purer, so that the keyword detection component 112 can perform keyword detection on the second speech segment after noise reduction more accurately and reliably, thereby improving the accuracy of keyword detection for the second speech segment.
[0079] In this application, the noise reduction component 113 can be understood as a device with an active noise reduction (NAC) algorithm.
[0080] and, Figure 4 The output of the noise reduction component 113 can also be connected to the first communication module 120. Therefore, the noise reduction component 113 can also be used to reduce the noise of the first speech segment output by the speech detection component 111, and send the noise-reduced first speech segment to the cloud server through the first communication module 120, so that the cloud server generates a first response signal based on the noise-reduced first speech segment. The advantage of this configuration is that by using the noise reduction component 113 to reduce the noise of the first speech segment output by the speech detection component 111, and then uploading the noise-reduced first speech segment to the cloud server through the first communication module 120, the cloud server can generate a first response signal based on the noise-reduced first speech segment. This improves the accuracy of speech recognition, reduces misrecognition, and generates more accurate response data. Furthermore, it increases the processing speed of speech data, enabling a real-time voice interaction experience.
[0081] In some optional embodiments, when the first communication module 120 in this application is a low-power Wi-Fi communication module, there may be a situation where there is no Wi-Fi network. Therefore, when the voice processor 110 determines that the second voice signal sent by any microphone 130 meets the second preset condition, the optional voice processor 110 can also send a wake-up event to the voice interaction system 140 to wake up the voice interaction system 140 and switch the voice interaction system 140 from standby state to working state, specifically as follows... Figure 5 .
[0082] The second input terminal of the voice processor 110 is connected to the output terminal of the voice interaction system 140, and the second output terminal of the voice processor 110 is connected to the input terminal of the voice interaction system 140.
[0083] In this application, the aforementioned second preset condition may refer to at least one of the following: detecting the presence of a preset wake-up word in the second voice signal for waking up the voice interaction system 140, and detecting the absence of a Wi-Fi network after receiving the second voice signal sent by any microphone 130.
[0084] The number of preset wake-up words is multiple, and can be flexibly set according to the type of voice interaction device 1000 and its application scenario. No restrictions are placed on this. For example, preset wake-up words could be "Hello, XX", "Hi, XXX", and "Zhang San, Zhang San", etc. It should be understood that XX in this application can be the identification information of the voice interaction device 1000, such as the name of the voice interaction device 1000.
[0085] That is, when the voice processor 110 determines that the second voice signal sent by any microphone 130 contains at least one preset wake-up word for waking up the voice interaction system 140, and / or detects that there is no Wi-Fi network after receiving the second voice signal sent by any microphone 130, the voice processor 110 sends a wake-up event to the voice interaction system 140 to wake up the voice interaction system 140, which is in standby mode. Then, the voice interaction system 140, which is in operation, performs voice recognition and other processing on the voice signal collected by the microphone 130 to generate a corresponding second response signal. The voice signal collected by the microphone 130 includes a first voice signal, or a first voice signal and at least one other voice signal.
[0086] Thus, when there is no Wi-Fi network and / or at least one wake-up word is present in the second voice signal, the voice processor 110 wakes up the voice interaction system 140 which is in standby mode, and sends the first voice signal sent by the microphone 130, or the first voice signal and at least one other voice signal, to the voice interaction system 140, so that the voice interaction system 140 generates a second response signal based on the first voice signal, or the first voice signal and at least one other voice signal, and performs voice interaction with the user based on the second response signal.
[0087] It is understood that when it is determined that at least one preset keyword exists in the second voice signal and a Wi-Fi network is present, the voice processor 110 of this application can send the first voice signal, or the first voice signal and at least one other voice signal, to the cloud server through the communication connection between the first communication module 120 and the cloud server, without waking up the voice interaction system 140. This can reduce the power consumption of the voice interaction device 1000 and further extend the standby time of the voice interaction device 1000. When it is determined that at least one preset wake-up word exists in the second voice signal, and / or when it is detected that there is no Wi-Fi network after receiving the second voice signal from any microphone 130, the voice processor 110 of this application can send a wake-up event to the voice interaction system 140 to bring the voice interaction system 140, which is in standby mode, into working mode. Furthermore, the first voice signal, or the first voice signal and at least one other voice signal, is sent to the voice interaction system 140, so that the voice interaction system 140 generates a second response signal based on the first voice signal, or the first voice signal and at least one other voice signal. This ensures that the system can respond to the user's interactive voice regardless of whether there is a Wi-Fi network or not, and regardless of whether the second voice signal carries a preset keyword or a preset wake-up word. This not only reduces the power consumption of the voice interaction device, but also improves the user's voice interaction experience.
[0088] In some alternative embodiments, such as Figure 6 As shown, the voice interaction system 140 of this application may further include: a first voice recognition module 143 and a first voice response module 144.
[0089] The input terminal of the first speech recognition module 143 is connected to the second output terminal of the speech processor 110. The first speech recognition module 143 is used to convert the speech signal output by the speech processor 110 into text information. The speech signal is any speech signal obtained by the microphone 130.
[0090] The input terminal of the first voice response module 144 is connected to the output terminal of the first voice recognition module 143. The first voice response module 144 is used to generate a second response signal based on the text information output by the first voice recognition module 143.
[0091] The interaction module 142 is connected to the output terminal of the first voice response module 144, and the interaction module 142 is also used to perform voice interaction according to the second response signal output by the first voice response module 144.
[0092] In this application, the first speech recognition module 143 can be understood as any functional module that supports an Automatic Speech Recognition (ASR) algorithm. This ASR algorithm is used to convert human speech into text.
[0093] The first speech response module 144 can be understood as any functional module that supports understanding speech signals and generating response signals corresponding to the speech signals. Optionally, in this application, the first speech response module 144 can be a Large Language Model (LLM) that supports understanding speech signals and generating response signals corresponding to the speech signals. This LLM is a model that uses massive amounts of training data to train a deep learning model to obtain a model that can recognize, understand, and generate corresponding text or speech content.
[0094] Optionally, if it is determined that at least one preset wake-up word exists in the second voice signal, and / or if it is detected that there is no Wi-Fi network after receiving the second voice signal from any microphone 130, the voice processor 110, after acquiring the first voice signal sent by the microphone 130, may send the first voice signal, or the first voice signal and at least one other voice signal, to the first voice recognition module 143 in the voice interaction system 140 that is in operation, so that the first voice recognition module 143 performs automatic voice recognition processing on the first voice signal, or the first voice signal and at least one other voice signal, to convert the first voice signal, or the first voice signal and at least one other voice signal, into corresponding text information. Then, the text information is output to the connected first voice response module 144, so that the first voice response module 144 performs semantic understanding and other processing on the acquired text information to generate a corresponding second response signal.
[0095] Next, the first voice response module 144 sends the generated second response signal to the interaction module 142, so that the interaction module 142 can interact with the user by voice based on the second response signal output by the first voice response module 144.
[0096] In this application, the second response signal may be response voice data, response text data, or response voice data and response text data.
[0097] Optionally, after receiving the second response signal sent by the first voice response module 144, the interaction module 142 can determine the type of the second response signal. If the second response signal is determined to be response voice data, the response voice data is played to the user through the speaker in the voice interaction system 140. If the second response signal is determined to be response text data, the response text data can be converted into response voice data and played to the user through the speaker in the voice interaction system 140. If the second response signal is determined to be both response voice data and response text data, the response voice data can be played to the user through the speaker in the voice interaction system 140 while the response text data is displayed to the user through the interactive screen provided by the interaction module 142, thus improving the diversity of voice interaction.
[0098] In some alternative embodiments, considering that after determining that the second voice signal sent by the microphone 130 meets the first preset condition, the voice processor 110 may determine that there is currently no wireless network, such as a Wi-Fi network, then the voice processor 110 can wake up the voice interaction system 140 and send the first voice signal sent by the microphone 130, or the first voice signal and at least one other voice signal, to the voice interaction system 140, so that the voice interaction system 140 processes the first voice signal, or the first voice signal and at least one other voice signal, to generate a second response signal. However, during or after the voice processor 110 sends the first voice signal sent by the microphone 130 to the voice interaction system 140, there may be a wireless network available. Therefore, this application can, when a wireless network is available, send the first voice signal, or the first voice signal and at least one other voice signal, to a cloud server through the voice interaction system 140, so that the cloud server can generate a first response signal based on the first voice signal, or the first voice signal and at least one other voice signal.
[0099] like Figure 7 As shown, the voice interaction system 140 of this application may further include: a voice storage module 145;
[0100] The input terminal of the voice storage module 145 is connected to the second output terminal of the voice processor 110 and the output terminal of the first voice recognition module 143, respectively. The voice storage module 145 is used to store the voice signal output by the voice processor 110 or the text information output by the first voice recognition module 143.
[0101] The output of the voice storage module 136 is connected to the cloud server through the second communication module 141. The voice storage module 145 is also used to transmit voice signals or text information to the cloud server through the second communication module 141, so that the cloud server generates a first response signal based on the voice signals or text information.
[0102] The first voice response module 144 is also used to receive the first response signal generated by the cloud server through the second communication module 141, and fuse the first response signal with the second response signal to obtain a fused response signal.
[0103] The interaction module 142 is also used to perform voice interaction based on the fused response signal output by the first voice response module 144.
[0104] In this application, the voice signal output by the aforementioned voice processor is any voice signal acquired by the microphone. Optionally, the voice signal can be a first voice signal, or a first voice signal and at least one other voice signal.
[0105] Optionally, when the voice processor 110 sends a voice signal to the voice interaction system 140, the voice storage module 145 in the voice interaction system 140 stores the voice signal sent by the voice processor 110. Simultaneously, the first voice recognition module 143 in the voice interaction system 140 automatically converts the voice signal sent by the voice processor 110 to obtain text information corresponding to the voice signal. Furthermore, after obtaining the text information, the first voice recognition module 143 can also send the text information to the voice storage module 145, laying the foundation for subsequent processing of the text information via a cloud server through a wireless network to generate a second response signal.
[0106] In some optional embodiments, when the voice storage module 145 stores the voice signal sent by the voice processor 110, and after storing the voice signal, if it is determined that a wireless network exists, the voice storage module 145 can send the voice signal it stores to the cloud server through the second communication module 141, so that the cloud server can generate a first response signal based on the received voice signal.
[0107] In some optional embodiments, when the voice storage module 145 stores the text signal sent by the first voice recognition module 143, and after storing the text signal, if it is determined that a wireless network exists, the voice storage module 145 can send the text information it stores to the cloud server through the second communication module 141, so that the cloud server can generate a first response signal based on the text information.
[0108] Optionally, while sending text information to the voice storage module 145, the first voice recognition module 143 of this application may also output text information to the first voice response module 144, so that the first voice response module 144 generates a second response signal based on the text information.
[0109] Because while the first voice response module 144 generates the second response signal based on the text information, the cloud server can generate the first response signal based on the received voice signal or text information, and also send the first response signal to the first voice response module 144. Therefore, when the first voice response module 144 receives the first response signal sent by the cloud server, it can fuse the first response signal and the second response signal to obtain a fused response signal. Then, the fused response signal is output to the interaction module 142, so that the interaction module 142 can perform voice interaction with the user based on the fused response signal. This can further improve the accuracy of voice interaction, thereby providing the user with more accurate interactive feedback.
[0110] The following is combined Figure 8 The cloud server provided in the embodiments of this application will be described. Figure 8 This is a schematic diagram of the structure of a cloud server provided in an embodiment of this application. Figure 8 As shown, the cloud server 2000 may include: a second speech recognition module 210 and a second speech response module 220.
[0111] The second speech recognition module 210 is used to convert the speech signal sent by the speech interaction device into text information.
[0112] The second voice response module 220 is used to generate a first response signal based on the text information output by the second voice recognition module 210, and send the first response signal to the voice interaction device so that the voice interaction device can perform voice interaction based on the first response signal.
[0113] In this application, the second speech recognition module 210 can be understood as any functional module that supports an Automatic Speech Recognition (ASR) algorithm. This ASR algorithm is used to convert human speech into text.
[0114] The second speech response module 220 can be understood as any functional module that supports understanding speech signals and generating response signals corresponding to the speech signals. Optionally, in this application, the second speech response module 220 can be a Large Language Model (LLM) that supports understanding speech signals and generating response signals corresponding to the speech signals. This LLM is a model that uses massive amounts of training data to train a deep learning model to obtain a model that can recognize, understand, and generate corresponding text or speech content.
[0115] Furthermore, the voice signal sent by the aforementioned voice interaction device includes a first voice signal, or a first voice signal and at least one other voice signal.
[0116] In some embodiments, after the cloud server 200 receives the voice signal sent by the voice interaction device, it can automatically perform voice recognition processing on the voice signal through the second voice recognition module 210 to convert the voice signal into corresponding text information. Then, the second voice recognition module 210 outputs the text information to the second voice response module 220, so that the second voice response module 220 performs semantic understanding and other processing on the acquired text information to generate a corresponding first response signal.
[0117] Next, the second voice response module 220 can return the first response signal to the voice interaction device so that the voice interaction device can interact with the user based on the first response signal.
[0118] Considering that the cloud server 2000 can obtain text information corresponding to the voice signal from the voice interaction device, the second voice response module 220 in the cloud server 2000 can directly perform semantic understanding and other processing on the text information to generate the corresponding first response signal, without the second voice recognition module 210 needing to perform any processing on the text information, thus further improving the data processing efficiency of the cloud server.
[0119] The technical solution provided in this application converts the voice signal sent by the voice interaction device into text information through a second voice recognition module in a cloud server. Then, a second voice response module in the cloud server performs semantic understanding and other processing on the text information output by the second voice recognition module to generate a first response signal. This first response signal is then sent to the voice interaction device, enabling the device to perform voice interaction based on the signal. This leverages the powerful data processing capabilities of the cloud server to quickly parse and process the voice signal uploaded by the voice interaction device, rapidly generating a response signal and achieving instant interaction between the user and the voice interaction device without prolonged waiting time, thus improving the user's voice interaction experience.
[0120] The following is a reference to the appendix. Figure 9 This application describes a voice interaction system proposed in its embodiments. Figure 9 As shown, the voice interaction system 10 includes: the voice interaction device 1000 described in the foregoing embodiments and the cloud server 2000 described in the foregoing embodiments.
[0121] In this device, the voice processor 110 determines that the microphone 130 has acquired a first voice signal, or the first voice signal and at least one other voice signal. Through the connection established between the first communication module 120 and the cloud server 2000, the first voice signal acquired by the microphone 130, or the first voice signal and at least one other voice signal, is sent to the second voice recognition module 210 in the cloud server 2000. The second voice recognition module 210 converts the first voice signal, or the first voice signal and at least one other voice signal, into text information and outputs the text information to the second voice response module 220. The second voice response module 220 generates a first response signal based on the text information and sends the first response signal to the voice interaction system 140 in the voice interaction device 1000, so that the voice interaction system 140 can perform voice interaction with the user based on the first response signal.
[0122] It should be understood that the specific implementation process of the voice interaction device 1000 and the cloud server 2000, as well as the interaction process between the voice interaction device 1000 and the cloud server 2000, in the above-mentioned embodiments of the voice interaction system 10 can be found in the above-mentioned embodiments of the voice interaction device 1000 and the above-mentioned embodiments of the cloud server 2000. To avoid repetition, they will not be described again here.
[0123] Figure 10 This is a flowchart illustrating a voice interaction method provided in an embodiment of this application. The voice interaction method provided in this application can be applied to any of the voice interaction devices described in the foregoing embodiments. The specific structure of the voice interaction device can be found in the foregoing embodiments, and will not be elaborated upon here. Figure 10 As shown, the voice interaction method may include the following steps:
[0124] S101: When the microphone acquires the first voice signal, it sends the first voice signal acquired by the microphone to the cloud server.
[0125] S102 performs voice interaction based on the first response signal sent by the cloud server.
[0126] In some alternative embodiments, the voice interaction method further includes:
[0127] When the microphone acquires the second voice signal, it determines whether the second voice signal meets the first preset condition;
[0128] If the second voice signal meets the first preset condition, a connection is established with the cloud server through the first communication module, and the first voice signal acquired by the microphone is sent to the cloud server. The second voice signal is earlier than the first voice signal.
[0129] In some alternative embodiments, the first response signal is generated based on any voice signal acquired by the microphone.
[0130] In some optional embodiments, determining whether the second voice signal meets a first preset condition includes:
[0131] The second speech signal acquired by the microphone is subjected to speech detection to obtain a second speech segment;
[0132] The second speech segment is subjected to keyword detection. When it is determined that the second speech segment includes at least one preset keyword, the second speech signal is determined to meet the first preset condition.
[0133] In some alternative embodiments, sending the first voice signal acquired by the microphone to a cloud server includes:
[0134] The first voice signal acquired by the microphone is subjected to voice detection to obtain a first voice segment. The first voice segment is then sent to the cloud server through the first communication module, so that the cloud server generates a first response signal based on the first voice segment.
[0135] In some optional embodiments, the voice interaction method further includes:
[0136] The second speech segment is subjected to noise reduction processing;
[0137] Keyword detection is performed on the second speech segment after noise reduction. When it is determined that the second speech segment after noise reduction includes at least one preset keyword, the second speech signal is determined to meet the first preset condition.
[0138] In some optional embodiments, the voice interaction method further includes:
[0139] The first speech segment is subjected to noise reduction processing, and the noise-reduced first speech segment is sent to the cloud server so that the cloud server generates a first response signal based on the noise-reduced first speech segment.
[0140] In some alternative embodiments, voice interaction is performed based on a first response signal sent by the cloud server, including:
[0141] Receive the first response signal sent by the cloud server, and perform voice interaction based on the first response signal.
[0142] In some optional embodiments, the voice interaction method further includes:
[0143] When it is determined that the second voice signal acquired by the microphone meets the second preset condition, a wake-up event is sent to the voice interaction system to wake up the voice interaction system.
[0144] In some optional embodiments, the voice interaction method further includes:
[0145] The speech signal is converted into text information, wherein the speech signal is any speech signal acquired by the microphone;
[0146] A second response signal is generated based on the text information;
[0147] Voice interaction is performed based on the second response signal.
[0148] In some optional embodiments, the voice interaction method further includes:
[0149] Store the voice signal, or store the text information corresponding to the voice signal;
[0150] The voice signal or the text information is transmitted to a cloud server, so that the cloud server generates a first response signal based on the voice signal or the text information;
[0151] Receive the first response signal generated by the cloud server, and fuse the first response signal with the second response signal to obtain a fused response signal;
[0152] Voice interaction is performed based on the fused response signal.
[0153] It should be understood that the embodiments of the voice interaction method in this application correspond to the aforementioned embodiments of the voice interaction device, and similar descriptions can be found in the aforementioned section on the embodiments of the voice interaction device. To avoid repetition, they will not be repeated here.
[0154] Figure 11 This is a flowchart illustrating another voice interaction method provided in an embodiment of this application. The voice interaction method provided in this application can be applied to any of the cloud servers described in the foregoing embodiments. The specific structure of the cloud server can be found in the foregoing embodiments, and will not be elaborated upon here. Figure 11 As shown, the voice interaction method may include the following steps:
[0155] S201, Receive a voice signal sent by a voice interaction device, wherein the voice signal includes a first voice signal.
[0156] S202, Generate a first response signal based on the voice signal.
[0157] S203, the first response signal is sent to the voice interaction device so that the voice interaction device performs voice interaction according to the first response signal.
[0158] In some optional embodiments, the voice interaction method further includes:
[0159] Receive text information sent by a voice interaction device, the text information corresponding to the voice signal;
[0160] A first response signal is generated based on the text information.
[0161] It should be understood that the voice interaction method embodiments of this application correspond to the aforementioned cloud server embodiments, and similar descriptions can be found in the aforementioned cloud server embodiment section. To avoid repetition, they will not be repeated here.
[0162] This application also provides a computer storage medium storing a computer program thereon, which, when executed by a computer, enables the computer to perform any of the above-described voice interaction methods.
[0163] This application also provides a computer program product containing program instructions that, when run on an electronic device, cause the electronic device to perform any of the above-described voice interaction methods.
[0164] When implemented using software, it can be implemented entirely or partially as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0165] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0166] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0167] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. For example, the functional modules in the various embodiments of this application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0168] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0169] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A voice interaction device, characterized in that, include: The system includes a voice processor, a first communication module, at least one microphone, and a voice interaction system, wherein the voice interaction system includes a second communication module. The first input terminal of the voice processor is connected to the microphone, and the first output terminal of the voice processor is connected to the first communication module. The voice processor is used to determine that the microphone has acquired a first voice signal, and to send the first voice signal acquired by the microphone to the cloud server through the connection established between the first communication module and the cloud server. The voice interaction system is connected to the cloud server through the second communication module, and the voice interaction system is used to receive the first response signal sent by the cloud server to perform voice interaction.
2. The voice interaction device according to claim 1, characterized in that, The power consumption of the first communication module is less than that of the second communication module.
3. The voice interaction device according to claim 1, characterized in that, The voice processor is further configured to: when it determines that the second voice signal acquired by the microphone meets the first preset condition, establish a connection with the cloud server through the first communication module, and send the first voice signal acquired by the microphone to the cloud server, wherein the second voice signal is earlier than the first voice signal.
4. The voice interaction device according to claim 1, characterized in that, The voice interaction system is also used to: perform voice interaction based on a first response signal sent by the cloud server; The first response signal sent by the cloud server is generated based on any voice signal acquired by the microphone.
5. The voice interaction device according to claim 3, characterized in that, The voice processor includes: a voice detection component and a keyword detection component; The input end of the voice detection component is connected to the microphone, and the voice detection component is used to perform voice detection on the second voice signal acquired by the microphone to obtain a second voice segment; The input end of the keyword detection component is connected to the output end of the speech detection component. The keyword detection component is used to perform keyword detection on the second speech segment output by the speech detection component. When it is determined that the second speech segment includes at least one preset keyword, it is determined that the second speech signal meets the first preset condition, and a connection is established with the cloud server through the first communication module.
6. The voice interaction device according to claim 5, characterized in that, The output of the voice detection component is connected to the first communication module. The voice detection component is also used to perform voice detection on the first voice signal acquired by the microphone to obtain a first voice segment, and send the first voice segment to the cloud server through the first communication module so that the cloud server generates a first response signal based on the first voice segment.
7. The voice interaction device according to claim 5, characterized in that, The voice processor also includes: a noise reduction component; The input terminal of the noise reduction component is connected to the output terminal of the speech detection component, and the noise reduction component is used to perform noise reduction processing on the second speech segment output by the speech detection component. The input of the keyword detection component is connected to the output of the noise reduction component, and the keyword detection component is used to detect keywords in the second speech segment after noise reduction output by the noise reduction component.
8. The voice interaction device according to claim 7, characterized in that, The output of the noise reduction component is connected to the first communication module. The noise reduction component is also used to perform noise reduction processing on the first speech segment output by the speech detection component, and send the noise-reduced first speech segment to the cloud server through the first communication module, so that the cloud server generates a first response signal based on the noise-reduced first speech segment.
9. The voice interaction device according to claim 1, characterized in that, The voice interaction system also includes: an interaction module; The interaction module is connected to the cloud server through the second communication module, and is used to receive the first response signal sent by the cloud server and perform voice interaction based on the first response signal.
10. The voice interaction device according to claim 1, characterized in that, The second input terminal of the voice processor is connected to the output terminal of the voice interaction system, and the second output terminal of the voice processor is connected to the input terminal of the voice interaction system. The voice processor is also used to send a wake-up event to the voice interaction system to wake up the voice interaction system when it determines that the second voice signal acquired by the microphone meets the second preset condition.
11. The voice interaction device according to claim 1, characterized in that, The voice interaction system further includes: a first voice recognition module and a first voice response module; The input terminal of the first speech recognition module is connected to the second output terminal of the speech processor. The first speech recognition module is used to convert the speech signal output by the speech processor into text information. The speech signal is any speech signal acquired by the microphone. The input terminal of the first voice response module is connected to the output terminal of the first voice recognition module, and the first voice response module is used to generate a second response signal based on the text information output by the first voice recognition module. The interaction module is connected to the output terminal of the first voice response module, and the interaction module is also used to perform voice interaction based on the second response signal output by the first voice response module.
12. The voice interaction device according to claim 11, characterized in that, The voice interaction system also includes: a voice storage module; The input terminal of the voice storage module is connected to the second output terminal of the voice processor and the output terminal of the first voice recognition module, respectively. The voice storage module is used to store the voice signal output by the voice processor or the text information output by the first voice recognition module. The output of the voice storage module is connected to the cloud server through the second communication module. The voice storage module is also used to transmit the voice signal or the text information to the cloud server through the second communication module, so that the cloud server generates a first response signal based on the voice signal or the text information. The first voice response module is further configured to receive a first response signal generated by the cloud server through the second communication module, and fuse the first response signal with the second response signal to obtain a fused response signal; The interaction module is also used to perform voice interaction based on the fused response signal output by the first voice response module.
13. A cloud server, characterized in that, include: Second speech recognition module and second speech response module; The second speech recognition module is used to convert the speech signal sent by the speech interaction device into text information; The second voice response module is used to generate a first response signal based on the text information output by the second voice recognition module, and send the first response signal to the voice interaction device so that the voice interaction device can perform voice interaction based on the first response signal.
14. A voice interaction system, characterized in that, It includes the voice interaction device as described in any one of claims 1-12, and the cloud server as described in claim 13.
15. A voice interaction method, characterized in that, Applied to the voice interaction device as described in any one of claims 1-12, the method comprises: When the microphone acquires the first voice signal, it sends the first voice signal acquired by the microphone to the cloud server; Voice interaction is performed based on the first response signal sent by the cloud server.
16. A voice interaction method, characterized in that, Applied to the cloud server as described in claim 13, the method includes: Receives voice signals sent by a voice interaction device, wherein the voice signals include a first voice signal; A first response signal is generated based on the speech signal; The first response signal is sent to the voice interaction device so that the voice interaction device can perform voice interaction based on the first response signal.
17. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the method as described in any one of claims 15-16.
18. A computer program product containing program instructions, characterized in that, When the program instructions are executed on the electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 15-16.