Speech interaction method and apparatus, and electronic device and storage medium
By recognizing user voice information in full-duplex mode and utilizing classifiers and large language models, the problem of responding when the user's voice is irrelevant to the application scenario is solved, and the quality of voice interaction and user experience are improved, especially in scenarios such as immersive cooking.
Patent Information
- Application Number
- PCT/CN2025/070169
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-12
- Filing Date
- 2025-01-02
- Publication Date
- 2025-07-17
AI Technical Summary
In full-duplex voice interaction scenarios, when the user's voice is irrelevant to the current application scenario, the electronic device still responds, resulting in a decrease in overall interaction quality and a poor user experience.
It adopts full-duplex mode, identifies whether the user's voice information belongs to a specific application scenario, responds only when it is relevant, and uses classifiers and large language models to identify specific intentions to achieve intelligent anti-interference.
It improves the quality of voice interaction and enhances the user experience, especially in scenarios such as immersive cooking, maintaining the immersive experience while enhancing intelligence.
Smart Images

Figure CN2025070169_17072025_PF_FP_ABST
Abstract
Description
Voice interaction method and device, electronic device and storage medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 12, 2024, with application number "2024100515181", the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of voice interaction technology, and in particular to a voice interaction method and device, an electronic device, and a storage medium. Background Art
[0003] In some voice interaction scenarios, non-full-duplex communication requires electronic devices to wake up each time, which is not intelligent enough and leads to a poor user experience. In full-duplex communication, Automatic Speech Recognition (ASR) monitors user voice, enabling a single wake-up and continuous conversation. However, if the user's voice is irrelevant to the current voice interaction scenario, the device will respond to every voice call, resulting in poor overall voice interaction quality and a degraded user experience. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a voice interaction method and device, an electronic device and a storage medium, which can solve the problem that in a full-duplex situation, when the user voice is irrelevant to the current voice interaction application scenario, if all user voices are responded to, the overall voice interaction quality is poor, which reduces the user experience.
[0005] In a first aspect, according to an embodiment of the present application, a voice interaction method is provided for an electronic device having a voice interaction function, comprising: controlling the electronic device to enter full-duplex mode; receiving a first voice message from a user; recognizing the first voice message; obtaining output information corresponding to the first voice message when the recognition result indicates that the first voice message belongs to a first application scenario; and controlling the electronic device to perform voice interaction with the user based on the output information.
[0006] In a second aspect, according to an embodiment of the present application, a voice interaction device is provided, comprising a first control module, a receiving module, a recognition module, an acquisition module, and a second control module. The first control module is used to control the electronic device to enter full-duplex mode. The receiving module is used to receive a first voice message from a user. The recognition module is used to recognize the first voice message. The acquisition module is used to obtain output information corresponding to the first voice message when the recognition result indicates that the first voice message belongs to voice information of a first application scenario. The second control module is used to control the electronic device to perform voice interaction with the user based on the output information.
[0007] In a third aspect, according to an embodiment of the present application, an electronic device is provided, which uses the voice interaction method of the first aspect to perform voice interaction with a user.
[0008] In a fourth aspect, according to an embodiment of the present application, an electronic device is provided, which includes a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the voice interaction method of the first aspect are implemented.
[0009] In a fifth aspect, according to an embodiment of the present application, a readable storage medium is provided, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the voice interaction method of the first aspect are implemented.
[0010] In a sixth aspect, according to an embodiment of the present application, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the voice interaction method of the first aspect are implemented.
[0011] According to an embodiment of the present application, in a full-duplex situation, when the user voice is related to the current first application scenario, responding to the user voice can improve the quality of the interaction between the electronic device and the user voice and enhance the user experience.
[0012] Additional aspects and advantages of the present application will become apparent in the following description or may be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0014] FIG1 shows a flow chart of a voice interaction method according to an embodiment of the present application;
[0015] FIG2 shows a second flow chart of a voice interaction method according to an embodiment of the present application;
[0016] FIG3 shows a third flow chart of a voice interaction method according to an embodiment of the present application;
[0017] FIG4 shows a fourth flow chart of a voice interaction method according to an embodiment of the present application;
[0018] FIG5 shows a structural block diagram of a voice interaction device provided according to an embodiment of the present application;
[0019] FIG6 shows a structural block diagram of an electronic device provided according to an embodiment of the present application;
[0020] FIG7 shows a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application;
[0021] FIG8 shows a fifth flow chart of a voice interaction method according to an embodiment of the present application;
[0022] FIG9 shows a sixth flow chart of a voice interaction method according to an embodiment of the present application;
[0023] FIG10 shows a schematic diagram of a central control screen of a cooking device provided according to an embodiment of the present application.
[0024] Among them, the correspondence between the figure marks and component names in Figures 5 to 7 is: 100: voice interaction device; 110: first control module; 120: receiving module; 130: identification module; 140: acquisition module; 150: second control module; 1000: electronic device; 1002: processor; 1004: memory; 1100: electronic device; 1101: radio frequency unit; 1102: network module; 1103: audio output unit; 1104: input unit; 11041: graphics processor; 11042: microphone; 1105: sensor; 1106: display unit; 11061: display panel; 1107: user input unit; 11071: touch panel; 11072: other input devices; 1108: interface unit; 1109: memory; 1110: processor. DETAILED DESCRIPTION
[0025] In order to more clearly understand the above-mentioned objects, features and advantages of the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features therein can be combined with each other in the absence of conflict.
[0026] In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the present application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present application is not limited to the specific embodiments disclosed below.
[0027] The voice interaction method and device, electronic device, and storage medium provided according to the embodiments of the present application are described in detail below with reference to Figures 1 to 10 through specific embodiments and their application scenarios.
[0028] According to an embodiment of the present application, a voice interaction method is provided. FIG1 shows one flow chart of the voice interaction method provided according to an embodiment of the present application. As shown in FIG1 , the voice interaction method is used in an electronic device having a voice interaction function, and includes:
[0029] Step S102: Control the electronic device to enter full-duplex mode.
[0030] Step S104: receiving a first voice message from a user.
[0031] Step S106: Recognize the first voice information.
[0032] Step S108: When the recognition result indicates that the first voice information belongs to the first application scenario voice information, output information corresponding to the first voice information is obtained.
[0033] Step S110: Based on the output information, control the electronic device to perform voice interaction with the user.
[0034] It is understandable that full-duplex refers to the ability to instantly transmit signals in both directions at the same time, that is, A transmits to B and B transmits to A, which is instantaneous synchronization. During the voice interaction process, full-duplex mode means that the electronic device is awakened once by the user, and the electronic device continues to monitor the user's voice. When the user performs voice interaction, it also continues to collect sound, and continues to process voice information. Multiple rounds of voice interaction can continue between the electronic device and the user. It is understandable that in voice interaction scenarios, such as immersive cooking, field question and answer, etc., a non-full-duplex mode is adopted, which requires multiple wake-ups, is not smart enough, and has a poor user experience. With full-duplex mode, only one wake-up is required, and the conversation can continue, avoiding the user having to wake up every time he speaks a sentence. However, when the user voice is irrelevant to the current voice interaction application scenario, the user voice is responded to, and the overall voice interaction quality is poor, which reduces the user experience.
[0035] In some embodiments of the present application, optionally, there is no restriction on the implementation method of controlling the electronic device to enter the full-duplex mode. For example, the electronic device can be pre-set to enter the full-duplex mode after being awakened, or directly enter the full-duplex mode when it recognizes that the user uses a specific wake-up word corresponding to the full-duplex mode.
[0036] In some embodiments of the present application, optionally, the recognition result indicates whether the first voice information belongs to the first application scenario voice information, which can also be understood as whether the recognition result of the first voice information is associated with the first application scenario. If so, the recognition result indicates that the first voice information belongs to the first application scenario voice information; if not, the recognition result indicates that the first voice information does not belong to the first application scenario voice information.
[0037] In some embodiments of the present application, optionally, for an electronic device, it is controlled to enter full-duplex mode and perform voice interaction with the user. The electronic device receives a first voice message from the user, then recognizes the first voice message to obtain a recognition result, and when the recognition result indicates that the first voice message belongs to the first application scenario voice message, obtains the corresponding output information of the first voice message. Based on the output information, the electronic device is controlled to perform voice interaction with the user. In the above manner, in the full-duplex case, when the user voice is related to the current first application scenario, responding to the user voice can improve the quality of the voice interaction between the electronic device and the user and enhance the user experience.
[0038] In some embodiments of the present application, optionally, when the recognition result indicates that the first voice information belongs to the first application scenario voice information, if the first voice information contains control instructions for other devices, the control instructions are sent directly or indirectly to the other devices to cause the other devices to perform corresponding operations, and combined with the execution results fed back by the other devices, output information corresponding to the first voice information is obtained. In some embodiments of the present application, the output information may optionally include content corresponding to different first voice information, corresponding to the first voice information, and the electronic device needs to interact with the user.
[0039] In some embodiments of the present application, the voice interaction method can optionally be applied to immersive cooking, laundry, domain question and answer, control, home appliance encyclopedia, etc. Specifically, the first application scenario can include immersive cooking, laundry, domain question and answer, control, home appliance encyclopedia, etc.
[0040] In some embodiments of the present application, optionally, an immersive cooking scenario is taken as an example to illustrate: when performing voice interaction, in the case of non-full-duplex, it is necessary to wake up each time, which is very unintelligent. In the case of full-duplex, only one wake-up is required, which avoids waking up every time the user speaks a sentence. Automatic Speech Recognition (ASR) is always monitoring the user's voice. Since cooking takes a certain amount of time, there is generally a certain time interval between different cooking steps, and the kitchen scene may be noisy, the ASR recognition accuracy is poor, and if a reply response is given to all ASR monitoring results, the overall conversation quality will be poor. In addition, when the user's current voice has nothing to do with cooking, irrelevant voice will disrupt the immersive cooking task and reduce the user experience.
[0041] In some embodiments of the present application, optionally, using an immersive cooking scenario as an example, a cooking device is controlled to enter full-duplex mode to interact with the user through voice. The electronic device receives a first voice message from the user, then recognizes the first voice message to obtain a recognition result. If the recognition result indicates that the first voice message belongs to an immersive cooking scenario, output information corresponding to the first voice message is obtained. Based on the output information, the cooking device is controlled to interact with the user through voice. In this manner, in a full-duplex mode, when the user's voice is relevant to the cooking device, the cooking device responds to the user's voice, improving the quality of the voice interaction between the cooking device and the user and enhancing the user experience.
[0042] In some embodiments of the present application, optionally, the voice interaction method further includes:
[0043] When the recognition result indicates that the first voice information does not belong to the first application scenario voice information, the electronic device is controlled to make a first response.
[0044] In some embodiments of the present application, optionally, the first voice information is recognized, and when the recognition result indicates that the first voice information does not belong to the first application scenario voice information, the electronic device is controlled to make a first response.
[0045] In some embodiments of the present application, optionally, the first response includes silence, that is, the electronic device does not perform any processing and remains silent.
[0046] In some embodiments of the present application, optionally, when the recognition result is not the current first application scenario, the text content corresponding to the first voice message received from the user may be simply presented on the interface without any other processing. For example, when the electronic device has a central control screen, the central control screen will display the text corresponding to the user's first voice message, and the electronic device may remain silent without any processing.
[0047] In some embodiments of the present application, optionally, the first response may also include a specific speech response to continue guiding the user to perform some other voice input.
[0048] In some embodiments of the present application, optionally, FIG2 shows a second flow chart of a voice interaction method provided according to an embodiment of the present application. As shown in FIG2 , the voice interaction method further includes:
[0049] Step S202: When the recognition result indicates that the first voice information does not belong to the first application scenario voice information, it is determined whether the first voice information includes the first keyword.
[0050] Step S204: When the first voice information includes the first keyword, obtain output information corresponding to the first voice information.
[0051] Step S206: Based on the output information, control the electronic device to perform voice interaction with the user.
[0052] It is understandable that in the first application scenario, if the user currently has other questions, such as questions about the weather, time, etc., if the user's other questions (i.e., the recognition result indicates that the voice information does not belong to the first application scenario) are always ignored, it is not smart enough and will also reduce the user experience. In real life, when people talk to each other, if the other party does not respond, they usually remind the other party by calling the other party's name. Through the above method, we can know that when the user interacts with the home device through voice, the home device can be recognized for specific intentions by setting keywords, thereby improving the intelligence and user experience.
[0053] In some embodiments of the present application, after recognizing the first voice message, if the recognition result indicates that the first voice message does not belong to the first application scenario, the first voice message is further determined to determine whether it includes the first keyword. If it does, output information corresponding to the first voice message is obtained, and based on the output information, the electronic device is controlled to perform voice interaction with the user. In some embodiments of the present application, identifying specific intents through the first keyword and selectively performing voice interaction can effectively enhance the user experience. Furthermore, the voice interaction process implements intelligent anti-interference, achieving a truly "immersive" experience. Responses to specific phrases are even more intelligent. In some embodiments of the present application, optionally, taking an immersive cooking scenario as an example, if a user asks other irrelevant questions during the cooking process, such as time, weather, laundry, and care issues, the cooking domain needs to filter out these questions and not respond. However, if the user insists on knowing the answers to such questions, or the cooking device's central control screen continues to ignore the user's intent, as shown in Figure 10, this will appear less intelligent and reduce the user experience. Therefore, if the user asks an unrelated question during immersive cooking, the user can add the first keyword in the question. After recognizing the first keyword, the cooking device obtains the output information corresponding to the user's voice information, and based on the output information, controls the electronic device to interact with the user through voice, and then continues to complete the subsequent question-and-answer mechanism, ensuring an immersive cooking experience while making the experience smarter.
[0054] In some embodiments of the present application, optionally, for example, the first keyword is X, and the user can say "What time is it now, X?", "X, is it raining outside?", "Can the washing machine wash down jackets? X, I am asking about our washing machine." As long as "X" is recognized in the text expressed by the user, the user's voice will be processed and replied to improve the user experience.
[0055] In some embodiments of the present application, optionally, when the first voice information includes the first keyword, obtaining output information corresponding to the first voice information includes:
[0056] In a case where the first voice information includes the first keyword, a target application scenario corresponding to the first voice information is obtained.
[0057] Based on the target application scenario, output information corresponding to the first voice information is obtained.
[0058] In some embodiments of the present application, optionally, when the first voice information includes the first keyword, the target application scenario corresponding to the first voice information can be determined first, and then the output information corresponding to the first voice information can be obtained according to the target application scenario.
[0059] Furthermore, according to the target application scenario, obtaining the output information corresponding to the first voice information may include: after determining the target application scenario corresponding to the first voice information, obtaining the second largest language model corresponding to the target application scenario, converting the first voice information into text, inputting the text into the second largest language model corresponding to the target application scenario, and then obtaining the output information corresponding to the first voice information.
[0060] Furthermore, it is also possible to first determine whether there is a primary language model corresponding to the application scenario. If so, the output information is obtained; if not, a fallback response is provided. For example, different vertical fields or application scenarios can correspond to different secondary language models. For example, in a cooking scenario, cooking-related questions and answers can be output using the large language model corresponding to the cooking scenario. If the speech is related to laundry but contains the first keyword, the text corresponding to the speech can be input into the large language model corresponding to the laundry scenario.
[0061] In some embodiments of the present application, optionally, when the first voice information includes keywords, the implementation method of obtaining output information corresponding to the first voice information includes converting the first voice information into text and inputting the text into the second largest language model.
[0062] Furthermore, when the recognition result indicates that the first voice information belongs to the first application scenario voice information, the first voice information is converted into text, the text is recognized, the text is sent to the first language model, and the output information corresponding to the first voice information is obtained. When the recognition result indicates that the first voice information does not belong to the first application scenario voice information, and the first voice information includes the first keyword, the method for obtaining the output information corresponding to the first voice information includes converting the first voice information into text, inputting the text into the second language model, and obtaining the output information corresponding to the second voice information. It should be noted that the first language model and the second language model may be the same or different. In some embodiments of the present application, optionally, the voice interaction method further includes:
[0063] When the first voice information does not include the first keyword, the electronic device is controlled to make a second response.
[0064] In some embodiments of the present application, optionally, the first voice information is recognized, and when the first voice information does not include the first keyword, the electronic device is controlled to make a second response.
[0065] In some embodiments of the present application, optionally, the second response includes silence, that is, the electronic device does not perform any processing and remains silent.
[0066] In some embodiments of the present application, optionally, when the first voice message does not include the first keyword, the text content corresponding to the first voice message received from the user may be simply presented on the interface without any other processing. For example, when the electronic device has a central control screen, the central control screen may display the text corresponding to the user's first voice message, and the electronic device may remain silent without any processing.
[0067] In some embodiments of the present application, optionally, the second response may also include a specific speech response to continue guiding the user to perform some other voice input.
[0068] In some embodiments of the present application, optionally, the first keyword includes at least one of the following:
[0069] Electronic device brand, electronic device type name, and custom content.
[0070] In some embodiments of the present application, the first keyword may optionally be the brand of the electronic device, including the Chinese name, English name, etc. The first keyword may also be the name of the type of electronic device, for example, oven, washing machine, etc. The first keyword may also be set as custom content.
[0071] In some embodiments of the present application, the set first keyword is simple and easy to remember, which can enhance user experience.
[0072] In some embodiments of the present application, optionally, recognizing the first voice information includes:
[0073] converting the first voice information into text;
[0074] Recognize text.
[0075] In some embodiments of the present application, optionally, the electronic device receives a first voice message from a user, converts the first voice message into text in real time and continuously, and subsequently recognizes the text to identify whether the first voice message belongs to a first application scenario, or whether the first voice message includes a first keyword, thereby providing a basis for subsequent voice interaction between the electronic device and the user.
[0076] In some embodiments of the present application, optionally, FIG3 shows a third flow chart of a voice interaction method provided according to an embodiment of the present application. As shown in FIG3 , recognizing text includes:
[0077] Step S302: input the text into the classifier to obtain the recognition result output by the classifier.
[0078] The classifier is used to determine whether the first voice information belongs to the first application scenario voice information.
[0079] In some embodiments of the present application, the text converted from the first voice information may optionally need to be recognized, and the recognition of the text may be accomplished by a classifier. Specifically, the text is input into the classifier, which determines whether the text belongs to the first application scenario based on the text, and then obtains a recognition result of the first voice information, the recognition result including: the first voice information belongs to the first application scenario or the first voice information does not belong to the first application scenario.
[0080] In some embodiments of the present application, optionally, when the first voice information belongs to the first application scenario, output information corresponding to the first voice information is obtained, and based on the output information, the electronic device is controlled to perform voice interaction with the user. When the first voice information does not belong to the first application scenario, in the first case: the electronic device is controlled to perform a first response. In the second case: whether the first voice information includes the first keyword is further determined. If the first voice information includes the first keyword, output information corresponding to the first voice information is obtained, and based on the output information, the electronic device is controlled to perform voice interaction with the user.
[0081] In some embodiments of the present application, whether the first voice information belongs to the first application scenario can be determined by a classifier, the judgment process is accurate, and the application scenarios are wide.
[0082] In some embodiments of the present application, optionally, when the recognition result indicates that the first voice information belongs to the first application scenario voice information, obtaining output information corresponding to the first voice information includes:
[0083] When the recognition result indicates that the first voice information belongs to the first application scenario voice information, the text is sent to the first large language model to obtain output information corresponding to the first voice information.
[0084] In some embodiments of the present application, optionally, the first voice information is recognized, and when the recognition result belongs to the first application scenario voice information, the text is sent to the first large language model, and the first large language model obtains the output information corresponding to the first voice information according to the text content, and then controls the electronic device to perform voice interaction with the user. It is understandable that a large language model (LLM), also known as a large language model, is a model based on machine learning and natural language processing technology, which learns to serve the ability of human language understanding and generation by training a large amount of text data. The core idea of LLM is to learn the patterns and language structures of natural language through large-scale unsupervised training, which can simulate the language cognition and generation process of human beings to a certain extent.
[0085] It is understandable that in addition to using the first language model, other technologies in the field of natural language processing (Neuro-Linguistic Programming, NLP) can also be adopted, including question-answering systems, etc.
[0086] In some embodiments of the present application, optionally, FIG4 shows a fourth flow chart of the voice interaction method provided according to an embodiment of the present application. As shown in FIG4 , before controlling the electronic device to enter the full-duplex mode, the method further includes:
[0087] Step S402: Acquire a first corpus applied to a first application scenario.
[0088] Step S404: Acquire a second corpus that is applied to a scenario other than the first application scenario.
[0089] Step S406: establishing a classifier.
[0090] Step S408: training a classifier using the first corpus and the second corpus to obtain a trained classifier.
[0091] In some embodiments of the present application, optionally, based on a first application scenario, a first corpus belonging to the first application scenario is prepared, and a second corpus that does not belong to the first application scenario is prepared. After a classifier is established, the classifier is trained using the first corpus and the second corpus. The input of the classifier is the first corpus, and the output of the classifier is 1. The input of the classifier is the second corpus, and the output of the classifier is 0, where 1 indicates that the input corpus belongs to the first application scenario, and 0 indicates that the input corpus does not belong to the first application scenario.
[0092] In some embodiments of the present application, optionally, taking the immersive cooking scenario as an example, some recipe-related question data can be used as the first corpus, and question data unrelated to cooking can be used as the second corpus. Then, the first corpus and the second corpus are used to train a classifier. The input of the classifier is a sentence (i.e., text converted from speech in full-duplex). When the input text belongs to the immersive cooking scenario (i.e., the first corpus), the output of the classifier is 1. When the input text does not belong to the immersive cooking scenario (i.e., the second corpus), the output of the classifier is 0, where 1 indicates related to cooking and 0 indicates unrelated to cooking.
[0093] For example, building a classifier can be divided into the following steps: First, prepare cooking-related corpus, the first corpus. Examples include "How to make chiffon cake?" and "Can corn oil be added to chiffon cake?" Second, prepare corresponding non-cooking corpus, the second corpus. Examples include "What's the weather like today?" and "Is the TV series X good?" Third, use the prepared corpus to train the classifier.
[0094] In some embodiments of the present application, optionally, classifiers are trained for different application scenarios. First, rules and boundaries need to be defined, and then the classifier learns the above rules and boundaries. For example, in an immersive cooking scenario, sentences related to cooking need to be identified. If in a cooking scenario, the user asks about the weather, the question is considered unrelated to cooking. In a laundry scenario, if a user interacts with a washing machine through voice, and the user asks about the weather, the user may be thinking about how to dry clothes. In this case, the question is considered to be related to laundry, that is, in different application scenarios, different application scenarios have different rules and boundaries. In some embodiments of the present application, optionally, the classifier includes a neural network model.
[0095] In some embodiments of the present application, optionally, the classifier can be a neural network model or a combination of multiple neural network models. In fact, it can be a combination of rules, keywords and neural network models.
[0096] In some embodiments of the present application, optionally, the neural network model can adopt a base model and fine-tune the training on this basis. The base model can be a general neural network model, including BERT (Bidirectional Encoder Representations from Transformers, language representation model), ERNIE (Wenxin model), FastText (fast text classifier), etc.
[0097] In some embodiments of the present application, optionally, FIG8 shows a fifth flow chart of a voice interaction method provided according to an embodiment of the present application. As shown in FIG8 , the voice interaction method, used in a cooking device, includes:
[0098] Step S502: Acquire user voice.
[0099] Among them, the cooking equipment is in full-duplex mode.
[0100] Step S504: Convert the user's voice into text.
[0101] Step S506: determine whether it is related to cooking.
[0102] If yes, go to step S508; otherwise, go to step S510.
[0103] Step S508: Use the first language model to process and reply to the text.
[0104] Step S510: The cooking device is silent.
[0105] In some embodiments of the present application, through the above method, in the full-duplex case, when the user voice is related to the current cooking scene, responding to the user voice can improve the quality of the interaction between the cooking device and the user voice and enhance the user experience.
[0106] In some embodiments of the present application, optionally, FIG9 shows a sixth flow chart of a voice interaction method provided according to an embodiment of the present application. As shown in FIG9 , the voice interaction method, used in a cooking device, includes:
[0107] Step S602: Acquire user voice.
[0108] Among them, the cooking equipment is in full-duplex mode.
[0109] Step S604: Convert the user's voice into text.
[0110] Step S606: determine whether it is related to cooking.
[0111] If yes, go to step S608; otherwise, go to step S610.
[0112] Step S608: Use the first language model to process and reply to the text.
[0113] Step S610: determine whether the current text includes keywords.
[0114] If yes, go to step S608, otherwise go to step S612.
[0115] Step S612: The cooking device is silent.
[0116] In some embodiments of this application, specific intent is identified through keywords, allowing for selective voice interaction, effectively enhancing the user experience. Furthermore, the voice interaction process is intelligently anti-interference, achieving a truly immersive experience. Responses to specific phrases are even more intelligent.
[0117] According to the voice interaction method provided in the embodiment of the present application, the execution subject may be a voice interaction device. According to the embodiment of the present application, the voice interaction device performing the voice interaction method is taken as an example to illustrate the voice interaction device provided in the embodiment of the present application.
[0118] According to an embodiment of the present application, a voice interaction device is provided. Optionally, Figure 5 shows a structural block diagram of the voice interaction device provided according to an embodiment of the present application. As shown in Figure 5, the voice interaction device 100 includes a first control module 110, a receiving module 120, an identification module 130, an acquisition module 140 and a second control module 150. The first control module 110 is used to control the electronic device to enter full-duplex mode. The receiving module 120 is used to receive a first voice message from a user. The identification module 130 is used to identify the first voice message. The acquisition module 140 is used to obtain output information corresponding to the first voice message when the recognition result indicates that the first voice message belongs to the first application scenario voice message. The second control module 150 is used to control the electronic device to perform voice interaction with the user based on the output information.
[0119] In some embodiments of the present application, optionally, for an electronic device, it is controlled to enter full-duplex mode and perform voice interaction with the user. The electronic device receives a first voice message from the user, then recognizes the first voice message to obtain a recognition result, and when the recognition result indicates that the first voice message belongs to the first application scenario voice message, obtains output information corresponding to the first voice message. Based on the output information, the electronic device is controlled to perform voice interaction with the user. In the above manner, in the case of full-duplex, when the user voice is related to the current first application scenario, responding to the user voice can improve the quality of the voice interaction between the electronic device and the user and enhance the user experience.
[0120] In some embodiments of the present application, the voice interaction device 100 can optionally be applied to immersive cooking, laundry, domain question and answer, control field, home appliance encyclopedia field, etc. Specifically, the first application scenario can include immersive cooking, laundry, domain question and answer, control, home appliance encyclopedia, etc.
[0121] The voice interaction device 100 provided according to the embodiment of the present application can implement each process of the above-mentioned voice interaction method embodiment and achieve the same technical effect. To avoid repetition, it will not be described here.
[0122] According to the embodiments of the present application, the voice interaction device can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or a device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a car-mounted electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., and is not specifically limited according to the embodiments of the present application.
[0123] According to the embodiments of the present application, the voice interaction device can be a device having an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited according to the embodiments of the present application.
[0124] The voice interaction device provided according to the embodiment of the present application can implement each process implemented in the above-mentioned voice interaction method embodiment. To avoid repetition, they will not be described here.
[0125] Optionally, as shown in Figure 6, an electronic device 1000 is also provided according to an embodiment of the present application. The electronic device 1000 includes a processor 1002 and a memory 1004. The memory 1004 stores programs or instructions that can be run on the processor 1002. When the program or instructions are executed by the processor 1002, the various steps of the above-mentioned voice interaction method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, they will not be repeated here.
[0126] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.
[0127] FIG7 is a schematic diagram of the hardware structure of an electronic device implementing an embodiment of the present application.
[0128] The electronic device 1100 includes but is not limited to: a radio frequency unit 1101, a network module 1102, an audio output unit 1103, an input unit 1104, a sensor 1105, a display unit 1106, a user input unit 1107, an interface unit 1108, a memory 1109 and a processor 1110.
[0129] Those skilled in the art will appreciate that the electronic device 1100 may further include a power source (e.g., a battery) to power various components. The power source may be logically connected to the processor 1110 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The electronic device structure shown in FIG7 does not limit the electronic device. The electronic device may include more or fewer components than shown, or may combine certain components, or have different component arrangements, which will not be described in detail here.
[0130] The processor 1110 is used to control the electronic device to enter the full-duplex mode.
[0131] The processor 1110 is configured to receive first voice information from a user.
[0132] The processor 1110 is configured to recognize the first voice information.
[0133] The processor 1110 is configured to obtain output information corresponding to the first voice information when the recognition result indicates that the first voice information belongs to the first application scenario voice information.
[0134] The processor 1110 is configured to control the electronic device to perform voice interaction with the user based on the output information.
[0135] The processor 1110 provided according to the embodiment of the present application can implement each process of the above-mentioned voice interaction method embodiment and achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0136] It should be understood that, according to an embodiment of the present application, the input unit 1104 may include a graphics processing unit (GPU) 11041 and a microphone 11042, and the graphics processor 11041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 1106 may include a display panel 11061, and the display panel 11061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1107 includes a touch panel 11071 and at least one of other input devices 11072. The touch panel 11071 is also called a touch screen. The touch panel 11071 may include two parts: a touch detection device and a touch controller. Other input devices 11072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an operating stick, which will not be repeated here.
[0137] The memory 1109 can be used to store software programs and various data. The memory 1109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1109 may include a volatile memory or a non-volatile memory, or the memory 1109 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct memory bus random access memory (DRRAM). The memory 1109 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memory.
[0138] Processor 1110 may include one or more processing units. Optionally, processor 1110 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 1110.
[0139] According to an embodiment of the present application, an electronic device is provided, having a voice interaction function, and optionally including:
[0140] Use the voice interaction method of any embodiment of the present application to perform voice interaction with the user.
[0141] In some embodiments of the present application, optionally, the electronic device includes one of the following:
[0142] Air conditioners, washing machines, refrigerators, dishwashers, water purifiers, range hoods, gas stoves, water heaters, cooking equipment, mobile devices, and central control screens.
[0143] According to an embodiment of the present application, a readable storage medium is also provided, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned voice interaction method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0144] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0145] According to an embodiment of the present application, a chip is further provided, which includes a processor and a communication interface. The communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned voice interaction method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0146] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0147] According to an embodiment of the present application, a computer program product is provided, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned voice interaction method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0148] The methods may be implemented in a variety of different ways depending on the specific features and / or example applications. For example, the methods may be implemented through a combination of hardware, firmware, and / or software. For example, in a hardware implementation, the processor may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, electronic devices, other device units for performing the above functions, and / or combinations thereof.
[0149] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium can be, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: a portable computer floppy disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory card, a floppy disk, an encoding mechanical device (such as a punched card or a groove with a raised structure on which instructions are recorded), and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be understood as a transmission signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium, or electrical signals transmitted through wires.
[0150] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0151] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
[0152] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A voice interaction method, wherein, For an electronic device with voice interaction function, including: Controlling the electronic device to enter the full-duplex mode; Receiving the first voice message from the user; Identifying the first voice message; When the recognition result indicates that the first voice message belongs to the first application scenario voice message, obtaining the output information corresponding to the first voice message; Based on the output information, controlling the electronic device to perform voice interaction with the user.
2. The voice interaction method according to claim 1, wherein, Also includes: When the recognition result indicates that the first voice message does not belong to the first application scenario voice message, controlling the electronic device to make a first response.
3. The voice interaction method according to claim 1, wherein, Also includes: When the recognition result indicates that the first voice message does not belong to the first application scenario voice message, determining whether the first voice message includes a first keyword; When the first voice message includes the first keyword, obtaining the output information corresponding to the first voice message; Based on the output information, controlling the electronic device to perform voice interaction with the user.
4. The voice interaction method according to claim 3, wherein, Also includes: When the first voice message does not include the first keyword, controlling the electronic device to make a second response.
5. The voice interaction method according to claim 3, wherein, The first keyword includes at least one of the following: Electronic device brand, electronic device type name, custom content.
6. The voice interaction method according to claim 3, wherein, When the first voice message includes the first keyword, obtaining the output information corresponding to the first voice message, including: When the first voice message includes the first keyword, obtaining the target application scenario corresponding to the first voice message; Based on the target application scenario, obtaining the output information corresponding to the first voice message.
7. The voice interaction method according to any one of claims 1 to 6, wherein, The identifying the first voice message includes: Converting the first voice message into text; Identifying the text.
8. The voice interaction method according to claim 7, wherein, The identifying the text includes: Inputting the text into a classifier to obtain the recognition result output by the classifier, where the classifier is used to determine whether the first voice message belongs to the first application scenario voice message.
9. The voice interaction method according to claim 8, wherein, When the recognition result indicates that the first voice message belongs to the first application scenario voice message, obtaining the output information corresponding to the first voice message, including: When the recognition result indicates that the first voice message belongs to the first application scenario voice message, sending the text to a first large language model to obtain the output information corresponding to the first voice message.
10. The voice interaction method according to claim 8 or 9, wherein, Before controlling the electronic device to enter the full-duplex mode, it also includes: Obtaining the first corpus applied to the first application scenario; Obtaining the second corpus applied outside the first application scenario; Establishing a classifier; Training the classifier with the first corpus and the second corpus to obtain the trained classifier.
11. The voice interaction method according to claim 10, wherein, The classifier includes a neural network model.
12. A voice interaction device, wherein, Includes: A first control module for controlling the electronic device to enter the full-duplex mode; A receiving module for receiving the first voice message from the user; An identifying module for identifying the first voice message; An acquisition module, configured to acquire output information corresponding to the first voice message when the recognition result indicates that the first voice message belongs to the first application scenario voice message; A second control module, configured to control the electronic device to perform voice interaction with the user based on the output information.
13. An electronic device, wherein, Having a voice interaction function, including: Performing voice interaction with the user by using the voice interaction method according to any one of claims 1 to 11.
14. An electronic device, wherein, Including: A memory, on which a program or instruction is stored; A processor, configured to implement the steps of the voice interaction method according to any one of claims 1 to 11 when executing the program or instruction.
15. A readable storage medium, on which a program or instructions are stored, wherein, The steps of the voice interaction method according to any one of claims 1 to 11 are implemented when the program or instruction is executed by the processor.
16. A computer program product, wherein, Including a computer program, wherein the steps of the voice interaction method according to any one of claims 1 to 11 are implemented when the computer program is executed by the processor.
Citation Information
Patent Citations
Speech control method, device and equipment and storage medium
CN109524010A
Voice control method and device, electrical equipment, storage medium and processor
CN112002315A
Voice control method and device of electronic equipment, computer equipment and storage medium
CN112017651A
Intelligent control method and device based on voice, electronic equipment and storage medium
CN112201246A
Voice interaction method and device, electronic equipment and storage medium
CN115565531A