Interaction method and apparatus, electronic device, and storage medium

By collecting user voice data in real time and combining it with various other information to determine whether to respond, the problem of cumbersome wake words in robot interaction has been solved, achieving natural and coherent human-computer interaction.

CN116028006BActive Publication Date: 2026-04-14HEFEI IFLYTEK TOYCLOUD TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-29
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing robot interactions require adding a wake word each time, resulting in an unnatural and cumbersome user experience.

Method used

By collecting user voice data in real time and combining it with the user's position relative to the device, voice interaction association information, and device interaction mode records, the system determines whether to respond to the user's voice, avoiding false responses to non-interactive voice messages and enabling multiple interactions with a single wake-up call.

Benefits of technology

It improves the convenience of human-computer interaction, making the interaction smoother and more natural, and reduces the reliance on wake words.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116028006B_ABST
    Figure CN116028006B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of human-computer interaction, and provides an interaction method, device, electronic equipment and storage medium, wherein the method comprises the following steps: after entering a current interaction mode, collecting user voice in real time; determining whether to respond to the user voice based on at least one of the following: the position information of the user relative to a local device corresponding to the user voice, the interaction association information of the user voice, and the interaction mode record of the local device; and in the case of determining to respond to the user voice, interacting based on the user voice. The method, device, electronic equipment and storage medium provided by the application collect user voice in real time in the current interaction mode, and provide conditions for one-time wake-up multi-time interaction. In the scene of realizing one-time wake-up multi-time interaction, the problem of false response of the local device to non-interaction voice is avoided, the human-computer interaction can be as smooth and natural as communication with people, and the convenience of human-computer interaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology, and in particular to an interaction method, device, electronic device, and storage medium. Background Technology

[0002] In existing robots, users need to add a wake word for each interaction before the robot can recognize that it is interacting with them. For example, when using the Alpha Egg robot, users need to add a wake word containing "Hello Egg" before each communication before proceeding with the subsequent interaction.

[0003] However, for users, they cannot interact with machines as smoothly and naturally as they would with people. The cumbersome operation of human-computer interaction reduces the user experience. Summary of the Invention

[0004] This invention provides an interaction method, device, electronic device, and storage medium to solve the problem that users need to add a wake word for each interaction to wake up the machine and perform the interaction, which is cumbersome.

[0005] This invention provides an interaction method, comprising:

[0006] After entering the current interaction mode, the user's voice is collected in real time;

[0007] Based on at least one of the following: the user's pose information relative to the local device corresponding to the user's voice, the interaction association information of the user's voice, and the interaction mode record of the local device, it is determined whether to respond to the user's voice.

[0008] If a response to the user's voice is determined, interaction is performed based on the user's voice.

[0009] According to an interaction method provided by the present invention, the step of determining the interaction-related information includes:

[0010] The user's speech is transcribed to obtain the transcribed text;

[0011] The interaction association information is determined based on the first correlation between the transcribed text and the wake word of the local device, and / or the second correlation between the transcribed text and the historical conversation text in the current interaction mode.

[0012] According to an interaction method provided by the present invention, the step of determining the first correlation between the transcribed text and the wake word of the local device includes:

[0013] Extract keywords from the transcribed text;

[0014] The first correlation is determined based on the correlation between the keyword and each wake-up word in the wake-up word.

[0015] According to an interaction method provided by the present invention, the step of determining the second correlation between the transcribed text and the historical conversation text in the current interaction mode includes:

[0016] Based on the historical conversation text, conversation prediction is performed to obtain the preset conversation text;

[0017] The second correlation is determined based on the correlation between the preset conversation text and the transcribed text.

[0018] According to an interaction method provided by the present invention, the step of determining the pose information includes:

[0019] Distance detection and / or orientation detection are performed on the user corresponding to the user's voice.

[0020] The pose information is determined based on the distance between the user corresponding to the user's voice and the local device, and / or the orientation of the user corresponding to the user's voice relative to the local device.

[0021] According to an interaction method provided by the present invention, the step of determining the interaction mode record includes:

[0022] The interaction mode record is determined based on the number of times the local device executes the interaction mode after the current power-on and the interaction sequence number of the current interaction mode after the current power-on.

[0023] According to an interaction method provided by the present invention, before the real-time acquisition of user voice, the method further includes:

[0024] Upon detecting a wake-up voice message carrying the wake-up word of the local device, the system enters the current interaction mode.

[0025] After entering the current interaction mode, the following is also included:

[0026] If no response is received from the user within the preset time period, the current interaction mode will be exited.

[0027] The present invention also provides an interactive device, comprising:

[0028] The acquisition unit is used to acquire user voice in real time after entering the current interaction mode;

[0029] The analysis unit is used to determine whether to respond to the user's voice based on at least one of the user's pose information relative to the local device, the interaction association information of the user's voice, and the interaction mode record of the local device.

[0030] An interaction unit is used to interact based on the user's voice when it is determined that a response to the user's voice is required.

[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the interaction method as described above.

[0032] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the interaction method as described above.

[0033] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the interaction method as described above.

[0034] This invention provides an interaction method, device, electronic device, and storage medium that, by real-time acquisition of user voice in the current interaction mode, provides conditions for multiple interactions with a single wake-up. Based on this, it determines whether to respond to the user voice by considering at least one of the following: the user's pose information relative to the local device, the interaction association information of the user's voice, and the interaction mode records of the local device. This distinguishes whether the user voice acquired in the interaction mode is interactive voice, thereby realizing human-computer interaction. In scenarios where multiple interactions are achieved with a single wake-up, it avoids the problem of the local device mistakenly responding to non-interactive voice, making human-computer interaction as smooth and natural as human communication, and improving the convenience of human-computer interaction. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0036] Figure 1 This is one of the flowcharts illustrating the interaction method provided by the present invention;

[0037] Figure 2 This is a schematic diagram of the process for determining the first correlation provided by the present invention;

[0038] Figure 3This is a schematic diagram of the process for determining the second correlation provided by the present invention;

[0039] Figure 4 This is a flowchart illustrating the process of determining pose information provided by the present invention;

[0040] Figure 5 This is the second flowchart illustrating the interaction method provided by the present invention;

[0041] Figure 6 This is a schematic diagram of the structure of the interactive device provided by the present invention;

[0042] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0044] In existing human-computer interaction processes, users need to add a wake word for each interaction in order to wake up the machine to interact. As a result, users cannot interact with the machine as smoothly and naturally as they would with a human, which reduces the user experience.

[0045] To address the aforementioned problems, this invention provides an interaction method to achieve activation during the initial interaction.

[0046] After waking up the machine, under certain conditions, subsequent interactions do not require the wake word, making it as convenient and natural as human-to-human interaction. Figure 1 This is one of the flowcharts illustrating the interaction method provided by the present invention, such as... Figure 1 As shown, this method can be applied to various intelligent terminals, including interactive robots, referred to here as the local device. The method includes:

[0047] Step 110: After entering the current interaction mode, collect the user's voice in real time;

[0048] Here, the current interaction mode refers to the interaction mode entered after the user wakes up the device using a wake word. It can be understood that one interaction mode corresponds to a complete interaction.

[0049] The conversation flow is as follows: After the local device is woken up and completes the user's verbal command, it remains in the current interaction mode and continues to collect the user's voice in real time. This prevents the local device from automatically exiting before the interaction is finished.

[0050] In interactive mode, users need to wake up the device again using a wake word. From 5 onwards, after one wake-up in interactive mode, multiple interactions can be performed, providing conditional support.

[0051] Step 120: Based on at least one of the following: the pose information of the user relative to the local device corresponding to the user's voice, the interaction association information of the user's voice, and the interaction mode record of the local device, determine whether to respond to the user's voice.

[0052] Specifically, the user voice collected in the current interaction mode may be the voice spoken by user 0 to continue interacting with this device, or it may be the voice spoken by the user in conjunction with other users.

[0053] Voice interaction between users. Therefore, after collecting user voice data, it is necessary to determine whether the user's voice is the voice used for interaction with the local device in the current interaction mode. That is, to determine whether the local device needs to respond to the user's voice.

[0054] To determine whether the user's voice is the voice used for interaction with the local device in the current interaction mode, the user's position relative to the local device can be considered.

[0055] The method includes at least one of the following: interactive information related to user voice, and interaction mode records of the local device. Specifically, when using two or more methods for judgment, the judgment results obtained from each method can be weighted, or the judgment results obtained from each method can be used for voting. This embodiment of the invention does not specifically limit this.

[0056] Here, the user corresponding to the user's voice can be the speaker obtained by voiceprint recognition of the user's voice.

[0057] User pose information relative to the local device refers to the user's position and posture relative to the local device. Specifically, pose information can be the distance between the user and the local device, or whether the user is facing the local device. It's understandable that, comparing a user being closer to the local device versus a user being farther away, the user's voice captured when the user is closer is more likely to be spoken speech used for human-computer interaction; similarly, comparing a user facing the local device versus a user with their back to the local device, the user's voice captured when facing the local device is more likely to be spoken speech used for human-computer interaction. Therefore, user pose information relative to the local device can be used as a basis for determining whether the local device needs to respond to the user's voice.

[0058] Interactive association information of user voice refers to information reflecting the relationship between the content of user voice and human-computer interaction. For example, interactive association information could be the association between the content of the transcribed text of user voice and the wake word, or the association between the transcribed text of user voice and the historical conversation text. It's understandable that when the collected user voice contains a wake word or information related to a wake word, it's similar to addressing someone by their name or nickname in a conversation. This allows us to infer whether the collected user voice is spoken by the user for human-computer interaction; the transcribed text of the collected user voice has a semantic or logical connection with the historical conversation text, similar to the contextual connection in human-computer communication. This allows us to infer whether the collected user voice is spoken by the user for human-computer interaction.

[0059] The interaction mode record of the local device refers to the historical record of the local device during human-computer interaction. Specifically, this interaction mode record can include the number of times the user enters an interaction mode, the specific number of interactions in the current interaction, or information such as the user's previous interaction patterns. It's understandable that the initial human-computer interaction generally requires multiple rounds of interaction to complete. Therefore, the number of times the user enters an interaction mode and the specific number of interactions in the current interaction determine whether the current interaction is the first interaction. Compared to non-first interactions, when the current interaction is the first interaction, the user's voice recorded is more likely to be dictated for human-computer interaction. When the local device records information such as the user's previous interaction patterns, such as the user's commonly used questioning tone and words, it reflects the user's general or personalized habits in human-computer interaction. For example, personalized habit information could be that the user's voice information always includes the phrase "I have a question" when interacting with the local device. It is understandable that when the collected user voice contains personalized habit information about the user's questions to the local device, it can be inferred that the collected user voice is more likely to be the user's spoken voice used for human-computer interaction.

[0060] Step 130: If it is determined that the user's voice is being responded to, then interaction is performed based on the user's voice.

[0061] Specifically, after determining whether to respond to the user's voice, the device can perform subsequent operations based on the result. If it determines to respond to the user's voice, the device responds via voice, video, or a combination of both, or it can determine the command corresponding to the user's voice and execute that command. For example, in the current interaction mode, if the user is one meter away from the device and facing it, asking "What's the weather like today?", the device, after determining that it needs to respond to the user's voice, replies "Today's weather is cloudy." Or, for example, if the user faces the device and issues the command "Egg, please turn it off immediately," the device, after determining that it needs to respond to the user's voice, performs a shutdown operation.

[0062] The method provided in this invention collects user voice in real time during the current interaction mode, providing conditions for multiple interactions with a single wake-up. Based on this, it determines whether to respond to the user voice by considering at least one of the following: the user's pose information relative to the local device, the interaction association information of the user's voice, and the interaction mode records of the local device. This distinguishes whether the user voice collected in the interaction mode is interactive voice, thereby realizing human-computer interaction. In scenarios where multiple interactions are achieved with a single wake-up, the problem of the local device mistakenly responding to non-interactive voice is avoided, making human-computer interaction as smooth and natural as human communication, thus improving the convenience of human-computer interaction.

[0063] Based on the above embodiments, step 120, the step of determining the interactive association information, includes:

[0064] The user's speech is transcribed to obtain the transcribed text;

[0065] The interaction association information is determined based on the first correlation between the transcribed text and the wake word of the local device, and / or the second correlation between the transcribed text and the historical conversation text in the current interaction mode.

[0066] Specifically, the wake word is a pre-defined word segment used to achieve voice wake-up. For example, the wake word could be "Dandan Nihao" (Hello Egg). After the local device recognizes the wake word, it is awakened and enters interactive mode. Generally, different local devices can correspond to the same or different wake words.

[0067] Considering that in human-computer interaction, user voice may carry segments similar to or related to the wake word, serving as an abbreviation or nickname for the local device to refer to it during the interaction, the need for a response can be determined based on the correlation between the transcribed text of the user's voice and the wake word. Here, the correlation between the transcribed text and the wake word is denoted as the first correlation. By determining the first correlation between the transcribed text and the wake word, interaction-related information can be identified, allowing the local device to recognize the user's use of the wake word's nickname or related segments during human-computer interaction, thus making the local device's interaction methods more flexible.

[0068] Furthermore, considering the semantic and / or logical connections between conversations in actual interactions, the correlation between the transcribed user speech and the historical conversation text in the current interaction mode can be determined based on their semantic and / or logical relevance. Here, the correlation between the transcribed text and the historical conversation text is denoted as the second correlation. The historical conversation text in the current interaction mode refers to the conversation text corresponding to the user's speech and the local device's speech during the human-computer interaction after entering this interaction mode. By leveraging the semantic and / or logical relevance between the transcribed user speech and the historical conversation text, the local device can enable multiple interactions with a single wake-up, making communication between the user and the local device smoother.

[0069] After obtaining the first correlation and the second correlation between the transcribed text of the user's speech and the historical conversation text, the interaction association information can be determined. Here, the interaction association information can be the first correlation and / or the second correlation themselves, or it can be the total correlation obtained by fusing the first correlation and the second correlation. This embodiment of the invention does not specifically limit this.

[0070] Based on any of the above embodiments Figure 2 This is a schematic diagram of the process for determining the first correlation provided by the present invention, as shown below. Figure 2 As shown, the step of determining the first correlation between the transcribed text and the wake word of the local device includes:

[0071] Step 210: Extract keywords from the transcribed text;

[0072] Step 220: Determine the first correlation based on the correlation between the keyword and each wake-up segment in the wake-up word.

[0073] Considering that the transcribed text may be quite long, performing correlation analysis between each word and the wake-up word would require significant computational resources. In this embodiment of the invention, keywords are extracted from the transcribed text before performing correlation analysis, and then correlation analysis is performed between these keywords and the wake-up word. This reduces the computational load on the local device's internal programs, enabling the device to respond quickly.

[0074] Specifically, when conducting relevance analysis, the relevance between keywords and the wake-up words in the wake-up phrase can be calculated. Here, wake-up words refer to the pre-defined words within the wake-up phrase. For example, if the wake-up phrase is "Egg Hello," then the wake-up words could be "Egg" and "Hello." The relevance between keywords and the wake-up words in the wake-up phrase can be determined based on the similarity between the semantic codes of the keywords and the semantic codes of the wake-up words. This similarity can be calculated using the cosine similarity algorithm. The higher the similarity, the greater the relevance between the keywords and the wake-up words, and the stronger the primary relevance. Alternatively, it can be determined by whether there are any identical words among the wake-up words of the keywords and the wake-up phrase; the more identical words there are, the stronger the primary relevance.

[0075] Based on any of the above embodiments Figure 3 This is a schematic diagram of the process for determining the second correlation provided by the present invention, such as... Figure 3 As shown, the step of determining the second correlation between the transcribed text and the historical conversation text in the current interaction mode includes:

[0076] Step 310: Based on the historical conversation text, perform conversation prediction to obtain the preset 5 conversation texts;

[0077] Specifically, the preset conversation text here refers to the text obtained by inferring what the user might communicate next based on the semantics and / or logic of the historical conversation text. For example, if the historical conversation text is "How's the weather today?", the preset conversation text could be "Today..."

[0078] Questions such as "What's the temperature today?", "What should I wear today?", and "What will the weather be like tomorrow?" are used. Here, conversation prediction can be achieved through a pre-trained conversation prediction model. This model can be trained on top of a language model. The conversation prediction model can be obtained through supervised learning, using the text of the first few sentences in a pre-collected conversation as training samples and the text of responses to the first few sentences in the conversation as training labels.

[0079] Step 320: Based on the correlation between the preset conversation text and the transcribed text,

[0080] Determine the second correlation.

[0081] Specifically, the correlation between the preset conversation text and the transcribed text is defined. This correlation can be determined by the similarity between the semantic codes of the sentences. The similarity can be determined by...

[0082] The cosine similarity algorithm calculates that the higher the similarity, the greater the correlation between them, i.e., the greater the correlation between them (0).

[0083] Based on any of the above embodiments Figure 4 This is a schematic diagram of the process for determining pose information provided by the present invention. For example... Figure 4 As shown, the steps for determining the pose information include:

[0084] Step 410: Perform distance detection and / or orientation detection on the user corresponding to the user's voice.

[0085] 5. Here, the distance detection refers to detecting the distance between the local device and the user corresponding to the user's voice.

[0086] Distance detection can be achieved through ultrasonic positioning, such as receiving reflected ultrasonic waves emitted by the device itself and calculating the distance between the device and the user based on the time difference between emission and reception of the ultrasonic waves and the propagation speed. Distance detection can also be achieved through visual positioning, such as capturing images of the surrounding environment of the device and identifying the user from the captured images, thereby calculating the distance between the device and the user based on the user's size in the image. Distance detection can also be achieved through sound source localization, that is, determining the distance between the device and the user by locating the sound source of the user's voice. No specific limitations are placed on the technology used for distance detection here.

[0087] In addition, orientation detection can be achieved through face detection. For example, image acquisition can be performed, and the user's face can be detected in the acquired image to determine whether the user is facing the device. Alternatively, it can be achieved through eye tracking, which tracks the user's gaze direction through the acquired image to determine whether the user is facing the device.

[0088] Step 420: Determine the pose information based on the distance between the user corresponding to the detected user voice and the local device, and / or the orientation of the user corresponding to the user voice relative to the local device.

[0089] After obtaining the distance between the user and the local device, and / or the user's orientation relative to the local device corresponding to the user's voice, the pose information can be determined. Specifically, the pose information can be determined using pre-set rules. For example, distance or orientation can be directly used as pose information, or a comprehensive judgment can be made based on distance and orientation to obtain the pose information. Based on the pose information, the current user interaction association information can be determined, allowing the local device to consider a wider range of interaction association information and providing a more natural interaction.

[0090] Based on any of the above embodiments, the step of determining the interaction mode record includes:

[0091] The interaction mode record is determined based on the number of times the local device executes the interaction mode after the current power-on and the interaction sequence number of the current interaction mode after the current power-on.

[0092] Specifically, considering common interaction habits, the first human-computer interaction after the device is powered on usually requires multiple rounds of interaction to complete. Therefore, an interaction mode record can be generated from the number of times the interaction mode is entered.

[0093] The number of times the user enters the interaction mode is accumulated after each power-on, and the situation after each power-on is independent of each other. The interaction sequence number of the current interaction mode after the current power-on reflects which interaction mode has been entered since this power-on. Based on these two factors, it can be determined whether the current interaction is the first human-computer interaction, and whether it is the first time is recorded in the interaction mode record. This is used to determine whether to respond to the user's voice. It is understandable that multiple rounds of interaction are generally required to complete the first human-computer interaction. Therefore, when the interaction mode record determines that it is the first time entering the human-computer interaction, it can be inferred that the collected user voice is more likely to be the user's spoken voice used for human-computer interaction.

[0094] Based on any of the above embodiments, before the real-time acquisition of user voice, the method further includes:

[0095] Upon detecting a wake-up voice message carrying the wake-up word of the local device, the system enters the current interaction mode.

[0096] After entering the current interaction mode, the method further includes:

[0097] If no response is received from the user within the preset time period, the current interaction mode will be exited.

[0098] Specifically, considering that users may interrupt the interaction in actual human-computer interaction, the device can exit the current interaction mode if it fails to collect any user voice responses within a preset time after entering the interaction mode. Here, the preset time refers to the duration of a single interaction. That is, the time from the completion of the last interaction until the device exits the current interaction mode if no user voice responses are collected. Typically, simple communication between people lasts about 10 minutes, and to efficiently utilize the device's energy, the preset time can be set to around 10-15 minutes. By exiting the current interaction mode if no user voice responses are collected within the preset time, the device's energy can be utilized more effectively, and the device's flexibility is improved.

[0099] The method provided in this embodiment of the invention sets conditions for entering and exiting the current interaction mode, thereby ensuring continuous and natural human-computer interaction as if communicating with a person, while also avoiding excessive consumption of the local device's energy.

[0100] Based on any of the above embodiments Figure 5 The second flowchart of the interaction method provided by the present invention is as follows: Figure 5 As shown, the interaction method may include the following steps:

[0101] Step 510: Enter the current interaction mode and collect user voice in real time.

[0102] The user voice collected in the current interaction mode may be spoken by the user to continue interacting with the local device, or it may be spoken by the user interacting with other people. Therefore, it is necessary to determine whether the collected user voice is the voice that the user needs a response from the local device.

[0103] Step 520: Based on at least one of the user's pose information relative to the local device, the interaction association information of the user's voice, and the interaction mode record of the local device, determine whether to respond to the user's voice. If a response is received, proceed to step 530; otherwise, proceed to step 540.

[0104] Step 530: If it is determined that the user's voice is being responded to, interaction is performed based on the user's voice.

[0105] Step 540: If it is determined that there is no response to the user's voice, further determine whether the waiting time for interaction in the current interaction mode exceeds the preset time. If it exceeds the preset time, proceed to step 550; otherwise, return to step 510 to continue collecting user voice.

[0106] Step 550: If it is determined that the preset time has been exceeded, that is, if no user voice to be responded to is collected within the preset time, exit the current interaction mode.

[0107] Based on any of the above embodiments Figure 6 This is a schematic diagram of the structure of the interactive device provided by the present invention, as shown below. Figure 6 As shown, the device includes:

[0108] The acquisition unit 610 is used to acquire user voice in real time after entering the current interaction mode;

[0109] The analysis unit 620 is used to determine whether to respond to the user's voice based on at least one of the user's pose information relative to the local device, the interaction association information of the user's voice, and the interaction mode record of the local device.

[0110] The interaction unit 630 is used to interact based on the user's voice when it is determined that the user's voice will be responded to.

[0111] The device provided in this embodiment of the invention collects user voice in real time during the current interaction mode, providing conditions for multiple interactions with a single wake-up. Based on this, it determines whether to respond to the user voice by considering at least one of the following: the user's pose information relative to the local device, the interaction association information of the user's voice, and the interaction mode records of the local device. This distinguishes whether the user voice collected in the interaction mode is interactive voice, thereby realizing human-computer interaction. In scenarios where multiple interactions are achieved with a single wake-up, the device avoids the problem of the local device mistakenly responding to non-interactive voice, making human-computer interaction as smooth and natural as human communication, and improving the convenience of human-computer interaction.

[0112] Based on any of the above embodiments, the analysis unit is further configured to:

[0113] The user's speech is transcribed to obtain the transcribed text;

[0114] The interaction association information is determined based on the first correlation between the transcribed text and the wake word of the local device, and / or the second correlation between the transcribed text and the historical conversation text in the current interaction mode.

[0115] Based on any of the above embodiments, the analysis unit is further configured to:

[0116] Extract keywords from the transcribed text;

[0117] The first correlation is determined based on the correlation between the keyword and each wake-up word in the wake-up word.

[0118] Based on any of the above embodiments, the analysis unit is further configured to:

[0119] Based on the historical conversation text, conversation prediction is performed to obtain the preset conversation text;

[0120] The second correlation is determined based on the correlation between the preset conversation text and the transcribed text.

[0121] Based on any of the above embodiments, the analysis unit is further configured to:

[0122] Distance detection and / or orientation detection are performed on the user corresponding to the user's voice.

[0123] The pose information is determined based on the distance between the user corresponding to the user's voice and the local device, and / or the orientation of the user corresponding to the user's voice relative to the local device.

[0124] Based on any of the above embodiments, the analysis unit is further configured to:

[0125] The interaction mode record is determined based on the number of times the local device executes the interaction mode after the current power-on and the interaction sequence number of the current interaction mode after the current power-on.

[0126] Based on any of the above embodiments, the acquisition unit is further configured to:

[0127] Upon detecting a wake-up voice message carrying the wake-up word of the local device, the system enters the current interaction mode.

[0128] After entering the current interaction mode, the following is also included:

[0129] If no response is received from the user within the preset time period, the current interaction mode will be exited.

[0130] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7As shown, the electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute an interaction method, which includes: after entering the current interaction mode, acquiring user voice in real time; determining whether to respond to the user voice based on at least one of the user's pose information relative to the local device, the interaction association information of the user's voice, and the interaction mode record of the local device; and if it is determined that the user's voice should be responded to, performing interaction based on the user's voice.

[0131] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0132] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the interaction methods provided by the above methods. The method includes: after entering the current interaction mode, acquiring user voice in real time; determining whether to respond to the user voice based on at least one of the user voice's pose information relative to the local device, the user voice's interaction association information, and the local device's interaction mode record; and if it is determined that the user voice should be responded to, performing interaction based on the user voice.

[0133] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the interaction methods provided by the above methods. The method includes: after entering the current interaction mode, acquiring user voice in real time; determining whether to respond to the user voice based on at least one of the user's pose information relative to the local device, the interaction association information of the user's voice, and the interaction mode record of the local device; and if it is determined that the user's voice should be responded to, performing interaction based on the user's voice.

[0134] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0135] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, etc.) to...

[0136] The server or network device, etc., performs the methods described in the various embodiments or some parts of the embodiments.

[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the techniques described in the foregoing embodiments can still be applied to other aspects.

[0138] The technical solution may be modified, or some of its technical features may be replaced with equivalent ones; however, these modifications or replacements do not cause the essence of the corresponding technical solution to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An interaction method, characterized in that, include: After entering the current interaction mode, the user's voice is collected in real time; Based on at least one of the user's pose information relative to the local device corresponding to the user's voice, the interaction mode record of the local device, and the interaction association information of the user's voice, or based on the interaction association information of the user's voice, determine whether to respond to the user's voice. If a response to the user's voice is determined, interaction is performed based on the user's voice. The interaction association information includes a first correlation between keywords extracted from the transcribed text of the user's voice and each wake-up word in the wake-up word of the local device; the first correlation is determined based on the correlation between the keywords and each wake-up word.

2. The interaction method according to claim 1, characterized in that, The steps for determining the interactive association information include: The user's speech is transcribed to obtain the transcribed text. The interaction association information is determined based on the first correlation, or based on the first correlation and the second correlation between the transcribed text and the historical conversation text in the current interaction mode.

3. The interaction method according to claim 2, characterized in that, The step of determining the second correlation between the transcribed text and the historical conversation text in the current interaction mode includes: Based on the historical conversation text, conversation prediction is performed to obtain the preset conversation text; The second correlation is determined based on the correlation between the preset conversation text and the transcribed text.

4. The interaction method according to claim 1, characterized in that, The steps for determining the pose information include: Distance detection and / or orientation detection are performed on the user corresponding to the user's voice. The pose information is determined based on the distance between the user corresponding to the user's voice and the local device, and / or the orientation of the user corresponding to the user's voice relative to the local device.

5. The interaction method according to claim 1, characterized in that, The steps for determining the interaction mode record include: The interaction mode record is determined based on the number of times the local device executes the interaction mode after the current power-on and the interaction sequence number of the current interaction mode after the current power-on.

6. The interaction method according to any one of claims 1 to 5, characterized in that, Prior to the real-time acquisition of user voice, the following also includes: Upon detecting a wake-up voice message carrying the wake-up word of the local device, the system enters the current interaction mode. After entering the current interaction mode, the following is also included: If no response is received from the user within the preset time period, the current interaction mode will be exited.

7. An interactive device, characterized in that, include: The acquisition unit is used to acquire user voice in real time after entering the current interaction mode; The analysis unit is used to determine whether to respond to the user's voice based on at least one of the user's pose information relative to the local device, the interaction mode record of the local device, and the interaction association information of the user's voice, or based on the interaction association information of the user's voice. An interaction unit is used to interact based on the user's voice when it is determined that a response to the user's voice is required; The interaction association information includes a first correlation between keywords extracted from the transcribed text of the user's voice and each wake-up word in the wake-up word of the local device; the first correlation is determined based on the correlation between the keywords and each wake-up word.

8. An electronic device comprising a microphone, a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The microphone is used to collect user voice in real time after entering the current interaction mode; The processor executes the program to determine whether to respond to the user's voice based on at least one of the user's pose information relative to the local device, the interaction mode record of the local device, and the interaction association information of the user's voice, or based on the interaction association information of the user's voice. And if it is determined that the user's voice is being responded to, interaction will be performed based on the user's voice; The interaction association information includes the first correlation between keywords extracted from the transcribed text of the user's voice and each wake-up word in the wake-up word of the local device.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the interaction method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice skill skipping method for man-machine conversation, electronic equipment and storage medium

    CN111506292A

  • Conversation state quitting method, terminal equipment and storage medium

    CN112328765A