Human-computer interaction method and device, electronic equipment and storage medium
By receiving and comparing the user's session request information and biometric information, the intelligent device can switch back to the corresponding session context when different users interact, solving the problem that continuous conversations cannot be conducted in a single-wheel dialogue scenario, and achieving more efficient human-computer interaction.
Patent Information
- Application Number
- CN202510100888.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
Existing smart devices cannot conduct continuous conversations in a single-wheel dialogue scenario, resulting in each human-computer interaction being regarded as a new session and unable to maintain the continuity of the interaction.
By receiving the session request information of the target user, extracting and comparing the biometric information to determine the user ID, establishing an active session set, and switching back to the corresponding session context when different users interact to maintain the continuity of the interaction.
It realizes that when different users interact with smart devices, they can switch back to the corresponding session context, maintain the continuity of human-computer interaction, save time, and improve the processing efficiency of smart devices.
Smart Images

Figure CN120010668A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of terminal technology, and in particular to a human-computer interaction method, device, electronic device and storage medium. Background Art
[0002] Smart devices have been widely used in the fields of smart customer service, smart speakers, smart transportation, etc. The main contents include: converting user commands from voice into text, then parsing the text of user commands into content that smart devices can understand, then executing user intentions, generating feedback signals, deciding on the dialogue strategy with the user, then converting the dialogue management strategy into fluent text that users can understand, and finally converting the text results into voice for the user.
[0003] At present, smart devices are mainly used in single-round dialogue scenarios, that is, dialogue scenarios in which a user inputs a request and responds immediately. In single-round dialogue scenarios, smart devices can accurately respond to the user's inquiry with the best output. However, after a question and an answer, the round of dialogue ends, and each human-computer interaction is considered a new conversation, and continuous dialogue is impossible. Summary of the invention
[0004] In view of this, the purpose of the present application is to provide a human-computer interaction method, device, electronic device and storage medium, which can establish corresponding session information for each user, so that when different users interact with smart devices, they can switch back to the corresponding session context to maintain the continuity of human-computer interaction, save time, and help improve the processing efficiency of smart devices.
[0005] In a first aspect, an embodiment of the present application provides a human-computer interaction method, the method comprising:
[0006] receiving a session request message sent by a target user, wherein the session request message does not carry a wake-up instruction, and an interval between a request time of the session request message and a session creation time after receiving the wake-up instruction sent by the target user is less than a preset time threshold;
[0007] Extracting the biometric information of the target user, and comparing the biometric information of the target user with the biometric information corresponding to a plurality of user IDs stored locally, to determine the target user ID of the target user;
[0008] According to the target user ID, an active session set is selected from the session sets corresponding to the multiple user IDs stored locally, each of the session sets including multiple single sessions;
[0009] An active single session is determined from the active session set, and session response information in the active single session is played, wherein the active single session indicates that an interval between a request time and a current time of a single session in the active session set is the smallest, and the session response information in the active single session is determined with reference to other single sessions in the active session set except the active single session.
[0010] In a possible implementation, the session set includes a single session stack and a last active time, the single session stack is used to store multiple single sessions, each single session includes a session request information sent by a user and session response information fed back by the cloud for the session request information; the last active time is used to be updated according to the request time of the last single session.
[0011] In a possible implementation manner, each of the session request information carries a request ID, and the request ID is stored in the single session stack. The playing of the session response information in the active single session includes:
[0012] Receive the session response information fed back by the cloud, where the session response information carries the request ID;
[0013] If the request ID carried in the session response information is consistent with the request ID carried in the session request information in the active single session, the session response information is played.
[0014] In a possible implementation, the method further includes: if the interval between the last active time of the inactive session set and the current time is not less than a preset duration threshold, canceling the inactive session set; wherein the inactive session set refers to a session set corresponding to other user IDs among the multiple user IDs except the target user ID.
[0015] In a possible implementation, extracting the biometric information of the target user, and comparing the biometric information of the target user with the biometric information corresponding to a plurality of locally stored user IDs to determine the target user ID of the target user includes:
[0016] Under the condition that the voice input volume is greater than a preset volume threshold, extracting biometric information of the target user, the biometric information including at least one of the following items: face information, body information and voiceprint information;
[0017] Comparing the target user's biometric information with the biometric information corresponding to multiple user IDs stored locally;
[0018] If the comparison is successful, the target user ID of the target user is determined from the multiple user IDs stored locally;
[0019] If the comparison fails, a new user ID is created, and the new user ID and the biometric information corresponding to the new user ID are saved locally.
[0020] In a possible implementation, the biometric information includes voiceprint information, and comparing the biometric information of the target user with biometric information corresponding to a plurality of locally stored user IDs includes:
[0021] Compare the target user's voiceprint information with the pre-recorded voiceprint model and calculate a similarity score;
[0022] According to the volume of the voice input, a weight coefficient is set for the similarity score to obtain a similarity output score;
[0023] If the similarity output score is greater than a first preset score threshold, a first score is output, and the first score indicates that the comparison is successful.
[0024] In a possible implementation, the biometric information includes voiceprint information and face information, and comparing the biometric information of the target user with the biometric information corresponding to the multiple user IDs stored locally includes:
[0025] Compare the target user's voiceprint information with the pre-recorded voiceprint model and calculate a similarity score;
[0026] According to the volume of the voice input, a weight coefficient is set for the similarity score to obtain a similarity output score;
[0027] Under the condition that the similarity output score is greater than the second preset score threshold, if the distance between the user's face and the camera is less than the preset distance threshold, and the facial recognition area is greater than the preset area threshold, then a first score is output; wherein the first score indicates that the comparison is successful;
[0028] Under the condition that the similarity output score is greater than the second preset score threshold, if the distance between the user's face and the camera is not less than the preset distance threshold, and the facial recognition area is not greater than the preset area threshold, a second score is output; wherein the second score indicates a comparison failure.
[0029] In a second aspect, an embodiment of the present application further provides a human-computer interaction device, the device comprising:
[0030] An information receiving module, configured to receive a session request message sent by a target user, wherein the session request message does not carry a wake-up instruction, and an interval between a request time of the session request message and a session creation time after receiving the wake-up instruction sent by the target user is less than a preset time threshold;
[0031] A feature extraction module, used to extract the biometric information of the target user, and compare the biometric information of the target user with the biometric information corresponding to a plurality of user IDs stored locally, to determine the target user ID of the target user;
[0032] A session selection module, configured to select an active session set from session sets corresponding to a plurality of user IDs stored locally according to the target user ID, each of the session sets including a plurality of single sessions;
[0033] The information playing module is used to determine an active single session from the active session set and play the session response information in the active single session, wherein the active single session indicates that the interval between the request time and the current time of the single session in the active session set is the smallest, and the session response information in the active single session is determined with reference to other single sessions in the active session set except the active single session.
[0034] In a third aspect, an embodiment of the present application further provides an electronic device, comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, the steps of the above-mentioned human-computer interaction method are performed.
[0035] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the human-computer interaction method as described above are executed.
[0036] The embodiment of the present application provides a human-computer interaction method, device, electronic device and storage medium, the method comprising: first receiving a session request message sent by a target user, the session request message does not carry a wake-up instruction, and the interval between the request time of the session request message and the session creation time after receiving the wake-up instruction sent by the target user is less than a preset time threshold; then extracting biometric information of the target user, and comparing the biometric information of the target user with the biometric information corresponding to multiple user IDs stored locally, to determine the target user ID of the target user; according to the target user ID, selecting an active session set from the session sets corresponding to the multiple user IDs stored locally, each session set including multiple single sessions; determining an active single session from the active session set, and playing session response information in the active single session, the active single session indicating that the interval between the request time of the single session in the active session set and the current time is the smallest, and the session response information in the active single session is determined with reference to other single sessions except the active single session in the active session set.
[0037] Compared with the prior art where smart devices are used in single-round conversation scenarios and cannot conduct continuous conversations, the embodiments of the present application can establish corresponding conversation information for each user, so that when different users interact with smart devices, they can switch back to the corresponding conversation context to maintain the continuity of human-computer interaction, save time, and help improve the processing efficiency of smart devices. In addition, without the need for complex AI capabilities in the cloud, the user's biometric information is extracted on the local side of the smart device and corresponding conversation information is established for each user, which can more accurately determine the user currently interacting with the smart device, helping to improve the processing accuracy of the smart device.
[0038] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0040] Figure 1 A flowchart of a human-computer interaction method provided in an embodiment of the present application;
[0041] Figure 2 A flowchart of another human-computer interaction method provided in an embodiment of the present application;
[0042] Figure 3 A schematic diagram of the structure of a human-computer interaction device provided in an embodiment of the present application;
[0043] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application usually described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, each other embodiment obtained by those skilled in the art without making creative work belongs to the scope of protection of the present application.
[0045] According to research, smart devices have been widely used in the fields of smart customer service, smart speakers, smart transportation, etc. The main contents include: converting user commands from voice into text, then parsing the text of user commands into content that smart devices can understand, then executing user intentions, generating feedback signals, deciding on the strategy for communicating with users, then converting the dialogue management strategy into fluent text that users can understand, and finally converting the text results into voice for users.
[0046] At present, smart devices are mainly used in single-round dialogue scenarios, that is, dialogue scenarios in which a user inputs a request and responds immediately. In single-round dialogue scenarios, smart devices can accurately respond to the user's inquiry with the best output. However, after a question and an answer, the round of dialogue ends, and each human-computer interaction is considered a new conversation, and continuous dialogue is impossible.
[0047] Based on this, the embodiments of the present application provide a human-computer interaction method, device, electronic device and storage medium, which can establish corresponding session information for each user, so that when different users interact with smart devices, they can switch back to the corresponding session context to maintain the continuity of human-computer interaction, save time, and help improve the processing efficiency of smart devices.
[0048] See also Figure 1 , Figure 1 This is a flow chart of a human-computer interaction method provided in an embodiment of the present application. Figure 1 As shown in , the human-computer interaction method provided by the embodiment of the present application includes:
[0049] S101, receiving a session request message sent by a target user, where the session request message does not carry a wake-up instruction, and the interval between the request time of the session request message and the session creation time after the target user sends the wake-up instruction is less than a preset time threshold;
[0050] S102, extracting biometric information of a target user, and comparing the biometric information of the target user with biometric information corresponding to a plurality of user IDs stored locally, to determine a target user ID of the target user;
[0051] S103, selecting an active session set from session sets corresponding to multiple user IDs stored locally according to the target user ID, each session set including multiple single sessions;
[0052] S104: Determine an active single session from the active session set, and play session response information in the active single session. The active single session indicates that the interval between the request time and the current time of the single session in the active session set is the smallest. The session response information in the active single session is determined by referring to other single sessions in the active session set except the active single session.
[0053] In the above steps S101 to S104, corresponding session information can be established for each user, so that when different users interact with the smart device, they can switch back to the corresponding session context to maintain the continuity of human-computer interaction, save time, and help improve the processing efficiency of the smart device. In addition, without the need for complex AI capabilities in the cloud, the user's biometric information is extracted on the local side of the smart device and corresponding session information is established for each user, which can more accurately determine the user currently interacting with the smart device, helping to improve the processing accuracy of the smart device.
[0054] The above steps S101 to S104 are described in detail below through a specific embodiment:
[0055] In step S101, a session request message sent by a target user is received, the session request message does not carry a wake-up instruction, and the interval between the request time of the session request message and the session creation time after receiving the wake-up instruction sent by the target user is less than a preset time threshold.
[0056] The embodiment of the present application sends a wake-up command to the smart device, so that the smart device enters the working state from the standby state to receive and process the user's voice command. Taking the smart speaker as an example, the smart speaker can be awakened by a wake-up word such as "Xiao X, Xiao X", so that the smart speaker enters the working state to receive the session request information sent by the target user.
[0057] The interval between the request time of the session request information in the above step S101 and the session creation time after receiving the wake-up instruction sent by the target user is less than the preset time threshold. This condition indicates that the current smart device is in a working state, so that the session request information without the wake-up instruction can be received and recognized by the smart device.
[0058] In which, within the preset time threshold after the target user sends the wake-up command, the target user can directly send a session request message to the smart device without repeatedly sending the wake-up command, making the voice interaction between the user and the smart device more convenient and rapid. For example, the preset time threshold can be set to 10 minutes.
[0059] In step S102, the biometric information of the target user is extracted, and the biometric information of the target user is compared with the biometric information corresponding to a plurality of user IDs stored locally, so as to determine the target user ID of the target user.
[0060] In this step, biometric information corresponding to multiple user IDs is pre-stored locally on the smart device, and biometric information of the locally stored user ID that is consistent with the biometric information of the target user is searched from the biometric information corresponding to the multiple user IDs stored locally, thereby determining the target user ID from the multiple user IDs.
[0061] Furthermore, if Figure 2 As shown, step S102 specifically includes:
[0062] Step S1021: Under the condition that the voice input volume is greater than a preset volume threshold, extract the biometric information of the target user, where the biometric information includes at least one of the following items: facial information, body shape information, and voiceprint information.
[0063] Here, extracting the target user's biometric information under the condition that the voice input volume is greater than a preset volume threshold can ensure that the biometric information extracted by the smart device is the biometric information of the target user, which helps to improve the accuracy of information extraction. For example, the average volume of the voice input can be checked. If it is lower than a preset volume threshold (such as 10 decibels), it is considered invalid input and will not be processed. Furthermore, by limiting the decibel threshold, invalid conversations can be effectively filtered out.
[0064] Exemplarily, when extracting the biometric information of the target user, the smart device can obtain it according to its hardware configuration, such as collecting facial information and body information through a camera, and collecting voiceprint information through a microphone. Taking the smart device as a smart speaker as an example, the speakers currently on the market that support human-computer dialogue have microphones, and some high-end speakers (especially speakers with screens) have cameras to achieve the purpose of extracting biometric information through microphones and / or cameras. For example, the processing module inside the smart speaker processes the video stream information input by the camera and the audio stream information input by the microphone, and mainly uses a plug-in design method to extract biometric information corresponding to different users from different input information. For example, facial information or body information is extracted from video stream information, and voiceprint features are extracted from audio stream information.
[0065] Step S1022: Compare the target user's biometric information with the biometric information corresponding to multiple user IDs stored locally.
[0066] Specifically, the biometric information extracted in step 1021 is saved so that the memory of the smart speaker stores all the biometric information extracted since the smart speaker is started and the user ID corresponding to each biometric information. Then, the biometric information of the target user is compared with the biometric information corresponding to the multiple user IDs stored locally to screen out the most likely user ID for this session.
[0067] Step S1023: If the comparison is successful, the user ID of the target user is determined from multiple user IDs stored locally.
[0068] Step S1024: If the comparison fails, a new user ID is created, and the new user ID and the biometric information corresponding to the new user ID are saved locally.
[0069] That is to say, for a new user, when the biometric information corresponding to the user ID is not stored locally, the smart device can extract the biometric information of the new user and store the user ID of the new user and the corresponding biometric information.
[0070] The embodiment of the present application compares the locally stored user biometric information with the biometric information of the user initiating the session, so as to quickly find the user ID of the target user, simplify the cloud call time, and respond more quickly.
[0071] In an optional embodiment, under the condition that the biometric information includes voiceprint information, step S1022 specifically includes:
[0072] The target user's voiceprint information is compared with the pre-recorded voiceprint model to calculate the similarity score; a weight coefficient is set for the similarity score according to the volume of the voice input to obtain a similarity output score; if the similarity output score is greater than a first preset score threshold, a first score is output, and the first score indicates that the comparison is successful.
[0073] For example: first check the average volume of the voice input. If it is lower than the preset volume threshold (for example, 10 decibels), it is considered as invalid input, and no processing is performed, and 0 points are output.
[0074] Under the condition that the voice input volume is greater than the preset volume threshold, the target user's voiceprint information and voice input volume are extracted and compared with the locally stored voiceprint information. The specific comparison method can be: using the voiceprint recognition algorithm, the target user's voiceprint information is compared with the pre-recorded voiceprint model to calculate the similarity score; according to the voice input volume, a weight coefficient is set for the similarity score to obtain a similarity output score; here, the weight coefficient is positively correlated with the voice input volume, but the sum of the weight coefficients does not exceed 1. If the similarity output score is greater than the first preset score threshold, the first score is output, and the first score indicates a successful comparison; if the similarity output score is not greater than the first preset score threshold, the second score is output, and the second score indicates a failed comparison. Exemplarily, the first preset score threshold can be 70 points.
[0075] If it is the voiceprint information of a new user, a unique new user ID is assigned to the voiceprint information, and the new user ID is the user ID of this session. In this case, 100 points are output.
[0076] The embodiment of the present application compares the user's biometric information by voiceprint information and the volume of the voice input, and can accurately identify the user's identity in a simple manner with a low error rate.
[0077] In another optional embodiment, under the condition that the biometric information includes voiceprint information and face information, step S1022 specifically includes:
[0078] Compare the target user's voiceprint information with the pre-recorded voiceprint model to calculate the similarity score; set a weight coefficient for the similarity score according to the volume of the voice input to obtain the similarity output score;
[0079] Under the condition that the similarity output score is greater than the second preset score threshold, if the distance between the user's face and the camera is less than the preset distance threshold, and the facial recognition area is greater than the preset area threshold, then a first score is output; wherein the first score indicates that the comparison is successful;
[0080] Under the condition that the similarity output score is greater than the second preset score threshold, if the distance between the user's face and the camera is not less than the preset distance threshold, and the facial recognition area is not larger than the preset area threshold, then the second score is output; wherein the second score indicates that the comparison failed.
[0081] In the above steps, if the similarity output score is greater than the second preset score threshold, it indicates that the voiceprint information comparison is successful. Here, the second preset score threshold can be equal to the first preset score threshold. Under the condition that the similarity output score is greater than the second preset score threshold, it is determined whether the distance between the user's face and the camera and the facial recognition area meet the requirements to determine whether the comparison is successful or not.
[0082] For example, the voice input volume of the microphone is first detected. If the voice input volume is lower than the preset volume threshold (such as 10 decibels), it is considered invalid input, not processed, and output 0 points. If the voice input volume is higher than the preset volume threshold, the next step is continued to extract the voiceprint information and the voice input volume in the microphone: using the voiceprint recognition algorithm, the target user's voiceprint information is compared with the pre-recorded voiceprint model, and the similarity score is calculated. According to the voice input volume, the weight coefficient is set for the similarity score to obtain the similarity output score; then, the video stream information input by the camera is detected: the face information of the camera and the size of the face recognition area are extracted. If the distance between the user's face and the camera is less than the preset distance threshold, and the face recognition area is larger than the preset area area threshold, the first score is output, and the first score indicates that the comparison is successful; that is, if the face is close to the camera and the face recognition area is larger, it can be considered that the face captured by the camera is the same person as the person heard by the microphone; further, if the face recognized by the camera matches the voiceprint information of the person heard by the microphone and the size of the face recognition area is as expected, the highest score (such as 100) is output. If the distance between the user's face and the camera is not less than the preset distance threshold, and the facial recognition area is not larger than the preset area threshold, a second score is output, and the second score indicates a failed comparison; that is, if the face is far away from the camera and the facial recognition area is small, it can be considered that the face captured by the camera and the person heard by the microphone are not the same person; further, if the face recognized by the camera does not match the voiceprint information of the person heard by the microphone or the size of the facial recognition area does not meet expectations, a lower score can be comprehensively calculated based on the degree of mismatch and the degree of difference in area size.
[0083] The embodiment of the present application compares the user's biometric information through voiceprint information and facial information, which can identify user characteristics in more detail and accurately with a lower error rate.
[0084] It should be added that the embodiments of the present application can also improve the accuracy of user identification by adding AI capabilities locally and utilizing AI capabilities.
[0085] In step S103, according to the target user ID, an active session set is selected from the session sets corresponding to the multiple user IDs stored locally, each session set including multiple single sessions.
[0086] Here, each user ID corresponds to a session set. Specifically, the active session set is selected according to the target user ID, that is, the session set corresponding to the target user ID is the active session set. Further, the active session set includes multiple single sessions.
[0087] Optionally, the session set includes a single session stack and a last active time. The single session stack is used to store multiple single sessions. Each single session includes a session request information sent by a user and session response information fed back by the cloud for the session request information. The last active time is used to update according to the request time of the last single session.
[0088] In an optional embodiment, step S103 specifically includes:
[0089] Step 1031: Receive a wake-up instruction sent by the user and create a session data structure, wherein each wake-up instruction corresponds to a session data structure, and each session data structure includes a user ID, a session ID, a session creation time, a last active time, and a single session stack.
[0090] Among them, the user ID is used as the unique identification of the user characteristics, and the session ID is used as the unique identification of the session data structure. The session data structure created after each wake-up instruction is received is uniquely marked, that is, the session generated each time the user sends a wake-up instruction is uniquely marked. The session creation time represents the time when the session data structure is created after receiving the wake-up instruction, and the last active time represents the time corresponding to the last session request information sent by the user. The single session stack is used to store multiple single sessions.
[0091] Step 1032: Under the same wake-up instruction, multiple session request information sent by the same user is saved in a single session stack, and session response information determined for each session request information fed back by the cloud is received and saved in the single session stack to obtain a session set; wherein the single session stack includes multiple single sessions and the request ID, request time, and response time corresponding to each single session.
[0092] Here, "multiple" means two or more, the request time means the time when the user sends the session request information, the response time means the time when the cloud feeds back the session response information, and the request ID is the unique identification identifier of a single session, which is set for the session request information when the user sends the session request information. Exemplarily, the request ID can be a digital identifier or a snowflake identifier.
[0093] According to the session request information after the user sends the wake-up command and the session response information determined for each session request information fed back by the cloud, the user ID, session ID and session creation time of the session data structure are determined, and the content of the single session stack is updated in real time according to the multiple session request information sent by the user and the session response information determined for each session request information fed back by the cloud, and the last active time is updated in real time according to the request time of the multiple session request information sent by the user, that is, the last active time is updated according to the request time of the last single session. Among them, the single session stack includes multiple single sessions and the request ID, request time and response time corresponding to each single session.
[0094] Specifically, each time the same user sends a wake-up command to the smart device, a session data structure is created according to the above steps 1031 and 1032; when multiple users send wake-up commands to the smart device, session data structures are created in the above manner.
[0095] In one possible implementation, when a user sends multiple wake-up instructions to a smart device within a preset time threshold, multiple session data structures will be created. If the previous session data structure is not destroyed after receiving a new wake-up instruction, the embodiment of the present application can determine the active session data structure from the active session set, and the interval between the last active time of the active session data structure and the current time is the smallest.
[0096] Here, the active session set includes multiple session data structures, and there is an active session data structure among the multiple session data structures. The last active time of the active session data structure is the request time of the session request information in the last single session. When the interval between the last active time of the session data structure and the current time is the smallest, it can be characterized that the session data structure is the most active, so as to determine the active single session from the active session data structure and play the session response information in the active single session.
[0097] In a possible implementation, if the interval between the last active time of an inactive session data structure in the active session set and the current time is not less than a preset time threshold, the inactive session data structure is deregistered; wherein the inactive session data structure refers to other session data structures in the active session set except the active session data structure.
[0098] Here, the active session set includes an active session data structure and an inactive session data structure, both of which include multiple rounds of single sessions, wherein the multiple rounds of single sessions included in the active session data structure include active single sessions, and the inactive session data structure is other session data structures in the active session set except the active session data structure.
[0099] The embodiment of the present application can destroy multiple inactive single sessions in the active session concentration by canceling the inactive session data structure, which can effectively release corresponding system resources (such as memory, etc.), saving resources and achieving faster response.
[0100] In step S104, an active single session is determined from the active session set, and session response information in the active single session is played. The active single session indicates that the interval between the request time and the current time of the single session in the active session set is the smallest. The session response information in the active single session is determined with reference to other single sessions in the active session set except the active single session.
[0101] Here, the single session with the smallest interval between the request time and the current time among multiple single sessions in the active session set is defined as the active single session. The active single session includes a session request information and a session response information corresponding to the session request information, and then the session response information in the active single session is played. The session request information of other single sessions in the active session set except the active single session can be responded by the cloud, but the session response information sent by the cloud is not played.
[0102] Specifically, the single conversation at the top of the single conversation stack is an active single conversation. The role of the active single conversation is as follows: Due to the responsiveness of the network and cloud AI, it is often seen that when the user initiates question 1, there is no reply for several seconds. In a smart device that does not support multi-round conversations, users can only wait passively. In the multi-round conversation of the embodiment of the present application, the user can continue to ask the next question 2 if he does not receive a reply to question 1. Assuming that the replies to question 2 and question 1 have reached the smart device side, the smart device will only play the reply to question 2, because the single conversation corresponding to question 2 is the current active single conversation, and the response to the inactive single conversation will be recorded, but not broadcast.
[0103] Further, the session response information in the active single session is determined by referring to other single sessions in the active session set except the active single session. Further, the session response information included in the active single session is determined by referring to the session request information in other single sessions in the active session set except the active single session. That is, the final session response information can only be obtained by comprehensively considering other session request information in the active session set. In other words, the session response information included in the active single session is obtained by contacting the associated content of the context, thereby ensuring the continuity of human-computer interaction.
[0104] Example 1: Target user A first asks "Xiao X, Xiao X, what's the weather like today?", and before receiving a reply from the smart device (because the smart device needs the AI capabilities of the cloud to respond and process, and depends on the network conditions, this reply time often takes 2-3 seconds or even longer), target user A can continue to ask "What is the oil price in Beijing today?" Then, the current active single session is "What is the oil price in Beijing today", and the smart device will eventually play the oil price in Beijing today.
[0105] Example 2: Target user B sends "Xiao X, Xiao X, play piano music" to the smart device. Two seconds later, target user C sends "Xiao X, Xiao X, what's the temperature today" to the smart device. Three seconds later, target user B sends "classical style" to the smart device. The smart device will eventually play a classical piano piece.
[0106] Example 3: Target user D sends "Xiao X, Xiao X, play piano music" to the smart device. Three seconds later, target user E sends "Xiao X, Xiao X, what's the temperature today" to the smart device. Two seconds later, target user D sends "classical style" to the smart device. Three seconds later, target user E sends "what time is it now" to the smart device. The smart device will eventually play the current time.
[0107] It should be noted that, within the preset time threshold after the target user sends the wake-up command, when the target user sends the session request information without the wake-up command multiple times, the cloud will associate the session request information of the same target user by default and feed back the session response information corresponding to each session request information to the smart device respectively, but the smart device only plays the session response information corresponding to the last session request information.
[0108] In a possible implementation, for the same user, each time a wake-up instruction is sent, it means that the previous session ends and a new session is reopened. Accordingly, subsequent session request information is also saved in the newly reopened session. For different users, within a preset time threshold after the first user sends the wake-up instruction, the session request information sent by the user is saved in the current session. When the second user sends the wake-up instruction, a new session for the second user will be created. If the first user sends a session request information without a wake-up instruction within the preset time threshold, the session request information is still saved in the session previously created by the first user, thereby achieving the purpose of switching back to the corresponding session context dialogue to maintain the continuity of human-computer interaction.
[0109] The above steps can save the user from answering questions that are no longer of interest to the user through active single-session judgment, and save the time of waiting for responses to inactive questions, making human-computer interaction smoother and more convenient. Furthermore, the human-computer interaction method in the embodiment of the present application can easily support multi-person multi-round conversations, greatly improving the user experience.
[0110] In an optional embodiment, each session request information carries a request ID, the request ID is stored in a single session stack, and the session response information in the active single session is played, including:
[0111] Receive the session response information fed back by the cloud, where the session response information carries the request ID; if the request ID carried in the session response information is consistent with the request ID carried in the session request information in the active single session, play the session response information.
[0112] The following description is made for the above embodiment:
[0113] a. According to the request ID carried in the session response information fed back by the cloud, determine whether the session request information corresponding to the session response information is the session request information in the active session set. If not, record the session response information but do not play it. Here, according to the request ID carried in the session response information, it can be determined that this is a question asked by another user. For some reason, the session corresponding to the user is no longer active, so the session can be abandoned.
[0114] b. If the session request information corresponding to the session response information is the session request information in the active session set, determine whether the session request information corresponding to the session response information is the session request information in the active single session according to the request ID carried by the session response information. If not, record the session response information but do not play it. The user may no longer care about the answer to the previous question for some reason, but continue to ask a new question. Then, there is no need to play the answer to the old question. If yes, play the session response information in the active single session.
[0115] In an optional embodiment, the embodiment of the present application also includes: if the interval between the last active time of the inactive session set and the current time is not less than a preset duration threshold, then canceling the inactive session set; wherein the inactive session set refers to a session set corresponding to other user IDs among multiple user IDs except the target user ID.
[0116] Here, each user ID corresponds to a session set, wherein the session set corresponding to the target user ID is an active session set, and the session sets corresponding to other user IDs are inactive session sets. The target user ID and other user IDs are user IDs saved on the smart device.
[0117] The embodiment of the present application can destroy inactive sessions by canceling inactive session sets, which can effectively release corresponding system resources (such as memory, etc.), thereby saving resources and achieving faster response.
[0118] The embodiment of the present application can establish corresponding session information for each user, so that when different users interact with smart devices, they can switch back to the corresponding session context to maintain the continuity of human-computer interaction, save time, and help improve the processing efficiency of smart devices. In addition, without the need for complex AI capabilities in the cloud, the user's biometric information is extracted on the local side of the smart device and corresponding session information is established for each user, which can more accurately determine the user currently interacting with the smart device, helping to improve the processing accuracy of the smart device.
[0119] Based on the same inventive concept, the embodiment of the present application also provides a human-computer interaction device corresponding to the human-computer interaction method. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the above-mentioned human-computer interaction method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0120] See also Figure 3 , Figure 3 This is a flow chart of a human-computer interaction device provided in an embodiment of the present application. Figure 3 As shown in , the human-computer interaction device provided in the embodiment of the present application includes:
[0121] The information receiving module 301 is used to receive a session request message sent by a target user, wherein the session request message does not carry a wake-up instruction, and the interval between the request time of the session request message and the session creation time after receiving the wake-up instruction sent by the target user is less than a preset time threshold;
[0122] A feature extraction module 302 is used to extract the biometric information of the target user, and compare the biometric information of the target user with the biometric information corresponding to multiple user IDs stored locally, to determine the target user ID of the target user;
[0123] A session selection module 303 is used to select an active session set from session sets corresponding to multiple user IDs stored locally according to the target user ID, each of which includes multiple single sessions;
[0124] The information playing module 304 is used to determine an active single session from the active session set, and play the session response information in the active single session, wherein the active single session indicates that the interval between the request time and the current time of the single session in the active session set is the smallest, and the session response information in the active single session is determined with reference to other single sessions in the active session set except the active single session.
[0125] In an optional embodiment, the session set includes a single session stack and a last active time, the single session stack is used to store multiple single sessions, each single session includes a session request information sent by a user and session response information fed back by the cloud for the session request information; the last active time is used to update according to the request time of the last single session.
[0126] In an optional embodiment, each session request message carries a request ID, and the request ID is stored in the single session stack. The information playing module 304 is specifically used for:
[0127] Receive the session response information fed back by the cloud, where the session response information carries the request ID;
[0128] If the request ID carried in the session response information is consistent with the request ID carried in the session request information in the active single session, the session response information is played.
[0129] In an optional embodiment, the embodiment of the present application further includes an information destruction module (not shown in the figure), and the information destruction module is used to:
[0130] If the interval between the last active time of the inactive session set and the current time is not less than a preset duration threshold, the inactive session set is cancelled; wherein the inactive session set refers to the session set corresponding to other user IDs among the multiple user IDs except the target user ID.
[0131] In an optional embodiment, the feature extraction module 302 is specifically used for:
[0132] Under the condition that the voice input volume is greater than a preset volume threshold, extracting biometric information of the target user, the biometric information including at least one of the following items: face information, body information and voiceprint information;
[0133] Comparing the target user's biometric information with the biometric information corresponding to multiple user IDs stored locally;
[0134] If the comparison is successful, the target user ID of the target user is determined from the multiple user IDs stored locally;
[0135] If the comparison fails, a new user ID is created, and the new user ID and the biometric information corresponding to the new user ID are saved locally.
[0136] In an optional embodiment, the feature extraction module 302 is further configured to:
[0137] Compare the target user's voiceprint information with the pre-recorded voiceprint model and calculate a similarity score;
[0138] According to the volume of the voice input, a weight coefficient is set for the similarity score to obtain a similarity output score;
[0139] If the similarity output score is greater than a first preset score threshold, a first score is output, and the first score indicates that the comparison is successful.
[0140] In an optional embodiment, the feature extraction module 302 is further configured to:
[0141] Compare the target user's voiceprint information with the pre-recorded voiceprint model and calculate a similarity score;
[0142] According to the volume of the voice input, a weight coefficient is set for the similarity score to obtain a similarity output score;
[0143] Under the condition that the similarity output score is greater than the second preset score threshold, if the distance between the user's face and the camera is less than the preset distance threshold, and the facial recognition area is greater than the preset area threshold, then a first score is output; wherein the first score indicates that the comparison is successful;
[0144] Under the condition that the similarity output score is greater than the second preset score threshold, if the distance between the user's face and the camera is not less than the preset distance threshold, and the facial recognition area is not greater than the preset area threshold, a second score is output; wherein the second score indicates a comparison failure.
[0145] Compared with the prior art where smart devices are used in single-round conversation scenarios and cannot conduct continuous conversations, the human-computer interaction device provided by the embodiment of the present application can establish corresponding conversation information for each user, so that when different users interact with the smart device, they can switch back to the corresponding conversation context to maintain the continuity of human-computer interaction, save time, and help improve the processing efficiency of the smart device. In addition, without the need for complex AI capabilities in the cloud, the user's biometric information is extracted on the local side of the smart device and corresponding conversation information is established for each user, which can more accurately determine the user currently interacting with the smart device, helping to improve the processing accuracy of the smart device.
[0146] See also Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 4 As shown in FIG. 4 , the electronic device 400 includes a processor 401 , a memory 402 and a bus 403 .
[0147] The memory 402 stores machine-readable instructions executable by the processor 401. When the electronic device 400 is running, the processor 401 communicates with the memory 402 via the bus 403. When the machine-readable instructions are executed by the processor 401, the above-mentioned Figure 1 to Figure 2 The steps of the human-computer interaction method in the method embodiment shown and the specific implementation methods can be found in the method embodiment, which will not be repeated here.
[0148] The present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the computer program can execute the above-mentioned Figure 1 to Figure 2 The steps of the human-computer interaction method in the method embodiment shown and the specific implementation methods can be found in the method embodiment, which will not be repeated here.
[0149] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0150] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0151] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0152] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0153] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application can essentially be embodied in the form of a software product, or in other words, the part that contributes to the prior art or the part of the technical solution. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0154] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application is described in detail with reference to the above-mentioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-mentioned embodiments within the technical scope disclosed in the present application, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A human-computer interaction method, characterized in that: The method comprises: receiving a session request message sent by a target user, wherein the session request message does not carry a wake-up instruction, and an interval between a request time of the session request message and a session creation time after receiving the wake-up instruction sent by the target user is less than a preset time threshold; Extracting the biometric information of the target user, and comparing the biometric information of the target user with the biometric information corresponding to a plurality of user IDs stored locally, to determine the target user ID of the target user; According to the target user ID, an active session set is selected from the session sets corresponding to the multiple user IDs stored locally, each of the session sets including multiple single sessions; An active single session is determined from the active session set, and session response information in the active single session is played, wherein the active single session indicates that an interval between a request time and a current time of a single session in the active session set is the smallest, and the session response information in the active single session is determined with reference to other single sessions in the active session set except the active single session.
2. The method according to claim 1, characterized in that The session set includes a single session stack and a last active time. The single session stack is used to store multiple single sessions. Each single session includes a session request information sent by a user and session response information fed back by the cloud for the session request information. The last active time is used to update according to the request time of the last single session.
3. The method according to claim 2, characterized in that Each of the session request information carries a request ID, and the request ID is stored in the single session stack. The playing of the session response information in the active single session includes: Receive the session response information fed back by the cloud, where the session response information carries the request ID; If the request ID carried in the session response information is consistent with the request ID carried in the session request information in the active single session, the session response information is played.
4. The method according to claim 2, characterized in that: The method further comprises: If the interval between the last active time of the inactive session set and the current time is not less than a preset duration threshold, the inactive session set is cancelled; wherein the inactive session set refers to the session set corresponding to other user IDs among the multiple user IDs except the target user ID.
5. The method according to claim 1, characterized in that The step of extracting the biometric information of the target user and comparing the biometric information of the target user with the biometric information corresponding to a plurality of user IDs stored locally to determine the target user ID of the target user includes: Under the condition that the voice input volume is greater than a preset volume threshold, extracting biometric information of the target user, the biometric information including at least one of the following items: face information, body information and voiceprint information; Comparing the target user's biometric information with the biometric information corresponding to multiple user IDs stored locally; If the comparison is successful, the target user ID of the target user is determined from the multiple user IDs stored locally; If the comparison fails, a new user ID is created, and the new user ID and the biometric information corresponding to the new user ID are saved locally.
6. The method according to claim 5, characterized in that The biometric information includes voiceprint information, and comparing the biometric information of the target user with biometric information corresponding to a plurality of locally stored user IDs respectively includes: Compare the target user's voiceprint information with the pre-recorded voiceprint model and calculate a similarity score; According to the volume of the voice input, a weight coefficient is set for the similarity score to obtain a similarity output score; If the similarity output score is greater than a first preset score threshold, a first score is output, and the first score indicates that the comparison is successful.
7. The method according to claim 5, characterized in that The biometric information includes voiceprint information and face information, and comparing the biometric information of the target user with the biometric information corresponding to the multiple user IDs stored locally includes: Compare the target user's voiceprint information with the pre-recorded voiceprint model and calculate a similarity score; According to the volume of the voice input, a weight coefficient is set for the similarity score to obtain a similarity output score; Under the condition that the similarity output score is greater than the second preset score threshold, if the distance between the user's face and the camera is less than the preset distance threshold, and the facial recognition area is greater than the preset area threshold, then a first score is output; wherein the first score indicates that the comparison is successful; Under the condition that the similarity output score is greater than the second preset score threshold, if the distance between the user's face and the camera is not less than the preset distance threshold, and the facial recognition area is not greater than the preset area threshold, a second score is output; wherein the second score indicates a comparison failure.
8. A human-computer interaction device, characterized in that: The device comprises: An information receiving module, configured to receive a session request message sent by a target user, wherein the session request message does not carry a wake-up instruction, and an interval between a request time of the session request message and a session creation time after receiving the wake-up instruction sent by the target user is less than a preset time threshold; A feature extraction module, used to extract the biometric information of the target user, and compare the biometric information of the target user with the biometric information corresponding to a plurality of user IDs stored locally, to determine the target user ID of the target user; A session selection module, configured to select an active session set from session sets corresponding to a plurality of user IDs stored locally according to the target user ID, each of the session sets including a plurality of single sessions; The information playing module is used to determine an active single session from the active session set and play the session response information in the active single session, wherein the active single session indicates that the interval between the request time and the current time of the single session in the active session set is the smallest, and the session response information in the active single session is determined with reference to other single sessions in the active session set except the active single session.
9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and the processor executes the machine-readable instructions to perform the steps of the human-computer interaction method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the human-computer interaction method according to any one of claims 1 to 7 are executed.