Robot active interaction method and device and computer equipment
By recognizing users' social status and generating decision-making scoring information, and combining this with memory-based interaction strategies to generate proactive interactive content, the problem of traditional robots being unable to dynamically adjust their behavior has been solved, achieving highly accurate proactive interaction and an optimized user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-03-27
AI Technical Summary
Traditional robots' active interaction systems cannot integrate vision, voice, behavior, and historical memory, making it impossible to determine when to actively interact and when to remain silent. They also lack growth memory and feedback learning capabilities, and cannot dynamically adjust their behavior according to the user's state, resulting in a poor user experience.
By acquiring the user's current contextual information, identifying social status, generating decision scoring information and action intent information, combining user memory interaction strategies to generate proactive interaction content, integrating multimodal contextual information for high-dimensional contextual awareness, dynamically determining whether to execute proactive or delayed behavior, and recording user profiles and feedback results to adaptively adjust interaction strategies.
It improves the accuracy of recognizing the intention of proactive interaction, avoids the lack of timely response at critical moments and frequent interruptions at inappropriate times, and achieves proactive interaction with low disturbance and high adaptability, thus enhancing the user experience.
Smart Images

Figure CN121744081A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus and computer device for active robot interaction. Background Technology
[0002] The proactive interaction capability of a robot refers to its ability to initiate communication or behavior proactively by sensing the environment and the user's state, without relying on explicit user commands. This capability significantly enhances the naturalness and fluency of human-computer interaction, making the robot more like an active partner rather than a passive tool. However, the proactive interaction capability of a robot relies heavily on its adjustment to various factors such as the user's state, feelings, and the context to improve the user experience. Therefore, improving the intelligence of a robot's proactive interaction is a current research focus.
[0003] Traditional robot intelligent interaction systems generally rely on passive command triggers, failing to integrate vision, voice, behavior, and historical memory simultaneously. This results in robots being unable to determine "when to be proactive and when to remain silent," leading to a lack of timely response at crucial moments and frequent disturbances to users at inappropriate times. Furthermore, the existing systems' event triggering and cyclical planning tasks are disconnected, failing to translate real-time events into subsequent planning or dynamically suppress or delay pre-defined tasks based on user status, thus unable to support a natural and continuous proactive behavior chain. In addition, robots lack growth memory and feedback learning capabilities, preventing them from adjusting proactive behavior based on user preferences and historical responses over long-term use. They remain stuck in fixed script-based interactions, lacking personalization and evolutionary capabilities, resulting in a poor user experience for the robot's proactive interaction functions. Summary of the Invention
[0004] Therefore, it is necessary to provide a method, apparatus, and computer device for active robot interaction to address the aforementioned technical problems.
[0005] In a first aspect, this application provides a method for active robot interaction, including:
[0006] Obtain the user's current context information and, based on the current context information, identify the user's current social status;
[0007] Based on the user's current social status, a social decision evaluation strategy is used to generate the robot's current decision score information, and based on the current decision score information, the robot's current action intention information is generated.
[0008] Based on the current action intent information, the robot generates the current active interaction content through a user memory interaction strategy, and based on the current active interaction content, the robot performs active interaction processing on the user.
[0009] Optionally, identifying the user's current social status based on the current fusion context information includes:
[0010] The current fusion scenario information is broken down into current modal data for each modality type;
[0011] Based on the current modality data of each modality type, the current user state data of each modality type is generated through the event detection model of each modality type;
[0012] Based on the current user status data for each modality type, the user's current social status is generated through a social status evaluation strategy.
[0013] Optionally, the step of generating the robot's current decision score information based on the user's current social status through a social decision evaluation strategy includes:
[0014] Based on the user's current social status, identify the arbitration scores for each key variable;
[0015] Based on the arbitration scores of each of the key variables, the current task execution status corresponding to the current social status is identified;
[0016] The current task execution status is used as the robot's current decision scoring information.
[0017] Optionally, generating the robot's current action intention information based on the current decision scoring information includes:
[0018] Based on the current task execution status, query the current intent generation strategy;
[0019] Based on the current modal data of each modal type, the robot's current action intention information is generated through the current intention recognition strategy.
[0020] Optionally, generating the robot's current active interaction content based on the current action intent information and through a user memory interaction strategy includes:
[0021] Obtain the user's current user profile and the user's historical interaction content;
[0022] Based on the user's current user profile, the user's historical interaction content, and the current action intent information, the robot's current task planning content is generated through a task decomposition strategy.
[0023] Based on the robot's current task planning content, the robot's current active interaction content is generated according to the language interaction model.
[0024] Optionally, the step of actively interacting with the user through the robot based on the current active interaction content includes:
[0025] Based on the robot's current active interaction content, the current active interaction task sequence of each functional component of the robot is identified, and each functional component is controlled to actively interact with the user through the current active interaction content of each functional component, so as to obtain the user's interaction feedback content.
[0026] Based on the user's interactive feedback, the current interactive content is adjusted to obtain new proactive interactive content. The new proactive interactive content replaces the current proactive interactive content, and the process returns to execute the steps of identifying the current proactive interactive task sequence of each functional component of the robot based on the robot's current proactive interactive content.
[0027] Secondly, this application also provides a robot active interaction device, comprising:
[0028] The acquisition module is used to acquire the user's current fusion context information and, based on the current fusion context information, identify the user's current social status;
[0029] The generation module is used to generate the robot's current decision score information based on the user's current social status and through a social decision evaluation strategy, and to generate the robot's current action intention information based on the current decision score information;
[0030] The interaction module is used to generate the robot's current active interaction content based on the current action intent information and through the user's memory interaction strategy, and to perform active interaction processing on the user through the robot based on the current active interaction content.
[0031] Optionally, the acquisition module is specifically used for:
[0032] The current fusion scenario information is broken down into current modal data for each modality type;
[0033] Based on the current modality data of each modality type, the current user state data of each modality type is generated through the event detection model of each modality type;
[0034] Based on the current user status data for each modality type, the user's current social status is generated through a social status evaluation strategy.
[0035] Optionally, the generation module is specifically used for:
[0036] Based on the user's current social status, identify the arbitration scores for each key variable;
[0037] Based on the arbitration scores of each of the key variables, the current task execution status corresponding to the current social status is identified;
[0038] The current task execution status is used as the robot's current decision scoring information.
[0039] Optionally, the generation module is specifically used for:
[0040] Based on the current task execution status, query the current intent generation strategy;
[0041] Based on the current modal data of each modal type, the robot's current action intention information is generated through the current intention recognition strategy.
[0042] Optionally, the interaction module is specifically used for:
[0043] Obtain the user's current user profile and the user's historical interaction content;
[0044] Based on the user's current user profile, the user's historical interaction content, and the current action intent information, the robot's current task planning content is generated through a task decomposition strategy.
[0045] Based on the robot's current task planning content, the robot's current active interaction content is generated according to the language interaction model.
[0046] Optionally, the interaction module is specifically used for:
[0047] Based on the robot's current active interaction content, the current active interaction task sequence of each functional component of the robot is identified, and each functional component is controlled to actively interact with the user through the current active interaction content of each functional component, so as to obtain the user's interaction feedback content.
[0048] Based on the user's interactive feedback, the current interactive content is adjusted to obtain new proactive interactive content. The new proactive interactive content replaces the current proactive interactive content, and the process returns to execute the steps of identifying the current proactive interactive task sequence of each functional component of the robot based on the robot's current proactive interactive content.
[0049] Thirdly, this application provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method described in any one of the first aspects.
[0050] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in any one of the first aspects.
[0051] Fifthly, this application provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the method described in any one of the first aspects.
[0052] The aforementioned robot proactive interaction method, apparatus, and computer equipment acquire the user's current fused context information and, based on this information, identify the user's current social state. Based on the user's current social state, a social decision evaluation strategy is used to generate the robot's current decision score information, and based on this score, the robot's current action intention information is generated. Based on this intention information, a user memory interaction strategy is used to generate the robot's current proactive interaction content, and based on this content, the robot performs proactive interaction processing with the user. This solution integrates multimodal contextual information to identify the user's current social state, thus providing reasonable high-dimensional contextual information for proactive interaction. Furthermore, the social decision evaluation strategy in this solution dynamically determines whether to execute proactive behavior, delay behavior, or inhibit behavior based on event priority, the user's current state (e.g., whether they are on a call or focused on work), environmental noise, and historical interaction preferences, achieving low-interference, highly adaptable proactive interaction. This solution improves the accuracy of recognizing the intent behind proactive interactions, avoiding the problems of lack of timely response at critical moments and frequent user interruptions at inappropriate times. Based on the current behavioral intent information, it generates the user's current proactive interaction content, transforming real-time events into subsequent planning. It supports the advancement, postponement, merging, and replacement of tasks, ensuring continuity and context sensitivity in proactive behavior. Furthermore, by recording data such as user profiles, behavioral habits, emotional changes, feedback results, and task execution effectiveness, and recalling relevant information in subsequent decision-making, the robot can adaptively adjust its proactive interaction strategy, effectively enhancing the user experience. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 This is a flowchart illustrating a robot's active interaction method in one embodiment;
[0055] Figure 2 This is a flowchart illustrating an example of robot proactive interaction in one embodiment;
[0056] Figure 3 This is a structural block diagram of a robot active interaction device in one embodiment;
[0057] Figure 4This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0059] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0060] The robot proactive interaction method provided in this application can be applied to robot proactive interaction systems. This system can be applied to a terminal, which is a decision engine agent located in the robot's control center. Its role is the system's "social brain," and it is the core technical point for achieving "perceptive awareness" in this invention. Its responsibility is to ultimately decide "whether to act," "when to act," and "how to act" based on the provided "fusion context." Technically, this is an agent based on a large language model that focuses on the "fusion context" data channel. Its core decision logic is a "social interruptibility judgment" mechanism, thereby executing the aforementioned robot proactive interaction method. This decision engine agent can be, but is not limited to, various personal computers, laptops, mid-range computers, etc. The terminal integrates multimodal contextual information to identify the user's current social state, thus providing reasonable high-dimensional contextual information for proactive interaction. Then, the social decision evaluation strategy of this solution includes dynamically judging whether to execute proactive behavior, delay behavior, or inhibit behavior based on event priority, the user's current state (such as whether they are on a call or focused on work), environmental noise, and historical interaction preferences, achieving low-interference, highly adaptable proactive interaction. This solution improves the accuracy of recognizing the intent behind proactive interactions, avoiding the problems of lack of timely response at critical moments and frequent user interruptions at inappropriate times. Based on the current behavioral intent information, it generates the user's current proactive interaction content, transforming real-time events into subsequent planning. It supports the advancement, postponement, merging, and replacement of tasks, ensuring continuity and context sensitivity in proactive behavior. Furthermore, by recording data such as user profiles, behavioral habits, emotional changes, feedback results, and task execution effectiveness, and recalling relevant information in subsequent decision-making, the robot can adaptively adjust its proactive interaction strategy, effectively enhancing the user experience.
[0061] In one exemplary embodiment, such as Figure 1 As shown, a method for active robot interaction is provided. Taking the application of this method to a terminal as an example, the method includes the following steps S101 to S103. Wherein:
[0062] Step S101: Obtain the user's current fusion context information, and identify the user's current social status based on the current fusion context information.
[0063] In this embodiment, the terminal acquires real-time data from various modalities through intelligent agents corresponding to the various sensing devices, camera devices, and audio acquisition devices installed on the robot, and uses all modalities of acquired data as the current fused context information. The intelligent agent is a sensor event agent, whose responsibility is to act as the system's "digital senses," converting unstructured raw sensor data (such as images and sounds) into structured "semantic events." Technically, this intelligent agent maintains a toolkit containing a series of pre-trained models for real-time analysis: the vision module calls a real-time pose estimation algorithm (e.g., a keypoint extraction-based model) to detect "falls" or "prolonged sitting"; it calls a face recognition model to confirm user identity; and it calls an object recognition model to analyze environmental changes (e.g., "a new coffee machine"). The hearing module calls a sound event classification model to detect "crying," "coughing," and "breaking glass" sounds. The ASR module calls an automatic speech recognition model to convert user speech into text. Its output is that when a valid event is detected (e.g., the posture estimation model detects a fall for two consecutive frames), the agent generates a new structured message (e.g., a message indicating that user Xiaoming has fallen) and publishes it to a new "semantic event" data channel. The agent can also connect to the internet to obtain real-time weather, temperature, and humidity data to assist in analyzing the user's current social state. Then, the terminal identifies the user's current social state based on the current fused contextual information. This current social state is quantified as an "interruptibility score" (e.g., a 0-4 scale), where different scores characterize the user's current disruptibility level. For example, a score of 0 (uninterruptible): the visual system detects the user is "on a call," "in an intense conversation," or "emotionally distressed" (crying). A score of 2 (low-priority interruption): the user is "watching TV" or "reading." A score of 4 (disruptive): the user is "idle" or "moving around." The specific identification process will be explained in detail later.
[0064] Step S102: Based on the user's current social status, the robot's current decision score information is generated through a social decision evaluation strategy, and the robot's current action intention information is generated based on the current decision score information.
[0065] In this embodiment, the terminal generates the robot's current decision score information based on the user's current social status through a social decision evaluation strategy, and then generates the robot's current action intention information based on the current decision score information. The social decision evaluation strategy involves the user evaluating the scores of two key variables: "task priority" (from the fused context) and "interruptibility score" (from real-time vision). If the task priority is high but the interruptibility score is low (e.g., the task is important, but the user is on a call), the decision is to suppress or delay. If the task priority is lower than or equal to the interruptibility score (e.g., the task is important, and the user is idle), the decision is to execute immediately. The task priority characterizes the priority of the robot's proactive interaction with the user, and the interruptibility score characterizes the degree of proactive interaction when the user is currently being proactively interacted with. The specific evaluation process will be described in detail later. The robot's current action intent is divided into two types: immediate execution and delayed execution. These represent the interaction time points where the robot will proactively interact. Immediate execution means the robot can immediately initiate an interaction, while delayed execution indicates that the robot is temporarily suspending execution, creating pending tasks, and setting execution trigger conditions. This ensures that the robot can trigger the pending tasks and proactively interact with the user when the trigger conditions are met. This current action intent information is generated by a planning event agent (perception layer). Its responsibility is to generate "semantic events" originating from "planning-driven" interactions. Technically, this is a time- and state machine-based agent. For periodic tasks, it acts as a scheduled job, triggering periodically (e.g., every morning at 7 AM), calling external APIs (such as weather, news, user calendars) or querying the internal knowledge base, and publishing the results (e.g., "rain in the afternoon," "meeting at 10 AM") as "semantic events." For long-term observation, it acts as a state machine, subscribing to "semantic event" channels for long-term pattern analysis. For example, it aggregates multiple "sitting posture" events published by agents, internally calculates "continuous sitting time," and only generates an event identifying "prolonged sitting reminder" when a threshold of "one consecutive hour" is met. This invention establishes a crucial "many-to-one" convergence point through a "semantic event" channel. All different types of driving sources—whether from real-time sensor events (such as "falling") or long-term planning events (such as "prolonged sitting")—are standardized and abstracted into the same data structure (i.e., "semantic events"). The advantage of this design is that it completely separates "event generation" from "event decision-making." Downstream "inference" layers can work from a unified, fused event stream without needing to know the source of the event (whether it's a real-time sensor or a long-term planner). This greatly simplifies the complexity of the decision-making model and unifies the underlying technical implementations of "event-driven" and "planning-driven" approaches in the user draft.
[0066] Step S103: Based on the current action intent information, the robot generates the current active interaction content through the user memory interaction strategy, and based on the current active interaction content, the robot performs active interaction processing on the user.
[0067] In this embodiment, the terminal generates the robot's current proactive interaction content based on the current action intent information and through a user memory interaction strategy. Based on this content, the robot then initiates proactive interaction with the user. This user memory interaction strategy generates the proactive interaction content needed to interact with the user by utilizing the "long-term memory" and "world model" recorded in the knowledge base agent. This knowledge base agent provides a unique read / write interface for persistent data to all other agents in the system (implemented via internal API or RPC, not through an event bus to ensure transaction consistency). Technically, it encapsulates a relational database for structured data and a vector database for semantic memory. To make the "context-aware" and "long-term learning" technical solutions of this invention reproducible, the core data structure of the knowledge base includes: User profile storage: maintaining information for each user, such as their name, relationship with family members, personal preferences (e.g., "prefers warm water when sick"), and current health status (e.g., "has a mild cold"). Long-term memory storage: recording past interactions. This includes the timing of the interaction, the event that triggered the interaction (e.g., "coughing"), the action taken by the robot (e.g., "providing warm water"), and the user's feedback (e.g., whether the user "accepted," "refused," or "interrupted" the action). World knowledge storage: mapping the physical environment. This includes the names of objects (e.g., "kettle") and their semantic locations (e.g., "kitchen countertop"), as well as a semantic map of the room layout. To-do task storage: storing actions that have been delayed. This includes the planned intent (e.g., "providing warm water") and a trigger condition (e.g., "only when the user's state becomes idle"). The specific generation process will be described in detail later.
[0068] Based on the above scheme, by integrating multimodal contextual information, the current social state of the user is identified, thus providing reasonable high-dimensional contextual information for proactive interaction. Then, the social decision-making evaluation strategy of this scheme includes dynamically judging whether to execute proactive behavior, delay behavior, or inhibit behavior based on event priority, the user's current state (e.g., whether on a call, focused on work), environmental noise, and historical interaction preferences, achieving low-interference and highly adaptable proactive interaction. This improves the accuracy of identifying the intent of the current proactive interaction behavior, avoiding the problems of lack of timely response at critical moments and frequent disturbance to the user at inappropriate times. Then, based on the current behavioral intent information, this scheme generates the user's current proactive interaction content, which can transform real-time events into subsequent planning, supporting the advance, postponement, merging, and replacement of tasks, making proactive behavior continuous and context-sensitive. Finally, this scheme records data such as user profiles, behavioral habits, emotional changes, feedback results, and task execution effects, and calls relevant information in subsequent decisions, enabling the robot to adaptively adjust the proactive interaction strategy, so that the proactive interaction content can effectively improve the user experience.
[0069] Optionally, based on the current fusion context information, identify the user's current social state, including: splitting the current fusion context information into current modality data of each modality type; generating current user state data of each modality type based on the current modality data of each modality type through event detection models of each modality type; and generating the user's current social state based on the current user state data of each modality type through a social state evaluation strategy.
[0070] In this embodiment, the terminal breaks down the current fused context information into current modal data of various modal types. These modal types include, but are not limited to, text, image, and audio types. The event detection models for each modal type are: a semantic recognition model based on a large language model for text, an image analysis model based on a convolutional neural network for images, and an acoustic analysis model based on an acoustic feature recognition network for audio.
[0071] Then, based on the current modal data of each modality type, the terminal generates current user status data for each modality type through the event detection model of each modality type. Among them, the current user status data is the fused context information of the user's environment obtained by analyzing the current monitoring data of different model types. For example, the image type is "night, user A is looking at a computer, wearing a blanket", the sound type is "coughing sound", and the text type is "temperature: -2°—0°, room temperature: 20°, weather: strong wind, snowfall".
[0072] Then, based on the current user state data of each modality, the terminal generates the user's current social state through a social state evaluation strategy. This social state evaluation strategy involves contextual fusion of the current user state data from each modality and then analyzing the user's current quantitative interruptibility score. For example, first, the context-aware agent decodes the "cough detected" event. Next, it performs knowledge base interaction: immediately initiating parallel queries to the knowledge base agent: query (user = "Alex") -> returns "Health condition: mild cold", "Preference: likes warm water"; query (event = "cough",...) -> returns "coughed 3 times in the past hour"; query (action = "provide warm water") -> returns "last feedback accepted"; query (user = "...", status = "social") -> simultaneously, the agent's vision module continuously analyzes the video stream (through the raw sensor information channel), identifies that the user is on a call, and has published "User status: on call". This information is acquired by the agent. Finally, the agent integrates all the above information and sends a message to the "fusion context" data channel ("Who: Alex, What: Cough (cold symptoms), Why is it important: High priority, Positive past feedback, Current status: In a call"). Then, the terminal identifies the user's interruptibility score using a state quantification scoring indicator. This indicator includes keywords for each user's state data corresponding to each score value. For example, a score of 0 (uninterruptible): visual detection indicates the user is "in a call," "having an intense conversation," or "feeling depressed" (crying). A score of 2 (low priority interruption): the user is "watching TV" or "reading." A score of 4 (interruptible): the user is "idle" or "moving around." Based on all the integrated information, the terminal identifies the matching between the keywords included in all the information and the keywords corresponding to each score value using the quantification scoring indicator. The score value corresponding to all the matched information is then used as the user's interruptibility score, i.e., the user's current social state.
[0073] Based on the above scheme, by extracting and identifying current user status data of multimodal types and then performing context fusion, the current social status of users can be quantitatively analyzed, thereby improving the accuracy of identifying the current social status of users.
[0074] Optionally, based on the user's current social status, a social decision evaluation strategy is used to generate the robot's current decision score information, including: identifying the arbitration scores of each key variable based on the user's current social status; identifying the current task execution status corresponding to the current social status based on the arbitration scores of each key variable; and using the current task execution status as the robot's current decision score information.
[0075] In this embodiment, the terminal identifies the arbitration scores of key variables based on the user's current social state. These key variables include "task priority" (from the fused context) and "interruptibility score" (from real-time vision). If the task priority is high but the interruptibility score is low (e.g., the task is important, but the user is on a call), the decision is to suppress or delay. If the task priority is lower than or equal to the interruptibility score (e.g., the task is important, and the user is idle), the decision is to execute immediately. Finally, based on the arbitration scores of each key variable, the terminal identifies the current task execution state corresponding to the current social state and uses this current task execution state as the robot's current decision score information. Specifically, the terminal performs task evaluation based on the acquired current social state and all fused information. Based on "coughing" and "high priority," it generates the candidate action "provide warm water" and sets the task priority to high. For social evaluation, based on "current state: on a call," it calls the "interruptibility model" and evaluates the interruptibility score to be extremely low.
[0076] Based on the above scheme, by scoring the user's current social status and all information in the integrated context from two key variables, the accuracy of the assessment of the user's current intrusiveness and the robot's current proactive interaction is improved.
[0077] Optionally, based on the current decision scoring information, the robot's current action intention information is generated, including: querying the current intention recognition generation strategy based on the current task execution status; and generating the robot's current action intention information based on the current modal data of each modality type through the current intention recognition strategy.
[0078] In this embodiment, the terminal queries the current intent generation strategy based on the current task execution state. Based on the current modal data of each modality type, the current intent recognition strategy generates the robot's current action intent information. Specifically, the current intent recognition strategy combines the scores of two key variables to identify the robot's current action intent and the intent content. Specifically: when the two key variables corresponding to the current task execution state are high task priority and low interruptibility score, the robot's current action intent information is a delayed execution of the robot's task intent; when the two key variables corresponding to the current task execution state are medium task priority and low interruptibility score, the robot's current action intent information is a delayed execution of the robot's task intent; when the two key variables corresponding to the current task execution state are low task priority and low interruptibility score, the robot's current action intent information is a delayed execution of the robot's task intent. When the two key variables corresponding to the current task execution state are high task priority and medium interruptibility score, the robot's current action intention is to execute immediately. When the two key variables corresponding to the current task execution state are high task priority and high interruptibility score, the robot's current action intention is to execute the robot's task intention immediately. When the two key variables corresponding to the current task execution state are medium task priority and medium interruptibility score, the robot's current action intention is to execute the robot's task intention with a delay. When the two key variables corresponding to the current task execution state are medium task priority and high interruptibility score, the robot's current action intention is to execute the robot's task intention immediately. The robot's task intention is an interactive task generated for the user through a large language model combined with contextual information.
[0079] For example, based on the current decision score information of user Alex obtained in the example above, i.e., the task priority is high and the interruptibility score is extremely low, then, because the task priority (high) is greater than the interruptibility score (extremely low), but the cost of interruption is too high, the decision is to delay execution (i.e., the robot's current action intention information). Finally, the output result: the agent does not send instructions to the executor, but writes a "to-do task" ("Task Intent: Provide warm water, Target User: Alex, Trigger Condition: User state becomes idle") to the to-do task storage of the knowledge base agent. The time of this delayed execution is T4, T4>T1+10 minutes. The triggering method of this delayed execution is that the terminal continuously collects the user's current action content through the camera device until it detects that the user hangs up the phone and waves goodbye. Then, the sensor event agent recognizes the state change and publishes ("Type: User state, User: Alex, State: Idle") to the "Semantic Event" data channel. Then, the decision engine agent listens for the "User State: Idle" event. This "Idle" state successfully matches the trigger condition in the "to-do task" storage. Finally, the output is: the decision engine releases the task and issues an "action intention" instruction ("execution intention: provide warm water") to the task execution agent.
[0080] Based on the above scheme, by combining the current task execution status, the robot's current decision-making scoring information can be quickly identified, thereby improving the efficiency of robot decision-making scoring.
[0081] Optionally, based on the current action intent information, the robot's current proactive interaction content is generated through a user memory interaction strategy, including: obtaining the user's current user profile and the user's historical interaction content; based on the user's current user profile, the user's historical interaction content, and the current action intent information, generating the robot's current task planning content through a task decomposition strategy; and based on the robot's current task planning content, generating the robot's current proactive interaction content according to a language interaction model.
[0082] In this embodiment, the terminal acquires the user's current user profile and the user's historical interaction content. Based on the user's current user profile, historical interaction content, and current action intent information, the terminal generates the robot's current task planning content through a task decomposition strategy. The task decomposition strategy involves the terminal performing task planning based on the user's current user profile, historical interaction content, and the aforementioned identified current action intent information, using a large language model. When the current action intent information is to immediately execute the intent to "provide warm water," the terminal accesses the "skill base" and queries the knowledge base agent for "world knowledge" to obtain the object's location ("cup: kitchen cabinet 1", "kettle: kitchen countertop"). Then, a task plan is generated: the large language model generates an ordered sequence of skills as the task plan: "Navigation ability" is invoked, target: "kitchen countertop"; "Visual search ability" is invoked, target: "cup"; "Manipulation ability" is invoked, action: "pick up"; object: "cup"; "Manipulation ability" is invoked, action: "fill"; substance: "warm water"; "Navigation ability" is invoked, target: "user Alex"; "Speech ability" is invoked, text: "Alex, I noticed you coughed several times just now. Your call is over now, drinking some warm water might make you feel better." Subsequently, on the event bus, the agent publishes this task plan sequence (or the first instruction after decomposition) to the action result data channel. Finally, the device consumes the instruction and executes the physical actions (navigation, grasping, speech) in sequence.
[0083] Based on the above scheme, by decomposing the task intent with the current action intent information, the current active interaction content of the robot is generated, thereby improving the efficiency of the robot's active interaction and the user's active interaction experience with the robot.
[0084] Optionally, based on the current active interaction content, the robot performs active interaction processing with the user, including: based on the robot's current active interaction content, identifying the current active interaction task sequence of each functional component of the robot, and controlling each functional component to actively interact with the user respectively through the current active interaction content of each functional component, obtaining the user's interaction feedback content; based on the user's interaction feedback content, performing interaction adjustment processing on the current interactive content to obtain new active interaction content, replacing the current active interaction content with the new active interaction content, and returning to execute the step of identifying the current active interaction task sequence of each functional component of the robot based on the robot's current active interaction content.
[0085] In this embodiment, the terminal identifies the current active interaction task sequence of each functional component of the robot based on the robot's current active interaction content. This current active interaction task sequence generates an ordered skill sequence for a large language model. For example, calling "navigation ability" with the target "kitchen countertop", calling "visual search ability" with the target "cup", calling "manipulation ability" with the action "pick up" and the object "cup", calling "manipulation ability" with the action "fill" and the substance "warm water", calling "navigation ability" with the target "user Alex", and calling "voice ability" with the text "Alex, I noticed you coughed several times just now. Your call is over now, drinking some warm water might make you feel better."
[0086] Then, the terminal controls each functional component to actively interact with the user based on the current active interaction content of each component, obtaining the user's interactive feedback. These functional components include, but are not limited to, language communication components (mouth) and motion support components (limbs). This interactive feedback content is the user's feedback information. Based on this feedback, the terminal uses a large language model to identify the user's feedback direction for the current active interaction, recording it in the knowledge base agent. This allows for adjustments when subsequent active interactions are triggered. Specifically, the terminal adjusts the current interactive content based on the user's feedback, obtaining new active interaction content, replacing the current active interaction content, and returning to execute the robot's current active interaction content, identifying the current active interaction task sequence steps of each functional component of the robot. For example, Alex on the device accepts water and says "thank you." The microphone captures the speech. Then, the Sensor Event Agent (ASR module) converts the speech into the text "thank you." Then, the Context Aware Agent (NLU module) interprets "thank you" as "user feedback: positive." Subsequently, the knowledge base agent records the "user feedback" for this interaction ("triggering event: cough") as "positive" in long-term memory (see above). Finally, the long-term impact is that this new record will increase the confidence of the next "cough" -> "water delivery" decision, achieving adaptive learning for the system.
[0087] Based on the above solution, by combining user feedback, the robot's subsequent actions are intelligently adjusted, thereby iteratively optimizing and improving the user experience when the robot actively interacts.
[0088] This application also provides an example of robot proactive interaction, such as... Figure 2 As shown, the specific processing procedure includes the following steps:
[0089] Step S201: Obtain the user's current fusion scenario information.
[0090] Step S202: The current fused scenario information is split into current modal data of each modal type.
[0091] Step S203: Based on the current modal data of each modal type, generate the current user state data of each modal type through the event detection model of each modal type.
[0092] Step S204: Based on the current user status data of each modality type, generate the user's current social status through a social status evaluation strategy.
[0093] Step S205: Based on the user's current social status, identify the arbitration scores of each key variable.
[0094] Step S206: Based on the arbitration scores of each key variable, identify the current task execution status corresponding to the current social status.
[0095] Step S207: Use the current task execution status as the robot's current decision scoring information.
[0096] Step S208: Based on the current task execution status, query the current intent generation strategy.
[0097] Step S209: Based on the current modal data of each modal type, the robot's current action intention information is generated through the current intention recognition strategy.
[0098] Step S210: Obtain the user's current user profile and the user's historical interaction content.
[0099] Step S211: Based on the user's current user profile, the user's historical interaction content, and current action intent information, generate the robot's current task planning content through a task decomposition strategy.
[0100] Step S212: Based on the robot's current task planning content, generate the robot's current active interaction content according to the language interaction model.
[0101] Step S213: Based on the robot's current active interaction content, identify the current active interaction task sequence of each functional component of the robot, and control each functional component to actively interact with the user through the current active interaction content of each functional component, so as to obtain the user's interaction feedback content.
[0102] Step S214: Based on the user's interactive feedback, adjust the current interactive content to obtain new proactive interactive content, replace the current proactive interactive content with the new proactive interactive content, and return to step S212.
[0103] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0104] Based on the same inventive concept, this application also provides a robot active interaction device for implementing the robot active interaction method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more robot active interaction device embodiments provided below can be found in the limitations of the robot active interaction method described above, and will not be repeated here.
[0105] In one exemplary embodiment, such as Figure 3 As shown, a robot active interaction device is provided, including: an acquisition module 310, a generation module 320, and an interaction module 330, wherein:
[0106] The acquisition module 310 is used to acquire the user's current fusion context information and, based on the current fusion context information, identify the user's current social status;
[0107] The generation module 320 is used to generate the robot's current decision score information based on the user's current social status and through a social decision evaluation strategy, and to generate the robot's current action intention information based on the current decision score information;
[0108] The interaction module 330 is used to generate the robot's current active interaction content based on the current action intent information and through the user's memory interaction strategy, and to perform active interaction processing on the user through the robot based on the current active interaction content.
[0109] Optionally, the acquisition module 310 is specifically used for:
[0110] The current fusion scenario information is broken down into current modal data for each modality type;
[0111] Based on the current modality data of each modality type, the current user state data of each modality type is generated through the event detection model of each modality type;
[0112] Based on the current user status data for each modality type, the user's current social status is generated through a social status evaluation strategy.
[0113] Optionally, the generation module 320 is specifically used for:
[0114] Based on the user's current social status, identify the arbitration scores for each key variable;
[0115] Based on the arbitration scores of each of the key variables, the current task execution status corresponding to the current social status is identified;
[0116] The current task execution status is used as the robot's current decision scoring information.
[0117] Optionally, the generation module 320 is specifically used for:
[0118] Based on the current task execution status, query the current intent generation strategy;
[0119] Based on the current modal data of each modal type, the robot's current action intention information is generated through the current intention recognition strategy.
[0120] Optionally, the interaction module 330 is specifically used for:
[0121] Obtain the user's current user profile and the user's historical interaction content;
[0122] Based on the user's current user profile, the user's historical interaction content, and the current action intent information, the robot's current task planning content is generated through a task decomposition strategy.
[0123] Based on the robot's current task planning content, the robot's current active interaction content is generated according to the language interaction model.
[0124] Optionally, the interaction module 330 is specifically used for:
[0125] Based on the robot's current active interaction content, the current active interaction task sequence of each functional component of the robot is identified, and each functional component is controlled to actively interact with the user through the current active interaction content of each functional component, so as to obtain the user's interaction feedback content.
[0126] Based on the user's interactive feedback, the current interactive content is adjusted to obtain new proactive interactive content. The new proactive interactive content replaces the current proactive interactive content, and the process returns to execute the steps of identifying the current proactive interactive task sequence of each functional component of the robot based on the robot's current proactive interactive content.
[0127] Each module in the aforementioned robot active interaction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0128] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for active robot interaction. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0129] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0130] In one exemplary embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of a beer warehouse inventory optimization method.
[0131] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of a beer warehouse inventory optimization method.
[0132] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of a beer warehouse inventory optimization method.
[0133] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0134] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0135] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0136] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for active robot interaction, characterized in that, The method includes: Obtain the user's current context information and, based on the current context information, identify the user's current social status; Based on the user's current social status, a social decision evaluation strategy is used to generate the robot's current decision score information, and based on the current decision score information, the robot's current action intention information is generated. Based on the current action intent information, the robot generates the current active interaction content through a user memory interaction strategy, and based on the current active interaction content, the robot performs active interaction processing on the user.
2. The method according to claim 1, characterized in that, The step of identifying the user's current social status based on the current fusion context information includes: The current fusion scenario information is broken down into current modal data for each modality type; Based on the current modality data of each modality type, the current user state data of each modality type is generated through the event detection model of each modality type; Based on the current user status data for each modality type, the user's current social status is generated through a social status evaluation strategy.
3. The method according to claim 1, characterized in that, The step of generating the robot's current decision score information based on the user's current social status through a social decision evaluation strategy includes: Based on the user's current social status, identify the arbitration scores for each key variable; Based on the arbitration scores of each of the key variables, the current task execution status corresponding to the current social status is identified; The current task execution status is used as the robot's current decision scoring information.
4. The method according to claim 3, characterized in that, The step of generating the robot's current action intention information based on the current decision scoring information includes: Based on the current task execution status, query the current intent generation strategy; Based on the current modal data of each modal type, the robot's current action intention information is generated through the current intention recognition strategy.
5. The method according to claim 1, characterized in that, The step of generating the robot's current active interaction content based on the current action intent information and through a user memory interaction strategy includes: Obtain the user's current user profile and the user's historical interaction content; Based on the user's current user profile, the user's historical interaction content, and the current action intent information, the robot's current task planning content is generated through a task decomposition strategy. Based on the robot's current task planning content, the robot's current active interaction content is generated according to the language interaction model.
6. The method according to claim 1, characterized in that, The step of actively interacting with the user through the robot based on the current active interaction content includes: Based on the robot's current active interaction content, the current active interaction task sequence of each functional component of the robot is identified, and each functional component is controlled to actively interact with the user through the current active interaction content of each functional component, so as to obtain the user's interaction feedback content. Based on the user's interactive feedback, the current interactive content is adjusted to obtain new proactive interactive content. The new proactive interactive content replaces the current proactive interactive content, and the process returns to execute the steps of identifying the current proactive interactive task sequence of each functional component of the robot based on the robot's current proactive interactive content.
7. A robot active interaction device, characterized in that, The device includes: The acquisition module is used to acquire the user's current fusion context information and, based on the current fusion context information, identify the user's current social status; The generation module is used to generate the robot's current decision score information based on the user's current social status and through a social decision evaluation strategy, and to generate the robot's current action intention information based on the current decision score information; The interaction module is used to generate the robot's current active interaction content based on the current action intent information and through the user's memory interaction strategy, and to perform active interaction processing on the user through the robot based on the current active interaction content.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.