Methods, devices, media, and program products for interaction

By dynamically cloning intelligent agents based on perceived information, robots can achieve personalized and adaptive interaction in multi-person interaction scenarios, solving the problem of inflexible humanoid robot interaction in existing technologies and improving the interaction experience.

CN122195252APending Publication Date: 2026-06-12SHANGHAI MATRIX SUPER INTELLIGENT SYSTEM INTEGRATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI MATRIX SUPER INTELLIGENT SYSTEM INTEGRATION CO LTD
Filing Date
2026-03-03
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

In existing technologies, humanoid robots lack flexibility and personalized adaptability, making it impossible to accurately distinguish the needs of different users in multi-user interaction scenarios, resulting in a poor interactive experience.

Method used

By dynamically cloning and adapting intelligent agents based on perceived information, the robot responds to interaction requests and generates personalized intelligent agents using visual and audio perception information, enabling personalized interaction for each user and supporting differentiated feedback and adaptive interaction in multi-user scenarios.

Benefits of technology

It significantly improves the robot's interactive experience in multi-person interaction scenarios, enabling personalized adaptive interaction. It can flexibly adjust the interaction method according to changes in the scenario and provide accurate interactive feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122195252A_ABST
    Figure CN122195252A_ABST
Patent Text Reader

Abstract

The purpose of the present application is to provide a method, device, medium and program product for interaction, the method comprising: in response to a wake-up interaction request initiated by a first interaction subject, obtaining perception information corresponding to the wake-up interaction request; cloning a target agent from a default agent as a template, determining a matching personalized capability according to the perception information, modifying the target agent according to the personalized capability, obtaining a modified first agent, and binding the first agent to the first interaction subject, wherein the target agent corresponds to the same basic capability as the default agent; starting the first agent, so that the robot interacts with the first interaction subject based on the first agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics, and more particularly to a technology for interaction. Background Technology

[0002] In existing technologies, humanoid robots often use near-field acoustic hardware and front-end acoustic solutions, resulting in a poor human-computer interaction experience. Humanoid robots typically achieve simple voice interaction through pre-programmed sounds and movements, lacking the ability to generalize to different scenarios. They often only support sound perception and voice interaction capabilities, frequently lacking visual perception and therefore unable to support human gesture interaction capabilities or autonomous finger manipulation. Furthermore, humanoid robots often have a pre-set single agent and skill list, unable to adapt in real-time based on the current environment and the interacting subject, leading to a lack of flexibility and personalized adaptability in multi-person interactions. To adapt to new scenarios, it is necessary to manually switch agents and re-bind the robot to a new agent. Summary of the Invention

[0003] One object of this application is to provide a method, device, medium, and program product for interaction.

[0004] According to one aspect of this application, a method for interaction is provided, applied to a robot, the method comprising:

[0005] In response to a wake-up interaction request initiated by the first interactive subject, obtain the perception information corresponding to the wake-up interaction request;

[0006] A target intelligent agent is cloned using a default intelligent agent as a template. Matching personalized capabilities are determined based on the perception information. The target intelligent agent is modified based on the personalized capabilities to obtain a modified first intelligent agent. The first intelligent agent is then bound to the first interactive subject. The target intelligent agent and the default intelligent agent have the same basic capabilities.

[0007] The first intelligent agent is activated, enabling the robot to interact with the first interactive subject based on the first intelligent agent.

[0008] According to one aspect of this application, a computer device for outputting a sequence of action instructions is provided, comprising a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of any of the methods described above.

[0009] According to one aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of any of the methods described above.

[0010] According to one aspect of this application, a computer program product is provided, comprising a computer program, characterized in that, when executed by a processor, the computer program implements the steps of any of the methods described above.

[0011] According to one aspect of this application, a robot for interaction is provided, the robot comprising:

[0012] The module is used to respond to a wake-up interaction request initiated by the first interactive subject and obtain the perception information corresponding to the wake-up interaction request.

[0013] The first and second modules are used to clone a target intelligent agent using a default intelligent agent as a template, determine matching personalized capabilities based on the perception information, modify the target intelligent agent based on the personalized capabilities to obtain a modified first intelligent agent, and bind the first intelligent agent to the first interactive subject, wherein the target intelligent agent and the default intelligent agent have the same basic capabilities.

[0014] The first and third modules are used to activate the first intelligent agent, enabling the robot to interact with the first interactive subject based on the first intelligent agent.

[0015] Compared with existing technologies, this application constructs an intelligent interaction scheme for robots in multi-user interaction scenarios based on perceptual information and dynamic intelligent agent technology. The robot's human-computer interaction process is triggered by a wake-up interaction request initiated by the interaction subject. Perceptual information comprehensively captures scene information (including environmental information and the characteristic information of the interaction subject). A new intelligent agent adapted to the perceptual information is dynamically cloned using a default intelligent agent as a template, and then the new intelligent agent interacts with the interaction subject. This enables the robot to have personalized interaction capabilities tailored to each user. The robot can flexibly adjust its interaction methods according to scene changes, significantly improving its scene generalization ability and making the interaction more aligned with actual scene needs. By dynamically switching intelligent agents in real time, the robot can accurately distinguish the needs of different users in multi-user interaction scenarios, providing differentiated interactive feedback, achieving personalized adaptive interaction, and realizing a unique interaction effect for each user, greatly improving the interactive experience in multi-user scenarios. Attached Figure Description

[0016] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0017] Figure 1 This diagram illustrates a method flowchart for interaction according to an embodiment of the present application;

[0018] Figure 2 This diagram illustrates a device structure diagram of an interactive robot according to an embodiment of this application;

[0019] Figure 3 Exemplary systems that can be used to implement the various embodiments described in this application are shown.

[0020] The same or similar reference numerals in the accompanying drawings represent the same or similar parts. Detailed Implementation

[0021] The present application will now be described in further detail with reference to the accompanying drawings.

[0022] In a typical configuration of this application, the terminal, the device of the service network, and the trusted party all include one or more processors (e.g., a central processing unit (CPU)), input / output interfaces, network interfaces, and memory.

[0023] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory. Memory is an example of computer-readable media.

[0024] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PCM), programmable random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0025] The devices referred to in this application include, but are not limited to, user equipment, network equipment, or devices composed of user equipment and network equipment integrated through a network. The user equipment includes, but is not limited to, any mobile electronic product capable of human-computer interaction (e.g., via a touchpad), such as smartphones and tablets. These mobile electronic products can use any operating system, such as Android or iOS. The network equipment includes an electronic device capable of automatically performing numerical calculations and information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and embedded devices. The network equipment includes, but is not limited to, computers, network hosts, single network servers, multiple network server clusters, or clouds composed of multiple servers. Here, a cloud consists of a large number of computers or network servers based on cloud computing, where cloud computing is a type of distributed computing, consisting of a virtual supercomputer composed of a group of loosely coupled computer clusters. The network includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, VPN network, wireless ad hoc network, etc. Preferably, the device can also be a program running on the user equipment, network device, or a device formed by integrating user equipment and network device, network device, touch terminal, or network device and touch terminal through a network.

[0026] Of course, those skilled in the art should understand that the above-described devices are merely examples, and other existing or future devices that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.

[0027] In the description of this application, "multiple" means two or more, unless otherwise expressly and specifically defined.

[0028] All data collected and processed in this application have been with the user's consent or permission, and have strictly complied with legal regulations, social ethics, and the public interest. This data processing includes, but is not limited to, tag management, rule setting, and recommendation decisions. The legal regulations mentioned include, but are not limited to: 1) relevant laws and regulations of various countries or organizations regarding the protection of personal information; 2) relevant laws and regulations of various countries or organizations regarding the protection of user information; and 3) relevant laws and regulations of various countries or organizations regarding the use of personal or user information by organizations.

[0029] Figure 1 The diagram illustrates a flowchart of an interaction method for a robot according to an embodiment of this application. The method includes steps S11, S12, and S13. In step S11, the robot responds to a wake-up interaction request initiated by a first interaction subject and obtains perception information corresponding to the wake-up interaction request. In step S12, the robot clones a target intelligent agent using a default intelligent agent as a template, determines matching personalized capabilities based on the perception information, modifies the target intelligent agent according to the personalized capabilities, obtains a modified first intelligent agent, and binds the first intelligent agent to the first interaction subject. The target intelligent agent and the default intelligent agent correspond to the same basic capabilities. In step S13, the robot activates the first intelligent agent, enabling the robot to interact with the first interaction subject based on the first intelligent agent.

[0030] In step S11, the robot responds to the wake-up interaction request initiated by the first interactive subject and obtains the perception information corresponding to the wake-up interaction request.

[0031] In some embodiments, the robot includes, but is not limited to, humanoid robots, bionic robots, wheeled robots, tracked robots, articulated robots, mobile operation robots, etc., and this example embodiment does not make any special limitation in this regard.

[0032] In some embodiments, the interaction subject includes, but is not limited to, any subject capable of interacting with the robot, such as a person, an animal, or other robots. This example embodiment does not impose any special limitations on this.

[0033] In some embodiments, the wake-up interaction request may be initiated by the first interactive subject via voice. The robot's microphone array (audio sensor) captures the voice information of the first interactive subject and identifies the wake-up interaction request from the voice information. In some embodiments, the wake-up interaction request may also be initiated by the first interactive subject via gesture. The wake-up interaction request is obtained by gesture recognition of the real-world image containing the first interactive subject captured by the robot's camera (visual sensor). In some embodiments, the robot only supports gesture wake-up to avoid accidental operation.

[0034] In some embodiments, the perception information includes at least one of visual perception information, environmental perception information, and audio perception information. The visual perception information is information associated with the interactive subject obtained from video information captured by the robot's camera. The environmental perception information is information associated with the robot's current environment obtained from video information captured by the robot's camera. The audio perception information is information associated with the interactive subject obtained from audio information captured by the robot's microphone array.

[0035] In step S12, the robot clones the target intelligent agent using the default intelligent agent as a template, determines the matching personalized capabilities based on the perception information, modifies the target intelligent agent according to the personalized capabilities, obtains the modified first intelligent agent, and binds the first intelligent agent to the first interactive subject. The target intelligent agent and the default intelligent agent have the same basic capabilities.

[0036] In some embodiments, personalized capabilities are determined based on perceived information. Personalized capabilities include at least one of a knowledge base, prompts, workflows, and embodied skills. The knowledge base is a structured / semi-structured data set that provides the agent with full-dimensional knowledge support. It is the "cognitive foundation" for the agent to understand the world and make reasoning decisions, and it is also the core carrier for avoiding repetitive learning and realizing knowledge reuse. Prompts are structured instructions that transmit task objectives, constraints, execution requirements, and scene contexts to the humanoid robot agent. They are the "communication interface" between humans and the agent, and they are also the "command signals" that guide the agent to retrieve knowledge from the knowledge base and plan workflows. Workflows are the "task execution framework" in which the agent breaks down complex tasks into a series of ordered and executable sub-tasks to complete the target task, and defines the execution logic, triggering conditions, connection methods, and exception handling rules of the sub-tasks. It is the core of the agent's decision-making and planning from "goal" to "steps". Embodied skills are a modular and reusable set of capabilities that the robot can use to perform specific basic actions / operations through physical execution. They are the final carrier for the agent's workflow to be implemented as actual physical behavior. In some embodiments, the robot's intelligent agent (i.e., artificial intelligence agent, AI Agent) is the core intelligence core that enables the robot to achieve autonomous perception, decision-making, action and interaction. In essence, it is an artificial intelligence system that integrates perception, cognition, planning, execution and learning capabilities. The intelligent agent transforms the robot from a "programmed mechanical actuator" into an "autonomous intelligent agent" that can understand the environment, make autonomous judgments and flexibly respond to complex scenarios.

[0037] In some embodiments, a target intelligent agent is cloned using a preset default intelligent agent as a template, so that the target intelligent agent and the default intelligent agent have the same basic capabilities. These basic capabilities include, but are not limited to, persona, language style, voice, and core functions, to avoid sudden changes in robot behavior. In some embodiments, the cloned target intelligent agent is modified according to the determined personalized capabilities, so that the modified first intelligent agent has the personalized capabilities, and then the first intelligent agent is bound to the first interaction subject. In this application, each interaction subject corresponds to a dedicated dynamic intelligent agent, and its personalized capabilities are matched to the wake-up interaction request, such as matching the attribute information of the interaction subject and / or matching the current scene type of the robot, to achieve personalized interaction, i.e., a unique interaction for each user.

[0038] In step S13, the robot activates the first intelligent agent, enabling the robot to interact with the first interactive subject based on the first intelligent agent.

[0039] In some embodiments, the first agent is activated, which makes the first agent effective, and then enables the robot to interact with the first interaction subject based on the first agent.

[0040] This application constructs an intelligent interaction scheme for robots in multi-user interaction scenarios based on perceptual information and dynamic intelligent agent technology. The system triggers the robot's human-computer interaction process by initiating a wake-up interaction request from the interacting subject. It comprehensively captures scene information (including environmental information and the characteristic information of the interacting subject) through perceptual information, dynamically clones a new intelligent agent adapted to the perceptual information using a default intelligent agent as a template, and initiates the interaction between the new intelligent agent and the interacting subject. This enables the robot to possess personalized interaction capabilities tailored to each user. The robot can flexibly adjust its interaction methods according to scene changes, significantly improving its scene generalization ability and making the interaction more aligned with actual scene needs. By dynamically switching intelligent agents in real time, the robot can accurately distinguish the needs of different users in multi-user interaction scenarios, providing differentiated interactive feedback, achieving personalized adaptive interaction, and realizing a unique interaction effect for each user, greatly improving the interactive experience in multi-user scenarios.

[0041] In some embodiments, the method further includes: identifying a hand gesture performed by a first interactive subject in a real-world scene captured by the robot's camera to obtain a wake-up interaction request initiated by the first interactive subject. In some embodiments, a real-world scene containing the first interactive subject is captured by the robot's camera, and the hand gesture performed by the first interactive subject in the real-world scene is identified. If the similarity between the hand gesture and at least one preset hand gesture is determined to be greater than or equal to a preset similarity threshold based on the identification result, it can be determined that the first interactive subject has initiated a wake-up interaction request.

[0042] In some embodiments, the perceived information includes at least one of visual perceived information, environmental perceived information, and audio perceived information. In some embodiments, the visual perceived information is information associated with the interactive subject obtained from video information captured by the robot's camera; the environmental perceived information is information associated with the robot's current environment obtained from video information captured by the robot's camera; and the audio perceived information is information associated with the interactive subject obtained from audio information captured by the robot's microphone array.

[0043] In some embodiments, the personalized capabilities include at least one of a knowledge base, prompts, workflows, and embodied skills. In some embodiments, the knowledge base provides the agent with a structured / semi-structured data set that supports all dimensions of knowledge. It serves as the "cognitive foundation" for the agent to understand the world and make reasoning decisions, and is also the core carrier for avoiding repetitive learning and achieving knowledge reuse. Prompts are structured instructions that transmit task objectives, constraints, execution requirements, and scene contexts to the humanoid robot agent. They are the "communication interface" between humans and the agent, and also the "command signal" that guides the agent to retrieve knowledge from the knowledge base and plan workflows. Workflows are the "task execution framework" by which the agent breaks down complex tasks into a series of ordered, executable sub-tasks to complete the target task, defining the execution logic, triggering conditions, connection methods, and exception handling rules of the sub-tasks. It is the core of the agent's decision-making and planning from "goal" to "steps." Embodied skills are a modular and reusable set of capabilities that allow the robot to complete specific basic actions / operations through physical execution. They are the final carrier for the agent's workflow to be implemented as actual physical behavior.

[0044] In some embodiments, the perceived information includes visual perceived information, and the personalization capability includes prompt information; wherein, determining the matching personalization capability based on the perceived information includes: determining prompt information that matches the visual perceived information based on the visual perceived information. In some embodiments, at least one visual feature associated with the first interactive subject can be obtained by performing feature extraction on the visual perceived information, and then selecting a preset prompt information that matches the at least one visual feature from a plurality of preset prompt information as the final prompt information, or generating a new prompt information that matches the at least one visual feature as the final prompt information.

[0045] In some embodiments, determining the prompt information matching the visual perception information includes: determining attribute information corresponding to the first interactive subject based on the visual perception information, and determining prompt information matching the attribute information. In some embodiments, determining the attribute information corresponding to the first interactive subject based on the visual perception information includes, but is not limited to, age, gender, occupation, etc. This example embodiment does not impose any special limitations on this; for example, the visual perception information can be input into a trained attribute prediction model to obtain the attribute information output by the model. In some embodiments, a preset prompt information matching the attribute information can be selected from multiple preset prompt information as the final prompt information, or a new prompt information matching the attribute information can be generated as the final prompt information; for example, a prompt information using a concise tone might be used for elderly individuals.

[0046] In some embodiments, the perceived information includes environmental perception information, and the personalized skill includes a workflow; wherein, determining the matching personalized ability based on the perceived information includes: determining a workflow that matches the environmental perception information. In some embodiments, at least one environmental feature associated with the robot's surrounding environment can be obtained by extracting features about the surrounding environment from the environmental perception information, and then a preset workflow that matches the at least one environmental feature can be selected as the final workflow from a plurality of preset workflows, or a new workflow that matches the at least one environmental feature can be generated as the final workflow.

[0047] In some embodiments, determining a workflow matching the environmental perception information includes: determining the current scene type of the robot based on the environmental perception information, and determining a workflow matching the scene type. In some embodiments, determining the current scene type based on the environmental perception information may include, but is not limited to, outdoor, exhibition hall, office, and home environments. This example embodiment does not impose any specific limitations on this. For example, the environmental perception information can be input into a trained scene prediction model to obtain the scene type output by the model. In some embodiments, a preset workflow matching the scene type can be selected from multiple preset workflows as the final workflow, or a new workflow matching the scene type can be generated as the final workflow. For example, a home scene might correspond to a workflow that prioritizes responding to life service instructions.

[0048] In some embodiments, the perceived information includes visual perceived information and environmental perceived information, and the personalization capability includes a knowledge base; wherein, determining the matching personalization capability based on the perceived information includes: determining a matching knowledge base based on the visual perceived information and the environmental perceived information. In some embodiments, feature extraction of the visual perceived information regarding the first interactive subject can be performed to obtain at least one visual feature associated with the first interactive subject, and feature extraction of the environmental perceived information regarding the surrounding environment can be performed to obtain at least one environmental feature associated with the robot's surrounding environment. The specific method has been detailed above and will not be repeated here. Then, a preset knowledge base that matches the at least one visual feature and the at least one environmental feature is selected from multiple preset knowledge bases as the final knowledge base, or a new knowledge base that matches the at least one visual feature and the at least one environmental feature is generated as the final knowledge base.

[0049] In some embodiments, determining a matching knowledge base based on the visual perception information and the environmental perception information includes: determining attribute information corresponding to the first interactive subject based on the visual perception information, and / or determining the scene type currently in which the robot is located based on the environmental perception information; and determining a matching knowledge base based on the attribute information and / or the scene type. In some embodiments, the specific methods for determining attribute information corresponding to the first interactive subject based on the visual perception information and / or determining the scene type currently in which the robot is located have been detailed above and will not be repeated here. In some embodiments, a preset knowledge base matching the attribute information and / or the scene type is selected from multiple preset knowledge bases as the final knowledge base, or a new knowledge base matching the attribute information and / or the scene type is generated as the final knowledge base, for example, a knowledge base corresponding to the exhibits in an exhibition hall scene.

[0050] In some embodiments, the perception information includes environmental perception information, and the personalization capability includes embodied skills; wherein, modifying the target agent based on the perception information to obtain a modified first agent includes: determining an embodied skill matching the environmental perception information based on the environmental perception information. In some embodiments, at least one environmental feature associated with the robot's surrounding environment can be obtained by extracting features about the surrounding environment from the environmental perception information, and then selecting a preset embodied skill matching the at least one environmental feature from a plurality of preset embodied skills as the final embodied skill, or generating a new embodied skill matching the at least one environmental feature as the final embodied skill.

[0051] In some embodiments, determining the embodied skill matching the environmental perception information includes: determining the current scene type of the robot based on the environmental perception information, and determining the embodied skill matching the scene type. In some embodiments, the specific method for determining the current scene type of the robot based on the environmental perception information has been detailed above and will not be repeated here. In some embodiments, a preset embodied skill matching the scene type is selected from multiple preset embodied skills as the final embodied skill, or a new embodied skill matching the scene type is generated as the final embodied skill; for example, an exhibition hall scene corresponds to the skill "exhibit explanation + directions".

[0052] In some embodiments, the perception information includes visual perception information and / or audio perception information; wherein, the method further includes: determining the position information of the first interactive subject based on the visual perception information and / or the audio perception information; causing the robot to turn towards the first interactive subject based on the position information. In some embodiments, feature extraction of the spatial position of the first interactive subject can be performed based on the visual perception information to obtain the position information of the first interactive subject (e.g., coordinate positioning). In some embodiments, based on the position information of the first interactive subject obtained based on the visual perception information, sound source localization of the first interactive subject can be performed based on the audio perception information, and the position information of the first interactive subject can be determined based on the localization result. In some embodiments, after the first intelligent agent is activated, the robot's head is turned towards the first interactive subject based on the position information of the first interactive subject, and after the turning is completed, interaction with the first interactive subject begins based on the first intelligent agent. In this application, the robot can quickly turn towards the current interactive subject and switch to the corresponding intelligent agent, supporting multi-person alternating interaction without delay, i.e., seamless switching between multiple users.

[0053] In some embodiments, the method further includes: in response to the activation of a second intelligent agent bound to a second interactive subject, placing the first intelligent agent in the background for standby, and pausing the robot's interaction with the first interactive subject. In some embodiments, after the first intelligent agent is activated, in response to a wake-up interaction request initiated by the second interactive subject, the second intelligent agent is bound to the second interactive subject and activated. The robot then begins interaction between the second interactive subjects based on the second intelligent agent, wherein the method of obtaining the second intelligent agent is the same as or similar to the method of obtaining the first intelligent agent described above. In some embodiments, the robot's head also turns from the first interactive subject bound to the first intelligent agent to the second interactive subject bound to the second intelligent agent. In some embodiments, in response to the activation of the second intelligent agent, the first intelligent agent bound to the first interactive subject is placed in the background for standby, rather than being deactivated; even if the first intelligent agent temporarily fails, the robot will pause its interaction with the first interactive subject.

[0054] In some embodiments, the method further includes: causing the second agent to copy the context information corresponding to the first agent. In some embodiments, the second agent also copies the context information of the first agent to maintain the continuity of the interaction. Context information refers to the set of all related information that the agent needs to acquire, store, reason about, and update in real time throughout the entire process of perception, decision-making, and execution of robot actions. Its core function is to enable the agent to understand "the current environment, its own state, task objectives, and interaction history," thereby making reasonable decisions that conform to the robot's physical characteristics and scenario requirements. Context information includes, but is not limited to, environmental perception information and the robot's current state. For example, context information includes the robot's perception that the interaction subject in front is a leader, the environment is a booth, the current state includes the leader's question, and the robot's recent response information. The agents in this application are all cloned from preset default agents, with only personalized modifications and adaptations, and the context information of the current agent is copied to avoid the robot exhibiting sudden behavioral changes in different user interactions, thus achieving behavioral consistency. For example, when the robot's head switches from the interaction subject "mother" to "child," the dialogue context is maintained, but the tone becomes more amiable.

[0055] In some embodiments, the method further includes: deleting the first intelligent agent in response to the first interactive subject leaving the robot's field of vision, and ending the robot's interaction with the first interactive subject. In some embodiments, if any part of the first interactive subject does not appear in the real-world image obtained by the robot's camera, it can be determined that the first interactive subject has left the robot's field of vision; or, if the duration for which any part of the first interactive subject does not appear in the real-world image obtained by the robot's camera is greater than or equal to a preset duration threshold, it can be determined that the first interactive subject has left the robot's field of vision. In some embodiments, in response to the first interactive subject leaving the robot's field of vision, the first intelligent agent bound to the first interactive subject is deleted, and the interaction between the robot and the first interactive subject ends. This application only terminates the corresponding intelligent agent when the interactive subject completely leaves the field of vision, preventing interaction interruption caused by brief occlusion. This application realizes real-time switching of intelligent agents, and the intelligent agent automatically terminates after the interactive subject leaves the field, which can ensure the continuity and personalization of interaction, avoid redundant resource occupation, and ensure the stable operation of the robot for a long time.

[0056] In some embodiments, the method further includes: activating a third agent waiting in the background, causing the robot to resume interaction with a third interactive subject bound to the third agent. In some embodiments, the third agent waiting in the background is also activated, causing the third agent to become active again, causing the robot to resume interaction with the third interactive subject bound to the third agent. In some embodiments, if there are multiple agents waiting in the background besides the first agent, an agent can be randomly selected from the multiple agents as the third agent, or the agent that has been waiting most recently (i.e., has the shortest corresponding waiting time) can be selected as the third agent.

[0057] In some embodiments, the method further includes: if there are no other intelligent agents waiting in the background besides the first intelligent agent, causing the robot to activate the default intelligent agent to return to a standby state. In some embodiments, if there are no other intelligent agents waiting in the background besides the first intelligent agent, then a preset default intelligent agent is activated, causing the robot to return to a standby state.

[0058] Figure 2 The diagram illustrates a device structure of an interactive robot according to an embodiment of this application. The robot includes a first module 11, a second module 12, and a third module 13. The first module 11 is used to obtain perception information corresponding to a wake-up interaction request initiated by a first interactive subject. The second module 12 is used to clone a target intelligent agent using a default intelligent agent as a template, determine matching personalized capabilities based on the perception information, modify the target intelligent agent according to the personalized capabilities to obtain a modified first intelligent agent, and bind the first intelligent agent to the first interactive subject. The target intelligent agent and the default intelligent agent correspond to the same basic capabilities. The third module 13 is used to activate the first intelligent agent, enabling the robot to interact with the first interactive subject based on the first intelligent agent.

[0059] Module 11 is used to respond to a wake-up interaction request initiated by the first interactive subject and obtain the perception information corresponding to the wake-up interaction request.

[0060] In some embodiments, the robot includes, but is not limited to, humanoid robots, bionic robots, wheeled robots, tracked robots, articulated robots, mobile operation robots, etc., and this example embodiment does not make any special limitation in this regard.

[0061] In some embodiments, the interaction subject includes, but is not limited to, any subject capable of interacting with the robot, such as a person, an animal, or other robots. This example embodiment does not impose any special limitations on this.

[0062] In some embodiments, the wake-up interaction request may be initiated by the first interactive subject via voice. The robot's microphone array (audio sensor) captures the voice information of the first interactive subject and identifies the wake-up interaction request from the voice information. In some embodiments, the wake-up interaction request may also be initiated by the first interactive subject via gesture. The wake-up interaction request is obtained by gesture recognition of the real-world image containing the first interactive subject captured by the robot's camera (visual sensor). In some embodiments, the robot only supports gesture wake-up to avoid accidental operation.

[0063] In some embodiments, the perception information includes at least one of visual perception information, environmental perception information, and audio perception information. The visual perception information is information associated with the interactive subject obtained from video information captured by the robot's camera. The environmental perception information is information associated with the robot's current environment obtained from video information captured by the robot's camera. The audio perception information is information associated with the interactive subject obtained from audio information captured by the robot's microphone array.

[0064] Module 12 is used to clone a target intelligent agent using a default intelligent agent as a template, determine matching personalized capabilities based on the perception information, modify the target intelligent agent based on the personalized capabilities to obtain a modified first intelligent agent, and bind the first intelligent agent to the first interactive subject, wherein the target intelligent agent and the default intelligent agent have the same basic capabilities.

[0065] In some embodiments, personalized capabilities are determined based on perceived information. Personalized capabilities include at least one of a knowledge base, prompts, workflows, and embodied skills. The knowledge base is a structured / semi-structured data set that provides the agent with full-dimensional knowledge support. It is the "cognitive foundation" for the agent to understand the world and make reasoning decisions, and it is also the core carrier for avoiding repetitive learning and realizing knowledge reuse. Prompts are structured instructions that transmit task objectives, constraints, execution requirements, and scene contexts to the humanoid robot agent. They are the "communication interface" between humans and the agent, and they are also the "command signals" that guide the agent to retrieve knowledge from the knowledge base and plan workflows. Workflows are the "task execution framework" in which the agent breaks down complex tasks into a series of ordered and executable sub-tasks to complete the target task, and defines the execution logic, triggering conditions, connection methods, and exception handling rules of the sub-tasks. It is the core of the agent's decision-making and planning from "goal" to "steps". Embodied skills are a modular and reusable set of capabilities that the robot can use to perform specific basic actions / operations through physical execution. They are the final carrier for the agent's workflow to be implemented as actual physical behavior. In some embodiments, the robot's intelligent agent (i.e., artificial intelligence agent, AI Agent) is the core intelligence core that enables the robot to achieve autonomous perception, decision-making, action and interaction. In essence, it is an artificial intelligence system that integrates perception, cognition, planning, execution and learning capabilities. The intelligent agent transforms the robot from a "programmed mechanical actuator" into an "autonomous intelligent agent" that can understand the environment, make autonomous judgments and flexibly respond to complex scenarios.

[0066] In some embodiments, a target intelligent agent is cloned using a preset default intelligent agent as a template, so that the target intelligent agent and the default intelligent agent have the same basic capabilities. These basic capabilities include, but are not limited to, persona, language style, voice, and core functions, to avoid sudden changes in robot behavior. In some embodiments, the cloned target intelligent agent is modified according to the determined personalized capabilities, so that the modified first intelligent agent has the personalized capabilities, and then the first intelligent agent is bound to the first interaction subject. In this application, each interaction subject corresponds to a dedicated dynamic intelligent agent, and its personalized capabilities are matched to the wake-up interaction request, such as matching the attribute information of the interaction subject and / or matching the current scene type of the robot, to achieve personalized interaction, i.e., a unique interaction for each user.

[0067] Module 13 is used to activate the first intelligent agent, enabling the robot to interact with the first interactive subject based on the first intelligent agent.

[0068] In some embodiments, the first agent is activated, which makes the first agent effective, and then enables the robot to interact with the first interaction subject based on the first agent.

[0069] In some embodiments, the robot is further configured to: recognize gestures performed by a first interactive subject in a real-world scene captured by the robot's camera, and obtain a wake-up interaction request initiated by the first interactive subject. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0070] In some embodiments, the perceived information includes at least one of visual perceived information, environmental perceived information, and audio perceived information. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0071] In some embodiments, the personalization capabilities include at least one of a knowledge base, prompts, workflows, and embodied skills. Here, related operations and... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0072] In some embodiments, the perceived information includes visual perceived information, and the personalization capability includes prompting information; wherein, determining the matching personalization capability based on the perceived information includes: determining prompting information matching the visual perceived information based on the visual perceived information. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0073] In some embodiments, determining the prompt information matching the visual perception information based on the visual perception information includes: determining attribute information corresponding to the first interactive subject based on the visual perception information, and determining prompt information matching the attribute information. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0074] In some embodiments, the perceived information includes environmental perception information, and the personalized skill includes a workflow; wherein, determining the matching personalized ability based on the perceived information includes: determining a workflow that matches the environmental perception information. Here, related operations and... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0075] In some embodiments, determining a workflow matching the environmental perception information includes: determining the current scene type of the robot based on the environmental perception information, and determining a workflow matching the scene type. Here, related operations are... Figure 1The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0076] In some embodiments, the perceived information includes visual perception information and environmental perception information, and the personalization capability includes a knowledge base; wherein, determining the matching personalization capability based on the perceived information includes: determining a matching knowledge base based on the visual perception information and the environmental perception information. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0077] In some embodiments, determining a matching knowledge base based on the visual perception information and the environmental perception information includes: determining attribute information corresponding to the first interactive subject based on the visual perception information; determining the scene type currently in which the robot is located based on the environmental perception information; and determining a matching knowledge base based on the attribute information and the scene type. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0078] In some embodiments, the perceived information includes environmental perception information, and the personalized capability includes embodied skills; wherein, modifying the target agent based on the perceived information to obtain a modified first agent includes: determining embodied skills matching the environmental perception information based on the environmental perception information. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0079] In some embodiments, determining the embodied skill matching the environmental perception information includes: determining the current scene type of the robot based on the environmental perception information, and determining the embodied skill matching the scene type. Here, related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0080] In some embodiments, the perception information includes visual perception information and / or audio perception information; wherein, the robot is further configured to: determine the position information of the first interactive subject based on the visual perception information and / or the audio perception information; and cause the robot to turn towards the first interactive subject based on the position information. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0081] In some embodiments, the robot is further configured to: in response to the activation of a second agent bound to a second interactive subject, place the first agent in the background for standby and pause the robot's interaction with the first interactive subject. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0082] In some embodiments, the robot is further configured to: cause the second agent to copy the context information corresponding to the first agent. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0083] In some embodiments, the robot is further configured to: delete the first agent and terminate the robot's interaction with the first agent in response to the first interactive subject leaving the robot's field of vision. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0084] In some embodiments, the robot is further configured to: activate a third intelligent agent waiting in the background, causing the robot to resume interaction with a third interactive subject bound to the third intelligent agent. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0085] In some embodiments, the robot is further configured to: if there are no other intelligent agents in the background besides the first intelligent agent, cause the robot to activate the default intelligent agent to return to a standby state. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.

[0086] Figure 3 Exemplary systems that can be used to implement the various embodiments described in this application are shown; such as Figure 3 As shown in some embodiments, system 300 can function as any of the devices described in each of the embodiments. In some embodiments, system 300 may include one or more computer-readable media having instructions (e.g., system memory or NVM / storage device 320) and one or more processors (e.g., one or more processors 305) coupled to the one or more computer-readable media and configured to execute the instructions to implement the module and thus perform the actions described in this application.

[0087] In one embodiment, the system control module 310 may include any suitable interface controller to provide any suitable interface to at least one of the processors 305 and / or any suitable device or component communicating with the system control module 310.

[0088] The system control module 310 may include a memory controller module 330 to provide an interface to the system memory 315. The memory controller module 330 may be a hardware module, a software module, and / or a firmware module.

[0089] System memory 315 can be used, for example, to load and store data and / or instructions for system 300. In one embodiment, system memory 315 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, system memory 315 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).

[0090] In one embodiment, the system control module 310 may include one or more input / output (I / O) controllers to provide interfaces to the NVM / storage device 320 and (one or more) communication interfaces 325.

[0091] For example, NVM / storage device 320 may be used to store data and / or instructions. NVM / storage device 320 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drive (HDD), one or more optical disc (CD) drives, and / or one or more digital universal optical disc (DVD) drives).

[0092] NVM / storage device 320 may include storage resources that are physically part of a device on which system 300 is mounted, or that can be accessed by the device without necessarily being part of it. For example, NVM / storage device 320 may be accessed via a network through one or more communication interfaces 325.

[0093] One or more communication interfaces 325 may provide the system 300 with an interface to communicate over one or more networks and / or with any other suitable device. The system 300 may wirelessly communicate with one or more components of a wireless network in accordance with any of one or more wireless network standards and / or protocols.

[0094] In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 (e.g., memory controller module 330). In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 to form a system-in-package (SiP). In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die. In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die to form a system-on-a-chip (SoC).

[0095] In various embodiments, system 300 may be, but is not limited to, a server, workstation, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). In various embodiments, system 300 may have more or fewer components and / or different architectures. For example, in some embodiments, system 300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0096] In addition to the methods and devices described in the above embodiments, this application also provides a computer-readable storage medium storing computer code that, when executed, performs the method described in any of the preceding embodiments.

[0097] This application also provides a computer program product that, when executed by a computer device, performs the method described in any of the preceding claims.

[0098] This application also provides a computer device, the computer device comprising:

[0099] One or more processors;

[0100] Memory, used to store one or more computer programs;

[0101] When the one or more computer programs are executed by the one or more processors, the one or more processors cause the one or more processors to perform the method as described in any of the preceding methods.

[0102] It should be noted that this application can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of this application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, a magnetic or optical drive, a floppy disk, or similar devices. Furthermore, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.

[0103] Furthermore, a portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0104] Communication media include media through which communication signals containing, for example, computer-readable instructions, data structures, program modules, or other data are transmitted from one system to another. Communication media can include guided transmission media (such as cables and wires (e.g., optical fibers, coaxial cables, etc.)) and wireless (unguided transmission) media capable of propagating energy waves, such as sound, electromagnetic, RF, microwave, and infrared. Computer-readable instructions, data structures, program modules, or other data can be embodied as modulated data signals in, for example, wireless media (such as carrier waves or similar mechanisms embodied as part of spread spectrum technology). The term "modulated data signal" refers to a signal whose one or more characteristics are altered or set in a manner that encodes information in the signal. Modulation can be analog, digital, or a hybrid modulation technique.

[0105] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. For example, computer-readable storage media include, but are not limited to, volatile memories such as random access memory (RAM, DRAM, SRAM); and non-volatile memories such as flash memory, various read-only memories (ROM, PROM, EPROM, EEPROM), magnetic and ferromagnetic / ferroelectric memories (MRAM, FeRAM); and magnetic and optical storage devices (hard disks, magnetic tapes, CDs, DVDs); or other media now known or hereafter developed capable of storing computer-readable information / data for use by a computer system.

[0106] Herein, one embodiment of this application includes an apparatus comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the apparatus is triggered to run a method and / or technical solution based on the foregoing embodiments of this application.

[0107] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within this application. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the apparatus claims may also be implemented by a single unit or device in software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.

Claims

1. A method for interaction, applied to a robot, wherein, The method includes: In response to a wake-up interaction request initiated by the first interactive subject, obtain the perception information corresponding to the wake-up interaction request; A target intelligent agent is cloned using a default intelligent agent as a template. Matching personalized capabilities are determined based on the perception information. The target intelligent agent is modified based on the personalized capabilities to obtain a modified first intelligent agent. The first intelligent agent is then bound to the first interactive subject. The target intelligent agent and the default intelligent agent have the same basic capabilities. The first intelligent agent is activated, enabling the robot to interact with the first interactive subject based on the first intelligent agent.

2. The method according to claim 1, wherein, The method further includes: By recognizing the gestures performed by the first interactive subject in the real-world scene captured by the robot's camera, a wake-up interaction request initiated by the first interactive subject can be obtained.

3. The method according to claim 1, wherein, The perceived information includes at least one of visual perceived information, environmental perceived information, and audio perceived information.

4. The method according to claim 3, wherein, The personalization capabilities include at least one of the following: knowledge base, prompts, workflow, and embodied skills.

5. The method according to claim 4, wherein, The perceived information includes visual perceived information, and the personalized capabilities include prompt information; The step of determining the matching personalized capabilities based on the perceived information includes: Based on the visual perception information, a prompt message matching the visual perception information is determined.

6. The method according to claim 5, wherein, The step of determining the prompt information matching the visual perception information based on the visual perception information includes: Based on the visual perception information, determine the attribute information corresponding to the first interactive subject, and determine the prompt information that matches the attribute information.

7. The method according to claim 4, wherein, The perceived information includes environmental perception information, and the personalized skills include workflow; The step of determining the matching personalized capabilities based on the perceived information includes: Based on the environmental perception information, a workflow matching the environmental perception information is determined.

8. The method according to claim 7, wherein, The step of determining the workflow that matches the environmental perception information based on the environmental perception information includes: Based on the environmental perception information, the current scene type of the robot is determined, and a workflow matching the scene type is determined.

9. The method according to claim 4, wherein, The perceived information includes visual perceived information and environmental perceived information, and the personalized capabilities include a knowledge base; The step of determining the matching personalized capabilities based on the perceived information includes: A matching knowledge base is determined based on the visual perception information and the environmental perception information.

10. The method according to claim 9, wherein, The step of determining a matching knowledge base based on the visual perception information and the environmental perception information includes: The attribute information corresponding to the first interactive subject is determined based on the visual perception information, and / or the scene type in which the robot is currently located is determined based on the environmental perception information; A matching knowledge base is determined based on the attribute information and / or the scenario type.

11. The method according to claim 4, wherein, The perceived information includes environmental perception information, and the personalized capabilities include embodied skills; The step of modifying the target agent based on the perceived information to obtain the modified first agent includes: Based on the environmental perception information, embodied skills that match the environmental perception information are determined.

12. The method according to claim 11, wherein, The step of determining the embodied skill that matches the environmental perception information based on the environmental perception information includes: Based on the environmental perception information, the robot determines the current scene type and identifies the embodied skill that matches the scene type.

13. The method according to claim 3, wherein, The perceived information includes visual perceived information and / or audio perceived information; The method further includes: The location information of the first interactive subject is determined based on the visual perception information and / or the audio perception information; This causes the robot to turn towards the first interactive subject based on the location information.

14. The method according to claim 1 or 13, wherein, The method further includes: In response to the activation of the second intelligent agent bound to the second interactive subject, the first intelligent agent is placed in the background for standby, and the robot's interaction with the first interactive subject is suspended.

15. The method according to claim 14, wherein, The method further includes: This causes the second agent to copy the context information corresponding to the first agent.

16. The method of claim 14, wherein, The method further includes: In response to the first interactive subject leaving the robot's field of vision, the first intelligent agent is deleted, and the robot's interaction with the first interactive subject ends.

17. The method according to claim 16, wherein, The method further includes: Activate the third intelligent agent that is waiting in the background, so that the robot can resume interaction with the third interactive subject that is bound to the third intelligent agent.

18. The method according to claim 17, wherein, The method further includes: If there are no other intelligent agents in the background besides the first intelligent agent, the robot will activate the default intelligent agent to return to the standby state.

19. A computer device for outputting a sequence of action instructions, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 18.

20. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 18.

21. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method as described in any one of claims 1 to 18.