Active interaction method, device and equipment of robot and medium
By combining perception modules and cognitive memory networks, intelligent robots can proactively determine interactive behaviors, solving the user experience problems of passive interaction modes and achieving more efficient user interaction.
Patent Information
- Application Number
- CN202610190082.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-15
AI Technical Summary
Existing intelligent robots in open service scenarios such as hospitals and nursing homes adopt a passive interaction mode, resulting in a poor user experience and an inability to proactively provide services.
By detecting the characteristics of the target user through the perception module, and using the gated memory enhancement recurrent unit in the long-term memory matrix and cognitive memory network, the current cognitive state and predicted state of the target user are determined, thereby proactively determining the interaction behavior and interacting with the user.
This improved the user experience of the robot, enabling it to proactively provide services and enhance its ability to interact with users.
Smart Images

Figure CN122047296A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robotics, and in particular to a method, apparatus, device, and medium for active interaction of robots. Background Technology
[0002] Intelligent robots deployed in open service settings such as hospitals and nursing homes currently generally adopt a passive "stimulus-response" interaction mode. That is, the robot strictly follows a pre-programmed process tree or waits for explicit voice commands before acting.
[0003] This robot interaction mode means that the robot will only perform corresponding operations after receiving instructions, and cannot provide services proactively, resulting in a poor user experience. Summary of the Invention
[0004] This invention provides a method, apparatus, device, and medium for active interaction of robots, enabling robots to actively interact and provide services, greatly improving the user experience of robots.
[0005] According to one aspect of the present invention, a method for active interaction of a robot is provided, the method comprising: Based on the characteristics of the target user detected by the perception module, the target state vector corresponding to the target user is determined, and based on the target state vector of the target user, the historical memory vector of the target user that matches the current scenario is determined in the long-term memory matrix. Based on the hidden state of the previous time step, the target state vector, and the historical memory vector of the gated memory enhancement recurrent unit in the cognitive memory network, the current cognitive state and the predicted state of the target user are determined. Based on the current cognitive state and the predicted state, the robot's proactive interaction behavior is determined, and the robot interacts with the target user based on the proactive interaction behavior.
[0006] According to another aspect of the present invention, an active interaction device for a robot is provided, comprising: The target state vector determination module is used to determine the target state vector corresponding to the target user based on the characteristics of the target user detected by the perception module, and to determine the historical memory vector of the target user that matches the current scenario in the long-term memory matrix based on the target state vector of the target user. The current cognitive state determination module is used to determine the current cognitive state and predicted state of the target user based on the internal hidden state of the previous time step of the gated memory enhancement recurrent unit in the cognitive memory network, the target state vector, and the historical memory vector. The active interaction behavior determination module is used to determine the robot's active interaction behavior based on the current cognitive state and the predicted state, and to enable the robot to interact with the target user based on the active interaction behavior.
[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the active interaction method of the robot according to any embodiment of the present invention.
[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the active interaction method of a robot according to any embodiment of the present invention.
[0009] The technical solution of this application includes: determining a target state vector corresponding to the target user based on the characteristics of the target user detected by the perception module; determining a historical memory vector of the target user that matches the current scenario in the long-term memory matrix based on the target state vector of the target user; determining the current cognitive state and predicted state of the target user based on the internal hidden state of the previous time step of the gated memory enhancement recurrent unit in the cognitive memory network, the target state vector, and the historical memory vector; determining the robot's active interaction behavior based on the current cognitive state and the predicted state; and enabling the robot to interact with the target user based on the active interaction behavior. This technical solution can greatly improve the user experience of the robot by interacting with the target user through the identified active interaction behavior of the robot.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1This is a flowchart of a robot's active interaction method according to Embodiment 1 of this application; Figure 2 This is a flowchart of a robot's active interaction method according to Embodiment 2 of this application; Figure 3 This is a data flow diagram provided according to the method described in Embodiment 2 of this application; Figure 4 This is a structural diagram of the cognitive memory network provided in Embodiment 2 of this application; Figure 5 This is a flowchart illustrating the collaborative workflow of the cognitive memory network and the active decision-making engine according to Embodiment 2 of this application; Figure 6 This is a structural schematic diagram of an active interaction device for a robot according to Embodiment 3 of this application; Figure 7 This is a schematic diagram of the structure of an electronic device that implements a robot active interaction method according to an embodiment of this application. Detailed Implementation
[0013] To enable those skilled in the art to better understand the present invention, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0014] It should be noted that the terms "first," "second," "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0015] Example 1 Figure 1This application provides a flowchart of a robot's active interaction method according to Embodiment 1. This embodiment is applicable to situations where a robot is controlled to perform active interaction. The method can be executed by the robot's active interaction device, which can be implemented in hardware and / or software. This active interaction device can be configured in an electronic device with data processing capabilities. Figure 1 As shown, the method includes: S110, based on the characteristics of the target user detected by the perception module, determine the target state vector corresponding to the target user, and based on the target state vector of the target user, determine the historical memory vector of the target user that matches the current scenario in the long-term memory matrix.
[0016] The perception module is used to detect user behavior, language, and other characteristics. It should be noted that the acquisition, storage, use, and processing of data in this application comply with relevant national laws and regulations and have been authorized by the user. The perception module can detect the language, behavior, and other characteristics of each user, and it can also detect user identity. Therefore, it can obtain the characteristics corresponding to each user, and one of these users can be the target user. This application embodiment uses the detection of the target user's characteristics as an example to illustrate how the robot determines proactive interaction behavior. The target state vector is a vector obtained by encoding the detected target user's characteristics. The long-term memory matrix can retain various standardized cognitive feature data accumulated by the user from historical interactions. Each vector in the matrix corresponds to a comprehensive feature record of the user in a specific state and scenario, covering multi-dimensional information such as the user's regular behavioral habits, the triggering and evolution characteristics of emotions and physiological states, feedback results to different robot interaction behaviors, and typical performance in special states.
[0017] Specifically, after the perception module on the robot detects the characteristics of the target user, it encodes the characteristics to obtain the target state vector corresponding to the target user; and then searches in the long-term memory matrix according to the target state vector and the target user's identifier to obtain the historical memory vector of the target user that matches the current situation.
[0018] S120, based on the hidden state of the previous time step, the target state vector, and the historical memory vector of the gated memory enhancement recurrent unit in the cognitive memory network, determines the current cognitive state and the predicted state of the target user.
[0019] The cognitive memory network refers to a network that analyzes a user's cognitive state. Its input data is the target user's target state vector. For example, if the target state vector reflects that the target user is dozing off, the cognitive memory network may analyze that the target user is sleepy. The cognitive memory network includes a gated memory enhancement loop unit to continuously update the user's cognitive state. It should be noted that the cognitive memory network may include a gated memory enhancement loop unit, a long-term memory matrix for storing the user's memory states, and a prediction network for predicting the user's next state.
[0020] Specifically, the gated memory enhancement loop unit can process the hidden state, target state vector, and historical memory vector of the previous time step to obtain the current cognitive state of the target user; the prediction network can then predict the cognitive state of the target user to obtain the predicted state.
[0021] S130, determine the robot's proactive interaction behavior based on the current cognitive state and the predicted state, and enable the robot to interact with the target user based on the proactive interaction behavior.
[0022] Specifically, after obtaining the current cognitive state and the predicted state, corresponding decisions can be made accordingly. For example, if the current cognitive state is that there is a risk of falling and the predicted state is that the probability of falling within 5 seconds is 80%, then the active interactive behavior can be protective following, voice prompts, etc.
[0023] For example, the current cognitive state and the predicted state can be processed by hierarchical reinforcement learning (HRL) to make decisions on the robot's proactive interaction behavior, thereby obtaining the robot's proactive interaction behavior and controlling the robot to interact with the target user based on the proactive interaction behavior.
[0024] The technical solution of this application includes: determining a target state vector corresponding to the target user based on the characteristics of the target user detected by the perception module; determining a historical memory vector of the target user that matches the current scenario in the long-term memory matrix based on the target state vector of the target user; determining the current cognitive state and predicted state of the target user based on the internal hidden state of the previous time step of the gated memory enhancement recurrent unit in the cognitive memory network, the target state vector, and the historical memory vector; determining the robot's active interaction behavior based on the current cognitive state and the predicted state; and enabling the robot to interact with the target user based on the active interaction behavior. This technical solution can greatly improve the user experience of the robot by interacting with the target user through the identified active interaction behavior of the robot.
[0025] Example 2 Figure 2This is a flowchart of a robot's active interaction method provided in Embodiment 2 of this application. This embodiment is an optimization based on the above embodiment.
[0026] like Figure 2 As shown, the method in this embodiment of the application specifically includes the following steps: S210, Based on the characteristics of the target user detected by the perception module, determine the target state vector corresponding to the target user.
[0027] For example, Figure 3 This is a data flow diagram of the method described in the embodiments of this application (which reflects the data transmission and reception between modules); in Figure 3 In this process, the user's features can be acquired through the multimodal perception module, and the target user identifier and target state vector can be output.
[0028] For example, based on the user identity ID and continuous tracking signal output by the perception module (voiceprint / visual fusion module), richer perception signals are further fused: Physiological and behavioral feature extraction: From the visual stream, a lightweight model is used to analyze the user's micro-expressions (such as pain, frowning), posture and gait (such as hunchback, shortened stride, swaying), movement frequency (such as repeatedly rubbing a certain part), and interactive attention (whether there is eye contact with the robot). Speech para-language feature extraction: From the audio stream, para-language features beyond the text content are extracted, including: speech energy changes, intonation (groaning, sighing), speech rate, and non-semantic vocalizations (coughing, wheezing). State vector encoding: The above features are fused with contextual information such as user identity, time, and robot position, and encoded into a unified temporal target user's target state vector. This serves as the input to the cognitive memory network. The cognitive memory network is used to process the target user's target state vector to obtain the current cognitive state and the predicted state; that is, steps S220-S260 can be executed within the cognitive memory network.
[0029] S220, Generate a query vector based on the target user's target state vector and the target user's identifier.
[0030] S230, determine the attention weight based on the query vector, and read the target user's historical memory vector from the long-term memory matrix based on the attention weight.
[0031] For example, the structure of the cognitive memory network is as follows: Figure 4 As shown, this cognitive memory network processes the following inputs at each time step t: : The target state vector of the target user in the current multimodal perception fusion time sequence. : The hidden internal state of the gated memory enhancement loop unit at the previous time step. ID: The target user identifier provided by the front-end system (i.e., multimodal user binding). And output the target user's current cognitive state and predicted state.
[0032] Specifically, after inputting the target user's target state vector and the target user's identifier into the cognitive memory network, a query vector can be generated to query historical memories in the long-term memory matrix. Specifically, this can be: Based on the current state and user identity, from the long-term memory matrix Retrieving relevant information from the user information stored in the vector database allows you to first generate a query vector during the search. : .
[0033] Specifically, obtain the query vector Then, calculate the attention weights: ; Based on the attention weights, the target user's historical memory vector is read from the long-term memory matrix. : .
[0034] in, It is an embedded representation of user identity. It is the dimension of the vector. It is the historical cognitive pattern most relevant to the current situation, that is, the historical memory vector.
[0035] S240, based on the gated memory enhancement recurrent unit in the cognitive memory network, the internal hidden state, target state vector and historical memory vector of the previous time step are processed to update the internal state and obtain the hidden state of the current time step.
[0036] For example, a cognitive memory network fuses the internal hidden state from the previous time step, the target state vector, and the historical memory vector to update the internal state and predict the future: Forget Gate (determines what information to discard from the internal state): ; Input gate (determines which new information to store in the state): ; Candidate memory cells: ; Renew memory cells: ; Output gate: ; Output the hidden state at the current time step: .
[0037] S250, based on the hidden state of the current time step and the historical memory vector, determine the current cognitive state of the target user.
[0038] For example, the hidden state at the current time step and the historical memory vector are concatenated to obtain the target user's current cognitive state. : .
[0039] S260 uses a prediction network in a cognitive memory network to predict the state of the target user and obtain the predicted state.
[0040] For example, the predicted state can specifically be the probability that the target user is in a certain state, such as the probability of falling down being 80%.
[0041] For example, this can be achieved through a small prediction network. Process the hidden state at the current time step and output the future. The state probability distribution of the step: .
[0042] in, Let P represent the probability of moving k steps into the future.
[0043] S270, determine the robot's behavior type based on the current cognitive state and the predicted state.
[0044] S280, determine the proactive interaction behavior based on the behavior type, and enable the robot to interact with the target user based on the proactive interaction behavior.
[0045] The behavior type can be a pre-set type. After determining the behavior type, specific active interaction behaviors are then determined within the scope of the behavior type.
[0046] For example, the steps of determining the robot's behavior type based on the current cognitive state and the predicted state, and determining the proactive interaction behavior based on the behavior type, can be implemented by a proactive decision engine based on hierarchical reinforcement learning (HRL). This decision engine uses a hierarchical reinforcement learning framework to decompose the complex proactive care task into learnable layers, specifically, a meta-policy layer and an execution policy layer. High-level (meta-strategy layer): Based on the current cognitive state Given a predicted state P, select a target from the macro-behavioral type space, such as: (maintaining observation, proactively approaching and inquiring, initiating gentle following, sending a remote reminder, performing reassuring behavior, etc.). The reward signal for this layer is designed to be sparse but long-term, for example, "successfully preventing a potential risk" would receive a high reward.
[0047] Bottom Layer (Execution Strategy Layer): After the upper layer selects the behavior type, the bottom layer strategy is responsible for generating specific, executable sequences of action parameters. For example, if the upper layer selects "actively approach and ask questions," the bottom layer plans the specific movement path, final stopping distance, head turning angle, and generates specific question statements. The reward signals at this layer are more concentrated, based on the smoothness of the action, user response, etc.
[0048] The Option (behavior type) output by the high-level strategy is a parameterized tuple: (behavior type, target user, expected intent, timeout parameter). The low-level strategy then generates specific Action sequences (active interaction behaviors) based on the Option: [(move to coordinates (x, y), speed v), (play voice text T, tone P), (robotic arm pose), etc.].
[0049] The data flow processing of the cognitive memory network described above corresponds to steps S220-S260; the data flow processing of the active decision engine based on hierarchical reinforcement learning corresponds to steps S270-S280; the cognitive memory network inputs the current cognitive state and the predicted state into the active decision engine. The two core modules form a tightly coupled, dynamically evolving cognitive decision-making closed loop through the following data flow. The collaborative workflow diagram of the cognitive memory network and the active decision engine is shown below. Figure 5 As shown, in Figure 5 In this process, the perception and cognition stage involves multimodal perception of the user to obtain the target user's target state vector, and the cognitive memory network processing the target user's target state vector to output the current cognitive state and the predicted state; the decision-making and planning stage is the stage where the active decision engine determines the active interaction behavior; based on Figure 5 The content, specifically the collaboration between cognitive memory networks and decision engines, is divided into the following parts: The depth of decision-making basis: the output of the gated memory enhancement loop unit in the cognitive memory network and Together, they constitute the enhanced state observation space of the HRL decision engine. For example, The combination of the "user mental state index" encoded in the code and the "probability of high fall risk" predicted in P significantly increases the probability of selecting Option (type='protective follow') in the higher-level strategy. The predicted uncertainty score is then used to adjust the exploration-exploitation tradeoff, choosing a more conservative option when there is high uncertainty.
[0050] Decision outcomes drive memory evolution: After the HRL engine executes proactive actions, the environment provides feedback (such as the user accepting care, explicitly refusing, or showing no response). This feedback signal, along with the final outcome (whether the risk was successfully prevented), is encapsulated into a feedback tuple. This is then passed back to the loop unit as an important event.
[0051] Instant Update: This feedback is immediately encoded as a state change. This may trigger a memory write-back process, transferring the experience of this interaction. Write to the corresponding user's long-term memory middle.
[0052] Long-term impact: In this way, users' personalized behavioral patterns (such as "responsive to gentle voice inquiries but uneasy about rapid approach") are continuously recorded and reinforced. The next time a similar situation occurs... When it appears, the retrieved memory This historical experience will be incorporated, enabling the decision engine to make more personalized and accurate judgments.
[0053] Closed-loop optimization and personalized adaptation: The entire system forms a closed loop of "perception -> cognitive modeling -> decision-making -> action -> feedback -> memory update -> optimization for the next perception and decision-making." As the number of interactions increases, As the matrix becomes increasingly rich, the system models the cognitive state of each user. The more accurate the decision-making process, the more personalized and efficient the strategies of the HRL decision engine become.
[0054] This collaborative design ensures that the system's proactive care behavior is not based on static rules, but rather stems from an evolving cognitive model that deeply understands the user, truly achieving a leap from "mechanical response" to "cognitive intelligence".
[0055] Optionally, in this embodiment of the application, after the robot interacts with the target user based on the proactive interactive behavior, the method further includes: obtaining environmental feedback information and determining a state change amount based on the environmental feedback information; updating the target user's memory in the long-term memory matrix based on the state change amount and the current cognitive state.
[0056] For example, the environmental feedback information can be information detected through multimodal sensing, and can be represented by feedback tuples. This means that encoding it yields the state change quantity.
[0057] For example, after the robot interacts with the target user based on the aforementioned proactive interaction behavior, it can update the long-term memory matrix according to environmental feedback. When a significant pattern deviation or important event (such as a successful proactive care) is detected, a write-back mechanism is triggered. Write gating, Represents the change in state: ; Generate new memory fragments: ; Update the memory matrix row for the corresponding user using gating: .
[0058] Optionally, after determining the proactive interaction behavior, the method further includes: quantifying and predicting risk indicators for the proactive interaction behavior based on a risk prediction network; if the quantification and prediction results of the risk indicators meet preset conditions, then controlling the robot to interact with the target user based on the proactive interaction behavior.
[0059] Specifically, after obtaining the proactive interaction behavior, it is quantitatively evaluated through a risk prediction network to avoid the proactive interaction behavior from causing undue impact on the user. After the evaluation is passed, the robot is controlled to interact with the target user based on the proactive interaction behavior.
[0060] For example, a validator can be set up that includes a rule base (such as minimum safe distance, prohibited vocabulary) and a trained risk prediction network. Through rapid simulation in the world model, the validator quantifies and evaluates indicators such as user discomfort and collision probability that the action sequence may cause. Only when all indicators are below the threshold can the action pass the test, and only actions that pass the test will be delivered for execution.
[0061] In this embodiment of the application, optionally, when determining the proactive interaction behavior according to the behavior type, the method further includes: obtaining the interaction template of the target user; and determining the proactive interaction behavior that conforms to the behavior type based on the interaction template.
[0062] For example, when generating the proactive interaction behavior, an interaction template matching the target user's identity is retrieved from the personalized policy library. For instance, for an elderly person with dementia who is known to resist strong interactions, the system may choose "observe from 2 meters away and send a remote reminder"; for a familiar postoperative patient, it may choose "approach to within 1 meter and ask in a gentle tone".
[0063] Furthermore, after obtaining the aforementioned proactive interaction behavior, proactive interactive operations are performed through the coordination of robot movement, lighting, screen expressions, and voice. For example, when performing "gentle following," the robot moves along an arc path 2 meters behind the user, displays a smiling expression on the screen, and plays soothing background music. It only asks a question via voice after the user has stayed for an extended period.
[0064] The HRL decision engine in this application embodiment can autonomously learn and optimize strategies through interaction with the environment, and the GMARU (Cognitive Memory Network) provides... P provides dynamic and predictive input for decision-making, enabling the determination of the entire active interaction to be adaptive and evolutionary, far exceeding static rules.
[0065] The embodiments of this application achieve the robot's proactive interaction through the steps of perception, cognition, decision-making, and execution. This hierarchical architecture and explicit memory matrix... This makes each step of the reasoning in the method described in the embodiments of this application traceable and auditable, and provides key behavioral safeguards through security verification.
[0066] The hierarchical reinforcement learning framework in this application, through the decomposition of "high-level goal setting (Option) and low-level fine-grained execution (Action)," greatly reduces the search space of each layer, improving learning efficiency and the modular reusability of the strategy. The meta-policy of "when to proactively care" learned by the high-level layer can guide various specific execution methods at the low level.
[0067] Because the internal memory of a typical LSTM is both "implicit" and "hybrid," all user memories are intertwined. The external memory matrix designed in this application embodiment... It enables long-term memory that is structured and explicitly stored according to the user, allowing for precise retrieval and editing through attention mechanisms, which directly supports the core goal of strong personalization.
[0068] The "layered" reinforcement learning in this application directly corresponds to the decision-making process of human caregivers: first, judging the intent ("He may need help"), and then considering the steps ("Go over first, then ask softly"). This layering not only makes learning more efficient, but also makes the final strategy more understandable and adjustable to human managers. High-level strategies can be interpreted as a series of "care intentions," facilitating human review and ethical alignment.
[0069] This technical solution pre-simulates the consequences of actions in a virtual environment and incorporates built-in security rule filters based on distance, tone, and content. This avoids the problems of potentially excessive intrusion or misjudgment of proactive behaviors.
[0070] Due to the initial Empty, HRL strategy is untrained. This technical solution adopts a strategy of "simulated data pre-training + human-supervised fine-tuning". First, HRL is trained using simulated patients in a digital twin environment, and then in a real-world scenario, staff members approve / reject the robot's proactive behaviors, providing initial... Signal.
[0071] This technical solution enables personalized processing for each user when multiple users simultaneously enter the sensing range by identifying user identifiers and storing data for different users in a long-term memory matrix. This solution highlights the importance of the front-end multimodal user binding and tracking module (i.e., the designed voiceprint / visual fusion). It provides clear and continuous user ID input for subsequent GMARU and HRL, and is a key prerequisite for the entire system to function.
[0072] In a specific instance, the characteristics of user A =[High gait instability, painful expression], retrieve their memory. =[Previous history of falls], generated =[High risk of falling], prediction =[Probability of falling within 5 seconds > 80%]. HRL executives made this selection accordingly. =(Protective following, User A, to prevent falls, 30 seconds), underlying planning =[Follow in an arc at a distance of 1 meter, preparing the robotic arm for a buffer posture]. After execution, if the user is safe, provide feedback. For positive, To alleviate the situation, trigger an update. User A's "prone to falling when sick" mode.
[0073] Example 3 Figure 6 This is a schematic diagram of the structure of a robot's active interaction device provided in Embodiment 3 of this application. This device can execute the robot's active interaction method provided in any embodiment of this invention, and possesses the corresponding functional modules and beneficial effects of the method. For example... Figure 6 As shown, the device includes: The target state vector determination module 310 is used to determine the target state vector corresponding to the target user based on the characteristics of the target user detected by the perception module, and to determine the historical memory vector of the target user that matches the current scenario in the long-term memory matrix based on the target state vector of the target user. The current cognitive state determination module 320 is used to determine the current cognitive state and predicted state of the target user based on the internal hidden state of the previous time step of the gated memory enhancement recurrent unit in the cognitive memory network, the target state vector, and the historical memory vector. The active interaction behavior determination module 330 is used to determine the robot's active interaction behavior based on the current cognitive state and the predicted state, and to enable the robot to interact with the target user based on the active interaction behavior.
[0074] The technical solution of this application embodiment includes: a target state vector determination module 310, used to determine the target state vector corresponding to the target user based on the characteristics of the target user detected by the perception module, and to determine the historical memory vector of the target user that matches the current scenario in the long-term memory matrix based on the target state vector of the target user; a current cognitive state determination module 320, used to determine the current cognitive state and predicted state of the target user based on the internal hidden state of the previous time step of the gated memory enhancement recurrent unit in the cognitive memory network, the target state vector, and the historical memory vector; and an active interaction behavior determination module 330, used to determine the robot's active interaction behavior based on the current cognitive state and the predicted state, and to enable the robot to interact with the target user based on the active interaction behavior. This technical solution can greatly improve the user experience of the robot by interacting with the target user through the identified active interaction behavior of the robot.
[0075] Optionally, in this embodiment of the application, the target state vector determination module 310 includes: A query vector generation unit is used to generate a query vector based on the target user's target state vector and the target user's identifier. The historical memory vector reading unit is used to determine the attention weight based on the query vector, and read the target user's historical memory vector from the long-term memory matrix based on the attention weight.
[0076] Optionally, in this embodiment of the application, the current cognitive state determination module 320 includes: The internal state update unit is used to process the internal hidden state, target state vector and historical memory vector of the previous time step based on the gated memory enhancement recurrent unit in the cognitive memory network, update the internal state, and obtain the hidden state of the current time step. The current cognitive state determination unit is used to determine the current cognitive state of the target user based on the hidden state of the current time step and the historical memory vector. The predictive state determination unit is used to predict the state of the target user based on the predictive network in the cognitive memory network, and obtain the predicted state.
[0077] Optionally, in this embodiment of the application, the active interaction behavior determination module 330 includes: A behavior type determination unit is used to determine the robot's behavior type based on the current cognitive state and the predicted state. The active interaction behavior determination unit is used to determine the active interaction behavior based on the behavior type.
[0078] Optionally, in this embodiment of the application, the device further includes: An interaction template acquisition unit is used to acquire the interaction template of the target user. Correspondingly, the active interaction behavior determination unit is specifically used to determine active interaction behaviors that conform to the behavior type based on the interaction template.
[0079] Optionally, in this embodiment of the application, the device further includes: An environmental feedback information acquisition module is used to acquire environmental feedback information and determine the state change amount based on the environmental feedback information. The memory update module is used to update the target user's memory in the long-term memory matrix based on the state change amount and the current cognitive state.
[0080] Optionally, in this embodiment of the application, the device further includes: The risk identification module is used to quantitatively predict risk indicators for the proactive interactive behavior based on a risk prediction network. The risk identification module is used to control the robot to interact with the target user based on the proactive interactive behavior if the quantitative prediction result of the risk indicator meets the preset conditions.
[0081] The robot active interaction device provided in this application embodiment can execute the robot active interaction method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0082] Example 4 Figure 7 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0083] like Figure 7As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0084] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0085] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the active interaction methods of a robot.
[0086] In some embodiments, the robot's active interaction method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded into and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the robot's active interaction method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the robot's active interaction method by any other suitable means (e.g., by means of firmware).
[0087] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0088] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0089] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0090] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0091] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0092] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0093] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0094] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for active interaction of a robot, characterized in that, include: Based on the characteristics of the target user detected by the perception module, the target state vector corresponding to the target user is determined, and based on the target state vector of the target user, the historical memory vector of the target user that matches the current scenario is determined in the long-term memory matrix. Based on the hidden state of the previous time step, the target state vector, and the historical memory vector of the gated memory enhancement recurrent unit in the cognitive memory network, the current cognitive state and the predicted state of the target user are determined. Based on the current cognitive state and the predicted state, the robot's proactive interaction behavior is determined, and the robot interacts with the target user based on the proactive interaction behavior.
2. The method according to claim 1, characterized in that, Based on the target user's target state vector, determine the target user's historical memory vector that matches the current scenario in the long-term memory matrix, including: Generate a query vector based on the target user's target state vector and the target user's identifier; Attention weights are determined based on the query vector, and the target user's historical memory vector is read from the long-term memory matrix based on the attention weights.
3. The method according to claim 1, characterized in that, Based on the hidden state of the previous time step, the target state vector, and the historical memory vector of the gated memory-enhancing recurrent unit in the cognitive memory network, the current cognitive state and predicted state of the target user are determined, including: The gated memory enhancement recurrent unit in the cognitive memory network processes the internal hidden state, target state vector, and historical memory vector of the previous time step, updates the internal state, and obtains the hidden state of the current time step. Based on the hidden state at the current time step and the historical memory vector, the current cognitive state of the target user is determined; The predictive network in the cognitive memory network is used to predict the state of the target user and obtain the predicted state.
4. The method according to claim 1, characterized in that, Determine the robot's proactive interaction behavior based on the current cognitive state and the predicted state, including: The robot's behavior type is determined based on the current cognitive state and the predicted state; Determine the proactive interaction behavior based on the behavior type.
5. The method according to claim 4, characterized in that, When determining the proactive interaction behavior based on the behavior type, the method further includes: Obtain the interaction template of the target user; Based on the interaction template, determine the proactive interaction behavior that matches the behavior type.
6. The method according to claim 1, characterized in that, After the robot interacts with the target user based on the proactive interactive behavior, the method further includes: Obtain environmental feedback information and determine the state change amount based on the environmental feedback information; The target user's memory is updated in the long-term memory matrix based on the state change and the current cognitive state.
7. The method according to claim 1, characterized in that, After determining the proactive interaction behavior, the method further includes: The risk indicators of the proactive interactive behavior are quantitatively predicted based on the risk prediction network. If the quantitative prediction result of the risk indicator meets the preset conditions, then the robot is controlled to interact with the target user based on the proactive interactive behavior.
8. An active interaction device for a robot, characterized in that, include: The target state vector determination module is used to determine the target state vector corresponding to the target user based on the characteristics of the target user detected by the perception module, and to determine the historical memory vector of the target user that matches the current scenario in the long-term memory matrix based on the target state vector of the target user. The current cognitive state determination module is used to determine the current cognitive state and predicted state of the target user based on the internal hidden state of the previous time step of the gated memory enhancement recurrent unit in the cognitive memory network, the target state vector, and the historical memory vector. The active interaction behavior determination module is used to determine the robot's active interaction behavior based on the current cognitive state and the predicted state, and to enable the robot to interact with the target user based on the active interaction behavior.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the active interaction method of the robot according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the active interaction method of the robot according to any one of claims 1-7.