An agent embodied interaction planning method based on active perception of environmental information

By employing multimodal perception and energy gradient methods, the agent proactively perceives the interactive characteristics of objects in complex environments and generates predictable embodied interactive actions. This solves the problems of unstable movement and insufficient safety of the agent in complex environments, achieving efficient interactive adaptation and security.

CN119077739BActive Publication Date: 2025-11-21TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411378786.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-11-21
Estimated Expiration
2044-09-30

AI Technical Summary

Technical Problem

When existing intelligent agents engage in embodied interactions in complex, dense, and human-inhabited environments, they struggle to understand the temporal characteristics of interaction forces and responses, leading to instability in movement and difficulty in ensuring safety.

Method used

By acquiring environmental information around the agent and perceptual signals from the object's interaction interface, multimodal information fusion is performed. Embody interactive actions are generated using energy gradients and interaction characteristic parameters, including visual and tactile perception. Interactive force stimuli are actively applied to observe the object's response, inferring the object's implicit interaction characteristics and generating predictable actions.

Benefits of technology

It improves the level of interaction and safety of intelligent agents in complex environments, overcomes the problems of limited visual perception and no solution in free movement space, and realizes the predictability and safety of robot actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119077739B_ABST
    Figure CN119077739B_ABST
Patent Text Reader

Abstract

The application relates to a kind of intelligent agent embodied interaction planning methods based on active sensing of environmental information, method includes the following steps: S1, obtaining aligned environmental information and the information at the interface of interaction between intelligent agent and object;S2, the semantic feature information, intelligent agent working environment and interaction characteristic parameter are carried out multi-modal information fusion representation, obtain environmental multidimensional information;S3, obtain embodied interaction sensing signal, based on the expectedness judgment of robot action based on the embodied interaction sensing signal, if the result is within the expectation, based on environmental multidimensional information and the specific task step of planning, a series of embodied interaction actions are generated based on the method of energy gradient, and the robot executes embodied interaction action;otherwise, a series of interaction actions with interaction force characteristics are generated based on the method of energy gradient of interaction characteristic parameter, and the robot executes interaction action.Compared with the prior art, the application has the advantages of improving the safety of machine action.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of embodied intelligence of collaborative agents, and in particular to an agent embodied interaction planning method based on active perception of environment information. BACKGROUND

[0002] In recent years, with the significant improvement in the depth and breadth of agent industry applications, its application scenarios have gradually moved from factories to community scenarios where it coexists with humans. This type of work scenario with unknown, complex, dense, and human living characteristics has a large number of interactive work tasks. This will put higher requirements on the adaptability of interactive actions in the agent work process. Taking a home service agent as an example, the indoor environment is highly unstructured and dynamic, and is a humanized environment, with various movable objects in the room, such as furniture, shoes, children's toys, and even humans (ourselves) living in these environments and performing daily activities, which pose a major challenge to the autonomous movement of agents. Agents not only need to consider the completion level and movement stability of tasks, but also need to consider the safety of movement. Ensuring the adaptability and safety of actions in the work scenario is a prerequisite for the application of service agents. Currently, the environment perception and representation method based on passive single visual modality often fails to meet the needs of agent motion planning constraints and embodied interaction explainability due to the neglect of the perception of implicit physical properties closely related to embodied interaction. This leads to uncoordinated or even movement failure of agents in such scenarios.

[0003] Currently, the main solution to agent action planning in such scenarios is to store the interaction properties of objects as priori in the hidden space of the neural network, use the trained neural network to infer the operation properties of the objects through visual observation, and represent them in a semantic form to constrain the embodied interaction actions of the agents.

[0004] However, the main problem with this method is the lack of predictability of the results of interactive actions, which often leads to unexpected movement failures. Agents have difficulty understanding the interaction force-interactive response timing characteristics of objects during embodied interaction, which makes it difficult to predict the expected results of agent interaction actions, increasing the risk of instability of agent interaction actions, and the safety of machine interaction actions cannot be fully guaranteed. SUMMARY

[0005] The purpose of the present application is to improve the understanding of the interaction force-interactive response timing signal of the matched pair by the agent, and to improve the adaptability of the agent's embodied interaction and the safety of the machine's action, and to provide an agent embodied interaction planning method based on active perception of environment information.

[0006] The purpose of the present application can be achieved by the following technical solutions:

[0007] The embodiment of the application discloses an agent embodied interaction planning method based on active perception of environment information, and the method comprises the following steps:

[0008] S1, obtaining environment information around an intelligent agent controlling robot movement and information at an intelligent agent and object interaction interface, the information at the intelligent agent and object interaction interface being embodied interaction perception signals, the environment information being a two-dimensional image and depth, extracting sampling data at the latest time from the obtained information, and aligning the sampling data to obtain aligned environment information and information at the intelligent agent and object interaction interface;

[0009] S2, performing target recognition and semantic segmentation on the aligned environment information to obtain semantic feature information, and performing fusion on the aligned environment information to obtain an intelligent agent working environment, inferring interaction characteristic parameters based on the information at the aligned intelligent agent and object interaction interface, and performing multi-modal information fusion representation on the semantic feature information, the intelligent agent working environment and the interaction characteristic parameters to obtain environment multi-dimensional information;

[0010] S3, obtaining embodied interaction perception signals, performing predictability judgment on robot movement based on the embodied interaction perception signals, if the judgment result is within the expectation, generating a series of first embodied interaction actions based on an energy gradient method based on the environment multi-dimensional information and a planned specific task step, and the robot performs the first embodied interaction action;

[0011] otherwise, generating a series of second interaction actions with interaction force characteristics based on an energy gradient method based on the interaction characteristic parameters, and the robot performs the second interaction action.

[0012] Further, the information at the intelligent agent and object interaction interface comprises interaction force stimulation, response displacement, temperature and vibration signals of the object in the interaction process, and the embodied interaction perception signals are obtained based on embodied interaction perception devices.

[0013] Further, the spatial position of the embodied interaction perception device is as follows:

[0014]

[0015] wherein, p i represents the Cartesian space position of the i th perception device relative to the parent joint, P i represents the position of the i th perception device relative to the robot center of mass coordinate system, represents the homogeneous transformation matrix of the parent joint corresponding to the perception device relative to the robot center of mass coordinate system.

[0016] Further, the specific steps of inferring the interaction characteristic parameters based on the aligned information at the intelligent agent and object interaction interface are as follows:

[0017] The information obtained from the interface between the aligned agent and the object is as follows:

[0018]

[0019] Where, Δx m It is an interactive displacement. It is the moving speed of the observation point at the interactive interface, F. m It refers to the interaction forces at the interface;

[0020] The information obtained above is used to infer the object's interaction characteristic parameters θ = [K, D, f]. T During the inference process, the interaction characteristics of objects are first modeled as follows:

[0021] Y(t)=H(t)θ+V(t)

[0022] It is the system superimposed output vector. It is a system overlay information matrix. It is system state noise;

[0023] Subsequently, the least squares parameter estimation method estimates the interaction characteristic parameters of the object:

[0024]

[0025] Assume J1(θ) in When the minimum value is obtained, let J1(θ) be relative to... The partial derivatives are zero, therefore we get

[0026]

[0027] Where y(j) represents the time-series input features observed at the interactive interface, v(t) represents noise, and ψ T J1(θ) represents the time-series input signal observed at the interactive interface, J1(θ) represents the system prediction error, K represents the elastic operation coefficient of the object, D represents the damping operation coefficient of the object, and f represents the operation loss constant of the object.

[0028] Furthermore, the specific steps for determining the predictability of robot actions based on the embodied interaction perception signals are as follows:

[0029] The planner first generates the target tracking point g* for the next time step based on the target position g:

[0030]

[0031] Where ε(x) is the effective solution space of the planner in Cartesian space, and x is the current position of the robot;

[0032] Determine whether the robot's speed and interaction force with the external environment, obtained from embodied interaction sensing signals, exceed the maximum safe speed v as the robot approaches the target point. safe and maximum interaction force F safe The steps for calculating the interaction force are: calculate the robot's acceleration when tracking the target point.

[0033]

[0034] Where M and C represent the inertia matrix and Coriolis matrix, respectively, q represents the robot joint space velocity, and G represents the gravity vector;

[0035] Therefore, the robot's speed when tracking the target point is Interactive power

[0036] The predictability of the robot's actions is determined, and the result of the robot action a is obtained.

[0037] Furthermore, the determination result of the robot action a is:

[0038]

[0039] Where a d This is a dangerous action, indicating something that was not expected. s For safety reasons, actions should be taken within the expected scope.

[0040] Furthermore, the system prediction error is:

[0041]

[0042] Where y(j) represents the time-series input features observed at the interactive interface, v(t) represents noise, and ψ T J1(θ) represents the time-series input signal observed at the interactive interface, and J1(θ) represents the system prediction error.

[0043] Furthermore, the energy gradient-based method generates a series of first embodied interaction actions, and the robot executes the first embodied interaction actions as follows:

[0044] First, the robot's workspace is topologically defined as the robot's body domain. Obstacle Domain and free motion space domain in Let r represent the robot's pose, h represent the robot's spatial configuration in the current pose, and x represent the robot's current position. This indicates the robot's workspace. Represents the obstacle domain. Represents the machine body domain. Represents the difference between the workspace and the obstacle domain and the robot body domain;

[0045] Subsequently, different spatial topologies were assigned different energy states, with the target domain being assigned a virtual energy potential field with elastic gravity. The free motion space domain is defined as having a linear damping field U. f (x), the obstacle domain is defined as U with energy operation cost. o,d The work region of W(θ, x) is (x)=W(θ,x).

[0046] By differentiating the energy state, the robot's driving commands are obtained as follows:

[0047] F* = -grad[U(x)]

[0048] The robot's driving command is the first embodied interactive action.

[0049] Furthermore, the energy state refers to the energy state of the robot when it is at a spatial position x, specifically:

[0050]

[0051] Among them, U o (x) represents U o,d The set of (x).

[0052] Furthermore, the method based on the energy gradient of the interaction characteristic parameter specifically generates a series of second interactive actions with interactive force characteristics as follows:

[0053] Calculate the specific F-phase second interactive action. act The specific characteristics of the second interactive action are determined by a sine function:

[0054] F act =F safe sin(ωt), t∈(0, 100)

[0055] Where w = π / 100 is the frequency of the interaction, and t is the number of samples.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] By actively applying a series of interactive force stimuli to an object, the agent can observe the temporal characteristics of the object's interactive force-interaction response after the stimulus through multimodal sensing devices. Based on these observed characteristics, the agent can infer the implicit interactive characteristics of the object and thus constrain the agent's interactive actions, improve the agent's understanding of the paired interactive force-interaction response temporal signals, thereby improving the agent's embodied interaction adaptation level and enhancing the safety of machine actions. Attached Figure Description

[0058] Figure 1 A diagram of the multimodal sensing system of this invention;

[0059] Figure 2 This is a flowchart of the multimodal sensing process of the present invention;

[0060] Figure 3 This is a flowchart illustrating the environmental information reasoning process of the present invention.

[0061] Figure 4 This is a flowchart of the embodied interaction planning action generation process of the present invention;

[0062] In the figure, there is a multimodal perception module 110, a perception communication and command issuing module 120, an environmental information reasoning module 130, an embodied interaction planning module 140, and a perception-motion coordination module 150. Detailed Implementation

[0063] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0064] This invention proposes an intelligent agent embodied interaction planning method based on active perception of environmental information, the method comprising the following steps:

[0065] S1. Acquire environmental information around the intelligent agent controlling the robot's movement and information at the interface between the intelligent agent and the object. The information at the interface between the intelligent agent and the object is an embodied interaction perception signal, and the environmental information is a two-dimensional image and depth. Extract the most recent sampling data from the acquired information and align the sampling data to obtain aligned environmental information and information at the interface between the intelligent agent and the object.

[0066] S2. Perform target recognition and semantic segmentation on the aligned environmental information to obtain semantic feature information, and fuse the aligned environmental information to obtain the agent's operating environment. Infer interaction characteristic parameters based on the information at the interface between the aligned agent and the object. Perform multimodal information fusion representation on the semantic feature information, the agent's operating environment and the interaction characteristic parameters to obtain multidimensional environmental information.

[0067] S3. Acquire embodied interaction perception signals, and make a predictability judgment on the robot's actions based on the embodied interaction perception signals. If the result is within expectations, then generate a series of embodied interaction actions based on the multidimensional information of the environment and the specific task steps of the plan, and the robot executes the embodied interaction actions.

[0068] Conversely, the method based on the energy gradient of the interaction characteristic parameters generates a series of interactive actions with interactive force characteristics, and the robot executes the interactive actions.

[0069] This invention enables intelligent agents to perceive physical signals at interactive interfaces like humans. By applying active interactive force stimuli to objects, the intelligent agent can observe the spatial state response of objects after stimulation through control methods such as electronic skin vision. The intelligent agent's understanding of the paired interactive force-interaction timing signals allows it to infer the implicit physical characteristics of objects. Through the active acquisition of interactive characteristic information of the surrounding environment, the robot possesses the ability to predict interactive actions. Based on this capability, the intelligent agent will overcome the challenges of perception and motion control brought about by unknown, complex, dense, and human-inhabited work environments, such as accidental touches, limited visual perception, and the inability to solve problems in free movement space.

[0070] The above method can be implemented using the following multimodal sensing system, such as... Figure 1 As shown, the system includes:

[0071] The multimodal perception module 110 is used to effectively collect local working environment information of the intelligent agent, so as to meet the needs of environmental perception and motion planning constraints of the intelligent agent.

[0072] The sensing communication and command issuing module 120 is responsible for the information acquisition and transmission of distributed multi-mode sensing devices and the transmission of intelligent agent motion control intelligence.

[0073] The environmental information reasoning module 130 processes the environmental information collected by the multimodal perception module to realize environmental modeling of the embodied intelligent agent and inference of environmental interaction characteristics.

[0074] The embodied interaction planning module 140 generates embodied interaction motion sequences based on the agent's task and environmental perception reasoning constraints to achieve active perception and ultimately complete the interactive task.

[0075] The perception-motion coordination module 150 determines the predictability of the current intelligent agent's embodied interaction actions, and coordinates and activates the corresponding active interaction perception and task operation planning units for embodied interaction planning.

[0076] Specifically, the multimodal perception module 110 can perceive the environment and the multidimensional information of the interaction between the intelligent agent and the environment. After the perceived data is calibrated for spatiotemporal consistency, it is transmitted to the environmental information reasoning module 130 and the perception-motion coordination module 150 via the perception communication and command issuing module 120 to meet the needs of generating the intelligent agent's embodied actions.

[0077] The perception communication and command issuing module 120 is responsible for the data path of the entire intelligent agent. It transmits the multimodal perception module 110 to the units that need to subscribe, and sends the action commands generated by the embodied interaction planning module 140 to the joint module to complete the embodied action.

[0078] The environmental information reasoning module 130 collects and fuses the collected environmental information to build a map, providing environmental constraints for the embodied interaction planning module 140 to generate actions and ensuring the predictability of the agent's actions.

[0079] The embodied interaction planning module 140 drives the robot to generate embodied interactive actions based on task planning requirements, environmental information constraints fused by the environmental information reasoning module 130, and the action generation unit activated by the perception-motion coordination module 150.

[0080] The perception-motion coordination module 150 evaluates the predictability of the agent's interactive actions based on the information features collected by the multimodal perception module 110 at the interactive interface. Based on the evaluation results, it activates the corresponding action generation unit of the embodied interaction planning module 140 to generate action commands.

[0081] In one embodiment of the present invention: the multimodal perception module 110 includes: an RGB-D visual perception unit 1102: used for geometric and image perception of the surrounding environment of the embodied intelligent agent, to meet the needs of large-scale rapid spatial recognition and mapping; an embodied interaction perception unit 1103: used for collecting object signals at the interaction interface, observing the spatial state response of the object after stimulation, and completing the pairing of interaction force-interaction response timing signals; and a spatiotemporal consistency alignment unit 1104: after the data collected by the multi-type sensing devices are time-aligned, spatiotemporal consistency calibration is completed, and the unified observation signal is transmitted to the environmental information reasoning module 120 through the perception communication and command issuance module.

[0082] Specifically, the specific implementation of the multimodal sensing module 110 ( Figure 2 )as follows:

[0083] The RGB-D visual perception unit 1102 acquires time-series signals of environmental information around the agent in the form of two-dimensional images and depth, and then prints the acquired visual data with timestamps and saves it to the corresponding queue.

[0084] The embodied interaction sensing unit 1103 collects information at the interface between the intelligent agent and the object. The collected information includes: the interaction force stimulation during the interaction process, the object's response displacement, temperature, vibration and other signals are collected in time sequence. Then, the collected multi-dimensional data is timestamped and stored in the corresponding queue in the form of a dictionary.

[0085] The spatiotemporal consistency alignment unit 1104 extracts the most recent sampled data from the data storage queue of the RGB-D visual perception unit 1102 and the data storage queue of the embodied interaction perception unit 1103 according to a specific sampling frequency, aligns the time dimension using the reference sensor alignment method, and sends the aligned data to the perception communication and command issuing module 120.

[0086] Considering the spatial consistency problem of distributed perception in the embodied interaction sensing unit 1103, a forward kinematic chain-based method is used to calculate the spatial position of each embodied interaction sensing device by acquiring the robot's joint position encoding information:

[0087]

[0088] In one embodiment of the present invention: the environmental information reasoning module 130 ( Figure 3 The system includes: a semantic reasoning unit 1302, which infers semantic information about the environment and identifies the semantic attributes of objects in the environment; an environmental geometry reconstruction unit 1303, which models the information of the environment around the agent based on geometric and appearance information from multimodal environmental perception to reconstruct a three-dimensional geometric map; an interaction characteristic reasoning unit 1304, which understands the interaction force-interaction response timing signals collected by the agent's interaction perception to infer the implicit characteristics of objects closely related to embodied interaction attributes that are difficult to observe visually; and a fusion mapping unit 1305, which fuses the constructed three-dimensional map, semantic map, and inferred object interaction features to construct a multi-dimensional environmental information map.

[0089] Specifically, the implementation of the environmental information reasoning module 130 is as follows:

[0090] Semantic reasoning unit 1302: Based on convolutional neural network technology, it performs target recognition and semantic segmentation on the acquired images, and then fuses the segmented semantic features with the object model reconstructed by the environment geometry reconstruction unit 1303 to realize semantic feature reasoning of the object.

[0091] Environmental geometry reconstruction unit 1303: Based on three-dimensional reconstruction technology (Nerf neural network, Gaussian sputtered surface, etc.), it fuses the acquired RGB images and depth point cloud information to perceive the geometric information of the intelligent agent's environment and complete the geometric reconstruction of the intelligent agent's working environment.

[0092] Interactive characteristic reasoning unit 1304: Based on the interactive force-interactive response timing features collected by the embodied interactive perception unit 1103, a data-driven system identification method is used to infer the implicit feature parameters of objects closely related to interactive characteristics.

[0093] The fusion mapping unit 1305: Based on the semantic feature information extracted by the semantic reasoning unit 1302, the environmental geometric information reconstructed by the environmental geometry reconstruction unit 1303, and the interaction characteristic parameters inferred by the interaction characteristic reasoning unit 1304, a multimodal information fusion representation method is used to perform environmental information fusion mapping. The resulting multidimensional environmental information will be used for action generation in the embodied interaction planning module 140.

[0094] In one embodiment of the present invention: the embodied interaction planning module 104 includes: an embodied interaction action generation unit 1402: generating a series of embodied interaction instructions to control the intelligent agent's body to apply a series of interactive actions to the interacting object, and completing the acquisition of interaction force-interaction response timing features; and a task operation action generation unit 1403: generating a series of intelligent agent operation actions with predictable interaction attributes based on environmental information constraints and the intelligent agent's task instructions.

[0095] Specifically, the implementation of the embodied interaction planning module 140 is as follows ( Figure 4 ):

[0096] To ensure coordination and consistency between the embodied interaction action generation unit 1402 and the task operation action generation unit 1403, the perception-motion coordination module 150 first uses data perceived by the multimodal perception module 110 and uses the judgment result to activate the corresponding embodied interaction action generation unit. When the agent's action is within a predictable range, the task operation action generation unit 1403 is activated; otherwise, the embodied interaction action generation unit 1402 is activated.

[0097] Embodied Interaction Action Generation Unit 1402: Based on the object interaction characteristic parameter inference of the interaction characteristic reasoning unit 1304, it generates a series of interactive actions with interactive force characteristics based on the energy gradient method, thereby guiding the intelligent agent to apply force stimulation to the object, and then the embodied interaction perception unit 1103 completes the perception of the object's interaction characteristic information.

[0098] Task action generation unit 1403: For intelligent agent task execution in complex scenarios, based on the multimodal information map and planned specific task steps generated by the environmental information reasoning module 130, a series of embodied interactive actions are generated based on the energy gradient method. By changing the spatial state of the intelligent agent itself and the surrounding environment, the problem of intelligent agent motion stability caused by unknown, complex, dense, and human-inhabited operation scenarios is overcome.

[0099] The specific steps for inferring interaction characteristic parameters from information at the interface between the agent and the object based on alignment are as follows:

[0100] The information obtained from the interface between the aligned agent and the object is as follows:

[0101]

[0102] Where, Δx m It is an interactive displacement. It is the moving speed of the observation point at the interactive interface, F. m It refers to the interaction forces at the interface;

[0103] The information obtained above is used to infer the object's interaction characteristic parameters θ = [K, D, f]. T During the inference process, it is assumed that the projection of the observed data onto the parameter vector is a linear model, and their parameters can be determined using linear fitting methods. The object interaction characteristics are modeled as follows:

[0104] Y(t)=H(t)θ+V(t)

[0105] It is the system superimposed output vector. It is a system overlay information matrix. It is system state noise.

[0106] Subsequently, the least squares parameter estimation method estimates the interaction characteristic parameters of the object:

[0107]

[0108] Assume J1(θ) in When the minimum value is obtained, let J1(θ) be relative to... The partial derivatives are zero, therefore we get

[0109]

[0110] When the action is within the range of unpredictability,

[0111] The energy gradient-based method generates a series of embodied interactive actions, specifically:

[0112] The goal of tactile-based active perception is to apply a series of force stimuli to obstacles and observe their specific shapes in order to obtain a sufficient set of sensory data. The input for active sensing is the interaction force F. act Its interaction force variation characteristics are determined by a sine function:

[0113] F act =F safe sin(ωt), t∈(0, 100)

[0114] Where w = π / 100 is the frequency of the interaction, and t is the number of samples.

[0115] When the action is within the expected range,

[0116] The method based on the energy gradient of interaction characteristic parameters generates a series of interactive actions with interactive force characteristics, specifically:

[0117] First, the robot's workspace is topologically defined as the robot's body domain. Obstacle Domain and free motion space domain in Let r represent the robot's pose, h represent the robot's spatial configuration in the current pose, and x represent the robot's current position. This indicates the robot's workspace. Represents the obstacle domain. Represents the machine body domain. Represents the difference between the workspace and the obstacle domain and the robot body domain;

[0118] Subsequently, different spatial topologies were assigned different energy states, with the target domain being assigned a virtual energy potential field with elastic gravity. The free motion space domain is defined as having a linear damping field U. f (x), the obstacle domain is defined as U with energy operation cost. o,d The work region of W(θ, x) is (x)=W(θ,x).

[0119] Therefore, the energy state of the robot at spatial position x is:

[0120]

[0121] Among them, U o (x) represents U o,d The set of (x);

[0122] By differentiating the energy state, the robot's driving commands are obtained as follows:

[0123] F * = -grad[U(x)]

[0124] The robot's driving commands are interactive actions with interactive force characteristics.

[0125] The specific steps for inferring interaction characteristic parameters from information at the interface between the agent and the object based on alignment are as follows:

[0126] The information obtained from the interface between the aligned agent and the object is as follows:

[0127]

[0128] Where, Δx m It is an interactive displacement. It is the moving speed of the observation point at the interactive interface, F.m It refers to the interaction forces at the interface;

[0129] The information obtained above is used to infer the object's interaction characteristic parameters θ = [K, D, f]. T During the inference process, the interaction characteristics of objects are first modeled as follows:

[0130] Y(t)=H(t)θ+V(t)

[0131] It is the system superimposed output vector. It is a system overlay information matrix. It is system state noise;

[0132] Subsequently, the least squares parameter estimation method estimates the interaction characteristic parameters of the object:

[0133]

[0134] Assume J1(θ) in When the minimum value is obtained, let J1(θ) be relative to... The partial derivatives are zero, therefore we get

[0135]

[0136] Where y(j) represents the time-series input features observed at the interactive interface, v(t) represents noise, and ψ T J1(θ) represents the time-series input signal observed at the interactive interface, J1(θ) represents the system prediction error, K represents the elastic operation coefficient of the object, D represents the damping operation coefficient of the object, and f represents the operation loss constant of the object.

[0137] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A method for planning embodied interaction of intelligent agents based on proactive perception of environmental information, characterized in that, The method includes the following steps: S1. Acquire environmental information around the intelligent agent controlling the robot's movement and information at the interface between the intelligent agent and the object. The information at the interface between the intelligent agent and the object is an embodied interaction perception signal, and the environmental information is a two-dimensional image and depth. Extract the most recent sampling data from the acquired information and align the sampling data to obtain aligned environmental information and information at the interface between the intelligent agent and the object. S2. Perform target recognition and semantic segmentation on the aligned environmental information to obtain semantic feature information, and fuse the aligned environmental information to obtain the agent's operating environment. Infer interaction characteristic parameters based on the information at the interface between the aligned agent and the object. Perform multimodal information fusion representation on the semantic feature information, the agent's operating environment and the interaction characteristic parameters to obtain multidimensional environmental information. S3. Obtain embodied interaction perception signals, and make a predictability judgment on the robot's actions based on the embodied interaction perception signals. If the judgment result is within expectations, then generate a series of first embodied interaction actions based on the multidimensional information of the environment and the specific task steps of the plan, and the robot executes the first embodied interaction actions. Conversely, a series of second interactive actions with interactive force characteristics are generated based on the energy gradient method of the interaction characteristic parameters, and the robot executes the second interactive actions; The information at the interface between the intelligent agent and the object includes the interactive force stimulation during the interaction process, the object's response displacement, temperature and vibration signals, and the embodied interaction perception signal is acquired based on the embodied interaction perception device. The specific steps for determining the predictability of robot actions based on the embodied interaction perception signals are as follows: The planner first bases its algorithm on the target location. Generate the target tracking point for the next moment. : in, It is the efficient solution space of the planner in Cartesian space. This is the robot's current position; Determine whether the robot's speed and interaction forces with the outside world, obtained from embodied interaction sensing signals, exceed the maximum safe speed as the robot approaches the target point. and maximum interaction force The steps for calculating the interaction force are: calculate the robot's acceleration when tracking the target point. : in and Let represent the inertia matrix and the Coriolis matrix, respectively, and q represent the robot's joint space velocity. Represents the gravity vector; Therefore, the robot's speed when tracking the target point is Interactive power Make a predictability judgment on the robot's actions to obtain the robot's actions. The judgment result; The robot's actions The judgment result is: in This is a dangerous action, indicating that it was not expected. For safety reasons, actions should be taken within the expected scope.

2. The intelligent agent embodied interaction planning method based on active perception of environmental information according to claim 1, characterized in that, The spatial location of the embodied interactive sensing device is: in, Indicates the first The Cartesian spatial position of each sensing device relative to the parent joint. Indicates the first The position of each sensing device relative to the robot's center of mass coordinate system. This represents the homogeneous transformation matrix of the parent joint corresponding to the sensing device relative to the robot's center of mass coordinate system.

3. The intelligent agent embodied interaction planning method based on active perception of environmental information according to claim 1, characterized in that, The specific steps for inferring interaction characteristic parameters from information at the interface between the agent and the object based on alignment are as follows: The information obtained from the interface between the aligned agent and the object is as follows: in, It is an interactive displacement. It is the movement speed of the observation point at the interactive interface. It refers to the interaction forces at the interface; The information obtained above is used to infer the interaction characteristic parameters of the object. During the inference process, the interaction characteristics of objects are first modeled as follows: It is the system superimposed output vector. It is a system overlay information matrix. It is system state noise; Subsequently, the least squares parameter estimation method is used to estimate the interaction characteristic parameters of the object, thus obtaining the system prediction error. , Assumption exist When the minimum value is obtained, let right The partial derivatives are zero, therefore we get in, This indicates the time-series input features observed at the interactive interface. v (t) represents noise. This indicates the timing input signal observed at the interactive interface. Indicates the system prediction error. This represents the elastic operating coefficient of an object. This represents the damping operating coefficient of the object. This represents the constant of the object's operational loss.

4. The intelligent agent embodied interaction planning method based on active perception of environmental information according to claim 1, characterized in that, The system prediction error is: in, This indicates the time-series input features observed at the interactive interface. v (t) represents noise. This indicates the timing input signal observed at the interactive interface. This represents the system prediction error.

5. The intelligent agent embodied interaction planning method based on active perception of environmental information according to claim 1, characterized in that, The energy gradient-based method generates a series of first embodied interactive actions, and the robot executes the first embodied interactive actions as follows: First, the robot's workspace is topologically defined as the robot's body domain. Obstacle domain and free motion space domain ,in Let r represent the robot's pose, h represent the robot's spatial configuration in the current pose, and x represent the robot's current position. This indicates the robot's workspace. Represents the obstacle domain. Represents the machine body domain. Represents the difference between the workspace and the obstacle domain and the robot body domain; Subsequently, different spatial topologies were assigned different energy states, with the target domain being assigned a virtual energy potential field with elastic gravity. The free motion space domain is defined as having a linear damping field. Obstacle domains are defined as having energy operation costs. Work domain, By differentiating the energy state, the robot's driving commands are obtained as follows: The robot's driving command is the first embodied interactive action.

6. The intelligent agent embodied interaction planning method based on active perception of environmental information according to claim 5, characterized in that, The energy state is when the robot is in a spatial position. The energy state at that time is specifically as follows: in, express A set of.

7. The intelligent agent embodied interaction planning method based on active perception of environmental information according to claim 1, characterized in that, The method based on the energy gradient of interaction characteristic parameters generates a series of second interactive actions with interactive force characteristics, specifically as follows: Calculate the specific second interactive action The specific characteristics of the second interactive action are determined by a sine function: in It is the frequency of interactive changes. It refers to the number of samples.

Citation Information

Patent Citations

  • Movement assistance device

    CN104582668A

  • Operation skill learning method based on agent active interaction

    CN116719409A