A high-dynamic target physical interaction control system and method for embodied agents and electronic equipment
Patent Information
- Application Number
- CN202610852079.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-09-25
AI Technical Summary
[0007]首先,解决现有强化学习环境采样效率低下的问题,避免智能体在目标处于动力学停滞或脱离有效交互空间等无效状态下,进行无意义的特征采集与算力浪费;
[0050]本发明通过构建多维环境快速重置机制,有效剔除了物理静滞与越界等无效特征,大幅提升了仿真环境的有效采样率并显著节约了算力消耗。
Smart Images

Figure CN122816002A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a highly dynamic target physical interaction control system, method, and electronic device for embodied intelligent agents, belonging to the field of artificial intelligence and robot control technology. Background Technology
[0002] As artificial intelligence extends into the physical world, embodied intelligence has become a core development trend. Driving embodied agents to perform instantaneous embodied perception and precise physical interaction with highly dynamic flying targets (such as high-dynamic ball games in sports, high-speed moving products on industrial assembly lines, and high-speed debris encountered by space stations) is currently a research hotspot in the field of embodied control. Reinforcement learning, due to its powerful end-to-end exploration capabilities, is widely used in policy training for such complex tasks. However, existing embodied reinforcement learning training schemes have the following significant drawbacks:
[0003] (1) Low efficiency and waste of computing power in embodied environment sampling: Existing round termination strategies usually rely only on global timeout settings, such as fixed maximum time step. When the interactive target becomes static in the physical environment, such as losing high speed and slowly rolling on the surface, or getting stuck in a dead corner, or seriously deviating from the effective embodied space of the agent, the system will still blindly collect a large number of static or invalid physical features, resulting in high simulation computing power waste.
[0004] (2) Computational Limitations and Reward Hacking Based on Virtual Coordinate Systems: Existing methods often design reward functions by calculating the absolute geometric distance between the end effector and the target in a global virtual coordinate system. However, due to calibration deviations in the local coordinate system of the underlying physical assets, spatial coordinate system misalignment errors are easily generated, causing the agent to learn false interaction positions that deviate from the real physical contact boundary. More seriously, the agent often exploits the loopholes in virtual distance rewards to generate locally optimal "reward hacking" behavior. For example, it may violate physical common sense by using its own body (rather than the end effector) to illegally intercept the target, or by outputting action commands with extremely high variance to cause high-frequency disordered oscillations in physical devices to increase the probability of random collisions. This strategy, which lacks strict physical contact constraints, generates action trajectories that are completely ineffective in the actual deployment stage, seriously hindering the smooth migration of embodied strategies to the real physical world (Sim-to-Real).
[0005] (3) Cold start difficulty under high dynamic sparse rewards: In highly dynamic embodied interaction tasks with global randomness, the success rate of the model's initial exploration is extremely low. The inability to obtain positive physical contact reward signals for a long time can easily lead to the degradation of the value network, making it impossible for the embodied strategy to converge. Summary of the Invention
[0006] The purpose of this invention is to address the problems existing in the prior art, and to provide a highly dynamic target physical interaction control system, method, and electronic device for embodied intelligent agents. The specific problems it solves are as follows:
[0007] First, address the problem of low sampling efficiency in existing reinforcement learning environments, and avoid meaningless feature collection and wasted computing power by agents when the target is in an ineffective state such as dynamic stagnation or out of the effective interaction space;
[0008] Secondly, it addresses the spatial boundary alignment error caused by calculations based on purely virtual coordinates, as well as the problem of "reward deception" behavior caused by agents exploiting rule loopholes. It also eliminates the high-frequency disordered oscillations of physical devices, forcing the model to learn action strategies that conform to real physical contact constraints, have smooth trajectories, and can be directly deployed on real machines.
[0009] Finally, the project addresses the difficulties of exploration divergence and cold start caused by the extremely low initial interaction fault tolerance in highly dynamic small target tasks, and achieves automated and smooth convergence of the model from a restricted local space to a global random scene.
[0010] The objective of this invention is achieved through the following technical solution:
[0011] A highly dynamic target physical interaction control system for embodied intelligent agents includes: an action execution and physical sensing module, a reinforcement learning strategy control module, and an environmental state monitoring module; the signal output terminals of the action execution and physical sensing module and the environmental state monitoring module are all connected to the signal input terminal of the reinforcement learning strategy control module.
[0012] The action execution and physical sensing module is used to simulate and collect real collision force feedback from various parts in real time through the physics engine, and transmit it as privileged information to the reinforcement learning strategy control module in one direction. At the same time, it receives instructions from the control brain to drive the core joints to perform highly dynamic interception actions.
[0013] The reinforcement learning strategy control module is used to receive signals from the action execution and physical sensing module and the environmental status monitoring module, thereby removing the dependence on complex sensor data in actual deployment and guiding the network to converge smoothly.
[0014] The environmental status monitoring module is used to process multi-dimensional status features in parallel in the background, and to judge in real time whether the current round has timed out, whether the target's absolute coordinates have escaped the compact physical boundary, and whether the target has fallen into a low-speed stuck state.
[0015] Preferably, the motion execution and physical sensing module includes: a basic support and torso linkage, an end effector, and a physical contact sensor;
[0016] The basic support and torso linkage are used to support the overall physical form of the multi-axis robotic arm and serve as the motion carrier for the end effector, providing the necessary degrees of freedom and workspace.
[0017] An end effector is used to directly physically interact with a highly dynamic target object, translating control commands into final embodied actions;
[0018] Physical contact sensors are used to monitor and collect force feedback between various parts of the robot and the target object in real time. During the simulation training phase, they serve as a privileged information source to distinguish between collision states of the end effector (legal interaction) and the torso (cheating), providing truth value basis for reward reshaping.
[0019] Preferably, the reinforcement learning strategy control module includes: an asymmetric network architecture and a reward and course learning control unit, wherein the signal output terminal of the reward and course learning control unit is connected to the signal input terminal of the asymmetric network architecture;
[0020] An asymmetric network architecture includes: an actor network and a commentator network, wherein:
[0021] The executor network is used to output action instructions based on regular observation information and to receive numerical scalars calculated by the reward and course learning control unit.
[0022] The commentator network, used for accurate value assessment, receives contact privilege information from the action execution and physical sensing module and status signals from the environmental status monitoring module;
[0023] The reward and course learning control unit is used to determine cheating and truncate reward calculations based on multidimensional feedback, and dynamically adjust the training difficulty and action variance to guide the network to converge smoothly.
[0024] Preferably, the environmental status monitoring module includes: time dimension truncation monitoring, spatial dimension bounding box monitoring, and physical kinetic energy anti-static monitoring;
[0025] The time-dimensional truncation monitoring is used to record and monitor the environmental simulation step size of a single round in real time. When the set maximum step threshold is reached, the round is terminated and a truncation signal is explicitly returned to the reinforcement learning policy control module to ensure the bootstrap evaluation accuracy of the value network and prevent the simulation from falling into an infinite loop.
[0026] Spatial dimension bounding box monitoring is used to construct a three-dimensional spatial boundary threshold around a set effective physical interaction area, compare the absolute coordinates of the target object in real time, and instantly determine the interaction failure and trigger the environment reset when the target escapes the set three-dimensional boundary, so as to avoid the intelligent agent from performing meaningless feature collection and wasting computing power in invalid space.
[0027] The physical kinetic energy anti-static monitoring is used to extract and calculate the three-dimensional composite linear velocity of interactive targets in real time. After a preset grace period, if the composite velocity continues to be lower than the static stagnation judgment threshold, the current physical environment is determined to be in a stagnant state and a reset is directly triggered to eliminate invalid physical states such as target stuck or landing still.
[0028] A control method based on a highly dynamic target physical interaction control system for embodied intelligent agents includes the following steps:
[0029] Step 1: The physical state of the agent and the interactive target in the simulation environment is obtained through the action execution and physical sensing module, and the force feedback is collected simultaneously as contact privilege information and input to the commentator network of the reinforcement learning policy control module; then the environment state monitoring module implements a multi-dimensional round termination and reset strategy, including setting spatial control boundaries, on the physical state, and outputs the effective training state after removing invalid samples.
[0030] Step 2: Based on the effective training state and privileged information output in Step 1, the reinforcement learning policy control module implements a reward reshaping method based on privileged information perception and stage truncation to calculate and output a reward value with physical common sense constraints.
[0031] Step 3: Combining the spatial control boundary from Step 1 with the reward value with physical common sense constraints calculated in Step 2, the reinforcement learning strategy control module implements a convergence-driven variance dynamic constraint and course learning control scheme, outputting a control model that achieves global generalization convergence.
[0032] Preferably, the specific steps of the multi-dimensional round termination and reset strategy described in step one are as follows:
[0033] Step 11: Time and Truncation Declaration: Set the maximum number of environment simulation steps, and explicitly declare truncation rather than task failure to the reinforcement learning policy optimization algorithm when timeout occurs, so as to ensure the accuracy of value network evaluation.
[0034] Steps 1 and 2: General criterion for the bounding box of the embodied feasible space: Let the absolute coordinates of the interactive target in the virtual environment be... , , The boundary thresholds for each axis of the three-dimensional bounding box for effective interaction of intelligent agents are set as follows: and , and , and When the value exceeds either the upper or lower limit, the interaction is immediately determined to have failed and the environment is reset.
[0035] Step 13: Reset Physical Kinetic Energy Detection: Let the three-dimensional linear velocity components of the interactive target object be respectively... In fact, the resultant linear velocity The general calculation formula is:
[0036]
[0037] After the initial grace period is set, if The hysteresis threshold remains consistently below the preset threshold. If so, the physical environment is determined to be stagnant and a reset is triggered.
[0038] Preferably, the specific steps of the reward reshaping method based on privileged information perception and stage truncation in step two are as follows:
[0039] Step 21: Simulation Privilege Anti-Cheat Interception: During the simulation training phase, virtual physical collision sensors are configured as privileged information sources for the illegal interaction parts of the embodied intelligent agent's main body. If force is detected, the round will be terminated immediately to eliminate the non-embodied cheating path of the intelligent agent using the main body's torso to illegally block.
[0040] Step 22: General Calculation of Truncated Approximation Reward: Let the spatial Euclidean distance between the end effector feature point and the target feature point be... The relative velocity component of the target object flying towards the agent is The reward function is constructed to continuously approximate the following:
[0041]
[0042] in, A distance scaling factor greater than zero. It is a natural constant, used in conjunction with the truncation indicator function logic: when The reward will be issued at the appropriate time. Once the target is successfully intercepted, the speed direction reverses. At that moment, a forced truncation was instantly triggered. Eliminate passive score-boosting behavior;
[0043] Steps 2 and 3: Privileged Contact Sparse Reward and Action Smoothness Penalty: A sparse reward for successful interaction is only issued when the end effector collision sensor, which serves as privileged information, detects a net contact force exceeding a set threshold. At the same time, negative weight penalties are introduced for abrupt acceleration / deceleration changes in embodied actions and joint motion states, forcing the network to converge into a smooth and deterministic motion trajectory without relying on physical contact sensors.
[0044] Preferably, the specific steps for implementing the convergence-driven variance dynamic constraint and course learning control scheme in step three are as follows:
[0045] Step 31: Spatial contraction and high-entropy exploration: In the early stage of model training, the random range of the interactive target's emission is contracted from the spatial dimension, while the action distribution exploration parameters are maintained at the bottom layer of the reinforcement learning algorithm, so as to enable the agent to conduct sufficient posture exploration in a limited and safe physical space and quickly establish the initial effective contact memory.
[0046] Step 32: Dynamic Entropy Decay and Variance Suppression: A dynamic decay strategy for the entropy coefficient of the action distribution based on the training process is introduced. After the system monitors that the end-contact reward score shows a steady upward trend, the variance dynamic decay mechanism is triggered. By smoothly reducing the entropy coefficient of the action distribution, the standard deviation of the action network output is forcibly reduced, the exploration noise is removed, and the high-frequency mechanical jitter of the embodied agent in the physical space is eliminated.
[0047] Step 33: Global Generalization and Deployment Convergence: The standard deviation of the action output is reduced to below the preset safe convergence threshold as the convergence mark of the current stage. Then, the emission random space of the interaction target is gradually expanded until global generalization is achieved. Finally, a physical interaction model with deterministic actions that can be directly deployed in the actual machine is output.
[0048] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements a highly dynamic target physical interaction control method for embodied intelligent agents.
[0049] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0050] This invention effectively eliminates invalid features such as physical stasis and out-of-bounds errors by constructing a multi-dimensional environment fast reset mechanism, which greatly improves the effective sampling rate of the simulation environment and significantly saves computing power consumption.
[0051] This invention completely blocks the cheating path of agents using virtual coordinate vulnerabilities to block the torso by introducing physical contact privilege information and truncated reward reconstruction strategy at the training end. It also eliminates high-frequency mechanical vibrations through smooth constraints and completely eliminates the dependence on expensive high-frequency force sensors in actual deployment by combining with asymmetric network architecture. The generated deterministic trajectory effectively bridges the gap between simulation and reality.
[0052] This invention introduces a course learning scheme based on spatial dynamic constraints and entropy decay. Through initial dimensionality reduction exploration, it successfully solves the cold start problem of highly dynamic tasks under extremely low fault tolerance. With the help of variance suppression guidance strategy in the later stage, it achieves a smooth transition and finally realizes automatic convergence in global random scenarios.
[0053] This invention effectively eliminates high-frequency mechanical oscillations and physical logic loopholes in the early stages of exploration, ultimately achieving precise interception of highly dynamic targets without reliance on physical force sensors and a smooth transition from simulation to real-world deployment. Furthermore, this invention is applicable to highly dynamic target interception and strike missions requiring instantaneous perception and precise physical interaction. Attached Figure Description
[0054] Figure 1 This is a schematic diagram of the architecture of a highly dynamic target physical interaction control system for embodied intelligent agents according to the present invention.
[0055] Figure 2 This is a flowchart of a highly dynamic target physical interaction control method for embodied intelligent agents according to the present invention. Detailed Implementation
[0056] The present invention will be further described in detail below with reference to the accompanying drawings: This embodiment is implemented under the premise of the technical solution of the present invention, and detailed implementation methods are given, but the protection scope of the present invention is not limited to the following embodiments.
[0057] like Figure 1 As shown in the figure, the high dynamic target physical interaction control system for embodied intelligent agents involved in this embodiment consists of three core parts: an action execution and physical sensing module, a reinforcement learning strategy control module, and an environmental state monitoring module, which ensures seamless migration from simulation training to actual deployment.
[0058] The motion execution and physical sensing module, serving as the physical interaction foundation and data acquisition terminal of the system, internally includes: basic support and torso links, end effectors, and physical contact sensors. In the simulation environment, this module is responsible for simulating the force sensors on the robotic arm links through the physics engine of the simulation software during the simulation training phase, and collecting real-time collision force feedback from various parts. This feedback is then transmitted unidirectionally as privileged information to the reinforcement learning policy control module. Simultaneously, it receives commands from the control brain to drive the core joints to perform highly dynamic interception actions.
[0059] The reinforcement learning strategy control module, serving as the system's core, integrates an asymmetric network architecture and a reward and curriculum learning control unit. The asymmetric network architecture is further divided into an actor network and a commentator network. The commentator network is specifically responsible for receiving contact privilege information from the action execution and physical sensing modules, as well as state signals from the environmental state monitoring module, for accurate value assessment; while the actor network outputs action commands based solely on conventional observation information. This asymmetric architecture fundamentally eliminates the reliance on complex sensor data in actual deployment. The reward and curriculum learning control unit is responsible for cheat detection and reward truncation calculations based on multi-dimensional feedback, and dynamically adjusts training difficulty and action variance to guide the network to smooth convergence.
[0060] The environmental status monitoring module, serving as the system's boundary guardian and status assessment arm, internally includes: time-dimensional truncation monitoring, spatial-dimensional bounding box monitoring, and physical kinetic anti-static monitoring. This module is responsible for parallel processing of multi-dimensional state features in the background, real-time analysis of whether the current round has timed out, whether the target's absolute coordinates have exceeded the compact physical boundary, and whether the target has entered a low-speed, stuck state. Once any of these conditions are triggered, the module sends an environmental reset and truncation signal to the control module, thereby significantly eliminating invalid feature samples and saving simulation computing power.
[0061] The motion execution and physical sensing module includes: a basic support and torso linkage, an end effector, and a physical contact sensor;
[0062] The basic support and torso linkage are used to support the overall physical form of the multi-axis robotic arm and serve as the motion carrier for the end effector, providing the necessary degrees of freedom and workspace.
[0063] An end effector is used to directly interact physically with highly dynamic target objects (such as striking, intercepting, or grabbing), translating control commands into final embodied actions. It is the core end unit for completing interactive tasks.
[0064] Physical contact sensors are used to monitor and collect force feedback between various parts of the robot and the target object in real time. During the simulation training phase, they serve as a privileged information source to distinguish the collision state between the end effector (legitimate interaction) and the torso (cheating), providing a truth value basis for reward reshaping.
[0065] The reinforcement learning strategy control module includes: an asymmetric network architecture and a reward and course learning control unit;
[0066] An asymmetric network architecture includes: an actor network and a commentator network, wherein:
[0067] An actor network is used to output action instructions based on regular observation information;
[0068] The commentator network, used for accurate value assessment, receives contact privilege information from the action execution and physical sensing module and status signals from the environmental status monitoring module;
[0069] The reward and course learning control unit is used to determine cheating and truncate reward calculations based on multidimensional feedback, and dynamically adjust the training difficulty and action variance to guide the network to converge smoothly.
[0070] The asymmetric network architecture and the reward and curriculum learning control unit form a control relationship based on parameter calculation and dynamic adjustment. Specifically, the approximation reward and sparse contact reward calculated by the reward and curriculum learning control unit are input into the commentator network in the asymmetric network architecture to calculate the value loss and update the network weights. Simultaneously, this unit dynamically adjusts the curriculum parameters based on the stage-by-stage task performance: on the one hand, the calculated dynamic variance coefficients are unidirectionally input into the executor network to constrain the output distribution range of action commands, thus eliminating mechanical oscillations; on the other hand, this unit simultaneously expands the random initial space of the interaction target to increase the task difficulty. Through the coordination of the above reward feedback and curriculum difficulty, the model is driven to smoothly transition from the initial limited exploration, ultimately achieving automatic convergence in a wide range of random scenarios.
[0071] The environmental status monitoring module includes: time dimension truncation monitoring, spatial dimension bounding box monitoring, and physical kinetic energy anti-static monitoring.
[0072] The time-dimensional truncation monitoring is used to record and monitor the environmental simulation step size of a single round in real time. When the set maximum step threshold is reached, the round is terminated and a truncation signal is explicitly returned to the control module to ensure the bootstrap evaluation accuracy of the value network and prevent the simulation from falling into an infinite loop.
[0073] Spatial dimension bounding box monitoring is used to construct a three-dimensional spatial boundary threshold around a set effective physical interaction area, compare the absolute coordinates of the target object in real time, and instantly determine the interaction failure and trigger the environment reset when the target escapes the set three-dimensional boundary, so as to avoid the intelligent agent from performing meaningless feature collection and wasting computing power in invalid space.
[0074] The physical kinetic energy anti-static monitoring is used to extract and calculate the three-dimensional composite linear velocity of interactive targets in real time. After a preset grace period, if the composite velocity continues to be lower than the static stagnation judgment threshold, the current physical environment is determined to be in a stagnant state and a reset is directly triggered to eliminate invalid physical states such as target stuck or landing still.
[0075] like Figure 2 The diagram illustrates the operational flow of a highly dynamic target physical interaction control system for embodied intelligent agents, as described in this invention. This flow encompasses the process from environment initialization and single-step control loops to curriculum evolution.
[0076] After system startup, the system first enters the environment initialization and dimensionality reduction exploration phase. In order to overcome the cold start problem of highly dynamic tasks, the system actively shrinks the launch space of the interactive target and simultaneously lowers the initial action standard deviation when initializing system parameters, so as to suppress the early ineffective violent oscillations of the agent and ensure that it establishes initial contact memory within the limited physical space.
[0077] Subsequently, the system enters a high-frequency single-step control loop, extracting the physical state and privileged information of the current control step, and performing a rigorous 3D environment reset judgment in parallel. First, it checks if the maximum simulation step size for a single round has been reached. If so, a truncation signal is returned to the reinforcement learning policy control module, and a bootstrap evaluation is performed to prevent value network distortion. Next, it checks if the target exceeds the 3D bounding box boundary. If so, the interaction is directly deemed a failure and a reset is triggered. Further, it checks if the target's linear velocity synthesis value remains continuously below a set hysteresis threshold. If so, the environment is deemed stuck and a reset is immediately initiated. These three lines of defense effectively filter out various invalid environment explorations.
[0078] After passing the environmental reset judgment, the system enters the evaluation and feedback phase. The system uses privileged information to strictly check whether physical collisions or forces have occurred at non-endpoint execution parts. If forces are found, it is judged as a violation or cheating behavior, and a high penalty is immediately imposed and the round is reset. If no violation occurs, the system continues to calculate the truncated approximation reward and the endpoint sparse contact reward, forcing the agent to learn the real physical interaction rules.
[0079] After completing the single-step reward settlement, the system enters the judgment stage of course training evolution. The system monitors in real time whether the phase contact score shows a steady upward trend. If no significant increase is observed, the current system parameters are maintained and the system returns to extracting the physical state of the current control step, continuing training at the current difficulty level. If the score steadily increases, the system triggers a variance dynamic decay mechanism to smoothly reduce the entropy coefficient of the action distribution to eliminate high-frequency spatial jitter, and gradually expands the target's launch space to improve generalization ability. If the global generalization convergence condition is not met, a reset is triggered, and the system continuously loops the above judgment and control process, returning to the step of "extracting the physical state and privileged information of the current control step" until the global generalization convergence condition is met, finally outputting a deployment model that is completely independent of privileged information and can be directly used for independent operation on a real machine.
[0080] This invention provides a highly dynamic target physical interaction control method for embodied intelligent agents. This method introduces real physical feedback as a core constraint into a reinforcement learning framework, and its specific steps are as follows:
[0081] Step 1: The physical state of the agent and the interactive target in the simulation environment is obtained through the action execution and physical sensing module, and the force feedback is collected simultaneously as contact privilege information and input to the commentator network of the reinforcement learning policy control module; then the environment state monitoring module implements a multi-dimensional round termination and reset strategy, including setting spatial control boundaries, on the physical state, and outputs the effective training state after removing invalid samples.
[0082] (a) Time and Truncation Declaration: Set the maximum number of environment simulation steps, and explicitly declare truncation to the reinforcement learning policy optimization algorithm when timeout occurs, instead of task failure, to ensure the accuracy of value network evaluation.
[0083] The reinforcement learning policy optimization algorithm adopts the optimizer part of algorithms such as PPO, SAC, or DQN, which are common modules in commonly used reinforcement learning algorithms; in this specific embodiment, the PPO algorithm is preferred.
[0084] (b) General criterion for embodied feasible bounding box: Let the absolute coordinates of the interactive target in the virtual environment be... , , The boundary thresholds for each axis of the three-dimensional bounding box for effective interaction of intelligent agents are set as follows: and , and , and When the interaction exceeds either the upper or lower limit, the interaction is immediately deemed to have failed and an environment reset is triggered.
[0085] (c) Physical kinetic energy detection reset: Let the three-dimensional linear velocity components of the interactive target object be respectively In fact, the resultant linear velocity The general calculation formula is:
[0086]
[0087] After the initial grace period is set, if The hysteresis threshold remains consistently below the preset threshold. If so, the physical environment is determined to be stagnant and a reset is triggered.
[0088] Step 2: Based on the effective training state and privileged information output in Step 1, the reinforcement learning policy control module implements a reward reshaping method based on privileged information perception and stage truncation to calculate and output a reward value with physical common sense constraints.
[0089] (a) Simulation privilege anti-cheating interception: During the simulation training phase, virtual physical collision sensors are configured as privileged information sources for the illegal interaction parts of the embodied intelligent agent body. If force is detected, the round will be terminated immediately to eliminate the non-embodied cheating path of the intelligent agent using the body to illegally block.
[0090] (b) General calculation of truncated approximation reward: Let the spatial Euclidean distance between the end effector feature point and the target feature point be... The relative velocity component of the target object flying towards the agent is The reward function that continuously approximates the reward is constructed as follows:
[0091]
[0092] in, A distance scaling factor greater than zero. It is a natural constant. In conjunction with the truncation indicator function logic: when The reward will be issued at the appropriate time. Once the target is successfully intercepted, the speed direction reverses. At that moment, a forced truncation was instantly triggered. We must eliminate any passive score-boosting behavior.
[0093] (c) Privileged contact sparsity reward and motion smoothness penalty: A sparse reward for successful interaction is only issued when the end effector collision sensor, which serves as privileged information, detects a net contact force higher than a set threshold. At the same time, negative weight penalties are introduced for abrupt acceleration and deceleration of embodied motion and joint motion state, forcing the network to converge a smooth and deterministic motion trajectory without relying on the actual contact sensor.
[0094] Step 3: Combining the spatial control boundary in Step 1 with the stage reward performance calculated in Step 2, the reinforcement learning strategy control module implements a convergence-driven variance dynamic constraint and course learning control scheme, outputting a control model that achieves global generalization convergence.
[0095] (a) Spatial contraction and high-entropy exploration: In the early stage of model training, the random range of the interactive target is contracted from the spatial dimension, while the action distribution exploration parameters (such as high-entropy coefficient) are maintained at the bottom layer of the reinforcement learning algorithm, so as to enable the agent to conduct sufficient posture exploration in a limited and safe physical space and quickly establish the initial effective contact memory.
[0096] (b) Dynamic Entropy Decay and Variance Suppression: A dynamic decay strategy for the entropy coefficient of the action distribution based on the training process is introduced. After the system monitors that the end-contact reward score shows a steady upward trend, the dynamic variance decay mechanism is triggered. By smoothly reducing the entropy coefficient of the action distribution, the standard deviation of the action network output is forcibly reduced, thus removing exploration noise and eliminating high-frequency mechanical jitter of the embodied agent in physical space.
[0097] (c) Global generalization and deployment convergence: The standard deviation of the action output is reduced to below the preset safe convergence threshold as the convergence mark of the current stage. Then, the emission random space of the interaction target is gradually expanded until global generalization is achieved. Finally, a physical interaction model with deterministic actions that can be directly deployed in the actual machine is output.
[0098] Example 1
[0099] This embodiment details the specific application of the high-dynamic target physical interaction control method and system for embodied intelligent agents described in this invention, which can be widely applied to high-dynamic interception scenarios such as sports robots and high-speed flying piece grasping in industrial production lines. The following describes the general implementation scheme of this invention in detail with specific method steps and system architecture.
[0100] 1. This embodiment uses the specific application scenario of a robotic arm intercepting a high-speed ping-pong ball as an example to explain in detail the general implementation scheme of the control method and electronic device described in this invention. The system implementation of this invention is strictly divided into two parts: a simulation training platform and a physical deployment electronic device. In the simulation training platform, the processor runs an asymmetric reinforcement learning network containing executors and commentators. It obtains the force data of the robotic arm linkage and the ping-pong paddle through the collision force sensor built into the physics engine. This force data is only used as privileged information input to the commentator network for reward reshaping and environment reset determination. In the physical deployment electronic device that completes the model migration to the real physical world, the hardware topology includes at least one processor, memory, a multi-axis robotic arm body, a ping-pong paddle as the end effector, and basic vision and joint position sensors. There is no need to install additional physical contact sensors. The physical processor only runs the executor network model after training convergence, receives the normal environmental state data stripped of privileged information in real time through the underlying communication bus, and issues high-frequency motion control commands to the robotic arm joint motors. In the algorithm motion mapping stage, only the core execution joints are exposed to the policy network.
[0101] 2. Implementation Scheme for Rapid Environment Reset Based on Multidimensional State Characteristics. The system performs three-dimensional reset determination in parallel within each control cycle. In the time dimension, a maximum simulation step size is set for each round; when this step size is reached, the round terminates with an explicit return cutoff signal, prompting the network to perform bootstrapping evaluation. In the spatial dimension, based on the aforementioned general bounding box determination logic, a compact three-dimensional spatial bounding box is constructed around the effective ping-pong table. If the absolute coordinates of the ping-pong ball escape this set boundary, the interaction is instantly determined to have failed and a reset is initiated. In the physical kinetic energy dimension, the three-dimensional linear velocity components of the ping-pong ball are extracted in real time. And through a general formula Calculate its velocity. After the initial grace period, if the velocity of the ping-pong ball... If the desktop remains below the preset static threshold, the environment will be determined to be frozen and reset.
[0102] 3. A truncated reward refactoring implementation plan with physical constraints. During the simulation evaluation phase, when the virtual collision sensor (which serves as privilege information) detects a non-paddle component, such as a robotic arm link, obstructing the ping-pong ball, the round is immediately terminated with a high penalty. The system extracts the velocity component of the ping-pong ball flying towards the robotic arm in real time as a general parameter. And calculate the spatial Euclidean distance between the center of the racket and the ping-pong ball. The system defines a continuously approximating reward using an exponential function. ,in This is the distance scaling factor. Only when... The reward is given when the ping-pong ball comes flying towards you. Once detected That is, when the ping-pong ball is hit back, the function instantly forces a cutoff. Finally, a fixed sparse bonus is awarded only when the virtual force sensor on the ping-pong paddle detects a net contact force higher than a preset threshold, and a negative penalty weight is applied to joint acceleration and deceleration abrupt changes in each control step, forcing the executor network to converge to a smooth force trajectory under the condition of stripping away privileged information.
[0103] A course learning control implementation scheme based on dynamic entropy decay. In the initial training phase, the system spatially shrinks the initial launch coordinates and velocity random range of the ping-pong ball, and lowers the logarithmic standard deviation of the underlying initial motion sampling. After monitoring a steady increase in the racket contact reward score, the system dynamically and smoothly reduces the entropy coefficient of the motion distribution, forcibly suppressing the motion output standard deviation to below a preset convergence threshold. Subsequently, it gradually expands the ping-pong ball launch random space to global generalization, outputting a real-world deployment model that can independently control a physical robotic arm.
[0104] The above description is merely a preferred embodiment of the present invention. These specific embodiments are different implementations based on the overall concept of the present invention, and the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A highly dynamic target physical interaction control system for embodied intelligent agents, characterized in that, include: The system includes an action execution and physical sensing module, a reinforcement learning policy control module, and an environmental state monitoring module. The signal output terminals of the action execution and physical sensing module and the environmental state monitoring module are all connected to the signal input terminal of the reinforcement learning policy control module. The action execution and physical sensing module is used to simulate and collect real collision force feedback from various parts in real time through the physics engine, and transmit it as privileged information to the reinforcement learning strategy control module in one direction. At the same time, it receives instructions from the control brain to drive the core joints to perform highly dynamic interception actions. The reinforcement learning strategy control module is used to receive signals from the action execution and physical sensing module and the environmental status monitoring module, thereby removing the dependence on complex sensor data in actual deployment and guiding the network to converge smoothly. The environmental status monitoring module is used to process multi-dimensional status features in parallel in the background, and to judge in real time whether the current round has timed out, whether the target's absolute coordinates have escaped the compact physical boundary, and whether the target has fallen into a low-speed stuck state.
2. The high dynamic target physical interaction control system for embodied intelligent agents according to claim 1, characterized in that, The motion execution and physical sensing module includes: a basic support and torso linkage, an end effector, and a physical contact sensor; The basic support and torso linkage are used to support the overall physical form of the multi-axis robotic arm and serve as the motion carrier for the end effector, providing the necessary degrees of freedom and workspace. An end effector is used to directly physically interact with a highly dynamic target object, translating control commands into final embodied actions; Physical contact sensors are used to monitor and collect force feedback between various parts of the robot and the target object in real time. During the simulation training phase, they serve as a privileged information source to distinguish between collision states of the end effector (legal interaction) and the torso (cheating), providing truth value basis for reward reshaping.
3. A high-dynamic target physical interaction control system for embodied intelligent agents according to claim 1, characterized in that, The reinforcement learning strategy control module includes: an asymmetric network architecture and a reward and course learning control unit, wherein the signal output terminal of the reward and course learning control unit is connected to the signal input terminal of the asymmetric network architecture. An asymmetric network architecture includes: an actor network and a commentator network, wherein: The executor network is used to output action instructions based on regular observation information and to receive numerical scalars calculated by the reward and course learning control unit. The commentator network, used for accurate value assessment, receives contact privilege information from the action execution and physical sensing module and status signals from the environmental status monitoring module; The reward and course learning control unit is used to determine cheating and truncate reward calculations based on multidimensional feedback, and dynamically adjust the training difficulty and action variance to guide the network to converge smoothly.
4. A high-dynamic target physical interaction control system for embodied intelligent agents according to claim 1, characterized in that, The environmental status monitoring module includes: time dimension truncation monitoring, spatial dimension bounding box monitoring, and physical kinetic energy anti-static monitoring. The time-dimensional truncation monitoring is used to record and monitor the environmental simulation step size of a single round in real time. When the set maximum step threshold is reached, the round is terminated and a truncation signal is explicitly returned to the reinforcement learning policy control module to ensure the bootstrap evaluation accuracy of the value network and prevent the simulation from falling into an infinite loop. Spatial dimension bounding box monitoring is used to construct a three-dimensional spatial boundary threshold around a set effective physical interaction area, compare the absolute coordinates of the target object in real time, and instantly determine the interaction failure and trigger the environment reset when the target escapes the set three-dimensional boundary, so as to avoid the intelligent agent from performing meaningless feature collection and wasting computing power in invalid space. The physical kinetic energy anti-static monitoring is used to extract and calculate the three-dimensional composite linear velocity of interactive targets in real time. After a preset grace period, if the composite velocity continues to be lower than the static stagnation judgment threshold, the current physical environment is determined to be in a stagnant state and a reset is directly triggered to eliminate invalid physical states such as target stuck or landing still.
5. A control method for a highly dynamic target physical interaction control system for embodied intelligent agents according to any one of claims 1-4, characterized in that, Includes the following steps: Step 1: The physical state of the agent and the interactive target in the simulation environment is obtained through the action execution and physical sensing module, and the force feedback is collected simultaneously as contact privilege information and input to the commentator network of the reinforcement learning policy control module; then the environment state monitoring module implements a multi-dimensional round termination and reset strategy, including setting spatial control boundaries, on the physical state, and outputs the effective training state after removing invalid samples. Step 2: Based on the effective training state and privileged information output in Step 1, the reinforcement learning policy control module implements a reward reshaping method based on privileged information perception and stage truncation to calculate and output a reward value with physical common sense constraints. Step 3: Combining the spatial control boundary from Step 1 with the reward value with physical common sense constraints calculated in Step 2, the reinforcement learning strategy control module implements a convergence-driven variance dynamic constraint and course learning control scheme, outputting a control model that achieves global generalization convergence.
6. The control method according to claim 5, characterized in that, The specific steps of the multi-dimensional round termination and reset strategy described in Step 1 are as follows: Step 11: Time and Truncation Declaration: Set the maximum number of environment simulation steps, and explicitly declare truncation rather than task failure to the reinforcement learning policy optimization algorithm when timeout occurs, so as to ensure the accuracy of value network evaluation. Steps 1 and 2: General criterion for the bounding box of the embodied feasible space: Let the absolute coordinates of the interactive target in the virtual environment be... , , The boundary thresholds for each axis of the three-dimensional bounding box for effective interaction of intelligent agents are set as follows: and , and , and When the value exceeds either the upper or lower limit, the interaction is immediately determined to have failed and the environment is reset. Step 13: Reset Physical Kinetic Energy Detection: Let the three-dimensional linear velocity components of the interactive target object be respectively... In fact, the resultant linear velocity The general calculation formula is: After the initial grace period is set, if The hysteresis threshold remains consistently below the preset threshold. If so, the physical environment is determined to be stagnant and a reset is triggered.
7. The control method according to claim 5, characterized in that, The specific steps of the reward reshaping method based on privileged information perception and stage truncation described in step two are as follows: Step 21: Simulation Privilege Anti-Cheat Interception: During the simulation training phase, virtual physical collision sensors are configured as privileged information sources for the illegal interaction parts of the embodied intelligent agent's main body. If force is detected, the round will be terminated immediately to eliminate the non-embodied cheating path of the intelligent agent using the main body's torso to illegally block. Step 22: General Calculation of Truncated Approximation Reward: Let the spatial Euclidean distance between the end effector feature point and the target feature point be... The relative velocity component of the target object flying towards the agent is The reward function is constructed to continuously approximate the following: in, A distance scaling factor greater than zero. It is a natural constant, used in conjunction with the truncation indicator function logic: when The reward will be issued at the appropriate time. Once the target is successfully intercepted, the speed direction reverses. At that moment, a forced truncation was instantly triggered. Eliminate passive score-boosting behavior; Steps 2 and 3: Privileged Contact Sparse Reward and Action Smoothness Penalty: A sparse reward for successful interaction is only issued when the end effector collision sensor, which serves as privileged information, detects a net contact force exceeding a set threshold. At the same time, negative weight penalties are introduced for abrupt acceleration / deceleration changes in embodied actions and joint motion states, forcing the network to converge into a smooth and deterministic motion trajectory without relying on physical contact sensors.
8. The control method according to claim 5, characterized in that, The specific steps for implementing the convergence-driven variance dynamic constraint and course learning control scheme described in step three are as follows: Step 31: Spatial contraction and high-entropy exploration: In the early stage of model training, the random range of the interactive target is contracted from the spatial dimension, while the action distribution exploration parameters are maintained in the reinforcement learning policy control module, so as to enable the agent to conduct sufficient posture exploration in a limited and safe physical space and quickly establish the initial effective contact memory. Step 32: Dynamic Entropy Decay and Variance Suppression: A dynamic decay strategy for the entropy coefficient of the action distribution based on the training process is introduced. After the system monitors that the end-contact reward score reaches the preset stage reward threshold, the variance dynamic decay mechanism is triggered. By smoothly reducing the entropy coefficient of the action distribution, the standard deviation of the executor network output is forcibly reduced, the exploration noise is removed, and the high-frequency mechanical jitter of the embodied agent in the physical space is eliminated. Step 33: Global Generalization and Deployment Convergence: The standard deviation of the action output is reduced to below the preset safe convergence threshold as the convergence mark of the current stage. Then, the physical interaction constraint boundary is gradually expanded outward, specifically by expanding the emission random space of the interaction target until global generalization is achieved. Finally, a physical interaction model with deterministic actions that can be directly deployed in the actual machine is output.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the highly dynamic target physical interaction control method for embodied intelligent agents as described in any one of claims 5 to 8.