Method for controlling action of embodied agent and electronic device
Patent Information
- Application Number
- CN202611257151.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-19
- Publication Date
- 2026-09-22
AI Technical Summary
然而,由于目标对象的形状、材质、空间位置以及周围环境状态可能存在差异,具身智能体在不同任务阶段下所面临的控制需求也可能不同
本发明实施例中,可以获取具身智能体执行目标任务时的当前执行阶段以及状态数据,并将这些数据输入至基于动态奖励函数训练得到的动作决策模型,生成与当前匹配的力度调整参数和位姿调整参数,以提高动作控制适配性;将力度调整参数和位姿调整参数转换为初始控制参数,并在控制过程中基于抓取力的波动程度以及碰撞风险预测结果对初始控制参数进行修正,可根据实际受力变化和运动安全风险调整控制参数,可提高任务执行过程的稳定性和安全性;最后基于修正后的目标控制参数执行当前执行阶段对应的动作,并在满足阶段完成条件时切换至下一执行阶段或结束目标任务,使得不同执行阶段之间能够连续衔接,进而有效提升具身智能体执行目标任务时的动作控制适配性、稳定性和安全性。
Smart Images

Figure CN122788005A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a method and electronic device for controlling the movement of an embodied intelligent agent. Background Technology
[0002] With the development of embodied intelligence technology, embodied intelligent agents are gradually being applied to task scenarios that require interaction with the real environment, such as object manipulation, assembly, transportation, and placement. In these task scenarios, embodied intelligent agents typically need to complete corresponding actions based on the state of the target object, the state of the actuator itself, and the state of the surrounding environment. The effectiveness of its action control directly affects the quality of the target task completion.
[0003] For tasks requiring high operational precision, embodied agents not only need to grasp, move, or place the target object, but also need to ensure the smoothness, accuracy, and safety of the action process. However, due to differences in the shape, material, spatial location, and surrounding environment of the target object, the control requirements faced by the embodied agent may vary at different stages of the task.
[0004] Existing embodied intelligent agents typically suffer from insufficient adaptability in motion control, poor stability during continuous operation, and untimely response to abnormal states when performing target tasks. Therefore, improving the adaptability, stability, and safety of motion control in embodied intelligent agents during target task execution has become a problem that needs to be solved. Summary of the Invention
[0005] To address the aforementioned problems, the present invention aims to provide a method and electronic device for controlling the action of an embodied intelligent agent, which can effectively improve the adaptability, stability, and security of action control of the embodied intelligent agent during the execution of a target task.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: On one hand, the present invention provides a method for controlling the action of an embodied intelligent agent, comprising: The current execution stage and state data of the embodied intelligent agent when performing the target task are obtained, and the state data includes multi-dimensional force data, visual data, pose data and obstacle data; The current execution stage and state data are input into the action decision model to obtain the force adjustment parameters and pose adjustment parameters. The action decision model is a model trained based on a dynamic reward function. The dynamic reward function is used to differentiate the evaluation of grasping force, pose accuracy, placement stability and collision risk under different execution stages. The force adjustment parameters and the pose adjustment parameters are converted into the initial control parameters of the embodied intelligent agent; During the process of controlling the embodied intelligent agent using the initial control parameters, the initial control parameters are corrected to target control parameters based on the fluctuation of the grasping force and the collision risk prediction results, and the grasping force is extracted from the multidimensional force data; Based on the target control parameters, the embodied intelligent agent is controlled to perform the action corresponding to the current execution stage, and when the current execution stage meets the stage completion conditions, it switches to the next execution stage or ends the target task.
[0007] On the other hand, the present invention also provides a motion control device for an embodied intelligent agent, comprising: The acquisition module is used to acquire the current execution stage and state data of the embodied intelligent agent when performing the target task. The state data includes multi-dimensional force data, visual data, pose data and obstacle data. The decision module is used to input the current execution stage and state data into the action decision model to obtain force adjustment parameters and pose adjustment parameters. The action decision model is a model trained based on a dynamic reward function. The dynamic reward function is used to differentiate the evaluation of grasping force, pose accuracy, placement stability and collision risk under different execution stages. The conversion module is used to convert the force adjustment parameters and the pose adjustment parameters into the initial control parameters of the embodied intelligent agent; The correction module is used to correct the initial control parameters to target control parameters based on the fluctuation of the grasping force and the collision risk prediction results during the process of controlling the embodied intelligent agent using the initial control parameters. The grasping force is extracted from the multidimensional force data. The control module is used to control the embodied intelligent agent to perform the action corresponding to the current execution stage based on the target control parameters, and to switch to the next execution stage or end the target task when the current execution stage meets the stage completion conditions.
[0008] On the other hand, the present invention also provides a server, including a processor and a memory, the memory storing a plurality of instructions; the processor loads instructions from the memory to execute steps in any of the embodied intelligent agent action control methods provided by the present invention.
[0009] On the other hand, the present invention also provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute steps in any of the embodied intelligent agent action control methods provided by the present invention.
[0010] On the other hand, the present invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps in any of the embodied intelligent agent action control methods provided by the present invention.
[0011] The beneficial effects of the technical solution provided by this invention include at least the following: In this embodiment of the invention, the current execution stage and state data of the embodied intelligent agent performing a target task can be obtained, and this data can be input into an action decision model trained based on a dynamic reward function to generate force adjustment parameters and pose adjustment parameters that match the current state, thereby improving the adaptability of action control. The force adjustment parameters and pose adjustment parameters are converted into initial control parameters, and the initial control parameters are corrected during the control process based on the fluctuation of the grasping force and the collision risk prediction results. The control parameters can be adjusted according to the actual force changes and motion safety risks, thereby improving the stability and safety of the task execution process. Finally, the action corresponding to the current execution stage is executed based on the corrected target control parameters, and the next execution stage is switched or the target task is ended when the stage completion condition is met, so that different execution stages can be continuously connected, thereby effectively improving the adaptability, stability and safety of action control when the embodied intelligent agent performs a target task. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram illustrating an application scenario of the action control method for an embodied intelligent agent provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating the action control method for an embodied intelligent agent provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the motion control device for an embodied intelligent agent provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] It is understood that in specific embodiments of the present invention, data involving user information and related data requires user permission or consent, and the collection, use and processing of such data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0016] See also Figure 1 This diagram illustrates an application scenario for a motion control method for embodied intelligent agents. This application scenario may include an embodied intelligent agent 101 and a server 102, which can exchange data via a network. The embodied intelligent agent 101 can be a single-arm collaborative robot, a dual-arm humanoid robot, a mobile manipulation composite robot, a high-precision specialized assembly robotic arm, etc.; the server 102 can be a single server or a server cluster consisting of multiple servers.
[0017] The embodied intelligent agent 101 can execute target tasks according to instructions. When the agent 101 executes the target task, the server 102 can obtain its current execution stage and state data. The state data may include multi-dimensional force data, visual data, pose data, and obstacle data. The current execution stage and state data are input into the action decision model to obtain force adjustment parameters and pose adjustment parameters. The action decision model is a model trained based on a dynamic reward function, which is used to differentiate the evaluation of grasping force, pose accuracy, placement stability, and collision risk at different execution stages. The force adjustment parameters and pose adjustment parameters are converted into the initial control parameters of the embodied intelligent agent. The server 102 can use the initial control parameters to control the embodied intelligent agent and, based on the fluctuation of the grasping force and the collision risk prediction results, correct the initial control parameters to target control parameters. Then, the server 102 uses the target control parameters to control the embodied intelligent agent 101 to execute the action corresponding to the current execution stage, and switches to the next execution stage or ends the target task when the current execution stage meets the stage completion conditions.
[0018] In this embodiment, a method for controlling the action of an embodied intelligent agent is provided, such as... Figure 2 As shown, the specific process of the action control method for this embodied intelligent agent can be as follows: S110. Obtain the current execution stage and state data of the embodied intelligent agent when it is performing the target task.
[0019] An embodied intelligent agent refers to a device with sensing, decision-making, and execution capabilities, such as an embodied robot, a robotic arm, a dexterous hand, or other intelligent devices capable of performing operations such as grasping, positioning, placing, and assembling. In this embodiment of the invention, an embodied robot is used as an example, wherein the specific robot has an end effector such as a dexterous hand or a gripper.
[0020] The target task refers to the task that the embodied intelligent agent needs to complete. This task can be a fine manipulation task targeting a target object. Fine manipulation tasks can include tasks such as grasping, placing, precision assembly, object handling, and handling of small materials. The target object is the object being grasped, placed, assembled, or moved. The target task can be manually set. For example, the target object may be a fragile piece of glass with a thickness of 0.5 mm, which needs to be grasped and moved to a precision slot 50 mm away to achieve stable placement.
[0021] The execution process of the target task can be divided into multiple stages according to actual needs. For example, the aforementioned task of moving the glass plate can be divided into three execution stages: grasping, localization, and placement. When the target intelligent agent performs the target task, it can obtain the execution stage it is in. State data refers to the multi-source data perceived by the embodied intelligent agent at the current moment, which may include multi-dimensional force data, visual data, pose data, and obstacle data.
[0022] Multidimensional force data can be obtained by force feedback sensors installed on the end effector of the embodied intelligent agent. The force feedback sensors can collect multidimensional force feedback signals in real time with high precision during the execution of the task, including key parameters such as the magnitude of the contact force, the distribution of the contact pressure, the coordinates of the contact position, the contact area, and the rate of change of the force, so as to comprehensively capture the contact state between the embodied intelligent agent and the target object.
[0023] Visual data is the data extracted from the target object and related environmental images collected by the visual sensors of the embodied intelligent agent. Specifically, it may include the target object image, target object position, outline, color, texture, size, orientation, target placement area, card slot position, obstacle image or three-dimensional spatial information, etc.
[0024] Pose data mainly refers to the pose data of the end effector of the embodied intelligent agent, which can be collected by the motion controller of the embodied intelligent agent. This pose data mainly includes position information and attitude information, where the position information is the three-dimensional coordinates of the end effector in the standard control, and the attitude information is the six joint angles and six joint angular velocities.
[0025] Obstacle data includes the obstacle's location coordinates, size, orientation, distance, bounding box information, minimum safe distance, and obstacle distribution information. This obstacle data can then be used for collision prediction, such as calculating the minimum distance between the predicted trajectory of the end effector and the obstacle to determine if a collision risk exists.
[0026] The current execution stage and state data of the embodied intelligent agent serve as the data foundation, upon which the action control of the embodied intelligent agent is achieved. The state data can be synchronously acquired by force feedback sensors, motion controllers, vision sensors, etc., with an acquisition frequency greater than or equal to 100Hz to ensure the real-time performance of the signals.
[0027] S120. Input the current execution stage and state data into the action decision model to obtain the force adjustment parameters and pose adjustment parameters.
[0028] The current execution stage and state data together constitute the model's input data, which is then fed into a pre-trained motion decision model. Through the analysis and processing of the motion decision model, the force adjustment parameters can be directly obtained. and pose adjustment parameters .
[0029] In some implementations, due to the heterogeneity and inconsistent dimensions of multi-source data such as force feedback data, visual data, pose data, and obstacle data, the multi-source states can be preprocessed and fused before being input into the action decision model. For example, when inputting the current execution stage and state data into the action decision model to obtain force adjustment parameters and pose adjustment parameters matching the current execution stage, the multi-dimensional force data, visual data, pose data, and obstacle data can be dimensionally aligned and fused to obtain a state vector; the state vector can be merged with the current execution stage to obtain input parameters; and the input parameters can be input into the action decision model to obtain force adjustment parameters and pose adjustment parameters matching the current execution stage.
[0030] Specifically, when performing dimensional alignment and fusion processing on multi-source data to obtain a state vector, principal component analysis can be performed on the visual data and the pose data to extract visual features and pose features; data related to grasping stability can be extracted from the multi-dimensional force data to form force features; the spatial coordinates, minimum safe distance, and size of obstacles can be extracted from the obstacle data to form obstacle features; and the visual features, pose features, force features, and obstacle features can be concatenated to obtain the state vector.
[0031] Visual data can be used to extract features such as target outline, color, and texture to form 256 original visual features. After dimensionality reduction by PCA, the first 24 dimensions are retained to obtain the visual features.
[0032] Attitude data includes position data and pose data. Here, only the pose data needs to be processed by PCA dimensionality reduction. The original attitude data includes 6 joint angles and 6 angular velocities, totaling 12 dimensions. After dimensionality reduction, 8 dimensions are retained to ensure the integrity of the attitude features. The position data is three-dimensional coordinates in a three-dimensional spatial coordinate system, totaling three dimensions, which do not require additional processing.
[0033] Multidimensional force data can be represented as It includes n-dimensional features such as contact force magnitude, pressure distribution, contact location, contact area, and force change rate; in order to eliminate the difference in dimensions and numerical scales, each feature is subjected to min-max normalization, mapping the value range to a unified dimensional scale of [0,1], thus obtaining multi-dimensional force features. .
[0034] Specifically, normalization can be performed using the following formula: ; For multidimensional force features, a two-step physics-first screening method can be used to reduce dimensionality, selecting 16 key features from n-dimensional multidimensional force features. First, interference signals unrelated to the contact state, such as ambient temperature drift and sensor baseline noise, are eliminated, retaining only features related to grasping stability, resulting in the first feature. For crane operators, features related to grasping stability can include contact force, pressure distribution, and contact position. Highly redundant and repetitive features, such as statistics that perfectly match the changing trend, are then removed from the first feature, resulting in the second feature. Finally, the second feature is sorted and screened according to its priority in influencing grasping stability, selecting 16 features strongly correlated with grasping stability, such as contact force components, pressure distribution peaks, contact position coordinates, and force change rate. This method eliminates invalid noise while ensuring a strong correlation between the features and the target task.
[0035] For obstacle data, the spatial coordinates (9-dimensional), minimum safe distance of the three obstacles closest to the embodied agent, and bounding box size information of the three obstacles (3-dimensional) can be obtained as a total of 13-dimensional obstacle features.
[0036] The state vector is obtained by concatenating visual features, pose features, force features, and obstacle features. One implementation method is to directly concatenate the individual features to obtain the state vector. Another implementation method, to ensure that force features dominate the state vector and to accommodate fine-grained operations, pre-assigned weights to each feature, and then concatenating the features in a specified order based on these weights, yields the state vector.
[0037] The weights corresponding to visual features, pose features, force features, and obstacle features are denoted as visual weight, pose weight, force weight, and obstacle weight, respectively. To ensure that the force feature dominates, the force weight is greater than any other weight. The specific weight values can be set according to actual needs, ensuring that the sum of the four weights is 1. For example, in this embodiment of the invention, the force weight is set to 0.4, the visual weight is set to 0.3, and the pose weight and obstacle weight are both set to 0.15.
[0038] Multiplying the weights by their corresponding features yields the weighted features. These features are then concatenated in a specified order to obtain the state vector. The specified order can be set according to actual needs. In this embodiment, the specified order can be force features, visual features, position features from pose features, posture features from pose features, and obstacle features. Concatenating according to this rule yields a 64-dimensional state vector, comprising 16 dimensions of force features, 24 dimensions of visual features, 3 dimensions of position features, 8 dimensions of posture features, and 13 dimensions of obstacle features.
[0039] By concatenating the state vector and the current execution stage, input parameters are obtained. These input parameters are then fed into the action decision model to obtain force adjustment parameters and pose adjustment parameters that match the current execution stage. The force adjustment parameters are primarily used to adjust the magnitude of the grasping force, while the pose adjustment parameters are primarily used to adjust the three-dimensional position of the end effector. Their coordinated output enables synchronized adjustment of force control and trajectory.
[0040] The action decision model can be pre-trained. During the training of the action decision model, a dynamic reward function can be used to differentiate the evaluation indicators related to grasping force, pose accuracy, placement stability, and collision risk in the action execution results at different execution stages.
[0041] In other words, the corresponding stage objectives differ at different execution stages of the target task. These different stage objectives mainly manifest in varying requirements for the grasping force, pose accuracy, placement stability, and collision safety of the embodied intelligent agent. For example, the grasping stage emphasizes flexible control of the grasping force, the positioning stage emphasizes precise adjustment of the end-effector pose, and the placement stage emphasizes stable placement of the target object. Simultaneously, collision risk must be considered at each execution stage. Therefore, when training the action decision model, by setting phased dynamic reward functions based on the stage objectives of different execution stages, the trained action decision model can learn the control emphasis at different execution stages. This allows it to output force and pose adjustment parameters that match the current execution stage when subsequently executing the target task, thereby improving the control accuracy, operational stability, and safety of the embodied intelligent agent when performing fine-grained tasks.
[0042] In some implementations, the motion decision model can be trained as follows: The sample execution phase and sample state data of the embodied agent performing a sample task are input into a basic decision model to obtain prediction force adjustment parameters and prediction pose adjustment parameters. The sample state data includes multi-dimensional force samples, visual samples, pose samples, and obstacle samples. Using the prediction force adjustment parameters and prediction pose adjustment parameters, the embodied agent is controlled to perform actions, and execution feedback data is obtained. The target reward value corresponding to the dynamic reward function is calculated based on the execution feedback data and the sample execution phase. The model parameters of the basic decision model are updated based on the target reward value until the training completion condition is met, thus obtaining the motion decision model.
[0043] Before training the model, training samples can be obtained first. These training samples can be the execution phase and state data of the embodied agent performing sample tasks. The state data can include multi-dimensional force samples, visual samples, pose samples, and obstacle samples. The execution phase and the current execution phase mentioned above have the same meaning, as do the state data mentioned above. For details, please refer to the corresponding explanations in the preceding embodiments. The difference is that the execution phase and state data are data used during the model training phase.
[0044] Before being input into the basic decision model, the sample state data still needs to be converted into a sample state vector. The specific processing method can be found in the corresponding content of the aforementioned embodiments, and will not be repeated here. After concatenating the sample state vector with the sample execution stage data, the sample input parameters can be obtained. Inputting these parameters into the basic decision model yields the corresponding prediction strength adjustment parameters and prediction pose adjustment parameters.
[0045] The embodied agent is controlled to perform actions using predicted force adjustment parameters and predicted pose adjustment parameters, and execution feedback data is acquired. Execution feedback data refers to the data actually collected under the control of the predicted force and predicted pose adjustment parameters.
[0046] When calculating the target reward value corresponding to the dynamic reward function using execution feedback data and the sample execution stage, the basic reward weight distribution corresponding to the sample execution stage can be obtained. The basic reward weight distribution includes weights corresponding to multiple dimensions, including grasping strength, pose accuracy, placement stability, and collision risk. Using the execution feedback data and the task objective of the sample task, a weight adjustment factor corresponding to each dimension is calculated. Based on the weight adjustment factor and the execution progress of the sample task, the basic reward weight distribution is adaptively adjusted to obtain the target reward weight distribution. The target reward value is calculated using the target reward weight distribution and the sub-reward function corresponding to each dimension.
[0047] In this embodiment of the invention, the dynamic reward function may include four sub-rewards. By adjusting the reward weights of these four sub-rewards, the reward function can be tailored to different stages. When calculating the target reward value, the basic reward weight distribution corresponding to the sample execution stage can be obtained. This basic reward weight distribution may include weights across multiple dimensions, such as grasping strength, pose accuracy, placement stability, and collision risk. The weight distribution corresponding to different execution stages can be pre-set, and the weight distribution corresponding to the sample execution stage can be directly used as the basic reward weight distribution.
[0048] For example, the weights corresponding to grasping force, pose accuracy, placement stability, and collision risk are denoted as follows: During the grasping phase, the objective is flexible force control, and its weight distribution can be set as follows: During the positioning phase, the goal is precise positioning, and the weight distribution can be set as follows: During the placement phase, the objective is stable placement, and the weight distribution can be set as follows: .
[0049] To adapt to actual execution conditions, the execution feedback data and the task objectives of the sample tasks can be used to calculate the weight adjustment factor corresponding to each dimension; based on the weight adjustment factor and the execution progress of the sample tasks, the basic reward weight distribution can be adaptively adjusted to obtain a target reward weight distribution that conforms to the actual situation.
[0050] Execution feedback data may include the collected real-time grasping force. Real-time spatial position of the end effector The variance of the real-time grasping force signal during the placement phase Real-time contact area .
[0051] Regarding the gripping force, the real-time gripping force can be calculated. With respect to the target's safe gripping force threshold The difference between them is used to obtain the force deviation, which is then used as a weighting adjustment factor; for pose accuracy, the real-time spatial position of the end effector can be calculated. Positioning coordinates of the target operation The difference between the values is used to obtain the positioning deviation, which is then used as a weighting adjustment factor. For placement stability, the real-time contact area can be calculated. Maximum contact area with the bottom of the target object The ratio between the two values is used to obtain the placement fit α. In addition, the degree of fluctuation in the gripping force can also be calculated, and the placement fit and force fluctuation are used together as weight adjustment factors.
[0052] In some implementations, the weight adjustment factor may include force deviation, positioning deviation, placement fit, and force fluctuation. Based on the weight adjustment factor and the execution progress of the sample task, the basic reward weight distribution is adaptively adjusted to obtain the target reward weight distribution. Specifically, if the force deviation is greater than a force threshold, the weight corresponding to the grasping force is increased, and the weight corresponding to the pose accuracy is decreased; if the positioning deviation is greater than a positioning threshold, the weight corresponding to the pose accuracy is increased, and the weight corresponding to the collision risk is decreased; if the placement fit or the force fluctuation meets the placement conditions, the weight corresponding to the placement stability is increased, and the weight corresponding to the grasping force is decreased; if the execution progress of the sample task reaches a preset progress, the weight corresponding to the collision risk is gradually increased to obtain the target reward weight distribution.
[0053] During adaptive adjustment, the sum of all weights remains 1, and the value of each individual weight ranges from [0,1]. During adaptive adjustment, if a force deviation greater than a force threshold is detected, the weight corresponding to the gripping force can be automatically increased. And automatically reduce the weight corresponding to the pose accuracy. The force threshold and specific adjustment data can be set according to actual needs. In this embodiment of the invention, the force threshold can be set to 0.05N. To increase the weight by 0.1 and ensure the sum of the weights is 1, we can... Decrease by 0.1.
[0054] If the positioning deviation is greater than the positioning threshold, the weight corresponding to the pose accuracy can be increased. And reduce the weight corresponding to the collision risk. The positioning threshold and specific adjustment data can be set according to actual needs. In this embodiment of the invention, the positioning threshold can be set to 0.5mm. To increase the weight by 0.1 and ensure the sum of the weights is 1, we can... Decrease by 0.1.
[0055] During the placement phase, if the placement fit α or force fluctuation meets the placement conditions, the weight corresponding to placement stability can be increased. And reduce the weight corresponding to collision risk. The placement conditions are that the fit degree α is less than the fit degree threshold, or the force fluctuation is greater than the force fluctuation threshold. The fit degree threshold, force fluctuation threshold, and specific adjustment values can be set according to actual needs. In this embodiment of the invention, the fit degree threshold can be 95%, and the force fluctuation threshold can be 0.03N. To increase the weight by 0.1 and ensure the sum of the weights is 1, we can... Decrease by 0.1.
[0056] If the execution progress of the sample task reaches the preset progress, the weight corresponding to the collision risk can be gradually increased. This ensures safe completion of the task. The preset progress can be set to 80%. Therefore, based on the actual execution situation, the target reward weight distribution can be adaptively adjusted from the basic reward weight distribution.
[0057] Each dimension has a corresponding sub-reward function. By fusing the target reward distribution and the sub-reward functions of each dimension, the target reward value can be calculated. As one implementation method, the target reward value can be calculated using the execution feedback data and the task objective of the sample task to calculate the sub-reward value corresponding to the sub-reward function of each dimension; the sub-reward values of all dimensions are then weighted based on the target reward weight distribution to obtain the target reward value.
[0058] The task objectives of the sample task may include the result to be achieved by the task, as well as some set values, such as the target safe grasping force threshold, maximum force deviation, maximum positioning deviation, maximum force fluctuation, discount factor and other parameters, which are used to calculate the target reward value in subsequent processes.
[0059] Regarding the grasping force dimension, the real-time grasping force can be acquired, and the difference between the real-time grasping force and the target safe grasping force threshold can be calculated to obtain the force deviation. The ratio between the force deviation and the maximum force deviation threshold can be calculated to obtain the force deviation ratio. Based on a first preset value and the force deviation ratio, the grasping force reward, i.e., the sub-reward value corresponding to the grasping force dimension, can be calculated. Specifically, the grasping force reward can be calculated according to the following formula: ; in, Sub-reward values corresponding to the dimension of grasping intensity; Characterizes the real-time grasping force collected; Characterizes the target's safe gripping force threshold; The maximum force deviation threshold is used to characterize the force. The smaller the force deviation, the closer the current grasping force is to the target safe grasping force, and the higher the grasping force reward. The larger the force deviation, the lower the grasping force reward. If the force deviation is greater than or equal to the maximum force deviation threshold, the grasping force reward is zero. This guides the basic decision model to output force adjustment parameters that enable the grasping force to approach the target safe grasping force.
[0060] Regarding the pose accuracy dimension, the real-time spatial position of the end effector can be obtained, and the difference between the real-time spatial position of the end effector and the target operation positioning coordinates can be calculated to obtain the positioning deviation. For positioning accuracy rewards, the real-time spatial position of the end effector can be obtained, and the spatial distance between the real-time spatial position and the target operation positioning coordinates can be calculated as the positioning deviation. The ratio between the positioning deviation and the maximum positioning deviation threshold can be calculated; the larger value between 0 and the ratio can be taken as the sub-reward value for this dimension. Specifically, the sub-reward value corresponding to the positioning accuracy dimension can be calculated according to the following formula: ; in, Sub-reward values corresponding to the dimension of positioning accuracy; Characterizes the real-time spatial position of the end effector; Characterizes the target's positioning coordinates; The maximum positioning deviation threshold is defined as follows: the smaller the positioning deviation, the closer the end effector is to the target operational positioning coordinates, and the higher the positioning accuracy reward; the larger the positioning deviation, the lower the positioning accuracy reward; if the positioning deviation is greater than or equal to the maximum positioning deviation threshold, the positioning accuracy reward is zero. This guides the basic decision model to output pose adjustment parameters that can reduce the distance between the end effector and the target operational positioning coordinates.
[0061] Regarding the placement stability dimension, the real-time grasping force acquired during the placement phase can be obtained, and force fluctuations can be determined based on multiple real-time grasping forces within a preset time window. The ratio between the force fluctuation and the maximum force fluctuation is calculated to obtain the fluctuation ratio. Based on a third preset value and the fluctuation ratio, the stable placement reward, i.e., the sub-reward value corresponding to the placement stability dimension, is calculated. Specifically, the stable placement reward can be calculated according to the following formula: ; in, Sub-reward values corresponding to the placement stability dimension; Fluctuations in characterization power; This represents the maximum force fluctuation. The smaller the fluctuation in gripping force, the more stable the target object is during placement, and the higher the stable placement reward; the larger the fluctuation in gripping force, the lower the stable placement reward; if the fluctuation in gripping force is greater than or equal to the maximum force fluctuation, the stable placement reward can be set to zero.
[0062] As one implementation method, to further improve the evaluation of placement stability, placement fit can be introduced; the stable placement bonus is calculated based on force fluctuation and placement fit together. Specifically, it can be calculated according to the following formula: ; Where α represents the placement fit, and is the real-time contact area. Maximum contact area with the bottom of the target object The ratio between them; The threshold representing the fit can be set according to actual needs. This reward can guide the basic decision-making model to output action adjustment parameters that reduce placement impact, decrease placement sway, and improve placement fit.
[0063] For the collision risk dimension, collision risk prediction can be performed, and the prediction result determines the sub-reward for this dimension. Details regarding collision risk prediction can be found in the preceding description and will not be repeated here. Specifically, the sub-reward for this dimension can be: ; In other words, if a collision risk is not predicted, a positive reward is given; if a collision risk is predicted, a negative incentive is given, guiding the basic decision-making model to take motion safety into account when outputting force adjustment parameters and pose adjustment parameters, thereby reducing the probability of collisions between the end effector, the target object, or the obstacle.
[0064] Finally, based on the target reward weight distribution and the sub-reward values of each dimension, the target reward value can be calculated. The specific calculation formula is as follows: ; In actual training, the training objective is to maximize the cumulative discount reward and optimize the strategy. The specific formula is as follows: ; Where π represents the reinforcement learning policy, τ represents the interaction trajectory between the robot and the environment, γ∈(0,1] is the discount factor (used to balance current rewards and future rewards), and T is the total step size in a single round of operation. Let be the standard state vector at time t. Let t be the control action at time t.
[0065] Using the total reward function as the evaluation criterion, the model parameters of the basic decision-making model are continuously iterated through autonomous trial and error interaction with the environment, so as to gradually converge to the optimal control strategy that adapts to fine operation and realize the autonomous optimization of the entire operation process.
[0066] It should be noted that when training the action decision model through reinforcement learning, a detailed scene covering targets of various shapes and materials, as well as complex working conditions, can be built in a virtual environment. The virtual environment is constructed using the Unity 3D engine, reproducing key parameters such as force feedback characteristics, friction coefficient, and gravity environment in the real scene; the training data used can also be collected in the virtual environment. In the virtual environment, training can be terminated when certain conditions are met, and the training can be transferred to the real physical environment for fine-tuning. The training termination conditions can be set according to actual needs. In this embodiment of the invention, the cumulative discount reward for 50 consecutive rounds of operation is ≥0.85, the operation error rate (object damage, excessive positioning deviation, collision) is ≤1%, and the force control error is ≤0.02N.
[0067] The model trained in the virtual environment is then transferred to a real physical environment for fine-tuning using a small amount of actual interaction data from the real environment. The training termination criteria in the real-world scenario are: cumulative discounted reward ≥ 0.9 over 30 consecutive rounds of operations, operation error rate ≤ 0.5%, force control error ≤ 0.01N, and trajectory deviation ≤ 0.05mm. Training can be terminated once these conditions are met. Of course, the training termination criteria in the real-world environment can also be set according to actual needs; this is just an example.
[0068] For example, the target object is a fragile glass shard with a thickness of 0.5mm (easily broken, requiring extremely high gripping force); the task is to grip the glass shard, move it into a precision slot 50mm away, and place it stably; the core parameter is the target's safe gripping force. Maximum allowable force deviation Maximum allowable positioning deviation Allowing maximum force fluctuation Discount factor For a scenario with a total step length of 70 steps in a single round of operations, the entire process of reinforcement learning optimization is as follows: Phase 1: Flexible grasping (t=0~20 steps, core objective: flexible grasping to avoid damage); State input: At t=0, the vision sensor detects the spatial position of the glass plate, and the force feedback sensor collects the initial contact force. (Far below the target safety force), no collision risk, standard state vector Includes all of the above information; Output: Based on the initial state input, the action decision model outputs the action. (Increase the gripping strength) ; Reward Feedback: Calculate rewards, =0 (significant deviation in intensity, moderate reward). =0.3 (Large positioning deviation, low reward). =0 (not yet in the placement phase), =+1 (No collision risk), Total Reward =0.36 (at this point, the weight) =0.5, =0.2, =0, =0.3); Trajectory optimization: During the t=0~20 steps, continuously receive force feedback signals and gradually adjust. This allows the gripping force to gradually approach 0.2N from 0.05N, while simultaneously based on real-time coordinate vectors and... Vector superposition optimizes the grasping trajectory, avoiding uneven force on the glass plate; at step t=20, , =0.8 The grasping trajectory fits the edge of the glass sheet without squeezing or breaking, completing the flexible grasping.
[0069] Phase 2: Precise Positioning (t=21~50 steps, core objective: precise movement, reducing positioning deviation); Status input: at t=21, the deviation between the real-time position of the end effector and the coordinates of the target slot. Grasping power (Stable), unobstructed, and no risk of collision; Action: Based on state input and reward feedback, the action decision model outputs... (Maintain the current gripping force to avoid changes in force that could cause the glass to slip.) ; Reward Feedback: =0.8 (stable force) =0.85 (deviation reduced, reward increased), =0, =+1, Total Reward =0.885 (at this point, the weight is adjusted to) =0.2, =0.5, =0, =0.3); Trajectory optimization: During the t=21~50 step period, the action decision model is continuously revised. The positioning deviation was gradually reduced, and the motion trajectory was smooth and jitter-free, avoiding glass plate wobbling; by t=50 steps, the positioning deviation was reduced to 0.2mm, meeting the requirements. Requirements =0.6.
[0070] Phase 3: Stable Placement (t=51~70 steps, core objective: stable placement, avoiding bumps and knocks); Status input: At t=51, the end effector is aligned with the slot, and the force feedback signal fluctuates. (Within permissible limits), no risk of collision; Action: Output (Release the force gradually to avoid the glass plate falling off due to sudden force release.) ); Reward Feedback: =0.8 (smooth unloading process). =0.6, =0.4, =+1, Total Reward =0.64 (at this point, the weight is adjusted to) =0.2, =0.2, =0.4, =0.2); Trajectory optimization: During the t=51~70 step period, the action decision model is adjusted slowly. and To ensure the glass slide enters the groove smoothly, the force signal fluctuation is always controlled within a certain range. Within t=70 steps, the glass plate is fully inserted into the groove, the gripping force drops to 0N, it is placed stably without bumps or damage, and the entire process is completed.
[0071] S130. The force adjustment parameters and the pose adjustment parameters are converted into the initial control parameters of the embodied intelligent agent.
[0072] The force adjustment parameters and pose adjustment parameters output by the action decision model are upper-level decision results, used to characterize how the grasping force, position and / or attitude of the end effector should be adjusted within the current control cycle; however, the actuator of the embodied agent usually cannot directly execute such abstract adjustment parameters, so the control center needs to convert them into control parameters that the actuator can recognize and execute.
[0073] Specifically, the force adjustment parameter can be used to characterize the direction and magnitude of the gripping force adjustment of the end effector. For example, when the current gripping force is less than the target safe gripping force, the force adjustment parameter can be a positive value to indicate an appropriate increase in gripping force; when the current gripping force is too large or the force fluctuation is large, the force adjustment parameter can be a negative value to indicate an appropriate decrease in gripping force. The control center can generate corresponding initial force control parameters based on this force adjustment parameter, such as torque control parameters, current control parameters, gripping force control parameters, or servo control parameters for the gripper motor.
[0074] Similarly, pose adjustment parameters can be used to characterize the direction and magnitude of position and / or attitude adjustment of the end effector. For example, when there is a deviation between the current position of the end effector and the target operating position, the pose adjustment parameters can include the position adjustment amount of the end effector in three-dimensional space, such as the displacement correction amount of the end effector in the x-axis, y-axis, and z-axis directions. The control center can convert the pose adjustment parameters into corresponding initial pose control parameters based on the robot's kinematic model or inverse kinematics calculation, such as the angle control parameters, angular velocity control parameters, displacement control parameters, or trajectory control parameters of each joint of the embodied agent's robotic arm. Thus, the initial control parameters can include initial force control parameters and initial pose control parameters, used to control the embodied agent to perform actions within the current control cycle.
[0075] S140. During the process of controlling the embodied intelligent agent using the initial control parameters, the initial control parameters are corrected to target control parameters based on the fluctuation of the grasping force and the collision risk prediction results.
[0076] Initial control parameters are preliminary control commands derived from the force adjustment parameters and pose adjustment parameters output by the action decision model. These parameters are used to control the embodied agent to begin executing the corresponding action. However, during the actual execution of the action, the contact state, force state of the target object, and the surrounding environment may change. Relying solely on the initial control parameters may lead to problems such as fluctuations in grasping force, object slippage, excessive local force, or collisions between the end effector and obstacles. Therefore, it is necessary to adjust the initial control parameters based on the actual situation during execution.
[0077] The target control parameters may include at least one of target force control parameters and target pose control parameters. The target force control parameters can be used to control the end effector to grasp the target object in a smoother manner, and the target pose control parameters can be used to control the end effector to complete position and / or attitude adjustment in a safer manner, thereby improving the flexibility, stability and safety of the embodied intelligent agent when performing the target task.
[0078] As one implementation method, when correcting the initial control parameters to target control parameters based on the fluctuation level of the grasping force and the collision risk prediction results, the fluctuation level of the grasping force can be determined according to multiple grasping forces within a preset time window during the process of controlling the embodied intelligent agent using the initial control parameters; collision risk prediction processing is performed based on the pose data of the target object, obstacles, and the embodied intelligent agent, where the target object is the processing object in the target task; if the fluctuation level of the grasping force is greater than the fluctuation threshold, the force adjustment parameter is corrected to obtain the target control parameters; if a collision risk is predicted, the force adjustment parameter and the pose adjustment parameter are corrected to obtain the target control parameters.
[0079] Among them, the grasping force is the contact force in the multidimensional force data. The fluctuation of the grasping force can be determined based on multiple real-time grasping force signals collected within a preset time window. For example, it can be determined based on the grasping force difference, grasping force range, grasping force variance, or grasping force standard deviation at adjacent sampling times.
[0080] If the fluctuation of the gripping force exceeds the preset force fluctuation threshold, it indicates that the contact state between the end effector and the target object is unstable, and there may be problems such as excessively tight gripping, excessively loose gripping, or uneven local force. In this case, the initial force control parameters can be corrected by feedback, such as reducing the increase of the gripping force, reducing the rate of change of the gripper motor torque, or performing reverse compensation on the gripping force to obtain the target control parameters, and making the actual gripping force tend to be stable based on the target control parameters.
[0081] When an embodied intelligent agent grasps a moving target object, the target object may collide with obstacles in the environment. To avoid collisions, the collision risk can be predicted, and collision avoidance can be actively achieved to improve operational safety.
[0082] As one implementation method, when performing collision risk prediction processing based on the pose data of the target object, obstacles, and embodied intelligent agent, the following steps can be taken: First, construct a spatial bounding volume for the target object and obstacles based on their pose data. Second, predict multiple predicted trajectory points of the embodied intelligent agent within a preset future time period based on the pose data of the embodied intelligent agent and the initial control parameters. Third, determine a target distance based on the distance between each predicted trajectory point and the spatial bounding volume. Fourth, if the target distance is less than a preset safety distance, a collision risk is predicted.
[0083] Based on the aforementioned collected data, three-dimensional pose data of the target object, obstacles, and the end effector of the embodied intelligent agent can be obtained, including center coordinates, dimensions, and orientation angle. For example, if the target object is a glass sheet on a table, the aforementioned visual sensor can identify that the center of the glass sheet is at (100mm, 50mm, 20mm), its dimensions are 40mm × 20mm × 0.5mm, and its orientation angle is 30 degrees. Based on the center coordinates, dimensions, and orientation angle, a spatial bounding volume can be directly generated for the target object and each obstacle using an online geometric fitting algorithm.
[0084] Using the pose data and initial control parameters of the embodied agent, multiple predicted trajectory points of the embodied agent within a preset future time period are predicted. In other words, the motion trajectory of the embodied agent in the future time period is predicted based on its current pose data and initial control parameters, and this motion trajectory may include multiple predicted trajectory points. Specifically, the joint motion of the embodied agent can be converted into the spatial motion state of the end effector based on the robot's kinematic model and Jacobian matrix. Then, using historical pose data from the most recent number of sampling periods, the current position, velocity, and acceleration of the end effector are estimated through Kalman filtering. Based on the current position, velocity, and acceleration, multiple spatial position points that the end effector may pass through within the next 5–10 ms are predicted using a motion extrapolation formula, forming multiple predicted trajectory points.
[0085] For each predicted trajectory point, the distance between that point and the enclosing space is calculated to obtain the target distance. If the target distance is less than the preset safety distance, a collision risk is predicted. This means that continuing to move according to the initial control parameters carries a high risk of subsequent collisions. The initial control parameters can be adjusted to achieve obstacle avoidance correction. For example, the direction of the end effector can be adjusted, the displacement adjustment range reduced, or a local obstacle avoidance path redefined. The force control parameters can be adjusted in conjunction with the obstacle avoidance correction results to obtain the target control parameters, thus preventing the target object from swaying, slipping, or experiencing abnormal forces due to changes in the end effector's direction or velocity.
[0086] If all target distances are greater than the preset safety distance, it can be determined that no collision risk has been predicted. That is, if the movement continues according to the initial control parameters, no collision will occur. No adjustments need to be made, and the initial control parameters can be directly determined as the target control parameters.
[0087] S150. Based on the target control parameters, control the embodied intelligent agent to perform the action corresponding to the current execution stage, and if the current execution stage meets the stage completion condition, switch to the next execution stage or end the target task.
[0088] The target control parameters are used to control the embodied agent to perform the actions corresponding to the current execution stage. For example, when the current execution stage is the grasping stage, the end effector can be controlled to adjust the grasping force based on the target force control parameters, and the contact position or contact posture between the end effector and the target object can be adjusted based on the target pose control parameters. When the current execution stage is the positioning stage, the end effector can be controlled to move towards the target operation positioning coordinates based on the target pose control parameters, and the grasping stability of the target object can be maintained based on the target force control parameters. When the current execution stage is the placement stage, the grasping force can be gradually reduced based on the target force control parameters, and the target object can be controlled to smoothly approach the placement area based on the target pose control parameters.
[0089] During the execution of the action corresponding to the current execution stage, it can be determined whether the current execution stage has been completed based on the corresponding stage completion conditions. Different execution stages can correspond to different stage completion conditions. For example, the stage completion conditions for the grasping stage may include at least one of the following: the real-time grasping force reaches the target safe grasping force range, the grasping force fluctuation is less than a preset force fluctuation threshold, and the target object has not slipped; the stage completion conditions for the positioning stage may include at least one of the following: the positioning deviation between the end effector and the target operation positioning coordinates is less than a preset positioning deviation threshold, and / or the collision risk prediction result indicates that there is no collision risk; the stage completion conditions for the placement stage may include at least one of the following: the fit between the target object and the target placement area reaches a preset fit threshold, the grasping force fluctuation is less than a preset force fluctuation threshold, and this remains true for a preset duration.
[0090] If the target task is not yet completed when the corresponding completion conditions are met in the current execution phase, the current execution phase will be switched to the next execution phase to continue the subsequent actions; if the current execution phase is the last execution phase of the target task, the target task will be terminated. Thus, the embodied intelligent agent can complete the target task in the order of grasping, locating, and placing, and automatically switch phases after each phase is completed, thereby improving the continuity and automation of the target task execution process.
[0091] Specifically, in this embodiment of the invention, if during the grasping phase, the grasping force Stable at Within the range, and for a duration It will automatically switch to the positioning stage (if positioning is not required, it will directly enter the placement stage).
[0092] If, during the positioning phase, the positioning deviation of the end effector meets the following requirements... If collision prediction shows no risk, the system automatically switches to the placement stage (if placement is not required, the weight is reset to zero). During the placement stage, the material fit... Force fluctuation And the duration Once the task is deemed complete, the weight is reset to zero.
[0093] The motion control scheme for embodied intelligent agents provided in this invention can be applied to various scenarios where embodied intelligent agents perform tasks. For example, taking the transfer of fragile glass shards by a robot as an example, the scheme provided in this invention can ensure that the robot grasps the glass shards with appropriate force, continuously adjusts the force and posture to avoid collisions, and ultimately transfers the glass shards safely.
[0094] The present invention can acquire the current execution stage and state data of an embodied intelligent agent when performing a target task. These two data points are input into a dynamic reward function to generate force adjustment parameters and pose adjustment parameters that match the current actual situation, thereby improving the adaptability of the embodied intelligent agent's motion control under different task stages and state conditions. The force adjustment parameters and pose adjustment parameters are converted into initial control parameters executable by the embodied intelligent agent. During control using these initial control parameters, the initial control parameters are corrected based on the fluctuation of the grasping force and the collision risk prediction results, which can reduce instability or safety issues caused by grasping force fluctuations, pose deviations, or collision risks. Based on the corrected target control parameters, the embodied intelligent agent is controlled to execute the action corresponding to the current execution stage. When the current execution stage meets the stage completion conditions, the system switches to the next execution stage or ends the target task. This achieves continuous connection between different execution stages in the target task, reduces control mismatch during stage switching, and effectively improves the adaptability, execution stability, and operational safety of the embodied intelligent agent when performing the target task.
[0095] To better implement the above methods, embodiments of the present invention also provide a motion control device for an embodied intelligent agent. This motion control device can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer; the server can be a single server or a server cluster composed of multiple servers.
[0096] For example, in this embodiment, the method of the present invention will be described in detail by taking the action control device of the embodied intelligent agent specifically integrated into the server as an example.
[0097] For example, such as Figure 3 As shown, the action control device 200 of the embodied intelligent agent may include an acquisition module 210, a decision module 220, a conversion module 230, a correction module 240, and a control module 250.
[0098] The acquisition module 210 is used to acquire the current execution stage and state data of the embodied intelligent agent when performing the target task. The state data includes multi-dimensional force data, visual data, pose data and obstacle data. The decision module 220 is used to input the current execution stage and state data into the action decision model to obtain the force adjustment parameters and the pose adjustment parameters. The action decision model is a model trained based on a dynamic reward function. The dynamic reward function is used to differentiate the evaluation of grasping force, pose accuracy, placement stability and collision risk under different execution stages. The conversion module 230 is used to convert the force adjustment parameters and the posture adjustment parameters into the initial control parameters of the embodied intelligent agent; The correction module 240 is used to correct the initial control parameters to target control parameters based on the fluctuation of the grasping force and the collision risk prediction results during the process of controlling the embodied intelligent agent using the initial control parameters. The grasping force is extracted from the multidimensional force data. The control module 250 is used to control the embodied intelligent agent to perform the action corresponding to the current execution stage based on the target control parameters, and to switch to the next execution stage or end the target task when the current execution stage meets the stage completion conditions.
[0099] In some embodiments, the decision module 220 is specifically used for: The multidimensional force data, visual data, pose data, and obstacle data are dimensionally aligned and fused to obtain a state vector. By merging the state vector with the current execution stage, the input parameters are obtained; The input parameters are input into the action decision model to obtain force adjustment parameters and pose adjustment parameters that match the current execution stage.
[0100] In some embodiments, the decision module 220 is specifically used for: Principal component analysis is performed on the visual data and the pose data to extract visual features and pose features. Data related to grasping stability is extracted from the multidimensional force data to form force features; The spatial coordinates, minimum safe distance, and size of the obstacles are extracted from the obstacle data to form obstacle features; The visual features, pose features, force features, and obstacle features are concatenated to obtain a state vector.
[0101] In some embodiments, the correction module 240 is specifically used for: During the process of controlling the embodied intelligent agent using the initial control parameters, the fluctuation degree of the grasping force is determined based on multiple grasping forces within a preset time window; Based on the pose data of the target object, obstacles, and the embodied intelligent agent, collision risk prediction is performed, where the target object is the processing object in the target task. If the fluctuation of the grasping force is greater than the fluctuation threshold, the force adjustment parameter is corrected to obtain the target control parameter; If a collision risk is predicted, the force adjustment parameters and pose adjustment parameters are corrected to obtain the target control parameters.
[0102] In some embodiments, the correction module 240 is specifically used for: Based on the pose data of the target object and obstacles, construct the spatial bounding volume of the target object and obstacles; Based on the pose data of the embodied intelligent agent and the initial control parameters, predict multiple predicted trajectory points of the embodied intelligent agent within a preset future time period; The target distance is determined based on the distance between each predicted trajectory point and the spatial bounding volume; If the target distance is less than the preset safety distance, a collision risk is predicted.
[0103] In some embodiments, the motion control device 200 of the embodied intelligent agent may further include a training module, which is specifically used for: The sample execution phase and sample state data when the embodied intelligent agent performs the sample task are input into the basic decision model to obtain the prediction force adjustment parameters and the prediction pose adjustment parameters. The sample state data includes multi-dimensional force samples, visual samples, pose samples and obstacle samples. Using the predicted force adjustment parameters and the predicted pose adjustment parameters, the embodied intelligent agent is controlled to perform actions, and execution feedback data is obtained; Calculate the target reward value corresponding to the dynamic reward function based on the execution feedback data and the sample execution stage; The model parameters of the basic decision model are updated based on the target reward value until the training completion condition is met, thus obtaining the action decision model.
[0104] In some embodiments, the training module is specifically used for: Obtain the basic reward weight distribution corresponding to the sample execution stage. The basic reward weight distribution includes weights corresponding to multiple dimensions, including grasping strength, pose accuracy, placement stability, and collision risk. Using the execution feedback data and the task objective of the sample task, calculate the weight adjustment factor corresponding to each dimension; Based on the weight adjustment factor and the execution progress of the sample task, the basic reward weight distribution is adaptively adjusted to obtain the target reward weight distribution; The target reward value is calculated using the target reward weight distribution and the sub-reward function corresponding to each dimension.
[0105] In some embodiments, the weight adjustment factor includes force deviation, positioning deviation, placement fit, and force fluctuation, and the training module is specifically used for: If the force deviation is greater than the force threshold, the weight corresponding to the grasping force is increased and the weight corresponding to the pose accuracy is decreased to obtain the target reward weight distribution. If the positioning deviation is greater than the positioning threshold, the weight corresponding to the pose accuracy is increased and the weight corresponding to the collision risk is decreased to obtain the target reward weight distribution. If the placement fit or the force fluctuation meets the placement conditions, increase the weight corresponding to the placement stability and decrease the weight corresponding to the gripping force to obtain the target reward weight distribution; If the execution progress of the sample task reaches the preset progress, the weight corresponding to the collision risk is gradually increased to obtain the target reward weight distribution.
[0106] In some embodiments, the training module is specifically used for: Using the execution feedback data and the task objective of the sample task, calculate the sub-reward value corresponding to the sub-reward function for each dimension; The target reward value is obtained by weighting the sub-reward values of all dimensions based on the target reward weight distribution.
[0107] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.
[0108] As can be seen from the above, the motion control device of the embodied intelligent agent in this embodiment can acquire the current execution stage and state data of the embodied intelligent agent when performing the target task. These two data points are input into force adjustment parameters and pose adjustment parameters generated based on a dynamic reward function to match the current actual situation, thereby improving the motion control adaptability of the embodied intelligent agent under different task stages and state conditions. The force adjustment parameters and pose adjustment parameters are converted into initial control parameters executable by the embodied intelligent agent. During control using these initial control parameters, the initial control parameters are corrected based on the fluctuation of the grasping force and the collision risk prediction results, which can reduce the instability or safety issues caused by grasping force fluctuations, pose deviations, or collision risks. Based on the corrected target control parameters, the embodied intelligent agent is controlled to perform the action corresponding to the current execution stage. When the current execution stage meets the stage completion conditions, the device switches to the next execution stage or ends the target task, achieving continuous connection between different execution stages in the target task, reducing control mismatch during stage switching, and thus effectively improving the motion control adaptability, execution stability, and operational safety of the embodied intelligent agent when performing the target task.
[0109] This invention also provides an electronic device, which can be a terminal, a server, or other similar device. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, personal computer, robot, embodied intelligent agent, etc.; the server can be a single server or a server cluster composed of multiple servers, etc.
[0110] In some embodiments, the device can also be integrated into multiple electronic devices. For example, the motion control device of the embodied intelligent agent can be integrated into multiple servers, and the motion control method of the embodied intelligent agent of the present invention can be implemented by multiple servers.
[0111] In this embodiment, a server will be used as an example for detailed description. For example, ... Figure 4 As shown, it illustrates a structural schematic diagram of the electronic device involved in an embodiment of the present invention, specifically: The server may include components such as a processor 310 with one or more processing cores, a memory 320 with one or more computer-readable storage media, a power supply 330, an input module 340, and a communication module 350. Those skilled in the art will understand that... Figure 4 The server architecture shown does not constitute a limitation on the server and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. Wherein: The processor 310 is the control center of the server, connecting various parts of the server via various interfaces and lines. It performs various server functions and processes data by running or executing software programs and / or modules stored in the memory 320, and by accessing data stored in the memory 320. In some embodiments, the processor 310 may include one or more processing cores; in some embodiments, the processor 310 may integrate an application processor and a modem processor, wherein the application processor primarily handles the operating system, user interface, and applications, while the modem processor primarily handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 310.
[0112] The memory 320 can be used to store software programs and modules. The processor 310 executes various functional applications and data processing by running the software programs and modules stored in the memory 320. The memory 320 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the server, etc. In addition, the memory 320 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 320 may also include a memory controller to provide the processor 310 with access to the memory 320.
[0113] The server also includes a power supply 330 that supplies power to the various components. In some embodiments, the power supply 330 can be logically connected to the processor 310 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 330 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0114] The server may also include an input module 340, which can be used to receive input numeric or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0115] The server may also include a communication module 350. In some embodiments, the communication module 350 may include a wireless module, through which the server can perform short-range wireless transmission, thereby providing users with wireless broadband internet access. For example, the communication module 350 can be used to help users send and receive emails, browse web pages, and access streaming media.
[0116] Although not shown, the server may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 310 in the server loads the executable files corresponding to the processes of one or more applications into the memory 320 according to the following instructions, and the processor 310 runs the applications stored in the memory 320, thereby implementing the steps in the methods of the various embodiments of the present invention.
[0117] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0118] As can be seen from the above, the electronic device provided in this embodiment of the invention can acquire the current execution stage and state data of the embodied intelligent agent when performing a target task. These two data points are input into a dynamic reward function to generate force adjustment parameters and pose adjustment parameters that match the current actual situation, thereby improving the adaptability of the embodied intelligent agent's motion control under different task stages and state conditions. The force adjustment parameters and pose adjustment parameters are converted into initial control parameters executable by the embodied intelligent agent. During control using these initial control parameters, the initial control parameters are corrected based on the fluctuation of the grasping force and the collision risk prediction results, which can reduce the instability or safety issues caused by grasping force fluctuations, pose deviations, or collision risks. Based on the corrected target control parameters, the embodied intelligent agent is controlled to execute the action corresponding to the current execution stage. When the current execution stage meets the stage completion conditions, the system switches to the next execution stage or ends the target task. This achieves continuous connection between different execution stages in the target task, reduces control mismatch during stage switching, and effectively improves the adaptability, execution stability, and operational safety of the embodied intelligent agent when performing the target task.
[0119] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0120] To this end, embodiments of the present invention provide a computer-readable storage medium storing a plurality of instructions which can be loaded by a processor to execute steps in any of the action control methods for an embodied intelligent agent provided in embodiments of the present invention.
[0121] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0122] According to one aspect of the present invention, a computer program product or computer program is provided, comprising a computer program / instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program / instructions from the computer-readable storage medium and executes the computer program / instructions, causing the electronic device to perform the methods provided in various optional implementations of the embodied intelligent agent's motion control or motion decision model training aspects provided in the above embodiments.
[0123] Since the instructions stored in the storage medium can execute the steps in any of the embodied intelligent agent action control methods provided in the embodiments of the present invention, the beneficial effects that any of the embodied intelligent agent action control methods provided in the embodiments of the present invention can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0124] The foregoing has provided a detailed description of the action control method and electronic device for an embodied intelligent agent provided by the embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for controlling the action of an embodied intelligent agent, characterized in that, The method includes: The current execution stage and state data of the embodied intelligent agent when performing the target task are obtained, and the state data includes multi-dimensional force data, visual data, pose data and obstacle data; The current execution stage and state data are input into the action decision model to obtain force adjustment parameters and pose adjustment parameters. The action decision model is a model trained based on a dynamic reward function, which is used to differentiate the evaluation of grasping force, pose accuracy, placement stability and collision risk at different execution stages. The step of inputting the current execution stage and state data into the action decision model to obtain force adjustment parameters and pose adjustment parameters includes: The multidimensional force data, visual data, pose data, and obstacle data are dimensionally aligned and fused to obtain a state vector; the state vector is then merged with the current execution stage to obtain input parameters; the input parameters are then input into the action decision model to obtain force adjustment parameters and pose adjustment parameters that match the current execution stage. The process of dimensional alignment and fusion of the multidimensional force data, visual data, pose data, and obstacle data to obtain a state vector includes: Principal component analysis is performed on the visual data and the pose data to extract visual features and pose features; data related to grasping stability is extracted from the multidimensional force data to form force features; spatial coordinates, minimum safe distance, and size of obstacles are extracted from the obstacle data to form obstacle features; the visual features, pose features, force features, and obstacle features are concatenated to obtain a state vector. The force adjustment parameters and the pose adjustment parameters are converted into the initial control parameters of the embodied intelligent agent; During the process of controlling the embodied intelligent agent using the initial control parameters, the initial control parameters are corrected to target control parameters based on the fluctuation of the grasping force and the collision risk prediction results, and the grasping force is extracted from the multidimensional force data; Based on the target control parameters, the embodied intelligent agent is controlled to perform the action corresponding to the current execution stage, and when the current execution stage meets the stage completion conditions, it switches to the next execution stage or ends the target task.
2. The method according to claim 1, characterized in that, In the process of controlling the embodied intelligent agent using the initial control parameters, the initial control parameters are corrected to target control parameters based on the fluctuation of the grasping force and the collision risk prediction results, including: During the process of controlling the embodied intelligent agent using the initial control parameters, the fluctuation degree of the grasping force is determined based on multiple grasping forces within a preset time window; Based on the pose data of the target object, obstacles, and the embodied intelligent agent, collision risk prediction is performed, where the target object is the processing object in the target task. If the fluctuation of the grasping force is greater than the fluctuation threshold, the force adjustment parameter is corrected to obtain the target control parameter; If a collision risk is predicted, the force adjustment parameters and pose adjustment parameters are corrected to obtain the target control parameters.
3. The method according to claim 2, characterized in that, The collision risk prediction process based on the pose data of the target object, obstacles, and the embodied intelligent agent includes: Based on the pose data of the target object and obstacles, construct the spatial bounding volume of the target object and obstacles; Based on the pose data of the embodied intelligent agent and the initial control parameters, predict multiple predicted trajectory points of the embodied intelligent agent within a preset future time period; The target distance is determined based on the distance between each predicted trajectory point and the spatial bounding volume; If the target distance is less than the preset safety distance, a collision risk is predicted.
4. The method according to any one of claims 1-3, characterized in that, The action decision model is trained in the following manner: The sample execution phase and sample state data when the embodied intelligent agent performs the sample task are input into the basic decision model to obtain the prediction force adjustment parameters and the prediction pose adjustment parameters. The sample state data includes multi-dimensional force samples, visual samples, pose samples and obstacle samples. Using the predicted force adjustment parameters and the predicted pose adjustment parameters, the embodied intelligent agent is controlled to perform actions, and execution feedback data is obtained; Calculate the target reward value corresponding to the dynamic reward function based on the execution feedback data and the sample execution stage; The model parameters of the basic decision model are updated based on the target reward value until the training completion condition is met, thus obtaining the action decision model.
5. The method according to claim 4, characterized in that, The step of calculating the target reward value corresponding to the dynamic reward function based on the execution feedback data and the sample execution stage includes: Obtain the basic reward weight distribution corresponding to the sample execution stage. The basic reward weight distribution includes weights corresponding to multiple dimensions, including grasping strength, pose accuracy, placement stability, and collision risk. Using the execution feedback data and the task objective of the sample task, calculate the weight adjustment factor corresponding to each dimension; Based on the weight adjustment factor and the execution progress of the sample task, the basic reward weight distribution is adaptively adjusted to obtain the target reward weight distribution; The target reward value is calculated using the target reward weight distribution and the sub-reward function corresponding to each dimension.
6. The method according to claim 5, characterized in that, The weight adjustment factors include force deviation, positioning deviation, placement fit, and force fluctuation. The adaptive adjustment of the basic reward weight distribution based on the weight adjustment factors and the execution progress of the sample task to obtain the target reward weight distribution includes: If the force deviation is greater than the force threshold, the weight corresponding to the grasping force is increased and the weight corresponding to the pose accuracy is decreased to obtain the target reward weight distribution. If the positioning deviation is greater than the positioning threshold, the weight corresponding to the pose accuracy is increased and the weight corresponding to the collision risk is decreased to obtain the target reward weight distribution. If the placement fit or the force fluctuation meets the placement conditions, increase the weight corresponding to the placement stability and decrease the weight corresponding to the gripping force to obtain the target reward weight distribution; If the execution progress of the sample task reaches the preset progress, the weight corresponding to the collision risk is gradually increased to obtain the target reward weight distribution.
7. The method according to claim 5, characterized in that, The calculation of the target reward value using the target reward weight distribution and the sub-reward function corresponding to each dimension includes: Using the execution feedback data and the task objective of the sample task, calculate the sub-reward value corresponding to the sub-reward function for each dimension; The target reward value is obtained by weighting the sub-reward values of all dimensions based on the target reward weight distribution.
8. An electronic device, characterized in that, It includes a processor and a memory, the memory storing multiple instructions; the processor loads instructions from the memory to execute the steps in the action control method for an embodied intelligent agent as described in any one of claims 1-7.