Mechanical arm control method, mechanical arm and storage medium

By acquiring the robot arm's task instructions and multimodal perception data to predict its state and determine the target joint configuration sequence, the robot arm's adaptability in complex dynamic environments is solved, achieving higher robustness and flexibility.

CN121315976APending Publication Date: 2026-01-13ANHUI KAIYANG TECHNOLOGY CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511752750.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Robotic arms have poor adaptability and insufficient robustness in complex dynamic environments, and cannot guarantee stable performance.

Method used

By acquiring the task instructions and multimodal perception data of the robotic arm, state prediction is performed to determine the target joint configuration sequence, and the robotic arm is controlled based on this. Visual, kinematic and force data are used for feature extraction and fusion to optimize joint motion to meet the constraints.

Benefits of technology

It improves the adaptability and robustness of the robotic arm in complex dynamic environments, and enhances the flexibility and stability of its movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121315976A_ABST
    Figure CN121315976A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a mechanical arm control method, a mechanical arm and a storage medium, and relates to the technical field of robot control. The method comprises the steps that task instructions and multi-modal sensing data of the mechanical arm are obtained, the task instructions are used for guiding operation behaviors of the mechanical arm, and the multi-modal sensing data are used for representing state characteristics and environment characteristics of the mechanical arm; state prediction is carried out on the basis of the task instruction and the multi-mode sensing data, future state features are obtained, and the future state features are used for guiding the motion trail of the mechanical arm; a target joint configuration sequence is determined based on the future state features, and the target joint configuration sequence is used for representing joint position information of the mechanical arm; and the mechanical arm is controlled based on the target joint configuration sequence. The technical problem that the adaptability of a mechanical arm in a complex dynamic scene is poor in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of robot control, in particular, to a robot arm control method, a robot arm and a storage medium. BACKGROUND

[0002] With the progress of sensing technology and artificial intelligence, a robot system can integrate multiple sensors to obtain more comprehensive environmental information and self-state. However, when the robot arm is executing a task, it often encounters external interference, such as changes in contact force, sudden changes in lighting conditions, and position deviations of the robot arm itself. The robot arm in the prior art has insufficient robustness when facing the above uncertainties and cannot guarantee stable performance in various complex dynamic environments.

[0003] There is currently no good solution to the above technical problems. SUMMARY

[0004] Embodiments of the present application provide a robot arm control method, a robot arm and a storage medium to at least solve the technical problem of poor adaptability of a robot arm in a complex dynamic scene in related technologies.

[0005] According to an aspect of embodiments of the present application, a robot arm control method is provided, the method comprising: obtaining a task instruction and multi-modal perception data of a robot arm, wherein the task instruction is used to guide the operation behavior of the robot arm, and the multi-modal perception data is used to represent the state features and environmental features of the robot arm; performing state prediction based on the task instruction and the multi-modal perception data to obtain future state features, wherein the future state features are used to guide the motion trajectory of the robot arm; determining a target joint configuration sequence based on the future state features, wherein the target joint configuration sequence is used to represent the joint position information of the robot arm; and controlling the robot arm based on the target joint configuration sequence.

[0006] Further, the multi-modal perception data at least includes: visual data, kinematics data and force sensation data, the visual data is used to represent the environmental visual features around the robot arm and the position of the target object, the kinematics data is used to represent the joint position and speed state of the robot arm, and the force sensation data is used to represent the contact force of the end effector of the robot arm.

[0007] Further, the state prediction based on the task instruction and the multi-modal perception data to obtain the future state feature comprises: fusing the task instruction and the multi-modal perception data based on the first feature vector, the second feature vector and the third feature vector to obtain a fusion feature, wherein the first feature vector comprises a visual feature vector, a joint state feature vector, a force feature vector and a language feature vector, the second feature vector is used to represent multi-modal prediction information at a future time and a constraint condition, and the third feature vector is used to guide a joint configuration sequence of the robot arm to meet the constraint condition; and performing state prediction based on the fusion feature to obtain the future state feature.

[0008] Further, the robot arm control method further comprises: extracting features from the visual data by using a visual encoder to obtain a visual feature vector; extracting features from the kinematics data by using a kinematics encoder to obtain a joint state feature vector; extracting features from the force perception data by using a force encoder to obtain a force feature vector; and extracting features from the task instruction by using an instruction parser to obtain a language feature vector.

[0009] Further, the state prediction based on the fusion feature to obtain the future state feature comprises: performing state prediction based on the fusion feature and historical data to obtain future visual features and future pose features; and performing state prediction based on the fusion feature and a current joint configuration sequence to obtain future matrix features and future joint torque features.

[0010] Further, the determination of the target joint configuration sequence based on the future state feature comprises: determining a target function and a target constraint condition based on the future state feature, wherein the target constraint condition is used to limit a joint movement range of the robot arm; and determining the target joint configuration sequence based on the target function and the target constraint.

[0011] Further, the target function comprises: a first function, a second function and a third function, the first function is used to reduce a trajectory tracking error of the robot arm in a task space, the second function is used to reduce joint vibration of the robot arm, and the third function is used to improve operability of the robot arm, a priority of the first function is higher than a priority of the second function, and a priority of the second function is higher than a priority of the third function.

[0012] Further, the control of the robot arm based on the target joint configuration sequence comprises: controlling the robot arm based on the target joint configuration sequence, and obtaining an effector pose of the robot arm; determining a target trajectory tracking error based on the target joint configuration sequence and the effector pose; updating the future state feature based on the target trajectory tracking error; determining a new joint configuration sequence based on the updated future state feature; and controlling the robot arm based on the new joint configuration sequence.

[0013] According to a further aspect of the embodiments of the present application, a mechanical arm control device is also provided, comprising: an acquisition module configured to acquire a task instruction of the mechanical arm and multi-modal perception data, wherein the task instruction is configured to guide an operation behavior of the mechanical arm, and the multi-modal perception data is configured to represent state features and environment features of the mechanical arm; a prediction module configured to perform state prediction based on the task instruction and the multi-modal perception data to obtain future state features, wherein the future state features are configured to guide a motion trajectory of the mechanical arm; a determination module configured to determine a target joint configuration sequence based on the future state features, wherein the target joint configuration sequence is configured to represent joint position information of the mechanical arm; and a control module configured to control the mechanical arm based on the target joint configuration sequence.

[0014] According to a further aspect of the embodiments of the present application, a mechanical arm is also provided, comprising: a memory configured to store an executable program; and a processor configured to run the program, wherein the program performs the mechanical arm control method in any of the above aspects when running.

[0015] According to a further aspect of the embodiments of the present application, a computer readable storage medium is also provided, comprising a stored executable program, wherein the executable program controls a device where the computer readable storage medium is located to perform the mechanical arm control method in the embodiments of the present application when running.

[0016] According to a further aspect of the embodiments of the present application, a computer program product is also provided, comprising a computer program, which implements the mechanical arm control method in the embodiments of the present application when executed by a processor.

[0017] In the embodiments of the present application, the task instruction of the mechanical arm and the multi-modal perception data are first acquired, then the state prediction is performed based on the task instruction and the multi-modal perception data to obtain the future state features for guiding the motion trajectory of the mechanical arm, and then the target joint configuration sequence for representing the joint position information of the mechanical arm is determined based on the future state features, and finally the mechanical arm is controlled based on the target joint configuration sequence. Thus, the adaptability of the mechanical arm in a complex dynamic environment is improved, thereby achieving the technical effect of significantly improving the robustness and flexibility of the motion of the mechanical arm, and further solving the technical problem of poor adaptability of the mechanical arm in a complex dynamic scene in the related art. BRIEF DESCRIPTION OF DRAWINGS

[0018] The accompanying drawings, which are included to provide a further understanding of the present application, form a part of the present application and illustrate the illustrative embodiments of the present application and its description, which are used to explain the present application, and do not constitute improper limitations on the present application. In the drawings:

[0019] Figure 1 is a flowchart of a mechanical arm control method according to an embodiment of the present application;

[0020] Figure 2 is a future state prediction branch diagram according to an embodiment of the present application;

[0021] Figure 3 is a cross-domain constraint optimization flow diagram according to an embodiment of the present application;

[0022] Figure 4 is a multi-modal fusion module structure diagram according to an embodiment of the present application;

[0023] Figure 5 is a schematic diagram of a predictive inverse kinematics system according to an embodiment of the present application;

[0024] Figure 6 is a structural block diagram of a robot arm control device according to an embodiment of the present application. DETAILED DESCRIPTION

[0025] In order to enable persons skilled in the art to better understand the scheme of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor should be within the scope of protection of the present application.

[0026] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] For the convenience of understanding, some descriptions of concepts related to the embodiments of the present application are exemplarily given for reference. As shown below:

[0028] Multi-modal fusion refers to the process of integrating and processing information from different perception channels or data sources into a unified and comprehensive representation. In this application, multi-modal fusion refers to the process of extracting, compressing, and interacting features from different modalities of data such as visual data, kinematic data, force sensation data, and task instructions through specific algorithms to form a fusion feature containing multiple types of information, in order to more comprehensively understand and predict the motion state of the robot arm and the environmental characteristics.

[0029] Inverse kinematics is a mathematical algorithm used in robot control to solve the configuration of robot joints in joint space from the desired position and attitude of the end effector in task space. In this application, the inverse kinematics method is extended to predictive inverse kinematics, which not only considers the current state but also predicts future state characteristics to optimize the joint configuration sequence, in order to improve the accuracy and robustness of trajectory tracking.

[0030] The Jacobian matrix is a matrix that describes the linear transformation relationship of the robot from joint space to task space. Future Jacobian prediction is based on the current and predicted future state of the robot arm to estimate the trend of the future Jacobian matrix, in order to more accurately predict the future motion state of the end effector.

[0031] According to an embodiment of the present application, a method embodiment of a robot arm control method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in a different order.

[0032] In this embodiment, a robot arm control method is provided, Figure 1 is a flowchart of a robot arm control method according to an embodiment of the present application, as Figure 1 shown, the flow includes the following steps:

[0033] Step S10, obtaining the task instruction and multi-modal perception data of the robot arm, wherein the task instruction is used to guide the operation behavior of the robot arm, and the multi-modal perception data is used to represent the state characteristics and environmental characteristics of the robot arm.

[0034] In the embodiment of the present application, the robot arm in this application can not only be a robot arm with redundant degrees of freedom, but also a conventional robot arm, which is not limited here.

[0035] The task instruction is information used to guide the robot arm to perform specific operations. The task instruction can take different forms, including natural language instructions or Cartesian coordinate instructions, which are not limited here.

[0036] The multi-modal perception data refers to data about the state of the robot arm and the environment acquired from different perception sources, mainly including the following types: visual data, kinematics data, force sensation data, and language instruction data, which are not limited here.

[0037] Exemplarily, a task space desired trajectory x_des(t:t+H) (i.e., a desired pose sequence within a prediction time domain H) is acquired through natural language instructions or Cartesian coordinate instructions, and multi-modal data of the robot arm at the current time is collected, including: visual data: RGB (Red Green Blue) images are acquired through Eye-on-Hand and Eye-on-Base cameras, visual features are extracted through a pre-trained visual encoder (such as MAE-ViT), and task-related features (such as an end effector, target objects, and obstacles) are compressed and screened through a perception resampler (Perceiver Resampler). Kinematics data: current joint positions , joint velocities are acquired through joint encoders. Force sensation data: contact forces of the end effector are acquired through force sensors.

[0038] It can be seen that the task instruction defines the target that the robot arm needs to reach, and the multi-modal perception data provides real-time environmental information and robot state for achieving this target.

[0039] In step S12, state prediction is performed based on the task instruction and the multi-modal perception data, and future state features are obtained, wherein the future state features are used to guide the motion trajectory of the robot arm.

[0040] In the embodiments of the present application, the future state features are a feature set representing the state of the robot arm and its environment within a certain time interval in the future, which is obtained through a prediction model and algorithm based on the given multi-modal perception data at the current time and the task instruction of the robot arm.

[0041] The state prediction based on the task instruction and the multi-modal perception data can be understood as that, by comprehensively analyzing the current task target of the robot arm and the multi-modal perception information of the environment where the robot arm is currently located, a prediction algorithm is used to infer and rehearse the potential motion state and environmental interaction of the robot arm within a future period of time.

[0042] It can be seen that the future state features allow the robot arm to evaluate potential impacts such as visual occlusion, force feedback, and physical constraints before action, so as to make more correct decisions, optimize the motion trajectory of the robot arm to achieve the task target, and avoid potential risks and limitations, thereby improving the adaptability of the robot arm in a dynamic environment.

[0043] At step S14, a target joint configuration sequence is determined based on the future state feature, where the target joint configuration sequence is used to represent joint position information of the robot arm.

[0044] In the embodiments of the present application, the target joint configuration sequence is a set of optimal joint position setting value sequences guiding the movement of the robot arm, which describes the expected positions of each joint of the robot arm in the future period of time, so that the robot arm can move according to the expected positions to complete the action specified by the task instruction, while meeting various constraint conditions such as kinematic constraint, dynamic constraint and safety constraint, etc., which are not limited here.

[0045] Determining the target joint configuration sequence based on the future state feature can be understood as determining a series of joint position setting values according to the future motion state of the robot arm and the environmental interaction information, through an optimization algorithm, to guide the robot arm to move according to the optimized trajectory and fully consider all constraint conditions in the future period of time.

[0046] As can be seen, determining the target joint configuration sequence based on the future state feature not only considers the requirements of the task instruction, but also deeply integrates the prediction results of multi-modal perception data, through complex optimization calculation, to obtain a joint position instruction sequence that can guide the robot arm to move stably under complex conditions, so as to improve the adaptability of the robot arm in complex dynamic scenarios.

[0047] At step S16, the robot arm is controlled based on the target joint configuration sequence.

[0048] In the embodiments of the present application, controlling the robot arm based on the target joint configuration sequence can be understood as using the determined optimal joint position setting value sequence to guide and command the movement of the robot arm, to ensure that the robot arm executes the task according to the predetermined trajectory and attitude requirements, while complying with all relevant kinematic, dynamic and safety constraints.

[0049] As can be seen, through the above control process based on the target joint configuration sequence, the robot arm not only can execute a complex motion trajectory, but also can make a quick response when encountering challenges, thereby enhancing the practicality and flexibility of the robot arm in various complex application scenarios.

[0050] By the above steps of the present application, firstly, the task instruction and the multi-modal perception data of the robot arm are acquired, then the state prediction is performed based on the task instruction and the multi-modal perception data, the future state feature for guiding the motion trajectory of the robot arm is obtained, then the target joint configuration sequence for representing the joint position information of the robot arm is determined based on the future state feature, and finally the robot arm is controlled based on the target joint configuration sequence. Thus, the purpose of improving the adaptability of the robot arm in a complex dynamic environment is achieved, thereby realizing the technical effect of significantly improving the robustness and flexibility of the motion of the robot arm, and further solving the technical problem of poor adaptability of the robot arm in a complex dynamic scene in the related art.

[0051] Optionally, in step S10, the multi-modal perception data at least includes: visual data, kinematics data and force sensation data, the visual data is used to represent the visual features of the environment around the robot arm and the position of the target object, the kinematics data is used to represent the joint position and speed state of the robot arm, and the force sensation data is used to represent the contact force of the end effector of the robot arm.

[0052] In the embodiment of the present application, an eye-on-hand and eye-on-base depth camera, such as RealSense D435i, is used to capture visual data within the workspace of the robot arm. The visual data includes but is not limited to RGB images and depth information, which can reflect the visual features of the environment around the robot arm, such as the layout of the environment, the shape, position and attitude of the target object, and the motion trajectory of any dynamic obstacle. Feature extraction is performed by a pre-trained visual encoder, such as Masked Autoencoder (MAE-ViT-Base), to compress and filter task-related features. The visual data is further processed by a perception resampler (Perceiver Resampler) to generate a visual vector, which is used for subsequent fusion and prediction.

[0053] The visual encoder can use a residual network (ResNet), an efficient network (EfficientNet), or other lighter or task-specific optimized convolutional neural network architectures instead of the Masked Autoencoder (MAE-ViT-Base). In terms of feature compression, in addition to the perception resampler (Perceiver Resampler), methods such as autoencoder (Autoencoder), principal component analysis or sparse coding can be used to reduce the dimensionality of high-dimensional visual features and extract task-related core features.

[0054] The joint encoder is used to collect kinematics data in real time, reflecting the internal state of the robot arm. The kinematics data includes but is not limited to the current joint position and joint speed Joint position data is a measurement of joint angle, and joint velocity reflects the instantaneous rate of change of the joint, both of which together describe the motion state of the robot arm.

[0055] Force sensation data is obtained through force sensors, and the force sensation data contains the magnitude and direction of the force generated when in contact with the environment (such as target objects, obstacles), as well as the numerical value of the moment.

[0056] In addition, in addition to visual data, kinematic data and force sensation data, other perception modalities can be introduced, such as a tactile sensor array (for more detailed contact perception), an acoustic sensor (for detecting abnormal sounds or environmental interactions), and an inertial measurement unit (for providing more accurate robot body pose and motion information). In terms of feature fusion, a sequence model such as a gated recurrent unit or a long short-term memory network can be used instead of a Transformer encoder, or a probabilistic graph model such as Bayesian fusion or Kalman filtering can be used for multi-sensor information fusion to adapt to different modal data characteristics and real-time requirements. A modal selection mechanism can also be introduced to dynamically select or weight different perception modalities according to task requirements, for example, focusing on tactile and force sensation in fine operations, and focusing on vision and inertial measurement units in navigation and obstacle avoidance.

[0057] As can be seen, the expanded perception modalities can provide richer and more comprehensive environment and robot state information, improving the system's perception ability and robustness in complex environments. The selective fusion mechanism can adaptively adjust the information processing strategy according to the task, optimize the allocation of computing resources, and at the same time maintain or improve the accuracy of trajectory tracking, constraint satisfaction and dynamic adaptability. In addition, using different visual encoders and feature compression methods can achieve different balances between computational efficiency, feature expression ability and generalization, possibly providing better performance in specific scenarios. Directly extracting three-dimensional environmental features helps to more accurately understand spatial geometric relationships, thereby improving the accuracy of collision detection and path planning.

[0058] Optionally, in step S12, a state prediction is performed based on the task instruction and the multi-modal perception data to obtain future state features, which can include the following execution steps:

[0059] Step S121, based on the first feature vector, the second feature vector and the third feature vector, the task instruction and the multi-modal perception data are fused to obtain a fusion feature, wherein the first feature vector includes a visual feature vector, a joint state feature vector, a force feature vector and a language feature vector, the second feature vector is used to represent multi-modal prediction information and constraint conditions at a future time, and the third feature vector is used to guide a joint configuration sequence of the robot arm to satisfy the constraint conditions;

[0060] Step S122, based on the fusion feature, a state prediction is performed to obtain future state features.

[0061] In this embodiment, the task instructions and multimodal perception data are fused based on the first feature vector, the second feature vector, and the third feature vector to obtain the fused feature. This can be understood as fusing visual feature vectors, joint state feature vectors, force feature vectors, language feature vectors, multimodal prediction information representing future moments, feature vectors representing constraints, and feature vectors guiding the joint configuration sequence of the robotic arm to satisfy the constraints to obtain the fused feature. .

[0062] State prediction based on fusion features can be understood as using fusion features that integrate multi-source information such as vision, kinematics, force perception, and task instructions to infer and predict various state changes of the robotic arm in a certain time domain in the future.

[0063] For example, to predict the future state based on fused features, one can first use the fused features... We construct future vision and pose prediction branches as well as future Jacobi and torque prediction branches. Figure 2 This is a schematic diagram of future state prediction branches according to one embodiment of this application, such as... Figure 2 As shown, in the future vision and pose prediction branch: using a conditional vision prediction model, with the task instruction as the target g and historical observations... ={ , , Given} (where m is the historical window length) as input, predict the visual features at the next H time points. H and the corresponding task space pose H, the loss function uses a weighted average of pixel-level mean square error and pose error. The loss function formula is as follows: Where α∈[0,1] are the visual-pose weight coefficients. Under the future Jacobian and moment prediction branch: based on the comprehensive motion transformation matrix, utilizing the current joint configuration... and future position H is the Jacobian matrix estimated for future times using Taylor expansion. The formula for estimating the Jacobian matrix is: ,in , The calculation is performed directly from the comprehensive motion transformation matrix (based on the relationship between joint velocity / acceleration and link kinematics), avoiding the noise amplification problem of traditional numerical differentiation. Then, it is combined with the predicted future joint configuration. H, predicting future joint moments by simplifying the inverse dynamics model. H, ensuring that the torque is within the hardware limit τ_min ≤ Within ≤τ_max.

[0064] Exemplarily, the state prediction of the fused features can also be performed using a unified end-to-end prediction model to obtain future state features, including: directly predicting the future H time visual features from the fused features using a large-scale transformation model or diffusion model End-to-end prediction of future H time visual features H, task space pose H, Jacobian matrix H, and joint torque H. The end-to-end prediction model can be trained through multi-task learning, sharing the underlying feature representation, thereby capturing the internal correlation between different future states. The end-to-end prediction model simplifies the system architecture, reduces the error accumulation between modules, and can obtain more consistent and accurate future state prediction through end-to-end optimization.

[0065] Exemplarily, in the future Jacobian and torque prediction branch, in addition to the comprehensive motion transformation matrix and the simplified inverse dynamics model, a simulation environment based on a physics engine (such as MuJoCo, Isaac Gym) can be introduced for model predictive control. Through rapid iteration and optimization in the simulation environment, combined with real-world data for domain adaptation, the shortcomings of purely data-driven models in physical consistency can be made up. For Jacobian estimation, numerical differentiation or finite difference method can be used, and combined with machine learning methods such as Gaussian process regression for noise suppression and smoothing processing to deal with the situation of complex comprehensive motion transformation matrix calculation or inaccurate model. By combining the physical model, the physical rationality and stability of the prediction results can be enhanced, especially when dealing with complex dynamics and contact tasks. The combination of numerical differentiation and Gaussian process regression provides another robust solution for Jacobian estimation, reducing the dependence on accurate comprehensive motion transformation matrix model.

[0066] Optionally, the robot arm control method further includes the following execution steps:

[0067] The visual data is extracted using a visual encoder to obtain a visual feature vector;

[0068] The kinematics data is extracted using a kinematics encoder to obtain a joint state feature vector;

[0069] The force sensation data is extracted using a force encoder to obtain a force feature vector;

[0070] The task instruction is extracted using an instruction parser to obtain a language feature vector.

[0071] In the embodiments of the present application, the visual encoder is used to extract features from the visual data to obtain the visual feature vector. It can be understood that the visual encoder (such as a pre-trained MAE-ViT-Base) receives the RGB images acquired from the eye-in-hand and eye-in-base cameras as input, automatically extracts key features in the images through a series of convolution, pooling, activation function and full connection layer or self-attention mechanism, which may include object edges, textures, colors, positions and postures, etc. The above key features are compressed into a fixed length vector, which is called a visual feature vector.

[0072] The kinematics encoder is used to extract features from the kinematics data to obtain the joint state feature vector. It can be understood that the kinematics encoder processes the kinematics data (such as joint position and joint velocity ) to obtain the joint state feature vector, which includes the current motion state of the robot arm, indicating the position, velocity, direction and other key information of the robot arm.

[0073] The force encoder is used to extract features from the force data to obtain the force feature vector. It can be understood that the force encoder receives the output of the force sensor (such as the end effector contact force ), and generates the force feature vector through data preprocessing (such as filtering, noise reduction), physical quantity calibration (such as force unit conversion) or feature learning (such as extracting force perception patterns through autoencoder).

[0074] The instruction parser is used to extract features from the task instructions to obtain the language feature vector. It can be understood that the instruction parser is responsible for parsing natural language instructions or Cartesian coordinate instructions into a set of numbers or vector representations that can be directly used to control the robot arm, obtaining the language feature vector.

[0075] As can be seen, through the conversion and generation of the above feature vectors, the present application can effectively convert the original data from different modalities into high-dimensional and abstract representations, facilitating efficient data fusion and analysis in the subsequent cross-modal attention module, thereby achieving accurate prediction and control of the future state of the robot arm.

[0076] Optionally, in step S122, the future state feature is obtained by predicting the state based on the fusion feature, including the following execution steps:

[0077] Step S1221, predicting the future visual feature and the future pose feature based on the fusion feature and the historical data;

[0078] Step S1222, predicting the future matrix feature and the future joint torque feature based on the fusion feature and the current joint configuration sequence.

[0079] In the embodiments of the present application, the historical data refers to the multimodal data records of the robot arm in the past period of time ={ , , } (wherein m is the length of the historical window). The historical data helps the prediction model to understand the motion trend of the robot arm and the change rule of the environment, thereby generating more accurate prediction results.

[0080] The future visual feature is the predicted RGB image feature in the field of view of the robot arm in the future H time points based on the fusion feature and the historical data, using a conditional visual prediction model or other prediction algorithm H.

[0081] The future pose feature is the task space pose feature corresponding to the future visual feature in the future H time points H.

[0082] The current joint configuration sequence refers to the joint state of the robot arm at the current time .

[0083] The future matrix feature is the Jacobian matrix at the future time .

[0084] The future joint torque feature is the torque obtained based on the predicted future joint configuration H. H.

[0085] The state prediction based on the fusion feature and the historical data to obtain the future visual feature and the future pose feature can be understood as using the generated fusion feature and the historical data of the robot arm in the past period of time as input to predict the image features that can be seen by the eye-in-hand and eye-in-base cameras of the robot arm in the future H time points, and the position and attitude of the robot arm end effector in the task space at the future H time points.

[0086] The state prediction based on the fusion feature and the current joint configuration sequence to obtain the future matrix feature and the future joint torque feature can be understood as using the fusion feature and the current joint configuration sequence to estimate the Jacobian matrix of the robot arm end effector at the future H time points and the torque required for each joint in the future H time points H.

[0087] It can be seen that the future state of the robot arm is predicted from the environment perception and the dynamic kinematics by the above state prediction steps, respectively, which provides comprehensive information for the predictive inverse kinematics of the robot arm, and significantly improves the motion control ability and work efficiency of the robot in a complex and dynamic scene.

[0088] Optionally, in step S14, determining the target joint configuration sequence based on the future state features comprises the following execution steps:

[0089] Step S141, determining the objective function and the target constraint condition based on the future state features, wherein the target constraint condition is used to limit the joint motion range of the robot arm;

[0090] Step S142, determining the target joint configuration sequence based on the objective function and the target constraint.

[0091] In the embodiments of the present application, the objective function aims to minimize the task space trajectory tracking error, the joint motion change amount of the robot arm, and maximize the operability of the robot arm, so as to ensure that the robot arm can efficiently, smoothly and safely complete the task.

[0092] The target constraint condition is a set of rules for limiting the optimization process, which ensures that the motion of the robot arm is both safe and feasible. Exemplarily, the target constraint condition includes but is not limited to the motion constraint: q_min≤ ≤q_max; _min≤( - ) / Δt≤ _max;ÿ_min≤( - ) / Δt≤ÿ_max. The dynamic constraint: τ_min≤ ≤τ_max, wherein =M( )× +C( , )× +G( ) (inertia, Coriolis, gravity term). Future collision constraint: based on the predicted visual feature , through a collision detection algorithm (such as a swept-sphere volume method), the distance between the link and the obstacle ≥d_safe (safety distance).

[0093] Determining the objective function and the target constraint condition based on the future state features can be understood as constructing a multi-objective function and a constraint condition for optimizing the future motion of the robot arm based on the predicted future visual feature, future pose feature, future matrix feature and future joint torque feature.

[0094] The determination of the target joint configuration sequence based on the target function and the target constraint can be understood as solving the above target function and target constraint condition by a hierarchical quadratic programming solver to determine a set of optimal joint configuration sequences that can meet all requirements within the prediction time domain H H .

[0095] Exemplarily, for a multi-objective optimization problem, a nonlinear optimization algorithm such as a sequential quadratic programming, an interior point method, or an augmented Lagrangian method can also be used instead of a traditional quadratic programming solver. By using a nonlinear optimization method, more complex optimization problems can be handled, especially when the robot arm approaches a singular point or a complex obstacle environment, a smoother and safer motion trajectory can be provided, and dynamic weight adjustment can better balance the conflicts between different optimization objectives.

[0096] In addition, the entire inverse kinematics solving process can also be regarded as a Markov decision process, and a deep reinforcement learning method such as a proximal policy optimization, a soft actor-critic, or a model-based reinforcement learning can be used to directly learn a policy from the current state to the optimal joint configuration sequence. The reward function can comprehensively consider factors such as trajectory tracking error, joint smoothness, operability, constraint satisfaction, and collision avoidance. The deep reinforcement learning method can eliminate the dependence on the construction and solver of an explicit quadratic programming problem, and autonomously learn an optimal control policy through interaction with the environment. The reinforcement learning method has strong adaptability and generalization ability, and can handle complex, high-dimensional, and nonlinear problems that are difficult to model by traditional optimization methods. Through the deep reinforcement learning method, the robot can autonomously discover more efficient and robust motion strategies, especially suitable for unknown or dynamically changing environments, and can achieve end-to-end optimization, reducing the complexity of manually designed rules.

[0097] As can be seen, by converting the predicted future state features into explicit target functions and constraint conditions, and then solving the target joint configuration sequence by an optimization algorithm, the robot arm can improve work efficiency and task success rate when facing complex and dynamic environments.

[0098] Optionally, the target function includes a first function, a second function, and a third function, the first function is used to reduce the trajectory tracking error of the robot arm in the task space, the second function is used to reduce the joint vibration of the robot arm, and the third function improves the operability of the robot arm, the priority of the first function is higher than the priority of the second function, and the priority of the second function is higher than the priority of the third function.

[0099] In the embodiment of the application, the first function can be represented as a first-level target function: , which is used to minimize the task space trajectory tracking error.

[0100] The second function can be represented as a second-level target function: , to minimize the amount of joint configuration change for smoothness.

[0101] The third function can be represented as a tertiary objective function: , to maximize the manipulability of the robot arm to avoid singularities.

[0102] It can be seen that by setting a hierarchical multi-objective function and its priority, the predictive inverse kinematics method of the robot arm can achieve high-precision trajectory tracking, smooth and stable motion, and good manipulability in a complex and variable environment, significantly improving the performance of the robot arm, especially in high dynamic response scenarios, it can more flexibly and safely perform tasks.

[0103] Figure 3 is a cross-domain constraint optimization flowchart according to an embodiment of the present application, as shown in Figure 3 First, based on the predicted future state features, a multi-objective quadratic programming problem is constructed, and the objective function priority is set to minimize the task space trajectory tracking error (the first objective function), minimize the amount of joint configuration change (the second objective function), and maximize the manipulability of the robot arm (the third objective function) as the main optimization objectives. The quadratic programming problem also contains a variety of constraint conditions, including: motion constraints (joint position, velocity, acceleration), dynamics constraints (joint torque), future collision constraints (safety distance). Subsequently, the quadratic programming problem is solved by a hierarchical quadratic programming solver, and the optimal joint configuration sequence in the prediction time domain is output, ensuring the accuracy, stability and safety of the robot arm in executing tasks in a complex environment.

[0104] Optionally, in step S16, the robot arm is controlled based on the target joint configuration sequence, including the following execution steps:

[0105] Step S161, control the robot arm based on the target joint configuration sequence, and obtain the pose of the actuator of the robot arm;

[0106] Step S162, determine the target trajectory tracking error based on the target joint configuration sequence and the pose of the actuator;

[0107] Step S163, update the future state features based on the target trajectory tracking error;

[0108] Step S164, determine a new joint configuration sequence based on the updated future state features;

[0109] Step S165, control the robot arm based on the new joint configuration sequence.

[0110] In this embodiment, controlling the robotic arm based on the target joint configuration sequence and obtaining the actuator pose of the robotic arm can be understood as receiving the optimized target joint configuration sequence. H Then, the system configures the first joint in the sequence. The controller sends commands to the robotic arm, guiding it to perform corresponding movements. Subsequently, the system needs to monitor the robotic arm's actual motion state in real time, collecting the actual pose information of the end effector through sensors such as joint encoders. This is used for subsequent error assessment and state prediction updates.

[0111] Determining the target trajectory tracking error based on the target joint configuration sequence and actuator pose can be understood as the system comparing the current actual actuator pose. The target trajectory tracking error is obtained by comparing the expected trajectory pose x_des(t+1) in the task space: =x_des(t+1)- .

[0112] Updating future state features based on target trajectory tracking error can be understood as, when a target trajectory tracking error is detected... At that time, the system will base its decisions on the target trajectory tracking error. Update the future state characteristics of the robotic arm.

[0113] Determining a new joint configuration sequence based on updated future state features can be understood as re-determining the joint configuration sequence in the future prediction time domain based on the updated future state features. H .

[0114] Controlling a robotic arm based on a new joint configuration sequence can be understood as obtaining a new joint configuration sequence through iterative optimization. H It will be used to guide the movement of the robotic arm in the next moment.

[0115] Exemplarily, in the error feedback mechanism, an adaptive control law can also be introduced to adjust the weight coefficients in the quadratic programming optimization, the parameters of the prediction model, or the constraint boundaries online according to the real-time tracking error and environmental changes. For example, when a persistent trajectory tracking error is detected, the weight of the primary objective function (trajectory tracking error) can be dynamically increased. In addition, online learning or meta-learning algorithms can be integrated to enable the prediction model and optimization strategy to be continuously updated and improved based on new observation data without the need for large-scale offline training. In this way, the adaptive ability of the system can be enhanced, enabling the robot arm to better cope with long-term changes such as unknown environmental disturbances, model uncertainties, or robot arm wear, thereby maintaining high performance for a longer period of time. The online learning mechanism enables the system to continuously accumulate experience from actual operations and continuously optimize performance.

[0116] In addition, for real-time requirements, computationally intensive modules such as multi-modal perception, feature fusion, future state prediction, and cross-domain constraint optimization can also be distributedly deployed. For example, visual processing can be pre-processed and feature extracted on an edge computing device, and quadratic programming solving can be implemented on a dedicated hardware accelerator. At the same time, lightweight communication protocols and data compression techniques can be used to optimize the data transmission efficiency between modules and reduce communication delay. In this way, through distributed deployment and hardware acceleration, the real-time performance of the entire system can be significantly improved to meet higher frequency control requirements, thereby supporting higher speed and more precise robot motion. The application of edge computing can also reduce the dependence on cloud resources and improve the autonomy and security of the system.

[0117] As can be seen, through the above steps, a closed-loop "perception-prediction-optimization-execution" cycle is formed to ensure that the robot arm can continuously adapt to constraint changes in a complex and dynamic environment and achieve high-precision and high-dynamic response motion control.

[0118] Figure 4 is a multi-modal fusion module structure diagram according to an embodiment of the present application, as Figure 4As shown, in the vector encoding stage, multiple feature vectors are obtained based on the multi-modal fusion module structure, including: visual feature vector, joint state feature vector, force feature vector, language feature vector, and special feature vector (future view feature and action optimization feature). Among them, the visual feature vector is obtained by the camera, and the visual feature vector is compressed and the task-related features are screened based on the visual encoder to obtain the visual feature. The joint state feature vector is obtained by the joint sensor, and the joint state feature vector is processed based on the multilayer perceptron encoder to obtain the state feature. The force feature vector is obtained by the force sensor, and the force feature vector is processed based on the linear encoder to obtain the force feature. The language feature vector is extracted according to the task instruction, and the language feature vector is processed based on the text encoder to obtain the language feature (task instruction). All the above features are fused based on the cross-modal encoder to output the fusion feature and the attention weight heat map. The future visual-pose prediction and the future Jacobian-torque prediction are obtained based on the fusion feature, and the robot arm is controlled by using the future visual-pose prediction and the future Jacobian-torque prediction.

[0119] Figure 5 is a schematic diagram of a predictive inverse kinematics system according to an embodiment of the present application, as Figure 5 shown, the predictive inverse kinematics system includes: a multi-modal perception layer: a vision module: eye-in-hand and eye-in-base cameras for obtaining environment and end effector visual data (RGB images); a kinematics module: joint encoders and force sensors for obtaining joint state and contact force data (joint position / speed contact force); an instruction module: converting task instructions (natural language or Cartesian coordinate instructions) into task space desired trajectories. A conversion encoder: respectively extracts features from visual, kinematic, force, and instruction data to generate corresponding feature vectors, and fuses multi-modal feature vectors through self-attention and cross-attention mechanisms to output fusion features.

[0120] A future state prediction layer: based on a conditional visual prediction model, outputs future visual features and future poses; based on a comprehensive motion transformation matrix, calculates a future Jacobian matrix, and combines a simplified inverse dynamics model to predict a future Jacobian matrix and a future torque matrix.

[0121] A cross-domain constraint optimization layer: a multi-objective quadratic programming optimization module: generates a multi-objective optimization problem according to the priority of the objective function and the constraint condition, and solves the optimization problem using a quadratic programming algorithm to output an optimal joint configuration sequence.

[0122] An execution feedback layer: an execution module: receives the optimal joint configuration sequence to drive the robot arm to perform motion; an error feedback module: calculates the tracking error of the actual pose and the desired pose, and feeds back to the future state prediction unit to correct the prediction result, controls the loop triggering of the perception-prediction-optimization-execution process, and ensures real-time performance.

[0123] A specific embodiment is provided, which selects a 7-DOF Franka Research 3 robot arm (redundancy 1), equipped with a Robotiq-2f-85 gripper, joint encoder resolution ≥ 16 bits, supporting joint position / velocity / torque feedback; the repeatability of the robot arm end is ±0.1mm (referring to Franka official technical manual “Franka Research 3 Technical Specification” section 5.2), which provides a hardware basis for high-precision trajectory tracking. Two RealSense D435i depth cameras are selected (installed at the end of the robot arm “eye on hand” and the base “eye on base” respectively), with a resolution of 1280×720 and a frame rate of 30fps; the above camera has a pose measurement error of ±0.05mm at a working distance of 1m (referring to Intel official product specification “Intel RealSense D435i Datasheet” section 3.2), which ensures the accuracy of visual pose measurement. A 6-axis force sensor is selected (accuracy ≤ 0.1N, referring to Robotiq official manual “Robotiq Force Torque Sensor Specification” section 4.1), which is used to obtain contact force feedback. Intel Core i7-13700K processor and NVIDIA RTX 4090 GPU are used to ensure real-time calculation of multi-modal feature fusion and quadratic programming.

[0124] Based on Ubuntu 22.04 LTS, ROS 2 Humble, and Franka Control Interface (FCI) to control the robot arm. The Transformer encoder is implemented based on PyTorch, the visual encoder uses the pre-trained MAE-ViT-Base, and the Perceiver Resampler compresses the visual features to 32 dimensions. The visual prediction model uses a U-Net structure to predict future images, and the Jacobian estimation is based on a comprehensive motion transformation matrix library (custom implementation, supporting 6-DOF pose derivative calculation). The qpOASES solver is used to solve the multi-objective optimization in layers, and the priority is realized through weight coefficients (for example, the weight of the first objective function is 1 , the weight of the second objective function is 10³, and the weight of the third objective is 10).

[0125] The parameter setting is the prediction time domain H=5 (corresponding to a 5ms future window, adapting to a 1kHz control frequency). The history window m=7 (using the last 7ms of history data to improve prediction stability). The visual-pose weight coefficient α=0.5, and the safety distance d_safe=5mm.

[0126] Under scenario 1: high-precision trajectory tracking (task: end effector moves along a circular trajectory with a radius of 25 cm, period 2 s). Through multi-modal fusion compensation and future Jacobian optimization, the average pose error is expected to be reduced to 0.25 mm (2.5 times the base accuracy), and the joint speed fluctuation can be controlled within ±0.08 rad / s (satisfying the precision scene specification). The core reason for the reduction of trajectory tracking error is that the future Jacobian estimation avoids the local optimization deviation caused by the time-varying Jacobian of the traditional method, and the real-time correction combined with visual feedback makes the error peak of the end effector in the high-speed motion section (such as the vertex of the circular trajectory) expected to be reduced from 1.2 mm to 0.35 mm.

[0127] Under scenario 2: joint limit avoidance (task: the robot arm moves from the initial pose to the target pose, and there are two joint positions close to the limit position in the path). Through joint trajectory optimization in the prediction time domain, the distribution of redundant degrees of freedom can be adjusted 200 ms in advance, so that the maximum position occupancy of joint 2 and joint 5 is controlled within 85%, and the trajectory can be completed without deceleration, and the time deviation from the planning time can be controlled within ≤5%.

[0128] Under scenario 3: dynamic obstacle avoidance (task: a 5cmx5cmx5cm obstacle suddenly appears in the robot arm motion path). Through the visual prediction model, the obstacle trajectory can be predicted 150 ms in advance, combined with collision constraint optimization, the obstacle avoidance response delay is reduced to 80 ms, the minimum obstacle avoidance distance can be stabilized at 8 mm (higher than the 5 mm safety distance), and the end effector trajectory offset can be controlled within ≤0.5 mm (without affecting the task accuracy).

[0129] Under scenario 4: load change robustness (task: the end load suddenly changes from 0 kg to 2 kg, and the trajectory tracking stability is tested). Through the future torque prediction branch, the load change can be perceived in advance, and the torque constraint is prioritized in the quadratic programming optimization, the error peak is controlled within 0.4 mm, and the convergence time is shortened to 100 ms.

[0130] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant national and regional laws, regulations and standards, and provide corresponding operation portal for user to choose authorization or refusal.

[0131] According to the embodiment of the present application, a robot arm control device embodiment is provided. It should be noted that the device can be used to execute the robot arm control method described above.

[0132] Figure 6This is a structural block diagram of a robotic arm control device according to one embodiment of this application, such as... Figure 6 As shown, taking a robotic arm control device 600 as an example, the device includes: an acquisition module 601, used to acquire task instructions and multimodal perception data of the robotic arm, wherein the task instructions are used to guide the operation behavior of the robotic arm, and the multimodal perception data are used to represent the state characteristics and environmental characteristics of the robotic arm; a prediction module 602, used to predict the state based on the task instructions and multimodal perception data to obtain future state characteristics, wherein the future state characteristics are used to guide the motion trajectory of the robotic arm; a determination module 603, used to determine a target joint configuration sequence based on the future state characteristics, wherein the target joint configuration sequence is used to represent the joint position information of the robotic arm; and a control module 604, used to control the robotic arm based on the target joint configuration sequence.

[0133] Embodiments of this application also provide a robotic arm, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.

[0134] Embodiments of this application also provide a computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.

[0135] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.

[0136] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0137] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.

[0138] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0139] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0140] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0141] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A robotic arm control method, characterized in that, The method includes: The task instructions and multimodal perception data of the robotic arm are acquired, wherein the task instructions are used to guide the operation behavior of the robotic arm, and the multimodal perception data are used to represent the state characteristics and environmental characteristics of the robotic arm; Based on the task instructions and the multimodal perception data, state prediction is performed to obtain future state features, wherein the future state features are used to guide the motion trajectory of the robotic arm. A target joint configuration sequence is determined based on the future state characteristics, wherein the target joint configuration sequence is used to represent the joint position information of the robotic arm; The robotic arm is controlled based on the target joint configuration sequence.

2. The method according to claim 1, characterized in that, The multimodal perception data includes at least: visual data, kinematic data, and force data. The visual data is used to represent the visual features of the environment around the robotic arm and the position of the target object. The kinematic data is used to represent the joint position and velocity state of the robotic arm. The force data is used to represent the contact force of the end effector of the robotic arm.

3. The method according to claim 2, characterized in that, The state prediction based on the task instructions and the multimodal perception data, to obtain future state features, includes: Based on the first feature vector, the second feature vector, and the third feature vector, the task instruction and the multimodal perception data are fused to obtain fused features. The first feature vector includes a visual feature vector, a joint state feature vector, a force feature vector, and a language feature vector. The second feature vector is used to represent multimodal prediction information and constraints at future moments. The third feature vector is used to guide the joint configuration sequence of the robotic arm to meet the constraints. Based on the fused features, state prediction is performed to obtain the future state features.

4. The method according to claim 3, characterized in that, The method further includes: The visual data is processed using a visual encoder to extract features, resulting in the visual feature vector. The kinematic encoder is used to extract features from the kinematic data to obtain the joint state feature vector; The force feature vector is obtained by extracting features from the force data using a force encoder. The language feature vector is obtained by extracting features from the task instructions using an instruction parser.

5. The method according to claim 3, characterized in that, The process of predicting the future state based on the fused features to obtain the future state features includes: Based on the fused features and historical data, state prediction is performed to obtain future visual features and future pose features; Based on the fused features and the current joint configuration sequence, state prediction is performed to obtain future matrix features and future joint torque features.

6. The method according to claim 1, characterized in that, The step of determining the target joint configuration sequence based on the future state features includes: The objective function and objective constraints are determined based on the future state characteristics, wherein the objective constraints are used to limit the range of motion of the robotic arm's joints; The target joint configuration sequence is determined based on the objective function and the objective constraints.

7. The method according to claim 6, characterized in that, The objective function includes a first function, a second function, and a third function. The first function is used to reduce the trajectory tracking error of the robotic arm in the task space. The second function is used to reduce the joint vibration of the robotic arm. The third function improves the operability of the robotic arm. The first function has a higher priority than the second function, and the second function has a higher priority than the third function.

8. The method according to any one of claims 1-7, characterized in that, The control of the robotic arm based on the target joint configuration sequence includes: The robotic arm is controlled based on the target joint configuration sequence, and the actuator pose of the robotic arm is obtained. Based on the target joint configuration sequence and the actuator pose, the target trajectory tracking error is determined; The future state features are updated based on the target trajectory tracking error; A new joint configuration sequence is determined based on the updated future state features; The robotic arm is controlled based on the new joint configuration sequence.

9. A robotic arm, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program executes the robotic arm control method according to any one of claims 1 to 8 when it runs.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the robotic arm control method according to any one of claims 1 to 8 when run on a computer or processor.

Citation Information

Cited By

  • Robot control method and device, electronic equipment and storage medium

    CN121832577A

  • Robot control method and device, electronic device, and storage medium

    CN121832577B