Method for mechanical arm real-time control technology combined with LLM
Through LLM parsing natural language instructions and combining hierarchical control and cache mechanisms, the problems of low human-computer interaction efficiency and programming complexity in robotic arm control are solved, and efficient and flexible robotic arm control is achieved, suitable for grabbing and placing fragile items.
Patent Information
- Application Number
- CN202510695070.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-08-22
AI Technical Summary
The lack of comprehensive solutions in the prior art that efficiently combines natural language instruction analysis, real-time motion planning and safety control has resulted in the need of cumbersome programming and debugging processes, making it difficult to achieve efficient and flexible human-computer interaction.
LLM is used to analyze natural language instructions to generate structured parameters, combine hierarchical control strategies, real-time trajectory planning and adaptive PID control, optimize the processing process through the cache mechanism, introduce a security pre-check layer for inverse kinematic accessibility and collision detection, and integrate position and force control parameters to achieve flexible contact and interaction.
It significantly improves the efficiency and convenience of human-computer interaction, reduces the technical threshold, improves the flexibility and adaptability of the robotic arm in dynamic environments, reduces the failure rate of motion planning, ensures efficient operation of the system, and is suitable for the grabbing and placement of fragile items.
Smart Images

Figure CN120516686A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robotic arms, and in particular relates to a method for real-time control technology of a robotic arm combined with LLM. Background Art
[0002] With the rapid development of artificial intelligence and robotics technology, robotic arms have been widely used in industrial production, logistics, services and other fields. Traditional robotic arm control technology mainly relies on preset programs and precise sensor feedback, which usually requires complex programming and debugging processes and has high requirements on the technical level of operators.
[0003] In recent years, natural language processing (NLP) technology has made significant progress, especially with the application of large-scale language models (LLMs), which enable machines to better understand and process human natural language commands. However, in practical applications, directly applying natural language commands to the real-time control of robotic arms still faces multiple challenges, including the ambiguity of natural language, the complexity of tasks, and the real-time requirements. As a result, robotic arm control often requires tedious programming and debugging, making efficient and flexible human-computer interaction difficult to achieve.
[0004] That is, the existing technology lacks a comprehensive solution that can efficiently combine natural language command parsing, real-time motion planning and safety control. Summary of the Invention
[0005] In one aspect, the present invention provides a method for real-time control of a robotic arm in combination with LLM, comprising the following steps:
[0006] S01 users give instructions through natural language;
[0007] The S02 user interaction layer receives natural language commands, and the LLM analyzes the intention and goal of the natural language commands and generates structured parameters;
[0008] S03: the structured parameters enter the code generation layer, and the structured parameters are converted into lightweight codes that can be executed by the real-time control layer;
[0009] The S04 trajectory planning layer simultaneously receives the motion constraints in the lightweight code and the real-time point cloud data of the environment perception layer to generate a smooth, collision-free joint space trajectory;
[0010] S05 receives the joint target index generated by the trajectory planning layer, and the RTOS (real-time control layer) calculates the control amount required by the motor through adaptive PID;
[0011] S06 PWM module converts the control quantity into a duty cycle signal to drive the servo motor to operate;
[0012] S07: The servo motor drives the robotic arm to move;
[0013] S08 integrates position and force control parameters in contact tasks and adjusts the target position in real time based on feedback;
[0014] The S09 feedback layer continuously monitors the robot arm status with a high sampling period of 0.5ms (2kHz) and immediately freezes execution when an abnormality is detected.
[0015] Preferably, the S02 user interaction layer receives natural language instructions, and the LLM analyzes the intention and goal of the natural language instructions to generate structured parameters, including:
[0016] S201: The user interaction layer receives a natural voice command and searches the cache for the same command.
[0017] S202 cache hit, directly obtain the corresponding structured parameters and lightweight code template from the cache, skip the LLM parsing step, and directly enter step S03;
[0018] If S203 cache misses, the instruction is passed to LLM for parsing. LLM parses the intent and goal of the natural language instruction and generates structured parameters.
[0019] S204 stores the generated structured parameters and lightweight code templates in a cache for subsequent use.
[0020] Preferably, the structured parameters include: task type, subtype, dynamic obstacle avoidance label, target object attributes, target position coordinates, and dynamic obstacle avoidance constraints.
[0021] Preferably, the structured parameters in S03 enter the code generation layer, and the structured parameters are converted into lightweight codes executable by the real-time control layer, including:
[0022] S301 code generation layer receives structured parameters from the user interaction layer;
[0023] The S302 safety pre-check layer performs inverse kinematics accessibility and collision detection verification on structural parameters;
[0024] If the verification in S303 succeeds, the parameters enter the code generation layer, which converts the structured parameters into lightweight code that can be executed by the real-time control layer;
[0025] S304 Verification failed, requesting the user to re-enter the instruction.
[0026] Preferably, the lightweight code includes: a motion instruction sequence for the robot arm joints, an obstacle avoidance detection and adjustment instruction sequence, and a control instruction sequence for grasping and placing actions.
[0027] Preferably, the S08 fuses the position and force control parameters in the contact task, adjusts the target position in real time according to the feedback, and fuses the position and force control parameters through an admittance control algorithm.
[0028] Another aspect of the present invention provides a real-time control system for a robotic arm in combination with LLM, characterized by comprising:
[0029] The user interaction layer is used to receive instructions from users in natural language and perform preliminary processing;
[0030] LLM (language model), used to parse the intent and goals of natural language instructions and generate structured parameters;
[0031] Cache layer, used to store parsing results of high-frequency instructions and lightweight code templates;
[0032] A code generation layer, configured to convert the structured parameters into lightweight codes executable by the real-time control layer;
[0033] A safety pre-check layer, used for performing inverse kinematics reachability and collision detection verification on the structural parameters;
[0034] Trajectory planning layer, used to generate smooth, collision-free joint space trajectories;
[0035] A real-time control layer (RTOS) receives the joint target position / speed generated by the trajectory planning layer and calculates the control amount required by the motor through adaptive PID;
[0036] A PWM module, configured to convert the control variable into a duty cycle signal;
[0037] The servo motor drives the robot arm to move according to the signal from the PWM module;
[0038] Feedback layer, used to monitor the status of the robotic arm;
[0039] The environmental perception layer is used to provide environmental information such as real-time point cloud data.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] Through the natural language processing capabilities of LLM, the present invention can directly receive natural language instructions from users and quickly parse and generate structured parameters, significantly improving the efficiency and convenience of human-computer interaction, reducing the dependence on professional programmers, and lowering the technical threshold and cost of robotic arm applications. Users do not need to have complex programming knowledge to control the robotic arm to complete various tasks through simple natural language instructions.
[0042] The present invention adopts a hierarchical control strategy, combined with real-time trajectory planning and adaptive PID control, which can quickly respond to user commands and dynamically adjust the motion trajectory and control parameters of the robot arm according to real-time environmental data. It shows high flexibility and adaptability in obstacle avoidance and task adjustment in dynamic environments.
[0043] The present invention introduces a safety pre-check layer to perform inverse kinematics reachability and collision detection verification on structural parameters, effectively reducing the motion planning failure rate and avoiding damage or mission failure of the robotic arm due to unreachable target positions or collision risks.
[0044] The present invention optimizes the processing flow of high-frequency instructions through a cache mechanism, reduces repeated calculations of LLM, and improves the response speed and efficiency of the system. The design of a multi-level cache structure can flexibly adjust the cache strategy according to different usage scenarios to ensure efficient operation of the system.
[0045] In the contact task of the present invention, the position and force control parameters are integrated through the admittance control algorithm to achieve smooth contact and interaction. The robotic arm can dynamically adjust the target position and speed according to the real-time sensed force control parameters, avoiding damage to the object caused by hard contact. It is particularly suitable for grasping and placing fragile objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a flow chart of the method of the present invention;
[0047] Figure 2 This is a relational operation diagram of the control system in the present invention.
[0048] Numbers in the figure: 1-user interaction layer; 2-LLM (language model), 3-cache layer, 4-code generation layer, 5-safety pre-check layer, 6-trajectory planning layer, 7-real-time control layer (RTOS), 8-PWM module, 9-servo motor, 10-feedback layer, 11-environmental perception layer. DETAILED DESCRIPTION
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0050] Example 1
[0051] like Figures 1 to 2 A method for real-time control of a robotic arm combined with LLM is shown, comprising the following steps:
[0052] S01 users give instructions through natural language;
[0053] S02 User Interaction Layer 1 receives natural speech commands, and LLM analyzes the intention and goal of the natural language commands and generates structured parameters;
[0054] S03 The structured parameters enter the code generation layer 4, which converts the structured parameters into lightweight codes that can be executed by the real-time control layer 7;
[0055] S04 trajectory planning layer 6 simultaneously receives the motion constraints in the lightweight code and the real-time point cloud data of the environment perception layer 11 to generate a smooth, collision-free joint space trajectory;
[0056] S05 receives the joint target position / speed generated by the trajectory planning layer 6, and the RTOS (real-time control layer 7) calculates the control amount required by the motor through adaptive PID;
[0057] S06 PWM module 8 converts the control quantity into a duty cycle signal to drive the servo motor 9 to operate;
[0058] S07 servo motor 9 drives the robot arm to move;
[0059] S08 integrates position and force control parameters (admittance) in contact tasks, and adjusts the target position in real time based on feedback (using a hierarchical strategy).
[0060] The S09 feedback layer 10 continuously monitors the robot arm status at a high sampling period of 0.5ms (2kHz) and immediately freezes execution when an abnormality is detected.
[0061] For ease of understanding, this embodiment defines the user's natural language instruction in step S01 as follows: "Place the white express box at the top of the car, avoiding the staff below." Furthermore, in a specific embodiment, in step S02, the user interaction layer 1 receives the natural language instruction, and the LLM analyzes the intent and goal of the natural language instruction to generate structured parameters, specifically including:
[0062] It should be noted that LLM refers to an artificial intelligence model that is trained based on massive text data and has powerful natural language processing capabilities. It is usually composed of billions or even trillions of parameters. It captures complex language patterns through deep learning and can handle multiple tasks such as question answering, writing, and code generation without the need to design a separate model for each task.
[0063] After receiving the natural language instruction "Put the white express box on the top of the car, and be careful to avoid the staff below", the user interaction layer 1 converts the instruction into text through automatic speech recognition, filters out irrelevant words such as "put", "of", and "pay attention", and outputs standardized text: Move the white express box to the top of the car to avoid obstacles.
[0064] LLM parses the intent and goals of natural language instructions, specifically: using a pre-trained classifier to determine the task type as "grasp-and-place" with a subtype of "high-altitude placement," while adding the task label "dynamic obstacle avoidance"; matching the object attributes of "white color and square shape" through the visual system to output the target object; combining the known 3D model of the car, calculating the center coordinates of the top safety area, such as [x=1.2m, y=0m, z=2.5m], and outputting the target position; real-time detection of the point cloud clustering results of the lower half of the car, generating dynamic obstacle avoidance parameters, maintaining the distance between the robotic arm and the staff, avoiding collisions with the staff, and extracting constraints; combining the above data to generate structured parameters.
[0065] In a specific embodiment, step S03, the structured parameters enter the code generation layer 4, and the structured parameters are converted into lightweight codes executable by the real-time control layer 7, specifically including:
[0066] The code generation layer 4 receives the structured parameters passed from the user interaction layer 1. These parameters include the task type (grasp-place), subtype (high-altitude placement), dynamic obstacle avoidance label, target object attributes (white color, square shape), target position coordinates ([x=1.2m, y=0m, z=2.5m]), and dynamic obstacle avoidance constraints extracted in step 02. The above structured parameters are then converted into lightweight code executable by the real-time control layer 7 through template matching.
[0067] Based on the task type "grasp-and-place" and the subtype "high-altitude placement", the code generation layer 4 selects a matching code template from the preset code template library. The code template contains the basic code framework of the robot arm's grasping action, moving to the target position, and placing action; then the target object attributes (white color, square shape) are embedded in the code of the grasping action, so that the system can control the code of the robot arm's end effector based on the code fragments recognized by the visual system, so that the robot arm can accurately grasp the target object; the target position coordinates ([x=1.2m, y=0m, z=2.5m]) are embedded in the code of moving to the target position, and the angles of each joint of the robot arm are calculated through the inverse kinematics algorithm to generate the corresponding joint motion instruction code; the constraints of dynamic obstacle avoidance are embedded in the motion control code, and so on.
[0068] The generated code is then syntax-checked and optimized, redundant code statements are removed, the loop structure is optimized, and the code is converted into a lightweight code format to ensure the correctness and efficiency of the code, so as to meet the requirements of the real-time control layer 7 for code execution efficiency and resource usage; the generated lightweight code includes but is not limited to: motion instruction sequences of the robotic arm joints, obstacle avoidance detection and adjustment instruction sequences, control instruction sequences for grasping and placing actions, etc. These codes will be passed to the trajectory planning layer 6 to provide input for subsequent robotic arm motion trajectory planning, etc.
[0069] In a specific embodiment, the S04 trajectory planning layer 6 simultaneously receives the motion constraints in the lightweight code and the real-time point cloud data of the environment perception layer 11 to generate a smooth, collision-free joint space trajectory, specifically including:
[0070] The trajectory planning layer 6 synchronously receives the subsequent required input data through a double buffering mechanism. One buffer is used to receive the static constraints from the lightweight code and the dynamic environment data from the environmental perception layer 11, and the other buffer is used for the actual trajectory planning calculation. The motion parameters are extracted from the lightweight code in step S03 to generate static constraints, including: the end effector target position of [x = 1.2m, y = 0m, z = 2.5m], the maximum joint acceleration limit of the manipulator, etc. The real-time point cloud data and dynamic obstacle motion prediction transmitted by the environmental perception layer 11 are obtained as dynamic environment data. The point cloud data can provide three-dimensional spatial information of the environment around the manipulator, including static obstacles (such as other objects in the car) and dynamic obstacles (such as the real-time position and motion trajectory of the staff). The dynamic obstacle motion prediction is based on the historical information and current state of the point cloud data, and estimates the position and velocity of dynamic obstacles in the future through a motion prediction algorithm.
[0071] The trajectory planning layer 6 uses the target pose of the end effector as the endpoint, combines the initial joint angles and velocities of the robot arm, and generates a smooth trajectory in the joint space based on a polynomial time parameterization method. The trajectory planning layer 6 then uses a real-time optimization algorithm to further optimize the generated trajectory, minimizing the movement time or energy consumption of the robot arm while satisfying all constraints.
[0072] The resulting trajectory will be a sequence of joint angles with time as the independent variable. Each time point corresponds to the target angle of each joint of the robotic arm. The above trajectory data will be passed to the real-time operating system (RTOS) to control the actual movement of the robotic arm.
[0073] In a specific embodiment, the target joint position / speed generated by the trajectory planning layer 6 is received, and the RTOS (real-time control layer 7) calculates the control quantity required by the motor through adaptive PID, specifically including:
[0074] The RTOS receives the joint target position and velocity generated by the trajectory planning layer 6 and adopts a hierarchical adaptive PID control strategy. It first calculates the expected joint angle based on the target trajectory, then adjusts the joint movement speed to ensure smooth tracking, and finally directly controls the torque of the servo motor 9 to achieve high dynamic response.
[0075] The RTOS receives the joint target positions and velocities generated by trajectory planning layer 6 and adopts an adaptive PID control strategy, combined with feedforward compensation and dynamic parameter adjustment, to ensure that the robot arm accurately tracks the trajectory. The control cycle is 0.5ms (2kHz). It obtains encoder feedback in real time through the EtherCAT bus, calculates position and velocity errors, and outputs optimized PWM control signals.
[0076] First, the system receives the desired position and velocity for each joint of the robot arm at the current moment from the trajectory planning layer 6. Simultaneously, it reads data from high-precision encoders installed on each joint in real time via the high-speed EtherCAT bus to obtain the actual position of the robot arm. It then quickly calculates the error between the target and actual positions and adjusts the control signals using an adaptive PID control algorithm. During this process, the system pre-calculates the inertial force generated by the robot arm during operation and compensates for the effects of gravity on the robot arm in different postures in real time to improve response speed. At the hardware level, an FPGA is used to accelerate PWM generation and encoder decoding, and critical computation tasks are bound to a dedicated CPU core to avoid scheduling delays. Safety mechanisms include limit protection, dynamic braking, and heartbeat detection. Limit protection prevents the robot arm from exceeding safety limits during operation, while dynamic braking ensures rapid shutdown in the event of an abnormality. The system continuously monitors its operating status and addresses any issues immediately. Ultimately, the system achieves a repeatable positioning accuracy of ±0.05°, a dynamic response time of less than 2ms, and can withstand sudden load fluctuations of ±15%, ensuring stable operation in complex environments.
[0077] In a specific embodiment, in steps S06 and S07, the PWM module 8 converts the control variable into a duty cycle signal to drive the servo motor 9 to operate, and the servo motor 9 drives the robot arm to move, specifically including:
[0078] The digital control quantity output by the RTOS is converted into a pulse width modulation (PWM) signal by the PWM module 8. The duty cycle is calculated based on the characteristics of the servo motor 9 and the movement requirements of the robot arm to ensure that the servo motor 9 can operate at an appropriate speed and torque; the PWM signal is amplified by the H-bridge drive circuit to drive the winding of the servo motor 9. After receiving the PWM signal, the servo motor 9 responds quickly and adjusts its speed and torque output according to the duty cycle of the signal. The encoder inside the motor feeds back position and speed information in real time, which is compared with the target value to form a closed-loop control.
[0079] The servo motor 9 actually reduces the speed and increases the torque through a harmonic reducer or planetary gear, thereby controlling the joints of the robotic arm. The control modes are divided into at least the following three types:
[0080] Position mode, used for high-precision positioning tasks, can directly follow the joint angle instructions sent by RTOS;
[0081] Speed mode, for continuous trajectory motion;
[0082] Torque mode, used for force control tasks, can achieve flexible interaction through current control.
[0083] During the operation of the PWM module 8 and the servo motor 9, once an abnormal situation such as overload, short circuit or overtemperature of the servo motor 9 is detected, the system will immediately trigger the protection mechanism, cut off the PWM signal, and stop the operation of the servo motor 9 to prevent equipment damage and safety accidents. At the same time, the fault information is fed back to the feedback layer 10 for subsequent fault diagnosis and processing.
[0084] In a specific embodiment, step S08 integrates position and force control parameters in the contact task and adjusts the target position in real time based on the feedback, specifically including:
[0085] During the grasping process, express boxes are mostly fragile objects that need to be handled with care. Therefore, in the contact task with the target object, the interaction between the robotic arm and the external environment needs to consider both position and force control, adjust the gripping force of the robotic arm, and avoid damage to the object.
[0086] Furthermore, through the six-dimensional force / torque sensor installed at the end of the robotic arm, the contact force and torque between the robotic arm and the target object or environment are sensed in real time, and these force control parameters are integrated with the position parameters (such as joint angles or the position of the end effector) through the admittance control algorithm. The admittance control algorithm defines the behavior of the robotic arm as an "admittance system", which converts the input force and torque into corresponding displacement and velocity. Through the real-time perception of force control parameters, the target position and velocity of the robotic arm are dynamically adjusted to achieve smooth contact and interaction. Then, when grasping fragile objects, the admittance control robotic arm automatically adjusts the grasping force according to the size of the contact force to avoid damage to the object.
[0087] In a specific embodiment, step S09, the feedback layer 10 continuously monitors the state of the robot arm at a high sampling period of 0.5ms (2kHz), and immediately freezes the execution when an abnormality is detected, specifically including:
[0088] In general, the present invention mainly adopts a hierarchical control strategy to achieve highly precise control of the robotic arm, and divides the control layer into a task layer and an execution layer. Among them, the task layer belongs to the high-level control and is responsible for the planning and decision-making of the overall task. The structured parameters generated by the LLM and the output of the trajectory planning layer 6 determine the motion target and contact strategy of the robotic arm; the execution layer belongs to the low-level control and is responsible for executing the instructions of the high-level control in real time and making rapid adjustments based on feedback. It mainly adjusts the control amount of the motor in real time through the adaptive PID controller and the admittance control algorithm to ensure that the movement of the robotic arm meets the task requirements.
[0089] During the contact task, the feedback layer 10 continuously monitors the status of the robot arm at a high sampling period of 0.5ms (2kHz), including information such as position, speed, force and torque. This feedback information is transmitted to the control algorithm in real time for dynamic adjustment of the target position and force control parameters. When it is detected that the contact force exceeds the preset threshold, the admittance control algorithm will automatically reduce the contact force of the robot arm. When the position deviation is large, the adaptive PID controller will adjust the control amount of the motor to quickly correct the position deviation.
[0090] Furthermore, the feedback layer 10 also monitors possible abnormal situations in real time, such as sudden changes in force control parameters, excessive position deviation, or sensor failure. Once an abnormally high contact force or other abnormal situation is detected, the system will immediately take measures, such as freezing execution, issuing an alarm, or switching to a safe mode, to prevent damage to the robotic arm or the target object.
[0091] When placing the express box, the admittance control algorithm will adjust the motion trajectory and strength of the robotic arm in real time according to the changes in contact force during the assembly process. When the robotic arm contacts the express box, it will use the feedback of force control parameters to gently clamp, move and place the express box, avoiding damage to the express box surface or even damage to the internal goods due to hard contact.
[0092] Example 2
[0093] Furthermore, in step S03, the structured parameters enter the code generation layer 4, that is, before the code is generated, the system is also provided with a security pre-check layer 5 to perform double verification on the parameters. Before the structured parameters are passed from the user interaction layer 1 to the code generation layer 4, they will first pass through the security pre-check layer 5 for reachability verification and collision detection verification to ensure that the robotic arm can perform tasks safely and reliably.
[0094] In the safety pre-check layer 5, inverse kinematics can determine whether the target position is within the working space of the robot arm based on the structural parameters and joint limitations of the robot arm. If the target position exceeds the reachable range of the robot arm, it is judged as unreachable. At the same time, the calculated joint angle needs to be within the physical limitation range of the robot arm. When the calculated joint angle exceeds the physical limitation range of the robot arm, it may cause damage to the robot arm or malfunction. Even if the target position is reachable, it is necessary to verify whether the path from the current position to the target position is feasible based on the kinematic model of the robot arm and the environmental constraints to ensure that the robot arm can reach the target position smoothly and without collision.
[0095] As mentioned above, if the inverse kinematics reachability check fails, the safety pre-check layer 5 will feedback an error message to the user interaction layer 1, prompting the user that the target position is unreachable or the joint angle is unreasonable, and requesting the user to re-enter the target position or adjust the task parameters.
[0096] In the safety pre-check layer 5, a collision detection check is also provided. The collision detection will construct a three-dimensional model of the environment around the robot arm in advance through the real-time point cloud data provided by the environmental perception layer 11, and use the collision detection algorithm to detect in real time whether there is a collision risk in the movement path of the robot arm. The above collision detection will be completed in a short time to ensure that the robot arm can adjust the movement trajectory in time. If a collision risk is detected, the safety pre-check layer 5 will immediately issue a warning to the user interaction layer 1 and request the user to re-plan the task path or adjust the target position. The user can also control the system to automatically adjust the movement trajectory of the robot arm.
[0097] That is, if the target position is beyond the reach of the robotic arm, the system will prompt the user "The target position is beyond the working range of the robotic arm, please re-enter the target position." If a collision risk is detected, the system will prompt the user "There is a collision risk in the robotic arm's motion path, please adjust the target position or path" to prompt the user to re-enter the task parameters or adjust the environment layout.
[0098] In actual application, the user inputs a natural language instruction: put the white express box on the carriage, and the system detects:
[0099] 1. The top of the carriage is a blind spot (inaccessible);
[0100] 2. Suggested correction: The recommended placement location is changed to the upper part of the carriage;
[0101] 3. Visually display the comparison diagram before and after the modification.
[0102] In summary, the security pre-check layer 5 in this embodiment can reduce the motion planning failure rate by 83%, shorten the average exception processing time to 1.2 seconds, and significantly improve the efficiency of human-computer interaction.
[0103] Example 3
[0104] Furthermore, the present invention also combines the LLM's real-time control technology for the robotic arm by introducing a cache mechanism to optimize the processing flow of high-frequency instructions and improve the response speed and efficiency of the system.
[0105] Specifically, in the actual application scenarios of the robotic arm, users may frequently issue certain similar instructions, such as repetitive grasping, handling or assembly tasks. The parsing results of such high-frequency instructions usually have a certain degree of repetitiveness. The system can cache the parsing results of these high-frequency instructions so that the user interaction layer 1 can directly call the pre-stored code template when receiving similar instructions, without the need to parse through the LLM every time, thereby significantly reducing the repeated calculations of the LLM and improving the real-time performance and response speed of the system.
[0106] In this embodiment, the cache mechanism is designed with a multi-level cache structure to store the parsing results of high-frequency instructions. The first-level cache is used to store the parsing results of the most recent 10 instructions. The second-level cache uses an SSD with a capacity of 1,000+ high-frequency instruction templates. The third-level cache is a cloud-based instruction library that supports cross-device synchronization. When the user interaction layer 1 receives a new instruction, it first searches the cache to see if the same instruction exists. If so, it is a cache hit. The pre-stored code template is directly called, skipping the LLM parsing step and entering the code generation layer 4. If the same instruction does not exist in the cache, the instruction is passed to the LLM for parsing. After the parsing is complete, the new parsing result is stored in the cache for subsequent use.
[0107] In practical applications, caches need to be updated regularly to ensure the accuracy and timeliness of their content. The update mechanism can be based on the following strategies:
[0108] Timestamp strategy: Set a timestamp for each cache item and regularly clean up expired cache items.
[0109] Frequency of use strategy: Update the cached items based on their frequency of use. If a cached item has not been used for a long time, it will be removed from the cache.
[0110] Dynamic adjustment strategy: Dynamically adjust the cache size according to the system load. When the system load is high, frequently used cache items are retained first; when the system load is low, the cache capacity can be appropriately expanded.
[0111] Example 4
[0112] Furthermore, a method for real-time control technology of a robotic arm combined with LLM, in step S02, the user interaction layer 1 receives a natural voice command, and the LLM analyzes the intention and goal of the natural language command and generates structured parameters, specifically including:
[0113] S201: User interaction layer 1 receives a natural voice command and searches the cache for the same command.
[0114] S202 cache hit, directly obtain the corresponding structured parameters and lightweight code template from the cache, skip the LLM parsing step, and directly enter step S03;
[0115] If S203 cache misses, the instruction is passed to LLM for parsing. LLM parses the intent and goal of the natural language instruction and generates structured parameters.
[0116] S204 stores the generated structured parameters and lightweight code templates in a cache for subsequent use.
[0117] Step S03: The structured parameters enter the code generation layer 4, and the structured parameters are converted into lightweight codes executable by the real-time control layer 7, which specifically includes:
[0118] S301 Code generation layer 4 receives structured parameters from user interaction layer 1;
[0119] S302 safety pre-check layer 5 performs inverse kinematics accessibility and collision detection verification on structural parameters;
[0120] If the verification in step S303 succeeds, the parameters enter the code generation layer 4, which converts the structured parameters into lightweight code that can be executed by the real-time control layer 7;
[0121] S304 Verification failed, requesting the user to re-enter the instruction.
[0122] Example 5
[0123] like Figure 2 The system used in the method of real-time control technology of a robotic arm combined with LLM includes:
[0124] User interaction layer 1: receives the user's natural language instructions, then queries cache layer 3 to determine whether there is a cache hit. If there is a cache hit, the structured parameters and lightweight code templates in cache layer 3 are directly called to enter code generation layer 4. If there is a cache miss, the instruction is passed to the LLM parsing layer.
[0125] Safety pre-check layer 5: is used to receive the structural parameters transmitted by the code generation layer 4 to perform inverse kinematics reachability and collision detection verification, and feed back the verification results to the code generation layer 4.
[0126] Code generation layer 4: is used to receive structured parameters passed by the user interaction layer 1 or the LLM parsing layer, pass the structured parameters to the security pre-check layer 5 for verification, receive the verification results of the security pre-check layer 5, generate lightweight code when the verification is successful, and pass it to the trajectory planning layer 6; if the verification fails, request the user to re-enter the command.
[0127] LLM parsing layer: used to receive instructions passed by the user interaction layer 1, parse the instruction intent and target, generate structured parameters, store the structured parameters and lightweight code templates in the cache layer 3, and pass the structured parameters to the code generation layer 4.
[0128] Cache layer 3: used to store the parsing results and lightweight code templates of high-frequency instructions, receive query requests from user interaction layer 1, return cached results, receive storage requests from the LLM parsing layer, and update cache content.
[0129] Trajectory planning layer 6: It is used to receive the motion constraints in the lightweight code passed by the code generation layer 4, receive the real-time point cloud data provided by the environment perception layer 11, generate smooth, collision-free joint space trajectories, and pass the trajectory planning results to the real-time control layer 7 (RTOS).
[0130] Real-time control layer 7 (RTOS): used to receive the joint target position / speed generated by the trajectory planning layer 6, and calculate the control quantity required by the motor through adaptive PID, and then pass the control quantity to the PWM module 8.
[0131] PWM module 8: used to receive the control quantity transmitted by the real-time control layer 7 (RTOS), convert the control quantity into a duty cycle signal, and then transmit the duty cycle signal to the servo motor 9.
[0132] The servo motor 9 is used to receive the duty cycle signal transmitted by the PWM module 8, thereby driving the robot arm to move.
[0133] Feedback layer 10: is used to continuously monitor the state of the robotic arm with a high sampling period. When an abnormal situation is detected, it sends a freeze execution instruction to the real-time control layer 7 and feeds back the abnormal information to the user interaction layer 1.
[0134] Environmental perception layer 11: used to provide environmental information such as real-time point cloud data, and pass the data to the trajectory planning layer 6 and the safety pre-inspection layer 5.
[0135] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0136] The above description is only used to illustrate the technical solution of the present invention and is not intended to limit it. Other modifications or equivalent substitutions made to the technical solution of the present invention by ordinary technicians in this field should be included in the scope of the claims of the present invention as long as they do not depart from the spirit and scope of the technical solution of the present invention.
Claims
1. A method for real-time control of a robotic arm in combination with LLM, characterized in that: The following steps are involved: S01 users give instructions through natural language; The S02 user interaction layer receives natural language commands, and the LLM analyzes the intention and goal of the natural language commands and generates structured parameters; S03: the structured parameters enter the code generation layer, and the structured parameters are converted into lightweight codes that can be executed by the real-time control layer; The S04 trajectory planning layer simultaneously receives the motion constraints in the lightweight code and the real-time point cloud data of the environment perception layer to generate a smooth, collision-free joint space trajectory; S05 receives the joint target index generated by the trajectory planning layer, and the RTOS (real-time control layer) calculates the control amount required by the motor through adaptive PID; S06 PWM module converts the control quantity into a duty cycle signal to drive the servo motor to operate; S07: The servo motor drives the robotic arm to move; S08 integrates position and force control parameters in contact tasks and adjusts the target position in real time based on feedback; The S09 feedback layer continuously monitors the robot arm status with a high sampling period of 0.5ms (2kHz) and immediately freezes execution when an abnormality is detected.
2. The method of a real-time control technology of a robotic arm combined with LLM according to claim 1, characterized in that: The S02 user interaction layer receives natural language instructions, and the LLM analyzes the intention and goal of the natural language instructions to generate structured parameters, including: S201: The user interaction layer receives a natural voice command and searches the cache for the same command. S202 cache hit, directly obtain the corresponding structured parameters and lightweight code template from the cache, skip the LLM parsing step, and directly enter step S03; If S203 cache misses, the instruction is passed to LLM for parsing. LLM parses the intent and goal of the natural language instruction and generates structured parameters. S204 stores the generated structured parameters and lightweight code templates in a cache for subsequent use.
3. The method of a real-time control technology of a robotic arm combined with LLM according to claim 2, characterized in that: The structured parameters include: task type, subtype, dynamic obstacle avoidance label, target object attributes, target position coordinates and dynamic obstacle avoidance constraints.
4. The method of a real-time control technology of a robotic arm combined with LLM according to claim 1, characterized in that: The S03 structured parameters enter the code generation layer, and the structured parameters are converted into lightweight codes executable by the real-time control layer, including: S301 code generation layer receives structured parameters from the user interaction layer; The S302 safety pre-check layer performs inverse kinematics accessibility and collision detection verification on structural parameters; If the verification in S303 succeeds, the parameters enter the code generation layer, which converts the structured parameters into lightweight code that can be executed by the real-time control layer; S304 Verification failed, requesting the user to re-enter the instruction.
5. The method of a real-time control technology of a robotic arm combined with LLM according to claim 4, characterized in that: The lightweight code includes: a motion instruction sequence for the robot arm joints, an obstacle avoidance detection and adjustment instruction sequence, and a control instruction sequence for grasping and placing actions.
6. The method of a real-time control technology of a robotic arm combined with LLM according to claim 1, characterized in that: The S08 fuses the position and force control parameters in the contact task, adjusts the target position in real time according to the feedback, and fuses the position and force control parameters through the admittance control algorithm.
7. A real-time control system for a robotic arm combined with LLM, characterized in that: include: The user interaction layer is used to receive instructions from users in natural language and perform preliminary processing; LLM (language model), used to parse the intent and goals of natural language instructions and generate structured parameters; Cache layer, used to store parsing results of high-frequency instructions and lightweight code templates; A code generation layer, configured to convert the structured parameters into lightweight codes executable by the real-time control layer; A safety pre-check layer, used for performing inverse kinematics reachability and collision detection verification on the structural parameters; Trajectory planning layer, used to generate smooth, collision-free joint space trajectories; A real-time control layer (RTOS) receives the joint target position / speed generated by the trajectory planning layer and calculates the control amount required by the motor through adaptive PID; A PWM module, configured to convert the control variable into a duty cycle signal; The servo motor drives the robot arm to move according to the signal from the PWM module; Feedback layer, used to monitor the status of the robotic arm; The environmental perception layer is used to provide environmental information such as real-time point cloud data.
Citation Information
Cited By
Hand-eye cooperative robot control system and method based on dynamic operator arrangement
CN121348922A