A robot control method and system based on a visual language action model, a terminal, and a storage medium

By combining visual language action models with depth images and camera intrinsic parameter matrices for 3D reconstruction and geometric fusion, the problem of VLA models being unable to accurately predict obstacle volume when visual occlusion or data is sparse is solved, thus achieving strict safety control and efficient execution of robot actions.

CN121696998BActive Publication Date: 2026-05-08GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)
Filing Date
2026-02-24
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing VLA models cannot accurately predict obstacle volume when there is visual occlusion or sparse data, making it difficult to achieve strict safety control and efficient execution of robot actions.

Method used

Risk assessment is performed by acquiring visual images and task instructions from the robot. 3D reconstruction and geometric fusion are then carried out by combining depth images and camera intrinsic parameter matrices to generate the target safety geometric boundary. The robot's actions are then corrected based on the risk identification results to ensure safety.

Benefits of technology

It enables strict safety control and efficient execution of robot actions in environments with visual occlusion or sparse data, thereby improving the safety and efficiency of robot tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121696998B_ABST
    Figure CN121696998B_ABST
Patent Text Reader

Abstract

The application discloses a robot control method and system based on a visual language action model, a terminal and a storage medium, and the method comprises the following steps: acquiring a visual image and a task instruction of a robot, performing risk assessment in a pre-trained visual language model according to the visual image and the task instruction, and obtaining an environment risk identification result; acquiring a depth image and a camera internal parameter matrix; if the environment risk identification result is a high-risk scene, performing three-dimensional reconstruction and geometric fusion on a candidate obstacle set according to the depth image and the camera internal parameter matrix to obtain a target safety geometric boundary; generating a nominal action according to the visual image and the task instruction; if the environment risk identification result is a high-risk scene, correcting the nominal action according to the target safety geometric boundary to obtain a control instruction, and sending the control instruction to a robot controller for execution. The application assesses risks based on vision and instructions, and realizes efficient execution and control of robot actions by correcting robot actions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot intelligent control technology, and in particular to a robot control method, system, terminal, and computer-readable storage medium based on a visual language action model. Background Technology

[0002] With the rapid development of embodied intelligence technology, the Visual-Language-Action (VLA) model has become the mainstream paradigm in the field of robot control. Through end-to-end training, the VLA model can directly map unstructured visual images and natural language commands into low-level control actions of the robot, demonstrating powerful semantic understanding and generalization capabilities.

[0003] Current VLA models typically employ "black box" strategies, with outputs that are probabilistic rather than based on deterministic physical constraints. To address safety issues, the mainstream approach utilizes reinforcement learning (RL) to integrate safety constraints into the training process. However, existing VLA models primarily focus on task completion rates, neglecting stringent safety guarantees. RL methods often treat safety as a "soft objective" (i.e., reward / penalty) rather than a "hard constraint," failing to provide strict, deterministic safety boundaries during inference. When faced with scenarios outside the training data distribution, VLA models are highly susceptible to unstable jitter or collision trajectories. Retraining-based methods are computationally expensive and difficult to apply directly to pre-trained large-scale VLA models. Adjustments to safety strategies often require model retraining, lacking plug-and-play flexibility. While pure end-to-end models possess semantic understanding capabilities, they lack prior geometric knowledge of the physical world, leading to an inability to accurately predict the physical volume of obstacles in cases of visual occlusion or data sparsity.

[0004] Traditional methods typically employ VLA models and safety constraint methods based on reinforcement learning (RL). However, due to shortcomings in safety (soft constraints rather than hard constraints), high computational costs and lack of flexibility when adjusting strategies, and the lack of physical geometric priors leading to the inability to accurately predict obstacle volumes in cases of visual occlusion or sparse data, it is difficult to achieve strict safety control and efficient execution of robot actions, which has become an urgent problem to be solved.

[0005] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0006] The main objective of this invention is to provide a robot control method, system, terminal, and computer-readable storage medium based on a visual language action model. This aims to solve the problem in the prior art that it is difficult to achieve strict safety control and efficient execution of robot actions when visual occlusion or data sparseness makes it impossible to accurately estimate the volume of obstacles.

[0007] To achieve the above objectives, the present invention provides a robot control method based on a visual language action model, the robot control method based on a visual language action model comprising the following steps:

[0008] The robot acquires visual images and task instructions, and performs risk assessment in a pre-trained visual language model based on the visual images and task instructions to obtain environmental risk identification results.

[0009] Acquire depth images and camera intrinsic parameter matrices. If the environmental risk identification result is a high-risk scene, then perform three-dimensional reconstruction and geometric fusion on the candidate obstacle set based on the depth images and the camera intrinsic parameter matrix to obtain the target safety geometric boundary.

[0010] A nominal action is generated based on the visual image and the task instruction. If the environmental risk identification result is a high-risk scenario, the nominal action is modified according to the target safety geometric boundary to obtain a control instruction. If the environmental risk identification result is a low-risk scenario, the control instruction is determined based on the nominal action and sent to the robot controller for execution.

[0011] Optionally, the robot control method based on a visual language action model further includes, after acquiring the robot's visual images and task instructions, performing risk assessment in a pre-trained visual language model based on the visual images and task instructions to obtain environmental risk identification results, the method further includes:

[0012] The visual image and the task instruction are input into the pre-trained visual language model for prediction to obtain the probability distribution;

[0013] Based on the probability distribution, the prompt words of the environmental risk identification results are filtered to obtain a set of candidate obstacles.

[0014] Optionally, in the robot control method based on a visual language action model, if the environmental risk identification result indicates a high-risk scene, the acquisition of depth images and camera intrinsic parameter matrices involves performing 3D reconstruction and geometric fusion of the candidate obstacle set based on the depth images and the camera intrinsic parameter matrix to obtain the target safety geometric boundary. Specifically, this includes:

[0015] If the environmental risk identification result is a high-risk scenario, then the location of the candidate obstacle set is extracted to obtain multiple spatial positioning points;

[0016] Obtain the depth image and camera intrinsic matrix of the candidate obstacle set, and transform all the spatial positioning points according to the depth image and the camera intrinsic matrix to form a three-dimensional point cloud set;

[0017] The three-dimensional point cloud set is subjected to outlier removal and spatial pass-through filtering to obtain the observation point cloud;

[0018] Based on the observed point cloud, the candidate obstacle set is reconstructed in three dimensions and geometrically fused to obtain the target safety geometric boundary.

[0019] Optionally, the robot control method based on a visual language action model, wherein the step of performing 3D reconstruction and geometric fusion of the candidate obstacle set based on the observed point cloud to obtain the target safety geometric boundary specifically includes:

[0020] Obtain the standard geometric shape parameters of the pre-built semantic geometry common library;

[0021] Based on the observed point cloud and the standard geometric shape parameters, geometric constraints are applied to the candidate obstacle set to obtain the initial safe geometric boundary:

[0022] ;

[0023] ;

[0024] in, Represents the observation shape matrix, Indicates the center of the observed shape. Represents the set of observation point clouds. This represents points in the observed point cloud. Represents the determinant of a matrix. Indicates transpose;

[0025] Obtain the confidence factor, and perform weighted fusion of the initial safe geometric boundary based on the confidence factor and the observed point cloud to obtain the target safe geometric boundary.

[0026] Optionally, in the robot control method based on visual language action model, the target safety geometric boundary includes a target safety center and a safety ellipsoid shape matrix;

[0027] The step of obtaining the confidence factor, and then weighting and fusing the initial safe geometric boundary based on the confidence factor and the observed point cloud to obtain the target safe geometric boundary, specifically includes:

[0028] Obtain the confidence factor, the number of points in the observed point cloud, the observation center of the standard object, and the prior center:

[0029] The initial safety geometric boundary is registered based on the observation center and the prior center to obtain the target safety center;

[0030] The initial safety geometric boundary is weighted and fused according to the standard geometric shape parameters to obtain the safety ellipsoid shape matrix:

[0031] ;

[0032] ;

[0033] in, This represents the confidence factor. Indicates the prior center, Indicates the target security center. The matrix representing the shape of the safe ellipsoid. This represents the prior shape matrix.

[0034] Optionally, the robot control method based on a visual language action model, wherein generating a nominal action based on the visual image and the task instruction, and if the environmental risk identification result is a high-risk scenario, then correcting the nominal action according to the target safety geometric boundary to obtain a control instruction; if the environmental risk identification result is a low-risk scenario, then determining the control instruction based on the nominal action and sending the control instruction to the robot controller for execution, specifically includes:

[0035] The visual image and the task instruction are input into the visual language action model for solving, and a nominal action is generated.

[0036] If the environmental risk identification result is a high-risk scenario, a control obstacle constraint function is constructed based on the target safety geometric boundary, and the nominal action is modified based on the control obstacle constraint function to obtain a control command. If the environmental risk identification result is a low-risk scenario, the control command is determined based on the nominal action and sent to the robot controller for execution.

[0037] Optionally, the robot control method based on a visual language action model, wherein if the environmental risk identification result is a high-risk scenario, constructing a control obstacle constraint function based on the target safety geometric boundary, and modifying the nominal action based on the control obstacle constraint function to obtain control commands, specifically includes:

[0038] If the environmental risk identification result is a high-risk scenario, the current environmental data is obtained, and the safety ellipsoid shape matrix is ​​constrained based on the current environmental data to obtain multiple obstacle constraints;

[0039] Construct a control obstacle constraint function based on the target safety geometric boundary and multiple obstacle constraints:

[0040] ;

[0041] in, Indicates the current state of the robot. Indicates a safety margin. Represents the minimum distance function. Represents an ellipsoid. Indicates the first An obstacle, Represents the control obstacle constraint function;

[0042] The optimal safety control quantity is determined based on the control obstacle constraint function.

[0043] The nominal action is corrected based on the optimal safety control quantity to obtain the control command:

[0044] ;

[0045] in, This represents the optimal safety control quantity. Represents the optimization variable. Indicates a nominal action.

[0046] Furthermore, to achieve the above objectives, the present invention also provides a robot control system based on a visual language action model, wherein the robot control system based on the visual language action model includes:

[0047] The obstacle recognition module is used to acquire the robot's visual images and task instructions, and to perform risk assessment in a pre-trained visual language model based on the visual images and task instructions to obtain environmental risk recognition results.

[0048] The safety boundary constraint module is used to acquire depth images and camera intrinsic parameter matrices. If the environmental risk identification result is a high-risk scene, the module performs 3D reconstruction and geometric fusion on the candidate obstacle set based on the depth images and the camera intrinsic parameter matrix to obtain the target safety geometric boundary.

[0049] The robot execution module is used to generate nominal actions based on the visual image and the task instructions. If the environmental risk identification result is a high-risk scenario, the nominal actions are modified according to the target safety geometric boundary to obtain control instructions. If the environmental risk identification result is a low-risk scenario, the control instructions are determined based on the nominal actions and sent to the robot controller for execution.

[0050] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a robot control program based on a visual language action model, and when the robot control program based on the visual language action model is executed by a processor, it implements the steps of the robot control method based on the visual language action model as described above.

[0051] This invention acquires visual images and task commands from a robot. Based on these images and commands, a risk assessment is performed in a pre-trained visual language model to obtain an environmental risk identification result. Depth images and camera intrinsic parameter matrices are also acquired. If the environmental risk identification result indicates a high-risk scenario, a 3D reconstruction and geometric fusion of a set of candidate obstacles is performed based on the depth images and camera intrinsic parameter matrices to obtain a target safety geometric boundary. A nominal action is generated based on the visual images and task commands. If the environmental risk identification result indicates a high-risk scenario, the nominal action is modified based on the target safety geometric boundary to obtain control commands. If the environmental risk identification result indicates a low-risk scenario, the control commands are determined based on the nominal actions and sent to the robot controller for execution. This invention assesses risk based on vision and commands, corrects robot actions to ensure safety, and achieves efficient execution and control of robot actions. Attached Figure Description

[0052] Figure 1 This is a flowchart of a preferred embodiment of the robot control method based on a visual language action model of the present invention;

[0053] Figure 2 This is a schematic diagram of the overall framework of a preferred embodiment of the robot control method based on a visual language action model of the present invention;

[0054] Figure 3 This is a structural diagram of a preferred embodiment of the robot control system based on a visual language action model of the present invention;

[0055] Figure 4 This is a structural diagram of a preferred embodiment of the terminal of the device of the present invention. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0057] Traditional methods typically employ VLA models and reinforcement learning-based safety constraint methods. However, these methods suffer from limitations in safety (soft constraints rather than hard constraints), high computational costs and lack of flexibility when adjusting strategies, and the inability to accurately predict obstacle volumes in cases of visual occlusion or data sparsity due to the lack of physical geometric priors. Consequently, they struggle to achieve strict safety control and efficient execution of robot actions. Therefore, a robot control method based on a visual language action model is needed. This method assesses risks based on vision and instructions, corrects robot actions to ensure safety, and achieves efficient execution and control of robot actions.

[0058] The preferred embodiment of the robot control method based on visual language action model described in this invention, such as... Figure 1 and Figure 2 As shown, the robot control method based on visual language action models includes the following steps:

[0059] Step S10: Obtain the robot's visual images and task instructions, and perform risk assessment in a pre-trained visual language model based on the visual images and task instructions to obtain environmental risk identification results.

[0060] Specifically, acquiring visual images of the robot ( ) and task instructions ( Based on the visual image and the task instructions, a risk assessment is performed in the pre-trained visual language model to obtain the environmental risk identification result, for example: "Based on the image and instructions, list all objects that may hinder the robot's movement and sort them by risk probability; if the environment is safe, output 'no risk'."

[0061] For example, such as Figure 2 As shown, the system first performs risk assessment and filters key information using a pre-trained Visual Language Model (VLM). If the risk confidence level is below a threshold (judged by the global gating unit), the "fast system" directly generates and executes action commands from the VLM. If a high-risk scenario is identified, the "slow system" is activated: it first identifies the target using a visual localization model, generates a scene point cloud using depth information, and then calculates precise "geometric constraints." Finally, the system inputs the initially generated nominal actions along with these safety constraints into a "quadratic programming" solver to calculate the final control commands that are both consistent with the task intent and absolutely safe, driving the robot to execute.

[0062] The process following step S10 also includes:

[0063] Step S11: Input the visual image and the task instruction into the pre-trained visual language model for prediction to obtain the probability distribution;

[0064] Step S12: Filter the prompt words of the environmental risk identification results according to the probability distribution to obtain a set of candidate obstacles.

[0065] Specifically, the visual image and the task instruction are input into the pre-trained visual language model for prediction to obtain a probability distribution ( The prompts for the environmental risk identification results are filtered based on the probability distribution to obtain a set of candidate obstacles. ).

[0066] In this embodiment, Top-p (kernel sampling) screening involves setting a cumulative probability threshold p (e.g., p=0.9). The probability is accumulated starting with the word with the highest probability until it exceeds p, forming a candidate obstacle set. A risk threshold is then set. . Let be the probability value of the obstacle word in the VLM output distribution. Case A (entering a fast system): If the set Empty, or risk confidence level of all objects < The environment is deemed "safe," and the task is completed efficiently in the fast system. Scenario B (Entering the slow system): If there is a risk confidence level for at least one object... ≥ If the environment is deemed "potentially dangerous," a slow system is entered to ensure the safety of the task execution process.

[0067] Furthermore, when environmental safety is determined, the system does not intervene in the model's output actions to maximize execution efficiency. VLA inference: and The input is a base VLA model (such as OpenVLA), and the VLA model directly outputs the nominal action. and directly As the final action Send to robot controller:

[0068] ;

[0069] in, Indicates a nominal action. This indicates the action that the robot ultimately performs.

[0070] Step S20: Obtain depth image and camera intrinsic parameter matrix. If the environmental risk identification result is a high-risk scene, then perform three-dimensional reconstruction and geometric fusion on the candidate obstacle set based on the depth image and the camera intrinsic parameter matrix to obtain the target safety geometric boundary.

[0071] Step S20 includes:

[0072] Step S21: If the environmental risk identification result is a high-risk scenario, then the location of the candidate obstacle set is extracted to obtain multiple spatial positioning points;

[0073] Step S22: Obtain the depth image and camera intrinsic matrix of the candidate obstacle set, and transform all the spatial positioning points according to the depth image and the camera intrinsic matrix to form a three-dimensional point cloud set;

[0074] Step S23: Perform outlier removal and spatial pass-through filtering on the three-dimensional point cloud set to obtain the observation point cloud;

[0075] Step S24: Perform three-dimensional reconstruction and geometric fusion of the candidate obstacle set based on the observed point cloud to obtain the target safety geometric boundary.

[0076] Specifically, if the environmental risk identification result is a high-risk scenario, then the location of the candidate obstacle set is extracted to obtain multiple spatial positioning points (for the set). Every obstacle in Spatial localization is performed to obtain the depth image and camera intrinsic matrix of the candidate obstacle set. Based on the depth image and camera intrinsic matrix, all spatial localization points are transformed (using coordinate transformation formulas) to form a 3D point cloud set (combined with the depth image). and camera intrinsic parameter matrix ,Will pixels within Converted to a 3D point cloud set in camera coordinate system The three-dimensional point cloud set is subjected to outlier removal and spatial pass-through filtering to obtain the observation point cloud (outlier removal and spatial pass-through filtering are performed to obtain the denoised observation point cloud). Based on the observed point cloud, the candidate obstacle set is reconstructed in three dimensions and geometrically fused to obtain the target safety geometric boundary.

[0077] As an example, firstly, VLM is used for scenario risk assessment (i.e., the split point of the "fast and slow system"): Prompt word construction: The system automatically... and Encapsulated in the process: "Analyze images and instructions. List possible fragile or dangerous obstacles on the path. If the environment is safe, output 'no risk'." Probability prediction and filtering: The pre-trained visual language model outputs the prediction results. In this embodiment, the model identifies "glass cup," and the confidence probability of this word in the output distribution... (0.92). Threshold determination: Set a risk threshold. = 0.8. Decision logic: Because (0.92)≥ (0.8) The system determines the environment as "high risk" and automatically activates the slow system channel. (Note: If there are no obstacles on the desktop, the pre-trained visual language model outputs "no risk", and the system will bypass subsequent geometric calculations and directly enter the fast system for execution.)

[0078] Furthermore, combining the depth map, the pixels within the bounding box are back-projected into 3D space to generate an observation point cloud. Prior fusion (semantic-geometric commonality prediction): Standard dimensions are retrieved from the "prior library". The number N of observation point clouds is calculated. Assuming the current view is obstructed and the point cloud is sparse, the system automatically increases the weight of the prior shape. By fusing the "observation-fitted ellipsoid" with the "standard prior ellipsoid", a safe geometric boundary for the target is generated. This ellipsoid can completely enclose the cup and is larger than the simply observed portion, providing additional safety margins. The observation point cloud is a collection of countless points with three-dimensional coordinates (X, Y, Z). It comes from depth images captured by an RGB-D camera and is calculated through backprojection. The data format is a list, where each element is a three-dimensional coordinate. These points represent the object surface that the camera can directly see at this moment.

[0079] In this embodiment, the coordinate transformation formula is as follows:

[0080] ;

[0081] in, Represented as 3D coordinates in the camera coordinate system, it indicates the spatial position of a point on the object's surface relative to the camera. This represents the depth value, which is the vertical distance between the object and the camera. This represents the inverse of the camera intrinsic matrix, used to transform the pixel coordinate system back to the physical imaging plane coordinate system. Represents the homogeneous pixel coordinates, where, Represents the pixel row and column coordinates of the object in the RGB image, with 1 indicating the augmented term of the homogeneous coordinates.

[0082] As an example, define categories: For application scenarios (such as home desktop organization), define semantic labels for common fragile or hazardous items, such as "stewed glass," "hot coffee," and "vase." Set geometric parameters: Define standard bounding box dimensions for each category and store the above mapping relationships as a JSON file. When the "slow system" identifies the object category but the point cloud is sparse, it directly calls the standard dimensions in the library to generate a safe bounding box. Application scenario: Desktop cleaning task. A target object (red block) and an obstacle (glass cup) are placed on the desktop. The robot's task is to pick up the red block and place it in the right-hand area.

[0083] Step S24 includes:

[0084] Step S241: Obtain the standard geometric shape parameters of the pre-built semantic geometry common library;

[0085] Step S242: Apply geometric constraints to the candidate obstacle set based on the observed point cloud and the standard geometric shape parameters to obtain the initial safe geometric boundary;

[0086] Step S243: Obtain the confidence factor, and perform weighted fusion of the initial safe geometric boundary based on the confidence factor and the observed point cloud to obtain the target safe geometric boundary.

[0087] Specifically, obtain the standard geometric shape parameters (e.g., standard, length:) of the pre-built semantic geometry common library. ,Width: ,high: Furthermore, according to Based on the semantic category, its corresponding standard shape matrix is ​​retrieved. ;

[0088] Based on the observed point cloud and the standard geometric shape parameters, geometric constraints are applied to the candidate obstacle set to obtain the initial safe geometric boundary:

[0089] ;

[0090] ;

[0091] in, Represents the observation shape matrix, Indicates the center of the observed shape. Represents the set of observation point clouds. This represents points in the observed point cloud. Represents the determinant of a matrix. Indicates transpose;

[0092] Obtain the confidence factor, and perform weighted fusion of the initial safe geometric boundary based on the confidence factor and the observed point cloud to obtain the target safe geometric boundary.

[0093] Step S243:

[0094] Step S2431: Obtain the confidence factor, the number of points in the observed point cloud, the observation center of the standard object, and the prior center;

[0095] Step S2432: Register the initial safety geometric boundary according to the observation center and the prior center to obtain the target safety center;

[0096] Step S2433: Weighted fusion of the initial safe geometric boundary according to the standard geometric shape parameters to obtain a safe ellipsoid shape matrix.

[0097] Specifically, the confidence factor, the number of points in the observed point cloud, the observation center of the standard object, and the prior center are obtained:

[0098] With the number of observed point clouds Proportional ( The larger the value, the more trusting the observations are. The smaller the value, the more it depends on prior values. Threshold for the number of saturated point clouds:

[0099] ;

[0100] in, This represents the confidence factor. This indicates the number of points in the observed point cloud. This represents the threshold for the number of points in the saturated point cloud.

[0101] The initial safety geometric boundary is registered based on the observation center and the prior center to obtain the target safety center. The initial safety geometric boundary is then weighted and fused according to the standard geometric shape parameters to obtain the safety ellipsoid shape matrix.

[0102] ;

[0103] ;

[0104] in, This represents the confidence factor. Indicates the prior center, Indicates the target security center. The matrix representing the shape of the safe ellipsoid. This represents the prior shape matrix.

[0105] In this embodiment, when observation data is missing ( When the value is at its minimum, the system automatically increases the proportion of prior weights, using "common sense" to complete the object shape. In this embodiment, to ensure the real-time performance of the system's calculations, common obstacles in the preset scene (such as water cups and boxes) are all placed upright. Therefore, the registration process of the prior model only includes position alignment in three-dimensional space, ignoring rotation transformations. That is, the geometric center of the standard object in the prior library is directly translated to the current observation center. The location serves as the prior center after registration. Subsequently, the observation center With the Priority Center (After registration) the same weighted fusion is performed to obtain the final security center.

[0106] Step S30: Generate a nominal action based on the visual image and the task instruction. If the environmental risk identification result is a high-risk scenario, the nominal action is modified according to the target safety geometric boundary to obtain a control instruction. If the environmental risk identification result is a low-risk scenario, the control instruction is determined based on the nominal action and sent to the robot controller for execution.

[0107] Step S30 includes:

[0108] Step S31: Input the visual image and the task instruction into the visual language action model for solving, and generate the nominal action;

[0109] Step S32: If the environmental risk identification result is a high-risk scenario, construct a control obstacle constraint function based on the target safety geometric boundary, modify the nominal action based on the control obstacle constraint function, and obtain a control command. If the environmental risk identification result is a low-risk scenario, determine the control command based on the nominal action and send the control command to the robot controller for execution.

[0110] Specifically, the visual image and the task instruction are input into the visual language action model for solving, and a nominal action is generated. If the environmental risk identification result is a high-risk scenario, a control obstacle constraint function is constructed based on the target safety geometric boundary, and the nominal action is modified based on the control obstacle constraint function to obtain a control command. If the environmental risk identification result is a low-risk scenario, the control command is determined based on the nominal action and sent to the robot controller for execution.

[0111] Step S32 includes:

[0112] Step S321: If the environmental risk identification result is a high-risk scenario, obtain the current environmental data, and constrain the safety ellipsoid shape matrix according to the current environmental data to obtain multiple obstacle constraints;

[0113] Step S322: Construct a control obstacle constraint function based on the target safety geometric boundary and multiple obstacle constraints;

[0114] Step S323: Determine the optimal safety control quantity based on the control obstacle constraint function;

[0115] Step S324: Modify the nominal action according to the optimal safety control quantity to obtain the control command.

[0116] Specifically, if the environmental risk identification result is a high-risk scenario, the current environmental data is acquired, and the safety ellipsoid shape matrix is ​​constrained based on the current environmental data to obtain multiple obstacle constraints, that is, the robot's end effector is modeled as an ellipsoid. , will be generated in the environment An obstacle ellipsoid As constraints, solve for safe actions (multiple obstacle constraints); construct control obstacle constraint functions based on the target safety geometric boundary and multiple obstacle constraints:

[0117] ;

[0118] in, Indicates the current state of the robot. Indicates a safety margin. Represents the minimum distance function. Represents an ellipsoid. Indicates the first An obstacle, This represents the control obstacle constraint function.

[0119] In this embodiment, a preset buffer distance (e.g., 0.05 meters) ensures that the robot does not get too close. Given the robot's current state (such as joint angles, end effector position, etc.), this distance can be analytically obtained from the geometric parameters of the two ellipsoids. Calculate the ellipsoid of the robot's end effector using the minimum distance function. With the An ellipsoid of obstacles The shortest Euclidean distance between them.

[0120] Further, an optimal safety control quantity is determined based on the control obstacle constraint function; the nominal action is then modified based on the optimal safety control quantity to obtain a control command:

[0121] ;

[0122] in, This represents the optimal safety control quantity. Represents the optimization variable. Indicates a nominal action.

[0123] Furthermore, all of them must be satisfied simultaneously. Safety forward invariance condition for each obstacle:

[0124] ;

[0125] in, The gradient of the barrier function is represented by a vector. This represents the system's drift term, for example, in the absence of control input ( When =0), the rate of change of the robot's state over time (usually determined by inertia, gravity, etc.). This represents the control input matrix, which describes the control quantity. How does it affect the rate of change of state? Represents extended class functions, This indicates the matrix transpose.

[0126] As an example, and Input a basic VLA model, and the model outputs the raw action. (For example: [0.5, 0.0, 0.0], meaning moving straight forward). Construct the control obstacle function: calculate the distance function between the robot's end effector and the safety ellipsoidal shape matrix of the water cup. Solve using quadratic programming; the solver calculates the optimal safety control quantity within 10ms. Result comparison: If This will cause a collision. It will be forcibly modified to move along the tangent direction of the ellipsoid (e.g., modified to [0.4, 0.2, 0.0], i.e., to move sideways); if the robot is far away from the obstacle, Then maintain with Consistent.

[0127] For example, the linear form is usually taken. (in >0). This term determines the degree of deceleration of the robot as it approaches the safety boundary. If this inequality is not satisfied, it means the robot will irreversibly crash into the obstacle. All If all are large enough (far from obstacles), then If approaching any obstacle, It will be forced to modify to tangentially avoid obstacles.

[0128] Furthermore, such as Figure 3As shown, based on the above-described robot control method based on visual language action models, this invention also provides a robot control system based on visual language action models, wherein the robot control system based on visual language action models includes:

[0129] The obstacle recognition module 51 is used to acquire the robot's visual images and task instructions, and to perform risk assessment in a pre-trained visual language model based on the visual images and task instructions to obtain environmental risk recognition results.

[0130] The safety boundary constraint module 52 is used to acquire depth images and camera intrinsic parameter matrices. If the environmental risk identification result is a high-risk scene, then the candidate obstacle set is reconstructed and geometrically fused according to the depth images and the camera intrinsic parameter matrix to obtain the target safety geometric boundary.

[0131] The robot execution module 53 is used to generate a nominal action based on the visual image and the task instruction. If the environmental risk identification result is a high-risk scenario, the nominal action is modified according to the target safety geometric boundary to obtain a control instruction. If the environmental risk identification result is a low-risk scenario, the control instruction is determined based on the nominal action and sent to the robot controller for execution.

[0132] Furthermore, such as Figure 4 As shown, based on the above-mentioned robot control method and system based on visual language action model, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 4 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0133] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a robot control program 40 based on a visual language action model, which can be executed by the processor 10 to implement the robot control method based on a visual language action model in this application.

[0134] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the robot control method based on the visual language action model.

[0135] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The terminals communicate with each other via a system bus.

[0136] In one embodiment, when the processor 10 executes the robot control program 40 based on the visual language action model stored in the memory 20, the following steps are performed:

[0137] The robot acquires visual images and task instructions, and performs risk assessment in a pre-trained visual language model based on the visual images and task instructions to obtain environmental risk identification results.

[0138] Acquire depth images and camera intrinsic parameter matrices. If the environmental risk identification result is a high-risk scene, then perform three-dimensional reconstruction and geometric fusion on the candidate obstacle set based on the depth images and the camera intrinsic parameter matrix to obtain the target safety geometric boundary.

[0139] A nominal action is generated based on the visual image and the task instruction. If the environmental risk identification result is a high-risk scenario, the nominal action is modified according to the target safety geometric boundary to obtain a control instruction. If the environmental risk identification result is a low-risk scenario, the control instruction is determined based on the nominal action and sent to the robot controller for execution.

[0140] The process of acquiring the robot's visual images and task instructions, performing risk assessment in a pre-trained visual language model based on the visual images and task instructions to obtain environmental risk identification results, further includes:

[0141] The visual image and the task instruction are input into the pre-trained visual language model for prediction to obtain the probability distribution;

[0142] Based on the probability distribution, the prompt words of the environmental risk identification results are filtered to obtain a set of candidate obstacles.

[0143] Specifically, the acquisition of depth images and camera intrinsic parameter matrices, if the environmental risk identification result is a high-risk scene, involves performing 3D reconstruction and geometric fusion on the candidate obstacle set based on the depth images and the camera intrinsic parameter matrix to obtain the target safety geometric boundary, including:

[0144] If the environmental risk identification result is a high-risk scenario, then the location of the candidate obstacle set is extracted to obtain multiple spatial positioning points;

[0145] Obtain the depth image and camera intrinsic matrix of the candidate obstacle set, and transform all the spatial positioning points according to the depth image and the camera intrinsic matrix to form a three-dimensional point cloud set;

[0146] The three-dimensional point cloud set is subjected to outlier removal and spatial pass-through filtering to obtain the observation point cloud;

[0147] Based on the observed point cloud, the candidate obstacle set is reconstructed in three dimensions and geometrically fused to obtain the target safety geometric boundary.

[0148] Specifically, the step of performing 3D reconstruction and geometric fusion of the candidate obstacle set based on the observed point cloud to obtain the target safety geometric boundary includes:

[0149] Obtain the standard geometric shape parameters of the pre-built semantic geometry common library;

[0150] Based on the observed point cloud and the standard geometric shape parameters, geometric constraints are applied to the candidate obstacle set to obtain the initial safe geometric boundary:

[0151] ;

[0152] ;

[0153] in, Represents the observation shape matrix, Indicates the center of the observed shape. Represents the set of observation point clouds. This represents points in the observed point cloud. Represents the determinant of a matrix. Indicates transpose;

[0154] Obtain the confidence factor, and perform weighted fusion of the initial safe geometric boundary based on the confidence factor and the observed point cloud to obtain the target safe geometric boundary.

[0155] The target safety geometric boundary includes the target safety center and a safety ellipsoid shape matrix;

[0156] The step of obtaining the confidence factor, and then weighting and fusing the initial safe geometric boundary based on the confidence factor and the observed point cloud to obtain the target safe geometric boundary, specifically includes:

[0157] Obtain the confidence factor, the number of points in the observed point cloud, the observation center of the standard object, and the prior center:

[0158] The initial safety geometric boundary is registered based on the observation center and the prior center to obtain the target safety center;

[0159] The initial safety geometric boundary is weighted and fused according to the standard geometric shape parameters to obtain the safety ellipsoid shape matrix:

[0160] ;

[0161] ;

[0162] in, This represents the confidence factor. Indicates the prior center, Indicates the target security center. The matrix representing the shape of the safe ellipsoid. This represents the prior shape matrix.

[0163] Specifically, the step of generating a nominal action based on the visual image and the task instruction, and if the environmental risk identification result is a high-risk scenario, then modifying the nominal action according to the target safety geometric boundary to obtain a control instruction; if the environmental risk identification result is a low-risk scenario, then determining the control instruction based on the nominal action and sending the control instruction to the robot controller for execution, includes:

[0164] The visual image and the task instruction are input into the visual language action model for solving, and a nominal action is generated.

[0165] If the environmental risk identification result is a high-risk scenario, a control obstacle constraint function is constructed based on the target safety geometric boundary, and the nominal action is modified based on the control obstacle constraint function to obtain a control command. If the environmental risk identification result is a low-risk scenario, the control command is determined based on the nominal action and sent to the robot controller for execution.

[0166] Specifically, if the environmental risk identification result is a high-risk scenario, a control obstacle constraint function is constructed based on the target safety geometric boundary, and the nominal action is modified based on the control obstacle constraint function to obtain a control command, which includes:

[0167] If the environmental risk identification result is a high-risk scenario, the current environmental data is obtained, and the safety ellipsoid shape matrix is ​​constrained based on the current environmental data to obtain multiple obstacle constraints;

[0168] Construct a control obstacle constraint function based on the target safety geometric boundary and multiple obstacle constraints:

[0169] ;

[0170] in, Indicates the current state of the robot. Indicates a safety margin. Represents the minimum distance function. Represents an ellipsoid. Indicates the first An obstacle, Represents the control obstacle constraint function;

[0171] The optimal safety control quantity is determined based on the control obstacle constraint function.

[0172] The nominal action is corrected based on the optimal safety control quantity to obtain the control command:

[0173] ;

[0174] in, This represents the optimal safety control quantity. Represents the optimization variable. Indicates a nominal action.

[0175] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a robot control program based on a visual language action model, and the robot control program based on the visual language action model, when executed by a processor, implements the steps of the robot control method based on the visual language action model as described above.

[0176] In summary, this invention provides a robot control method, system, terminal, and storage medium based on a visual language action model. The method includes: acquiring visual images and task instructions from the robot; performing risk assessment in a pre-trained visual language model based on the visual images and task instructions to obtain an environmental risk identification result; acquiring depth images and a camera intrinsic parameter matrix; if the environmental risk identification result indicates a high-risk scenario, performing 3D reconstruction and geometric fusion on a set of candidate obstacles based on the depth images and camera intrinsic parameter matrix to obtain a target safety geometric boundary; generating nominal actions based on the visual images and task instructions; correcting the nominal actions based on the target safety geometric boundary to obtain control instructions; and sending the control instructions to the robot controller for execution. This invention assesses risk based on vision and instructions, corrects robot actions to ensure safety, and achieves efficient execution and control of robot actions.

[0177] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal system that includes that element.

[0178] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.

[0179] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A robot control method based on a visual language action model, characterized in that, The robot control method based on visual language action models includes: The robot acquires visual images and task instructions, and performs risk assessment in a pre-trained visual language model based on the visual images and task instructions to obtain environmental risk identification results. Acquire depth images and camera intrinsic parameter matrices. If the environmental risk identification result is a high-risk scene, then perform three-dimensional reconstruction and geometric fusion on the candidate obstacle set based on the depth images and the camera intrinsic parameter matrix to obtain the target safety geometric boundary. A nominal action is generated based on the visual image and the task instruction. If the environmental risk identification result is a high-risk scenario, the nominal action is modified according to the target safety geometric boundary to obtain a control instruction. If the environmental risk identification result is a low-risk scenario, the control instruction is determined based on the nominal action and sent to the robot controller for execution. The process of acquiring depth images and camera intrinsic parameter matrices, and if the environmental risk identification result indicates a high-risk scene, involves performing 3D reconstruction and geometric fusion on the candidate obstacle set based on the depth images and camera intrinsic parameter matrices to obtain the target safety geometric boundary. Specifically, this includes: If the environmental risk identification result is a high-risk scenario, then the location of the candidate obstacle set is extracted to obtain multiple spatial positioning points; Obtain the depth image and camera intrinsic matrix of the candidate obstacle set, and transform all the spatial positioning points according to the depth image and the camera intrinsic matrix to form a three-dimensional point cloud set; The three-dimensional point cloud set is subjected to outlier removal and spatial pass-through filtering to obtain the observation point cloud; Based on the observed point cloud, the candidate obstacle set is reconstructed in three dimensions and geometrically fused to obtain the target safety geometric boundary. The step of performing three-dimensional reconstruction and geometric fusion of the candidate obstacle set based on the observed point cloud to obtain the target safety geometric boundary specifically includes: Obtain the standard geometric shape parameters of the pre-built semantic geometry common library; Based on the observed point cloud and the standard geometric shape parameters, geometric constraints are applied to the candidate obstacle set to obtain the initial safe geometric boundary: ; ; in, Represents the observation shape matrix, Indicates the center of the observed shape. Represents the set of observation point clouds. This represents points in the observed point cloud. Represents the determinant of a matrix. Indicates transpose; Obtain the confidence factor, and perform weighted fusion of the initial safe geometric boundary based on the confidence factor and the observed point cloud to obtain the target safe geometric boundary.

2. The robot control method based on a visual language action model according to claim 1, characterized in that, The process of acquiring the robot's visual images and task instructions, performing risk assessment in a pre-trained visual language model based on the visual images and task instructions to obtain environmental risk identification results, further includes: The visual image and the task instruction are input into the pre-trained visual language model for prediction to obtain the probability distribution; Based on the probability distribution, the prompt words of the environmental risk identification results are filtered to obtain a set of candidate obstacles.

3. The robot control method based on a visual language action model according to claim 1, characterized in that, The target security geometric boundary includes the target security center and a security ellipsoid shape matrix; The step of obtaining the confidence factor, and then weighting and fusing the initial safe geometric boundary based on the confidence factor and the observed point cloud to obtain the target safe geometric boundary, specifically includes: Obtain the confidence factor, the number of points in the observed point cloud, the observation center of the standard object, and the prior center: The initial safety geometric boundary is registered based on the observation center and the prior center to obtain the target safety center; The initial safety geometric boundary is weighted and fused according to the standard geometric shape parameters to obtain the safety ellipsoid shape matrix: ; ; in, This represents the confidence factor. Indicates the prior center, Indicates the target security center. The matrix representing the shape of the safe ellipsoid. This represents the prior shape matrix.

4. The robot control method based on a visual language action model according to claim 3, characterized in that, The process of generating a nominal action based on the visual image and the task instruction, and modifying the nominal action according to the target safety geometric boundary if the environmental risk identification result is a high-risk scenario to obtain a control instruction if the environmental risk identification result is a low-risk scenario, determining the control instruction based on the nominal action and sending the control instruction to the robot controller for execution, specifically includes: The visual image and the task instruction are input into the visual language action model for solving, and a nominal action is generated. If the environmental risk identification result is a high-risk scenario, a control obstacle constraint function is constructed based on the target safety geometric boundary, and the nominal action is modified based on the control obstacle constraint function to obtain a control command. If the environmental risk identification result is a low-risk scenario, the control command is determined based on the nominal action and sent to the robot controller for execution.

5. The robot control method based on a visual language action model according to claim 4, characterized in that, If the environmental risk identification result is a high-risk scenario, a control obstacle constraint function is constructed based on the target safety geometric boundary, and the nominal action is modified based on the control obstacle constraint function to obtain a control command, specifically including: If the environmental risk identification result is a high-risk scenario, the current environmental data is obtained, and the safety ellipsoid shape matrix is ​​constrained based on the current environmental data to obtain multiple obstacle constraints; Construct a control obstacle constraint function based on the target safety geometric boundary and multiple obstacle constraints: ; in, Indicates the current state of the robot. Indicates a safety margin. Represents the minimum distance function. Represents an ellipsoid. Indicates the first An obstacle, Represents the control obstacle constraint function; The optimal safety control quantity is determined based on the control obstacle constraint function. The nominal action is corrected based on the optimal safety control quantity to obtain the control command: ; in, This represents the optimal safety control quantity. Represents the optimization variable. Indicates a nominal action.

6. A robot control system based on a visual language action model, characterized in that, The robot control system based on the visual language action model is applied to the robot control method based on the visual language action model according to any one of claims 1-5, wherein the robot control system based on the visual language action model includes: The obstacle recognition module is used to acquire the robot's visual images and task instructions, and to perform risk assessment in a pre-trained visual language model based on the visual images and task instructions to obtain environmental risk recognition results. The safety boundary constraint module is used to acquire depth images and camera intrinsic parameter matrices. If the environmental risk identification result is a high-risk scene, the module performs 3D reconstruction and geometric fusion on the candidate obstacle set based on the depth images and the camera intrinsic parameter matrix to obtain the target safety geometric boundary. The robot execution module is used to generate nominal actions based on the visual image and the task instructions. If the environmental risk identification result is a high-risk scenario, the nominal actions are modified according to the target safety geometric boundary to obtain control instructions. If the environmental risk identification result is a low-risk scenario, the control instructions are determined based on the nominal actions and sent to the robot controller for execution.

7. A terminal, characterized in that, The terminal includes: a memory, a processor, and a robot control program based on a visual language action model stored in the memory and executable on the processor. When the robot control program based on the visual language action model is executed by the processor, it implements the steps of the robot control method based on a visual language action model as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a robot control program based on a visual language action model, which, when executed by a processor, implements the steps of the robot control method based on a visual language action model as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Method and device for securely executing motion trajectory of robot in working environment

    CN120862714A