Robot operation method, system and terminal based on visual language model

By using visual language models and tactile fusion technology, the robot generates initial planning strategies and corrects them in real time in unstructured environments, solving the problems of lack of physical attribute perception and difficulty in long-range fine operation, and improving operational accuracy and stability.

CN121696989AActive Publication Date: 2026-03-20GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)

Patent Information

Application Number
CN202610200293.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-11
Publication Date
2026-03-20
Estimated Expiration
2046-02-11

AI Technical Summary

Technical Problem

In unstructured environments, robots lack the ability to perceive physical properties and struggle with long-range precision operations, leading to deviations in operational results.

Method used

Visual images and natural language instructions are acquired through a visual language model to generate an initial planning strategy. Tactile signals during the operation are collected in real time using tactile sensors and input into a tactile language translation model for semantic mapping to generate a tactile semantic description. When the tactile semantic description indicates an operation abnormality, it is fed back to the visual language model for re-reasoning and planning correction.

Benefits of technology

It improves the robot's operational accuracy and stability in unstructured environments, enabling it to cope with instantaneous physical disturbances and autonomously correct cognitive biases, thus achieving highly robust operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121696989A_ABST
    Figure CN121696989A_ABST
Patent Text Reader

Abstract

The invention relates to the field of robot intelligent control, and discloses a robot operation method, system and terminal based on a visual language model, and the method comprises the steps: obtaining a visual image and a natural language instruction, carrying out the combined analysis of the visual image and the natural language instruction through a visual language large model, and generating an initial planning strategy; the robot is controlled to execute operation actions according to the initial planning strategy, and original tactile signals in the operation process are collected; inputting the original tactile signal into a tactile language translation model for feature extraction and semantic mapping to obtain tactile semantic description; when the prompt operation of the tactile semantic description is abnormal or the physical attribute does not accord with the visual expectation, the tactile semantic description is fed back to the visual language large model as an enhanced prompt word; and according to the visual image and the enhanced prompt word, the current task scene is reasoned again, a corrected planning strategy is generated, and the robot is controlled to execute an operation action. The operation precision of the robot is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot intelligent control technology, and in particular to a robot operation method, system, terminal, and computer-readable storage medium based on a visual language model. Background Technology

[0002] Robots are increasingly operating in complex, unstructured environments such as factories, homes, healthcare facilities, and logistics hubs. In these scenarios, robots need dexterous grasping and manipulation capabilities similar to human hands to handle a wide variety of objects with diverse physical properties (e.g., fragile, slippery, deformable). Vision-Language Models (VLMs), with their powerful semantic understanding, common-sense reasoning, and open-vocabulary scene parsing capabilities, are widely used in robot task planning and decision-making. By translating high-level natural language instructions into sequences of robot actions, VLMs endow robots with the potential to handle unseen objects and complex tasks. Furthermore, to compensate for the limitations of a single visual modality in physical interaction, tactile perception, as an important supplement to vision, has been introduced to provide local geometric information and mechanical feedback of the contact surface. Vision-tactile fusion has become a key technological direction for improving the operational stability of robots.

[0003] However, traditional visual-tactile fusion often combines visual and tactile data through simple signal splicing or Kalman filtering. Traditional fusion methods remain at the signal or feature level, failing to transform tactile signals into semantic information understandable by the VLM (Virtual Model). When physical interaction anomalies occur at the lower level (such as slippage due to an overly smooth object), the lower-level signals cannot be transmitted to the higher-level decision-making brain. This prevents the VLM from using its common-sense reasoning ability to correct higher-level policies. If all tactile feedback is processed by a large model, it leads to excessively high inference latency, making it unable to handle millisecond-level physical slippage. Conversely, relying solely on the lower-level controller cannot address policy errors.

[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0005] The main objective of this invention is to provide a robot operation method, system, terminal, and computer-readable storage medium based on a visual language model, aiming to solve the problem in the prior art that robots lack perception of physical attributes and have difficulty in long-range fine operation in unstructured environments, resulting in deviations in the final operation results.

[0006] To achieve the above objectives, the present invention provides a robot operation method based on a visual language model, the robot operation method based on a visual language model comprising the following steps: The visual images of the target scene and the natural language instructions for the operation task are obtained. The visual images and natural language instructions are jointly parsed using a large visual language model to generate an initial planning strategy. The robot is controlled to perform operations according to the initial planning strategy, and the tactile sensors set on the robot's end effector are used to collect raw tactile signals in real time during the operation. The original tactile signal is input into a pre-trained tactile language translation model for feature extraction and semantic mapping to obtain a tactile semantic description in natural language form. When the tactile semantic description indicates an operational anomaly or that the physical properties do not match the visual expectations, the tactile semantic description is fed back as an enhanced prompt word to the visual language big model. Based on the visual image and the enhanced prompts, the current task scenario is re-reasoned using the visual language big data model to generate a revised planning strategy, and the robot is controlled to perform operations according to the revised planning strategy.

[0007] Optionally, the robot operation method based on a visual language model, wherein the step of jointly parsing the visual image and the natural language command through a large visual language model to generate an initial planning strategy specifically includes: Based on the natural language instruction, the visual image is parsed using the visual language big model to identify and locate the target object indicated in the instruction, and the two-dimensional semantic bounding box of the target object in the visual image is output. The two-dimensional semantic bounding box is used as a prompt word and input into the segmentation model for segmentation to obtain the pixel-level accurate binary mask of the target object; Based on the pixel-level precise binary mask, the depth point cloud corresponding to the visual image is filtered to extract the local point cloud of the target object. The local point cloud is input into the geometric grasping generation model for calculation to generate a candidate set of six-degree-of-freedom grasping poses, and the set of six-degree-of-freedom grasping poses is used as the initial planning strategy.

[0008] Optionally, the robot operation method based on a visual language model further includes, after generating the initial planning strategy: For operational tasks that require tools or specific operations, the functional key points of the target object are identified through the large visual language model. Based on the natural language instructions of the operation task, the functional key points are transformed into spatial geometric constraints; The initial planning strategy is filtered or optimized based on the spatial geometric constraints to generate an optimized planning strategy that meets the functional requirements, and the operation actions are executed according to the optimized planning strategy.

[0009] Optionally, the robot operation method based on a visual language model further includes: When the robot performs an operation according to the initial planning strategy, it continuously acquires the original tactile signals; The original tactile signal is analyzed to determine whether it contains characteristics of slippage or abnormal contact force. If the original tactile signal contains the characteristics representing slippage or the abnormal contact force, then based on a preset reflexive control law, the robot's grasping action is corrected at the millisecond level to maintain grasping stability.

[0010] Optionally, in the robot operation method based on the visual language model, the reflexive control law includes a slip suppression strategy, a gripping force maintenance strategy, or a flexible protection strategy. The slip suppression strategy is used to automatically trigger a step increase in grip force when high-frequency micro-slip features are detected in the tactile signal, until the slip signal disappears; The gripping force maintenance strategy is used to adjust the normal gripping force in real time based on the PID controller so that the normal gripping force meets the preset gripping stability criterion. The flexible protection strategy is used to estimate the stiffness of an object based on tactile signals, and to dynamically limit the robot's maximum output grip force when the stiffness of the object is lower than a first preset threshold.

[0011] Optionally, the robot operation method based on a visual language model, wherein the step of inputting the original tactile signal into a pre-trained tactile language translation model for feature extraction and semantic mapping to obtain a tactile semantic description in natural language form, specifically includes: The original tactile signals are input as time-series data into a pre-trained tactile language translation model; The tactile language translation model extracts high-dimensional tactile features from the original tactile signal, maps the high-dimensional tactile features into natural language text, and forms the tactile semantic description. The tactile semantic description is used to describe the physical properties of an object or the contact state between a robot and an object in natural language. The physical properties include texture, hardness, or friction characteristics, and the contact state includes contact stability, slippage tendency, or force distribution state.

[0012] Optionally, the robot operation method based on a visual language model, wherein the step of re-reasoning about the current task scenario based on the visual image and the enhanced cue words, using the visual language model to generate a revised planning strategy, specifically includes: The visual image, the enhanced prompt words, and relevant historical case knowledge retrieved from the vectorized operation experience repository are fused using multimodal information to obtain fused information; The fused information is comprehensively analyzed by the visual language big model to re-evaluate the current task's state, physical constraints, and reasons for failure, and to obtain the re-evaluation result. Based on the reassessment results, a revised planning strategy is generated that conforms to the current physical reality and task objectives.

[0013] Optionally, in the robot operation method based on a visual language model, the step of generating the initial planning strategy further includes: Construct a vectorized operation experience repository, which is used to store historical success cases; Extract the visual semantic features of the current task, and perform similarity retrieval based on the visual semantic features in the vectorized operation experience repository; Historical successful cases with similarity higher than a second preset threshold are used as target cases. Prior knowledge of the target cases is extracted, and the initial planning strategy is generated based on the prior knowledge.

[0014] Furthermore, to achieve the above objectives, the present invention also provides a robot operating system based on a visual language model, wherein the robot operating system based on the visual language model includes: The multimodal semantic perception and planning module is used to acquire visual images of the target scene and natural language instructions for the operation task. It uses a large visual language model to jointly parse the visual images and natural language instructions to generate an initial planning strategy. The motion execution and tactile acquisition module is used to control the robot to perform operation actions according to the initial planning strategy, and to use tactile sensors set on the robot's end effector to collect raw tactile signals in real time during the operation process; The tactile semantic translation module is used to input the original tactile signal into a pre-trained tactile language translation model for feature extraction and semantic mapping to obtain a tactile semantic description in natural language form. The visual-touch fusion judgment module is used to feed back the tactile semantic description as an enhanced prompt word to the visual language big model when the tactile semantic description indicates an abnormal operation or the physical attributes do not match the visual expectations. The closed-loop replanning module is used to re-reason about the current task scenario based on the visual image and the enhanced prompt words through the visual language big model, generate a revised planning strategy, and control the robot to perform operation actions according to the revised planning strategy.

[0015] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a robot operation program based on a visual language model stored in the memory and executable on the processor, wherein when the robot operation program based on the visual language model is executed by the processor, it implements the steps of the robot operation method based on the visual language model as described above.

[0016] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a robot operation program based on a visual language model, and when the robot operation program based on the visual language model is executed by a processor, it implements the steps of the robot operation method based on the visual language model as described above.

[0017] In this invention, visual images and natural language commands are acquired, and a large-scale visual language model is used to jointly analyze the visual images and natural language commands to generate an initial planning strategy. The robot is then controlled to execute actions according to the initial planning strategy, and raw tactile signals are collected during the operation. These raw tactile signals are input into a tactile language translation model for feature extraction and semantic mapping to obtain a tactile semantic description. When the tactile semantic description indicates an operational anomaly or a discrepancy between physical attributes and visual expectations, the tactile semantic description is fed back as an enhanced cue word to the large-scale visual language model. Based on the visual image and the enhanced cue word, the current task scenario is re-reasoned to generate a revised planning strategy, and the robot is then controlled to execute the actions. This invention solves the problems of lack of physical attribute perception and difficulty in long-range fine-grained operations in unstructured environments, thus improving the robot's operational accuracy. Attached Figure Description

[0018] Figure 1 This is a flowchart of a preferred embodiment of the robot operation method based on a visual language model of the present invention; Figure 2 This is a technical schematic diagram of the robot operation method based on a visual language model according to the present invention; Figure 3 This is a structural diagram of a preferred embodiment of the robot operating system based on a visual language model of the present invention; Figure 4 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation

[0019] This application provides a robot operation method, system, and terminal based on a visual language model. To make the objectives, technical solutions, and effects of this application clearer and more explicit, the following detailed description is provided with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining this application and are not intended to limit this application.

[0020] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0021] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0022] The preferred embodiment of the robot operation method based on a visual language model described in this invention, such as... Figure 1 and Figure 2 As shown, the robot operation method based on the visual language model includes the following steps: Step S10: Obtain the visual image of the target scene and the natural language instructions for the operation task. Use the visual language big data model to jointly parse the visual image and the natural language instructions to generate an initial planning strategy.

[0023] Specifically, the robot acquires color images and corresponding depth images of the target scene using its RGB-D (RGB-Depth, color depth image) camera, which serve as the visual images. It also receives text or voice commands from the user defining the operation task, converting the voice commands into text as the natural language commands. It is understood that this embodiment clarifies the specific technical means and sources of data acquisition, concretizing abstract acquisition steps into executable technical actions. The visual images come from the robot's own RGB-D camera, and the data includes color and depth images. The natural language commands come from user input, which can be direct text or voice that needs to be converted.

[0024] Furthermore, the step of jointly parsing the visual image and the natural language instruction using a large visual language model to generate an initial planning strategy specifically includes: Based on the natural language instruction, the visual image is parsed using the visual language big model to identify and locate the target object indicated in the instruction, and the two-dimensional semantic bounding box of the target object in the visual image is output. The two-dimensional semantic bounding box is used as a prompt word and input into the segmentation model for segmentation to obtain the pixel-level accurate binary mask of the target object; Based on the pixel-level precise binary mask, the depth point cloud corresponding to the visual image is filtered to extract the local point cloud of the target object. The local point cloud is input into the geometric grasping generation model for calculation to generate a candidate set of six-degree-of-freedom grasping poses, and the set of six-degree-of-freedom grasping poses is used as the initial planning strategy.

[0025] In this embodiment, a pre-trained Visual Language Model (VLM) is first used as the core inference engine. The VLM directly parses the input image and text instructions end-to-end, understands the task intent, and outputs a two-dimensional semantic bounding box of the target object in the image. Next, to avoid information loss due to abstract semantic transmission in traditional cascaded detection, this embodiment explicitly couples a large segmentation model, such as the Segment Anything Model (SAM). The bounding box output by the VLM is used as input to the SAM to generate a pixel-level precise binary mask of the target object.

[0026] Furthermore, the generated mask is used to filter the depth point cloud data, retaining only the point cloud of the target region. A geometric grasping generation model, such as AnyGrasp (a six-DOF general object grasping algorithm), is invoked and combined with camera intrinsics to generate a set of candidate 6-DoF (six-DOF) grasping poses in 3D space.

[0027] Specifically, based on depth image data acquired by an RGB-D camera, back-projection is performed using the camera's intrinsic parameter matrix to reconstruct the original 3D depth point cloud data of the scene. Next, the semantic mask generated in the preceding steps is used to perform point-to-point indexing filtering on the original point cloud, retaining only the local point cloud of the target object. This local point cloud serves as input to a geometric grasping generation model (such as AnyGrasp), aiming to eliminate interference from the background environment (such as desktops and clutter), allowing the model to focus on calculating the geometric surface features of the target object. Finally, the geometric grasping generation model is invoked to analyze this local point cloud, outputting a set of candidate 6-DoF grasping poses (including spatial position coordinates and rotational orientation). This pose set is a concrete parameterized representation of the "initial planning strategy" at the geometric level, transforming the abstract semantic instructions generated by the VLM into specific end effector target points required for solving the inverse kinematics of the robotic arm, for subsequent motion planning modules to select and execute.

[0028] As can be seen, compared with traditional detection methods based on fixed category labels, this embodiment utilizes VLM to achieve open-vocabulary object detection, which can identify and locate objects not seen in the training set. Furthermore, it combines SAM to achieve a refined transition from semantic boxes to pixel-level masks, effectively eliminating interference from background point clouds and improving the accuracy of subsequent pose generation.

[0029] Furthermore, the generation of the initial planning strategy also includes, prior to: Construct a vectorized operation experience repository, which is used to store historical success cases; Extract the visual semantic features of the current task, and perform similarity retrieval based on the visual semantic features in the vectorized operation experience repository; Historical successful cases with similarity higher than a second preset threshold are used as target cases. Prior knowledge of the target cases is extracted, and the initial planning strategy is generated based on the prior knowledge.

[0030] Understandably, this invention utilizes a large vision-language model to perform end-to-end parsing of natural language instructions and environmental images. By combining a semantic segmentation model and a grasping detection model, it generates the grasping pose of the target object and introduces a vectorized long-term experience repository to assist in task planning by retrieving historical similar cases.

[0031] Specifically, the system maintains a vectorized operational experience repository, storing historically successful task cases (including visual features, task descriptions, action sequences, etc.). Before generating actions, the system extracts the visual semantic features of the current scene and performs a similarity search in the experience repository. If a similar historical successful case is matched, its prior knowledge is extracted to assist AnyGrasp in sampling candidate poses. By introducing a long-term memory mechanism and retrieving historical experience to provide heuristic priors for the current task, the system significantly reduces the online inference overhead of large models in complex scenes, achieving a "the more you use it, the more proficient you become" effect.

[0032] It should be noted that the aforementioned experience base can be a retrieval system based on a vector database or a structured storage format based on a knowledge graph, as long as it can achieve feature matching and prior retrieval for similar tasks.

[0033] Furthermore, the generation of the initial planning strategy further includes: For operational tasks that require tools or specific operations, the functional key points of the target object are identified through the large visual language model. Based on the natural language instructions of the operation task, the functional key points are transformed into spatial geometric constraints; The initial planning strategy is filtered or optimized based on the spatial geometric constraints to generate an optimized planning strategy that meets the functional requirements, and the operation actions are executed according to the optimized planning strategy.

[0034] In this embodiment, the functional key points of an object (such as handles and openings) are predicted based on a large model, and the abstract semantics are transformed into dynamic spatial geometric constraints based on the key points to guide refined operations.

[0035] Specifically, for tasks requiring tool use or specific operations (such as "using a hammer" which requires holding the handle), VLM further predicts the functional key points of the object (such as "tool handle" or "container center"). The system transforms the task requirements into a dynamic spatial constraint problem based on the key point locations, filters and optimizes the geometric grasping pose to meet the functional operation requirements.

[0036] As can be seen, traditional grasping methods only focus on the geometric stability of the object's center of mass (whether it can be grasped). This invention introduces semantic functional key points (such as recognizing the "hammer handle" rather than the "hammer head") and a vectorized experience repository. The advantage of this is that it endows the robot with the ability to "use tools" (not just carry), enabling it to perform fine operations such as alignment and straightening; at the same time, by utilizing historical success or failure experience retrieval, it significantly improves reasoning efficiency and planning success rate in complex scenarios.

[0037] Step S20: Control the robot to perform operation actions according to the initial planning strategy, and use the tactile sensor set on the robot's end effector to collect the original tactile signals in real time during the operation process.

[0038] Specifically, the high-level task planning is translated into the robot's low-level physical actions. The system takes the grasping pose (including spatial position and orientation) determined in the planning strategy as the target and inputs it into the robot arm's motion planning module to generate a joint space trajectory that smoothly and without collisions moves the robot arm from its current position to the target pose. Subsequently, this trajectory is parsed in real time into a series of continuous joint angle commands and sent to the robot's servo drive system. Finally, each joint motor receives the commands and executes them precisely, driving the end effector (such as a gripper) to complete predetermined operations such as approaching and grasping.

[0039] Furthermore, the system continuously and frequently collects raw physical signals related to the contact interface throughout the entire process of the robot performing grasping, lifting, placing, or manipulating actions according to the planned strategy by using tactile sensors integrated in specific parts of the robot's end effector (such as the inner side of the gripper's fingertips, the contact surface, or the joints). These tactile sensors (the tactile sensors mentioned in this invention can be flexible sensors based on the piezoresistive principle, or capacitive, visual-tactile sensors, or other sensors that can provide contact force and texture information)

[0040] These raw tactile signals are low-level data streams that have not been interpreted by higher-level semantics. Their types and sources include distributed pressure signals, high-frequency vibration and texture signals, body deformation signals, and torque signals.

[0041] The system includes: Distributed pressure signal: Provided by a high-density flexible piezoresistive or capacitive sensor array, it captures the three-dimensional force distribution (normal pressure and shear force) and its dynamic changes on the contact surface in real time, forming a tactile image. High-frequency vibration and texture signal: Captures micro-vibrations generated at the contact interface through accelerometers or dynamic force sensors; its frequency and amplitude characteristics contain information such as the surface texture and the moment of initiation of sliding. Body deformation signal: Sensing the degree of deformation of the end effector itself (such as a soft gripper) through strain gauges embedded in the flexible coating or drive structure, indirectly calculating the stiffness and shape adaptability of the contact object. Torque signal: Provided by a torque sensor installed on the end effector base, it accurately measures the resultant force and torque applied to the object during operation, used for macroscopic force control and stability assessment.

[0042] The data acquisition process is completed by a dedicated signal conditioning circuit and a high-speed data acquisition card. This circuit amplifies, filters, and digitizes the analog output of the sensor, generating a high-dimensional time-series signal sequence with millisecond-level or even higher time resolution. These raw data streams are uploaded to the system's tactile processing unit via a real-time communication interface, providing the most basic, unmodified physical world information input for subsequent reflection control and semantic translation.

[0043] Furthermore, the robot operation method based on the visual language model also includes: When the robot performs an operation according to the initial planning strategy, it continuously acquires the original tactile signals; The original tactile signal is analyzed to determine whether it contains characteristics of slippage or abnormal contact force. If the original tactile signal contains the characteristics representing slippage or the abnormal contact force, then based on a preset reflexive control law, the robot's grasping action is corrected at the millisecond level to maintain grasping stability.

[0044] Understandably, this invention innovatively constructs a hierarchical, progressive, closed-loop control architecture for visual-tactile fusion, comprising "lower-level reflection" and "upper-level cognition." When performing the initial operation based on the upper-level cognition loop (replanning based on tactile-language translation), this invention simultaneously initiates a lower-level reflection loop (action refinement based on local tactile reflections) that does not rely on large-model reasoning, and designs a fast-reflexive force controller that does not undergo VLM reasoning. The original tactile signals are continuously acquired by this fast-reflexive force controller.

[0045] The fast-reflective force controller performs online analysis of the continuously input raw tactile signals. Its analysis objectives are twofold: first, to identify whether specific characteristic patterns representing micro-slip appear in the signal (e.g., a sudden increase in high-frequency vibration energy in a specific frequency band, or periodic fluctuations in shear force); and second, to determine whether the signal deviates from the expected contact force range (e.g., a sharp drop in normal force indicates instability, or an overload of normal force indicates compression).

[0046] Furthermore, once the fast-reflex force controller outputs a positive judgment result (i.e., confirming the existence of slippage characteristics or abnormal contact force characteristics), the system will immediately trigger a preset reflexive control law without waiting for the decision of the upper-level cognitive loop. This control law directly acts on the robot's grasping action control loop. For example, it dynamically adjusts the torque output of the servo motor through a PID controller (Proportional-Integral-Derivative Controller) to increase the normal gripping force to suppress slippage, or it dynamically limits the amplitude through the force control interface to protect fragile objects. The entire "detection-trigger-execution" closed loop is completed in milliseconds, aiming to maintain the physical stability of the grasp before the intervention of higher-level planning. This solves the problem that large model inference latency cannot cope with millisecond-level physical disturbances, ensuring the real-time stability and safety of the operation and preventing objects from slipping or being damaged.

[0047] Furthermore, the reflexive control law includes a slip suppression strategy, a gripping force maintenance strategy, or a flexible protection strategy; The slip suppression strategy is used to automatically trigger a step increase in grip force when high-frequency micro-slip features are detected in the tactile signal, until the slip signal disappears; The gripping force maintenance strategy is used to adjust the normal gripping force in real time based on the PID controller so that the normal gripping force meets the preset gripping stability criterion. The flexible protection strategy is used to estimate the stiffness of an object based on tactile signals, and to dynamically limit the robot's maximum output grip force when the stiffness of the object is lower than a first preset threshold.

[0048] In this embodiment, the reflective control law mainly includes the following three aspects: Slip detection and suppression: Real-time monitoring of high-frequency micro-slip characteristics of tactile signals. Once slip is detected, an automatic grip force increment strategy is triggered until the slip signal disappears.

[0049] Maintaining grip stability: Real-time adjustment of normal grip force using a PID algorithm. To ensure the following gripping stability criteria are met and prevent objects from slipping: ; in, For normal grip force, The coefficient of friction, The external force acting on an object This is a dynamic safety margin.

[0050] Flexible object protection: For fragile or easily deformable objects, the deformation gradient of the tactile image is calculated in real time to estimate stiffness. If extremely low stiffness is detected, a force limiting mechanism is triggered to dynamically truncate excessive displacement commands issued by the VLM, limiting the grip force to within a first preset threshold.

[0051] Step S30: Input the original tactile signal into the pre-trained tactile language translation model for feature extraction and semantic mapping to obtain a tactile semantic description in natural language form.

[0052] Specifically, the original tactile signals are input as time-series data into a pre-trained tactile language translation model; The tactile language translation model extracts high-dimensional tactile features from the original tactile signal, maps the high-dimensional tactile features into natural language text, and forms the tactile semantic description. The tactile semantic description is used to describe the physical properties of an object or the contact state between a robot and an object in natural language. The physical properties include texture, hardness, or friction characteristics, and the contact state includes contact stability, slippage tendency, or force distribution state.

[0053] In this embodiment, the raw tactile signals typically originate from a high-density array of sensors on the end effector, outputting high-frequency, multi-channel time-series voltage or digital signals. In this step, the system performs necessary preprocessing on these signals (such as normalization) and organizes them into standardized time-series data frames, which serve as direct input to the tactile-language translation model. This ensures that the model can receive and process signal streams that are consistent with the training data format and reflect the dynamic changes in physical interaction.

[0054] Furthermore, the pre-trained tactile language translation model receives the aforementioned temporal data. The model first automatically extracts discriminative high-dimensional tactile features from the raw, high-dimensional temporal signal using its deep encoder network. Subsequently, the model's decoder uses these high-dimensional features as conditions to drive a language generation head, which generates a coherent text description according to the grammatical and semantic rules of natural language. This completes the end-to-end mapping from numerical features to natural language text, forming the final tactile semantic description.

[0055] It is understood that the tactile semantic description generated by this invention includes two core dimensions: description of the object's physical properties and description of the contact state.

[0056] Object physical property description: This aims to characterize the material properties of the object itself. For example, the model might output "rough surface", "soft material" or "very slippery", which correspond to the evaluation of texture, hardness and friction properties, respectively.

[0057] Contact state description: This aims to characterize the dynamic state of the interaction process. For example, the model may output "stable grip", "lateral sliding is occurring", or "uneven fingertip pressure distribution", which correspond to the evaluation of contact stability, sliding tendency, and force distribution state, respectively.

[0058] Step S40: When the tactile semantic description indicates an operational abnormality or that the physical properties do not match the visual expectations, the tactile semantic description is fed back as an enhanced prompt word to the visual language big model.

[0059] Specifically, the semantic descriptions generated in real time by the tactile language translation model are analyzed and judged instantly. The judgment is based on two clear logical dimensions: the first is operation anomaly detection, which matches the descriptive text with a predefined library of anomaly keywords (such as "slippage," "looseness," "insufficient grip strength"). If the text contains such words, it is determined that a physical anomaly requiring attention has occurred. The second is physical attribute conflict detection, which directly compares the assertions in the description about the characteristics of the object itself (such as "very soft," "slippery surface") with the attribute inferences made by the visual language model based solely on visual images of the same object (such as inferences of "hard object," "dry surface"). If the two are contradictory, it is determined that the physical attributes do not match the visual expectations.

[0060] Furthermore, if any of the above conditions are met, the system will initiate the feedback process. The system will integrate the tactile semantic description, the specific reason for the trigger judgment (e.g., "slippage detected" or "conflict between tactile feedback material and visual judgment"), and relevant historical case knowledge retrieved in real time from the long-term experience base (if any) into a structured information package.

[0061] The information packet is formatted into a clear, enhanced cue word that can be directly processed by the large model, and is fed back to the decision input loop of the vision-language large model in real time through a dedicated interface. This enables the large model to initiate a new round of reasoning based on enhanced cognition that integrates visual scene, task intent, and physical reality feedback. The system can automatically initiate a replanning process by combining failure records in the experience base with the current tactile feedback, achieving self-healing from failures.

[0062] Step S50: Based on the visual image and the enhanced prompt words, the current task scenario is re-reasoned through the visual language big model to generate a revised planning strategy, and the robot is controlled to perform operation actions according to the revised planning strategy.

[0063] The step of re-reasoning about the current task scenario based on the visual image and the enhanced cue words, using the visual language big data model, and generating a revised planning strategy specifically includes: The visual image, the enhanced prompt words, and relevant historical case knowledge retrieved from the vectorized operation experience repository are fused using multimodal information to obtain fused information; The fused information is comprehensively analyzed by the visual language big model to re-evaluate the current task's state, physical constraints, and reasons for failure, and to obtain the re-evaluation result. Based on the reassessment results, a revised planning strategy is generated that conforms to the current physical reality and task objectives.

[0064] In this embodiment, when the system initiates re-inference, it fuses three types of key information: First, the original visual image. Second, enhanced cue words containing physical descriptions, such as "the object surface is slippery" or "the grip has slipped." Third, relevant historical case knowledge retrieved in real time from a vectorized experience base. Through feature alignment and context stitching techniques, this information is integrated into a unified fused information representation, providing a data foundation for deep analysis.

[0065] Furthermore, the visual language big data model comprehensively analyzes the fused information to achieve cognitive deepening. First, it reassesses the current task state, such as determining whether the object has shifted or is in an unstable grasp. Second, it re-identifies physical constraints, confirming real attributes that are difficult to judge visually based on tactile feedback, such as the object's hardness and surface friction coefficient. Finally, it diagnoses potential causes of failure, inferring whether the anomaly stems from insufficient grip strength, an overly slippery surface, or an improper movement trajectory, thus outputting a comprehensively updated reassessment result.

[0066] Furthermore, based on the reassessment results, the visual language big data model is re-planned. The final revised planning strategy is an executable sequence of specific actions that integrates multimodal cognition, reflecting the adaptive decision-making capability that the system acquires through closed-loop feedback.

[0067] The present invention has the following beneficial effects: (1) Solved the problem of the disconnect between high-level semantics and low-level control: By guiding the key points of semantic functions, the robot is given the ability to understand the functions of object parts and perform fine long-range operations such as alignment and straightening.

[0068] (2) It bridges the modal gap between tactile signals and large model cognition: It is the first to translate tactile signals into natural language, enabling large models to "understand" the physical world (such as softness and hardness, friction characteristics), thereby achieving intelligent decision-making in accordance with physical laws under zero-sample conditions.

[0069] (3) Achieved highly robust operation with self-healing ability: The unique dual-loop mechanism of "bottom-level physical reflection + upper-level semantic replanning" can not only cope with instantaneous physical disturbances, but also autonomously correct cognitive biases through closed-loop feedback, which significantly improves the success rate and adaptability of the robot in unstructured environments when facing unknown, fragile or slippery objects.

[0070] Furthermore, such as Figure 3 As shown, based on the above-described robot operation method based on a visual language model, this invention also provides a robot operating system based on a visual language model, wherein the robot operating system based on a visual language model includes: The multimodal semantic perception and planning module 51 is used to acquire visual images of the target scene and natural language instructions for the operation task, and to jointly parse the visual images and natural language instructions through a large visual language model to generate an initial planning strategy. The motion execution and tactile acquisition module 52 is used to control the robot to perform operation actions according to the initial planning strategy, and to use the tactile sensor set on the robot's end effector to collect the original tactile signals in real time during the operation process. The tactile semantic translation module 53 is used to input the original tactile signal into a pre-trained tactile language translation model for feature extraction and semantic mapping to obtain a tactile semantic description in natural language form. The visual-touch fusion judgment module 54 is used to feed back the tactile semantic description as an enhanced prompt word to the visual language big model when the tactile semantic description prompts an operation abnormality or the physical attributes do not match the visual expectations. The closed-loop replanning module 55 is used to re-reason about the current task scenario based on the visual image and the enhanced prompt words through the visual language big model, generate a revised planning strategy, and control the robot to perform operation actions according to the revised planning strategy.

[0071] Furthermore, such as Figure 4 As shown, based on the above-mentioned robot operation method and system based on visual language model, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 4 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0072] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a robot operation program 40 based on a visual language model, which can be executed by the processor 10 to implement the robot operation method based on a visual language model in this application.

[0073] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the robot operation method based on the visual language model.

[0074] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.

[0075] In one embodiment, when the processor 10 executes the robot manipulation program 40 based on the visual language model in the memory 20, the following steps are performed: The visual images of the target scene and the natural language instructions for the operation task are obtained. The visual images and natural language instructions are jointly parsed using a large visual language model to generate an initial planning strategy. The robot is controlled to perform operations according to the initial planning strategy, and the tactile sensors set on the robot's end effector are used to collect raw tactile signals in real time during the operation. The original tactile signal is input into a pre-trained tactile language translation model for feature extraction and semantic mapping to obtain a tactile semantic description in natural language form. When the tactile semantic description indicates an operational anomaly or that the physical properties do not match the visual expectations, the tactile semantic description is fed back as an enhanced prompt word to the visual language big model. Based on the visual image and the enhanced prompts, the current task scenario is re-reasoned using the visual language big data model to generate a revised planning strategy, and the robot is controlled to perform operations according to the revised planning strategy.

[0076] Specifically, the step of jointly parsing the visual image and the natural language instruction using a large visual language model to generate an initial planning strategy includes: Based on the natural language instruction, the visual image is parsed using the visual language big model to identify and locate the target object indicated in the instruction, and the two-dimensional semantic bounding box of the target object in the visual image is output. The two-dimensional semantic bounding box is used as a prompt word and input into the segmentation model for segmentation to obtain the pixel-level accurate binary mask of the target object; Based on the pixel-level precise binary mask, the depth point cloud corresponding to the visual image is filtered to extract the local point cloud of the target object. The local point cloud is input into the geometric grasping generation model for calculation to generate a candidate set of six-degree-of-freedom grasping poses, and the set of six-degree-of-freedom grasping poses is used as the initial planning strategy.

[0077] The generation of the initial planning strategy further includes: For operational tasks that require tools or specific operations, the functional key points of the target object are identified through the large visual language model. Based on the natural language instructions of the operation task, the functional key points are transformed into spatial geometric constraints; The initial planning strategy is filtered or optimized based on the spatial geometric constraints to generate an optimized planning strategy that meets the functional requirements, and the operation actions are executed according to the optimized planning strategy.

[0078] The robot operation method based on the visual language model further includes: When the robot performs an operation according to the initial planning strategy, it continuously acquires the original tactile signals; The original tactile signal is analyzed to determine whether it contains characteristics of slippage or abnormal contact force. If the original tactile signal contains the characteristics representing slippage or the abnormal contact force, then based on a preset reflexive control law, the robot's grasping action is corrected at the millisecond level to maintain grasping stability.

[0079] The reflexive control law includes a slip suppression strategy, a gripping force maintenance strategy, or a flexible protection strategy. The slip suppression strategy is used to automatically trigger a step increase in grip force when high-frequency micro-slip features are detected in the tactile signal, until the slip signal disappears; The gripping force maintenance strategy is used to adjust the normal gripping force in real time based on the PID controller so that the normal gripping force meets the preset gripping stability criterion. The flexible protection strategy is used to estimate the stiffness of an object based on tactile signals, and to dynamically limit the robot's maximum output grip force when the stiffness of the object is lower than a first preset threshold.

[0080] Specifically, the step of inputting the original tactile signal into a pre-trained tactile language translation model for feature extraction and semantic mapping to obtain a tactile semantic description in natural language form includes: The original tactile signals are input as time-series data into a pre-trained tactile language translation model; The tactile language translation model extracts high-dimensional tactile features from the original tactile signal, maps the high-dimensional tactile features into natural language text, and forms the tactile semantic description. The tactile semantic description is used to describe the physical properties of an object or the contact state between a robot and an object in natural language. The physical properties include texture, hardness, or friction characteristics, and the contact state includes contact stability, slippage tendency, or force distribution state.

[0081] Specifically, the step of re-reasoning about the current task scenario based on the visual image and the enhanced cue words, using the visual language big data model, to generate a revised planning strategy includes: The visual image, the enhanced prompt words, and relevant historical case knowledge retrieved from the vectorized operation experience repository are fused using multimodal information to obtain fused information; The fused information is comprehensively analyzed by the visual language big model to re-evaluate the current task's state, physical constraints, and reasons for failure, and to obtain the re-evaluation result. Based on the reassessment results, a revised planning strategy is generated that conforms to the current physical reality and task objectives.

[0082] The generation of the initial planning strategy further includes, prior to: Construct a vectorized operation experience repository, which is used to store historical success cases; Extract the visual semantic features of the current task, and perform similarity retrieval based on the visual semantic features in the vectorized operation experience repository; Historical successful cases with similarity higher than a second preset threshold are used as target cases. Prior knowledge of the target cases is extracted, and the initial planning strategy is generated based on the prior knowledge.

[0083] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a robot operation program based on a visual language model, and the robot operation program based on the visual language model, when executed by a processor, implements the steps of the robot operation method based on the visual language model as described above.

[0084] In summary, this invention provides a robot operation method, system, and terminal based on a visual language model. The method includes: acquiring visual images and natural language commands; jointly parsing the visual images and natural language commands using a large visual language model to generate an initial planning strategy; controlling the robot to execute operational actions according to the initial planning strategy and collecting raw tactile signals during the operation; inputting the raw tactile signals into a tactile language translation model for feature extraction and semantic mapping to obtain a tactile semantic description; when the tactile semantic description indicates an operational anomaly or a discrepancy between physical attributes and visual expectations, feeding the tactile semantic description back to the large visual language model as an enhanced cue word; and re-reasoning about the current task scenario based on the visual image and the enhanced cue word to generate a revised planning strategy and control the robot to execute operational actions. This invention solves the problems of lack of physical attribute perception and difficulty in long-range fine-grained operation in unstructured environments, thereby improving the robot's operational accuracy.

[0085] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.

[0086] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0087] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A robot manipulation method based on a visual language model, characterized in that, The robot operation method based on the visual language model includes: The visual images of the target scene and the natural language instructions for the operation task are obtained. The visual images and natural language instructions are jointly parsed using a large visual language model to generate an initial planning strategy. The robot is controlled to perform operations according to the initial planning strategy, and the tactile sensors set on the robot's end effector are used to collect raw tactile signals in real time during the operation. The original tactile signal is input into a pre-trained tactile language translation model for feature extraction and semantic mapping to obtain a tactile semantic description in natural language form. When the tactile semantic description indicates an operational anomaly or that the physical properties do not match the visual expectations, the tactile semantic description is fed back as an enhanced prompt word to the visual language big model. Based on the visual image and the enhanced prompts, the current task scenario is re-reasoned using the visual language big data model to generate a revised planning strategy, and the robot is controlled to perform operations according to the revised planning strategy.

2. The robot operation method based on a visual language model according to claim 1, characterized in that, The step of jointly parsing the visual image and the natural language instruction using a large visual language model to generate an initial planning strategy specifically includes: Based on the natural language instruction, the visual image is parsed using the visual language big model to identify and locate the target object indicated in the instruction, and the two-dimensional semantic bounding box of the target object in the visual image is output. The two-dimensional semantic bounding box is used as a prompt word and input into the segmentation model for segmentation to obtain the pixel-level accurate binary mask of the target object; Based on the pixel-level precise binary mask, the depth point cloud corresponding to the visual image is filtered to extract the local point cloud of the target object. The local point cloud is input into the geometric grasping generation model for calculation to generate a candidate set of six-degree-of-freedom grasping poses, and the set of six-degree-of-freedom grasping poses is used as the initial planning strategy.

3. The robot operation method based on a visual language model according to claim 2, characterized in that, The generation of the initial planning strategy is followed by: For operational tasks that require tools or specific operations, the functional key points of the target object are identified through the large visual language model. Based on the natural language instructions of the operation task, the functional key points are transformed into spatial geometric constraints; The initial planning strategy is filtered or optimized based on the spatial geometric constraints to generate an optimized planning strategy that meets the functional requirements, and the operation actions are executed according to the optimized planning strategy.

4. The robot operation method based on a visual language model according to claim 1, characterized in that, The robot operation method based on the visual language model further includes: When the robot performs an operation according to the initial planning strategy, it continuously acquires the original tactile signals; The original tactile signal is analyzed to determine whether it contains characteristics of slippage or abnormal contact force. If the original tactile signal contains the characteristics representing slippage or the abnormal contact force, then based on a preset reflexive control law, the robot's grasping action is corrected at the millisecond level to maintain grasping stability.

5. The robot operation method based on a visual language model according to claim 4, characterized in that, The reflexive control law includes a slip inhibition strategy, a gripping force maintenance strategy, or a flexible protection strategy. The slip suppression strategy is used to automatically trigger a step increase in grip force when high-frequency micro-slip features are detected in the tactile signal, until the slip signal disappears; The gripping force maintenance strategy is used to adjust the normal gripping force in real time based on the PID controller so that the normal gripping force meets the preset gripping stability criterion. The flexible protection strategy is used to estimate the stiffness of an object based on tactile signals, and to dynamically limit the robot's maximum output grip force when the stiffness of the object is lower than a first preset threshold.

6. The robot operation method based on a visual language model according to claim 1, characterized in that, The step of inputting the original tactile signal into a pre-trained tactile language translation model for feature extraction and semantic mapping to obtain a tactile semantic description in natural language form specifically includes: The original tactile signals are input as time-series data into a pre-trained tactile language translation model; The tactile language translation model extracts high-dimensional tactile features from the original tactile signal, maps the high-dimensional tactile features into natural language text, and forms the tactile semantic description. The tactile semantic description is used to describe the physical properties of an object or the contact state between a robot and an object in natural language. The physical properties include texture, hardness, or friction characteristics, and the contact state includes contact stability, slippage tendency, or force distribution state.

7. The robot operation method based on a visual language model according to claim 1, characterized in that, The step of re-reasoning about the current task scenario based on the visual image and the enhanced cue words, using the visual language big data model, and generating a revised planning strategy specifically includes: The visual image, the enhanced prompt words, and relevant historical case knowledge retrieved from the vectorized operation experience repository are fused using multimodal information to obtain fused information; The fused information is comprehensively analyzed by the visual language big model to re-evaluate the current task's state, physical constraints, and reasons for failure, and to obtain the re-evaluation result. Based on the reassessment results, a revised planning strategy is generated that conforms to the current physical reality and task objectives.

8. The robot operation method based on a visual language model according to claim 1, characterized in that, The strategy for generating the initial planning step also includes: Construct a vectorized operation experience repository, which is used to store historical success cases; Extract the visual semantic features of the current task, and perform similarity retrieval based on the visual semantic features in the vectorized operation experience repository; Historical successful cases with similarity higher than a second preset threshold are used as target cases. Prior knowledge of the target cases is extracted, and the initial planning strategy is generated based on the prior knowledge.

9. A robot operating system based on a visual language model, characterized in that, The robot operating system based on the visual language model includes: The multimodal semantic perception and planning module is used to acquire visual images of the target scene and natural language instructions for the operation task. It uses a large visual language model to jointly parse the visual images and natural language instructions to generate an initial planning strategy. The motion execution and tactile acquisition module is used to control the robot to perform operation actions according to the initial planning strategy, and to use tactile sensors set on the robot's end effector to collect raw tactile signals in real time during the operation process; The tactile semantic translation module is used to input the original tactile signal into a pre-trained tactile language translation model for feature extraction and semantic mapping to obtain a tactile semantic description in natural language form. The visual-touch fusion judgment module is used to feed back the tactile semantic description as an enhanced prompt word to the visual language big model when the tactile semantic description indicates an abnormal operation or the physical attributes do not match the visual expectations. The closed-loop replanning module is used to re-reason about the current task scenario based on the visual image and the enhanced prompt words through the visual language big model, generate a revised planning strategy, and control the robot to perform operation actions according to the revised planning strategy.

10. A terminal, characterized in that, The terminal includes: a memory, a processor, and a robot operation program based on a visual language model stored in the memory and executable on the processor. When the robot operation program based on the visual language model is executed by the processor, it implements the steps of the robot operation method based on a visual language model as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Deformable object interactive operation control method based on visual touch-language-action multi-mode model

    CN119526422A

  • Digital exhibition product display method based on deep learning

    CN120215704A

  • Multi-mode-based body robot control method and device, electronic equipment, readable storage medium and program product

    CN121043159A

Cited By

  • Prompt word optimization method and device, electronic equipment and storage medium

    CN121920385A

  • Robot vision positioning calibration method and system based on dexterous hand touch sense

    CN122008256A