A method and apparatus for deformable object manipulation based on vision and language models
Patent Information
- Application Number
- CN202411640054.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-11-18
AI Technical Summary
此外,现有技术中的机器人操作多依赖于预编程的动作,缺乏灵活应对实时变化情况的能力
[0033](1)本发明通过整合先进的视觉处理技术和CLIP语言模型,利用视觉处理技术捕捉操作环境的视觉数据,并通过深度学习分析该数据以识别物体的空间位置及形态变化;利用CLIP语言模型解析操作任务的自然语言指令,提取出具体的动作要求和操作目标;将识别的视觉信息与语言指令的解析结果经过深度融合,生成精准详细的操作策略,从而显著提高了机器人在复杂环境中执行拾取、移动和放置任务的能力,提高了机器人处理柔性物体的精度与效率,为智能制造、家居和服务等领域的应用提供了有效的技术支持。
Smart Images

Figure CN119501933B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deformable object manipulation technology, and in particular to a method and apparatus for manipulating deformable objects based on vision and language models. Background Technology
[0002] In the fields of intelligent robotics and automation, the demand for handling fabrics, ropes, and other flexible materials continues to grow, particularly in industries such as apparel manufacturing, healthcare, home furnishings, and services. The flexibility and deformability of these materials pose significant challenges to traditional automation and robotics technologies. Conventional robotic systems are typically designed to handle rigid, fixed-shape objects, relying on precise robotic arms and pre-programmed procedures to perform tasks. However, the unpredictable shapes and variable physical properties of flexible objects demand that robotic systems possess more advanced perception and adaptation capabilities to achieve high-precision manipulation.
[0003] Traditional methods typically involve using machine vision systems to identify the position and orientation of objects. However, these systems often struggle to accurately recognize and process the complexity of visual information caused by irregular shapes. For example, automatically folding clothing on a garment production line or flattening gauze in surgery requires robots to understand the specific state and needs of the objects in order to adjust their operational strategies. Furthermore, existing robot operations largely rely on pre-programmed actions, lacking the ability to flexibly respond to real-time changes. In addition, although natural language processing technology has been introduced into robot control systems to provide a more intuitive way to input commands, most existing systems have failed to effectively integrate language commands with machine vision data. This results in robots still exhibiting limitations in understanding and execution when performing tasks involving complex language descriptions. For example, when the instruction involves "precisely folding a corner on an irregularly shaped piece of fabric," relying solely on language or visual information often makes it difficult to complete the precise operation.
[0004] Therefore, there is an urgent need to develop a new robot manipulation technology that can effectively combine advanced vision processing capabilities and deep language understanding to improve the accuracy and efficiency of handling flexible objects. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the prior art by providing a method and apparatus for manipulating deformable objects based on vision and language models, thereby improving the accuracy and efficiency of handling flexible objects.
[0006] The objective of this invention can be achieved through the following technical solutions:
[0007] A method for manipulating deformable objects based on vision and language models includes the following steps:
[0008] An operating environment is set up for the deformable object to be manipulated, which includes a camera for acquiring visual data of the deformable object and a robotic arm for manipulating the deformable object.
[0009] In the operating environment, visual information of deformable objects is captured to obtain image data, and corresponding language commands are recorded.
[0010] Based on the language instructions, the language model extracts key actions and target objects to generate operation instructions; the visual processing model extracts spatial features from the image data; the spatial features and operation instructions are fused, and the final operation strategy is generated through machine learning algorithms.
[0011] The operation strategy is converted into execution instructions for the robotic arm to manipulate deformable objects;
[0012] The training process for the language model, visual processing model, and machine learning algorithm includes:
[0013] The process involves simulating the motion and deformation of deformable objects to generate simulation training data; capturing visual information of deformable objects in a real environment and recording corresponding language commands to generate real training data; combining the simulation training data with the real training data to form a sample set including image data and language commands; and training the language model, visual processing model, and machine learning algorithm based on the sample set.
[0014] Furthermore, the operating environment also includes an operating platform, on which the deformable object to be manipulated is located, and multiple cameras are located around the operating platform to capture visual information of the deformable object from different angles.
[0015] The robotic arm's workspace covers the entire operating platform.
[0016] Furthermore, the language model is a CLIP language model, which takes natural language-based language instructions as input and outputs the task's operation objectives and steps to extract key actions and target objects.
[0017] Furthermore, the visual processing model is a Transformer network;
[0018] The camera is an RGB-D camera, and the extracted image data includes RGB images and corresponding depth images. The visual processing model takes RGB images and corresponding depth images as input and outputs the spatial features of the object.
[0019] Furthermore, the machine learning algorithm is a deep reinforcement learning algorithm, and the generation expression of the operation policy is:
[0020] S=f(F i Task
[0021] In the formula, S represents the operation strategy, and F... i The spatial features are defined as Task, the operation instructions are defined as f, and the policy generation function is defined as f. The algorithm is optimized and adjusted based on the deep reinforcement learning algorithm.
[0022] Furthermore, the motion planning of the robotic arm follows the following model:
[0023] q t+1 =q t +Δq
[0024] In the formula, q t Let q be the joint angle vector of the robotic arm at time t, and Δq be the change in joint angle under the current operation strategy. t+1 Let be the joint angle vector of the robotic arm at time t+2.
[0025] Furthermore, during the manipulation of deformable objects, the robotic arm monitors its execution process in real time via a camera and dynamically adjusts its operation strategy based on feedback.
[0026] Furthermore, the deformable object includes fabric and rope.
[0027] Furthermore, the robotic arm performs pick-up, move, and / or place operations based on an operational strategy.
[0028] The present invention also provides a deformable object manipulation device based on vision and language models, comprising:
[0029] A camera is used to collect visual data from deformable objects.
[0030] A robotic arm is used to manipulate deformable objects;
[0031] The processor, connected to both the camera and the robotic arm, is used to execute the steps described above.
[0032] Compared with the prior art, the present invention has the following advantages:
[0033] (1) This invention integrates advanced visual processing technology and CLIP language model. It uses visual processing technology to capture visual data of the operating environment and analyzes the data through deep learning to identify the spatial position and shape changes of objects. It uses CLIP language model to parse the natural language instructions of the operation task and extract the specific action requirements and operation goals. It deeply integrates the identified visual information with the parsing results of the language instructions to generate accurate and detailed operation strategies, thereby significantly improving the robot's ability to perform picking, moving and placing tasks in complex environments, improving the accuracy and efficiency of the robot in handling flexible objects, and providing effective technical support for applications in intelligent manufacturing, home and service fields.
[0034] (2) This invention simultaneously collects and processes data from both the simulation environment and the real environment. By combining the simulation data with the data collected from the real environment, the gap between the simulation and the real environment can be narrowed, and the optimization performance of the visual processing and natural language processing models can be improved.
[0035] (3) The CLIP language model of this invention can accurately understand and transform complex language expressions through multimodal deep learning capabilities; the operation strategy generation adopts Transformer... r The network combines the local perception capabilities of convolutional layers with Transformer... r Its global context capture capability can effectively extract detailed spatial features of deformable objects (such as cloth). Attached Figure Description
[0036] Figure 1 This is a flowchart illustrating a method for manipulating deformable objects based on a vision and language model, as provided in an embodiment of the present invention.
[0037] Figure 2 This is a schematic diagram illustrating offline training and real-world prediction of a fabric according to an embodiment of the present invention;
[0038] Figure 3 This is a schematic diagram of the state of an operating environment provided in an embodiment of the present invention. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0040] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0041] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0042] Example 1
[0043] like Figure 1 As shown, this embodiment provides a method for manipulating deformable objects based on vision and language models, including the following steps:
[0044] S1: Establish an operating environment for the deformable object to be manipulated. The operating environment includes a camera for acquiring visual data of the deformable object and a robotic arm for manipulating the deformable object.
[0045] S2: As Figure 2 As shown, a deformable object is simulated to simulate its motion and deformation, generating simulation training data; visual information of deformable objects in the real environment is captured and corresponding language commands are recorded to generate real training data; the simulation training data and real training data are combined to form a sample set including image data and language commands.
[0046] S3: Based on language instructions, extract key actions and target objects through the language model to generate operation instructions;
[0047] S4: Extract spatial features from image data using a visual processing model; fuse spatial features with operation instructions; and generate the final operation strategy using machine learning algorithms.
[0048] S5: Transforms the operation strategy into execution instructions for the robotic arm to manipulate deformable objects.
[0049] Specifically, in step S1, such as Figure 3 As shown, the operating environment also includes an operating platform, on which the deformable object to be manipulated is located. There are multiple cameras located around the operating platform to capture visual information of the deformable object from different angles.
[0050] The robotic arm's workspace covers the entire operating platform.
[0051] In this embodiment, the setup of the operating environment includes installing two RGB-D cameras, positioned opposite each other on the operating platform, to capture visual data of the operating scene from different angles, improving the accuracy and reliability of visual recognition. At least one precision robotic arm is installed to perform complex object manipulation tasks; and a 1m x 1m operating platform is used to provide sufficient workspace to ensure the flexibility and efficiency of the robotic arm operation. The hardware configuration of the operating environment ensures that the system can capture high-quality visual data and provide precise robotic arm control. The image data captured by the RGB-D cameras can be represented as:
[0052] (I i D i ), i = 1, 2, ..., N
[0053] Among them, I i It is an RGB image, D i It is the corresponding depth image.
[0054] The software configuration includes installing an operating system to support the basic operation of the robot system, deploying a vision processing program to process image data acquired from RGB-D cameras to identify the shape and position of objects; deploying the CLIP language model to parse natural language instructions and translate them into specific action instructions that the robotic arm can execute; and tools for collecting and analyzing operational data to support real-time feedback and subsequent optimization of the system.
[0055] In step S2, PyFlex software is used in a simulation environment to simulate the physical properties of fabric and other deformable objects, collecting language commands and visual data to generate high-quality training data. This data is combined with data collected in the real environment to narrow the gap between simulation and reality, optimizing the performance of visual processing and natural language processing models. In the real environment, an RGB-D camera system captures visual information of the operation scene and records the corresponding language commands. The collected dataset is processed to form a sample set. This sample set contains image data (I i D i ) and corresponding language instructions The corresponding expression is:
[0056]
[0057] in, These are the corresponding natural language instructions.
[0058] In step S3, the language model is the CLIP language model. The input of the CLIP language model is language instructions based on natural language, and the output is the operation goal and steps of the task, so as to extract key actions and target objects.
[0059] In other words, the CLIP language model, through its multimodal understanding capabilities, translates natural language instructions... Transforming complex natural language expressions into specific operational goals and steps, extracting key actions and target objects, and ensuring that the system can accurately parse complex natural language expressions and transform them into operational instructions that the robot can execute.
[0060] The CLIP model, through multimodal deep learning capabilities, accurately understands and transforms complex language expressions, recognizing key actions and target objects (such as "grasp," "move," and "place") in instructions, along with their corresponding positional and attribute descriptions. Using the CLIP language model, the system can generate precise operational strategies, optimizing the accuracy and efficiency of task execution. Furthermore, the CLIP language model enhances the system's responsiveness to complex and ambiguous instructions, significantly improving the robot's operational flexibility and adaptability. The application of this language processing technology significantly improves the robot's performance in tasks involving deformable objects, providing advanced language understanding support for intelligent robotics technology.
[0061] That is, the input to the CLIP language model is natural language instructions. The output consists of the specific operational objectives and steps of the task:
[0062]
[0063] In step S4, the visual processing model uses a lightweight Transformer network (such as MobileViT) to process the RGB-D image and extract the spatial features F of the object. i :
[0064] F i =MobileViT(I i D i )
[0065] These features are fused with the instruction results parsed by the CLIP language model to generate the final operation strategy S:
[0066] S=f(F i Task
[0067] Here, f is the policy generation function, which is optimized and adjusted using a deep reinforcement learning algorithm.
[0068] The policy generation function f is optimized and adjusted using machine learning algorithms (such as deep reinforcement learning) to ensure the accuracy and effectiveness of the operation.
[0069] The Transformer network combines the local perception capabilities of convolutional layers with the global context capture capabilities of the Transformer to effectively extract detailed spatial features of deformable objects (such as fabric), including the object's position, shape, size, and possible deformation states. Simultaneously, this method employs machine learning algorithms, particularly deep reinforcement learning, to analyze and optimize the fused visual and language parsing results. These algorithms learn from historical operational data, automatically adjusting and optimizing operational parameters and strategies to more accurately execute complex tasks, and dynamically adjusting operational strategies based on real-time feedback, thereby ensuring both efficiency and accuracy.
[0070] In step S5, the robotic arm is guided to perform actions including picking, moving, and placing according to the generated operation strategy. The operation strategy is converted into execution instructions for the robotic arm through the controller module, and the motion planning of the robotic arm follows the following model:
[0071] q t+1 =q t +Δq
[0072] In the formula, qt is the joint angle vector of the robotic arm at time t, Δq is the change in joint angle under the current operation strategy, and q t+1 Let be the joint angle vector of the robotic arm at time t+2.
[0073] Preferably, the system monitors the execution process in real time and dynamically adjusts the operation strategy based on feedback to ensure efficient operation results.
[0074] This step ensures that the robotic arm can accurately follow the strategy generated based on visual data analysis and CLIP language model parsing results, performing tasks involving complex spatial positioning and object shape adjustments. The system monitors the execution process in real time and adjusts the robot's operations accordingly based on feedback to adapt to real-time changes in the object's state, thereby achieving high-precision operation.
[0075] The above is an introduction to the method embodiments. The following describes the present invention further through device embodiments.
[0076] The present invention also provides a deformable object manipulation device based on vision and language models, comprising:
[0077] A camera, used to collect visual data of deformable objects, preferably an RGB-D camera;
[0078] A robotic arm is used to manipulate deformable objects;
[0079] The processor, connected to both the camera and the robotic arm, is used to execute the steps of the vision and language model-based deformable object manipulation method described above.
[0080] It should be noted that the specific details and beneficial effects of the device in this application can be found in the above-described method embodiments, and will not be repeated here.
[0081] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A method for manipulating deformable objects based on vision and language models, characterized in that, Includes the following steps: An operating environment is set up for the deformable object to be manipulated, which includes a camera for acquiring visual data of the deformable object and a robotic arm for manipulating the deformable object. In the operating environment, visual information of deformable objects is captured to obtain image data, and corresponding language commands are recorded. Based on the language instructions, the language model extracts key actions and target objects to generate operation instructions. Spatial features are extracted from the image data using a visual processing model; the spatial features and operation instructions are fused together, and a final operation strategy is generated using a machine learning algorithm. The operation strategy is converted into execution instructions for the robotic arm to manipulate deformable objects; The training process for the language model, visual processing model, and machine learning algorithm includes: Simulate deformable objects to simulate their motion and deformation, and generate simulation training data. Visual information of deformable objects in a real environment is captured and corresponding language instructions are recorded to generate real training data; the simulation training data is combined with the real training data to form a sample set including image data and language instructions; the language model, visual processing model and machine learning algorithm are trained based on the sample set. The language model is the CLIP language model. The input of the CLIP language model is language instructions based on natural language, and the output is the operation goal and steps of the task, so as to extract key actions and target objects. The visual processing model is a Transformer network; The camera is an RGB-D camera, and the extracted image data includes RGB images and corresponding depth images. The input of the visual processing model is the RGB image and the corresponding depth image, and the output is the spatial features of the object. The operating environment also includes an operating platform, on which the deformable object to be manipulated is located. There are multiple cameras located around the operating platform to capture visual information of the deformable object from different angles. The workspace of the robotic arm covers the entire operating platform; The machine learning algorithm is a deep reinforcement learning algorithm, and the generative expression of the operation policy is: In the formula, As an operational strategy, For spatial features, For operation instructions, The policy generation function is optimized and adjusted based on the deep reinforcement learning algorithm. The deformable objects include fabric and rope.
2. The method for manipulating deformable objects based on vision and language models according to claim 1, characterized in that, The motion planning of the robotic arm follows the following model: In the formula, Let be the joint angle vector of the robotic arm at time t. This represents the change in joint angle under the current operating strategy. Let be the joint angle vector of the robotic arm at time t+1.
3. The method for manipulating deformable objects based on vision and language models according to claim 2, characterized in that, During the manipulation of deformable objects, the robotic arm monitors its execution process in real time via a camera and dynamically adjusts its operation strategy based on feedback.
4. The method for manipulating deformable objects based on vision and language models according to claim 1, characterized in that, The robotic arm performs pick-up, move, and / or place operations based on an operational strategy.
5. A deformable object manipulation device based on vision and language models, characterized in that, include: A camera is used to collect visual data from deformable objects. A robotic arm is used to manipulate deformable objects; The processor, connected to both the camera and the robotic arm, is used to execute the steps of the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Robot control method and device based on visual language pre-training model and medium
CN115933387A
Visual language navigation method combining image description and text generation image
CN117571014A
Robot manipulation method based on visual language large model
CN118559711A