Deformable Object Shape Control Method Based on Visual Language Model and Historical Data Learning
By introducing visual language models and historical data learning into the shape control method of deformable object, the problems of multimodal information fusion and insufficient utilization of historical data in the existing methods are solved, and efficient processing and precise control of complex tasks are achieved.
Patent Information
- Application Number
- CN202510293736.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The existing deformable object shape control methods lack the ability to fusion multi-modal object information and the ability to learn historical data, making it difficult to deal with complex multi-step tasks.
Using a method based on visual language model and historical data learning, we use parameterized polygon models and historical databases to plan and optimize actions by obtaining user language target instructions and visual sub-object sequences, decomposing language targets and combining visual targets.
It realizes the effective fusion of multimodal target information, uses historical data to improve task performance, can handle complex deformable object shape control tasks, and improves the accuracy and stability of task execution.
Smart Images

Figure CN119820579B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent robot operation, and particularly to a deformable object shape control method based on a vision-language model and historical data learning. Background Art
[0002] Currently, the shape control of deformable objects is an important task in the field of intelligent robot operation, and has broad application value and application prospects in both daily life and industrial production. For example, in a home environment, an assistive robot can help humans complete tasks such as sorting clothes; in an industrial assembly line, deformable object shape control technology can be applied to multiple links such as assembly and packing. This requires effective perception of deformable objects and planning of reasonable deformable object shape control actions according to different types and modalities of goals provided by humans.
[0003] Existing deformable object shape control methods can be divided into vision-based target methods and language-based target methods. The vision-based target method uses a sequence of visual sub-goals pre-provided by human experts as a reference and guidance to calculate the action sequence required to achieve the target shape of the deformable object. The language-based target method uses the language target instructions provided by human users as a guidance, and by decomposing the language instructions, gradually calculates the action sequence required to reach the target shape of the deformable object. However, the existing methods have the following problems:
[0004] 1) Lack of the ability to fuse multi-modal target information: The target in the visual modality can provide fine-grained pixel-level guidance during the deformable object shape control process, and the target in the language modality can provide rich semantic information guidance. The existing methods only rely on single-modal target information and do not fully utilize and effectively fuse multi-modal target information.
[0005] 2) Lack of the ability to learn from historical data: During the execution of the method, a large amount of historical execution data will accumulate over time. The existing methods do not have the ability to retrieve, evaluate, and learn from historical data, and do not fully utilize valuable historical data to help improve the performance of the current task.
[0006] 3) Poor performance in complex tasks: The existing methods can only handle relatively simple deformable object shape control tasks, such as square cloth, and perform poorly when facing more complex deformable objects and more complex multi-step tasks. Summary of the Invention
[0007] In view of the problems existing in the prior art, the present invention provides a deformable object shape control method based on a vision-language model and historical data learning. The specific technical solutions adopted by the present invention are as follows:
[0008] The present invention discloses a deformable object shape control method based on visual language model and historical data learning, including:
[0009] Obtain historical execution data, preprocess the data and construct a historical database;
[0010] Obtain the user's language target instruction and visual sub-goal sequence;
[0011] Process the user's language target instruction through the visual language model and decompose it into a language sub-goal sequence;
[0012] For each language sub-goal and its corresponding visual sub-goal in the language sub-goal sequence, loop through the following steps until all sub-goals are completed;
[0013] Obtain the image of the current deformable object and fit the corresponding parameterized polygon model;
[0014] Fit the corresponding parameterized polygon model according to the current visual demonstration sub-goal;
[0015] Calculate the task matching score between each data in the historical database and the current task;
[0016] Calculate the object matching score between each data in the historical database and the current object;
[0017] Screen out a specific number of historical data from the historical database according to the task matching score and the object matching score;
[0018] Calculate the folding symmetry axis required for the deformable object shape control action through the visual language model;
[0019] Calculate the first grasping point, the first placement point, the second grasping point and the second placement point of the shape control action through the operability optimization algorithm;
[0020] Execute the shape control action of the deformable object.
[0021] As a further improvement, the preprocessing of the data and the construction of the historical database in the present invention specifically include:
[0022] The historical execution data of the deformable object shape control is divided into historical data based on language goals and historical data based on visual goals according to the target type;
[0023] For each piece of historical data based on language goals, the preprocessing process includes: using a pre-trained text embedding encoder for the user's language instruction and language sub-goal in the historical data based on language goals, and mapping them to a dimension of The eigenvector, while fitting the corresponding parameterized polygon model according to the images of the deformable object before and after deformation in the historical data, and scoring the execution effect of the historical data through the vision-language model, denoted as , combining the historical data, the eigenvector, the parameterized polygon model, and the evaluation score obtained through preprocessing to construct a historical database based on the language objective;
[0024] For each piece of historical data based on the vision objective, fitting the parameterized polygon models of the deformable object before and after shape change in the historical data and the parameterized polygon model of the vision demonstration sub-objective sequence in the historical data. In addition, the calculation process of the evaluation score of the historical data based on the language objective is the same. Similarly, using the vision-language model to score the execution effect of the historical data, and combining the historical data, the parameterized polygon model, and the evaluation score obtained through preprocessing to construct a historical database based on the vision objective;
[0025] Integrating the historical databases based on the language objective and the vision objective to construct a historical database.
[0026] As a further improvement, the processing of the user's language objective instruction by the vision-language model in the present invention, which decomposes it into a sequence of language sub-objectives, specifically includes:
[0027] Inputting the user's language objective instruction and the vision demonstration sub-objective sequence into the vision-language model, analyzing the user's language objective instruction at the task level through the vision-language model, and decomposing the user's language instruction into a series of language sub-objectives. Each language sub-objective represents an independent single-step task and corresponds one-to-one with the vision demonstration sub-objective sequence.
[0028] As a further improvement, the calculation of the task matching score for each data in the historical database and the current task in the present invention specifically includes:
[0029] For the historical data based on the language objective in the historical database, if the user's language objective instruction of a set of data is , and the language sub-objective is , and the corresponding eigenvectors are respectively and ; for the user's language objective instruction and the current language sub-objective in the current task, as well as the corresponding eigenvectors and , calculating the task matching score , and the formula is as follows:
[0030] ;
[0031] where Represents the cosine similarity between two feature vectors, is the evaluation score calculated by the vision-language model, is the weighting coefficient;
[0032] For the historical data based on visual objects in the historical database, if the parameterized polygon models corresponding to the visual demonstration sub-goals of the two frames before and after the current step in a set of data are and , and the parameterized polygon models corresponding to the visual demonstration sub-goals of the two frames before and after the current step of the current task are and ; Normalize the coordinates of all polygon models to the range [0, 1], and calculate the task matching score :
[0033] ;
[0034] where is the evaluation score calculated by the vision-language model, is the weighting coefficient, is the Euclidean distance between polygon models.
[0035] As a further improvement, calculating the object matching score between each data in the historical database and the current object in the present invention specifically includes:
[0036] If the image before the shape change of the deformable object in a set of historical data is , and the corresponding parameterized polygon model is ; If the parameterized polygon model of the current deformable object is ; Rotate the parameterized polygon model and the parameterized polygon model to the standard orthogonal position, and calculate the Euclidean distance between the polygon models as the object matching score , which can be expressed by the formula as follows:
[0037] .
[0038] As a further improvement, calculating the folding symmetry axis required for the shape control action of the deformable object by the vision-language model in the present invention specifically includes:
[0039] If the current language sub-goal is , the visual demonstration sub-goal is , the parameterized polygon model corresponding to the visual demonstration sub-goal is , the current image is , and the corresponding parameterized polygon model is , selected from the historical database A set of historical data based on the language objective is , A set of historical data based on the visual objective is . Taking this as the input of multi-modal, through the visual language model Perform inference calculation on the current task to generate the folding symmetry axis required for the shape control action of the deformable object , which is expressed by the formula as follows:
[0040] .
[0041] As a further improvement, the operability optimization algorithm described in the present invention specifically includes:
[0042] The operability optimization algorithm determines the first grasping point and the second grasping point by maximizing the weighted sum of the quadrilateral area and the distance from the point to the folding symmetry axis. The specific steps are as follows: If the first grasping point at the current step , the second grasping point , the starting point of the folding symmetry axis is , and the ending point is ; calculate , , and The area of the quadrilateral enclosed by ; calculate the distance from the first grasping point to the folding symmetry axis ; calculate the distance from the second grasping point to the folding symmetry axis , and the specific formula is as follows:
[0043]
[0044] ;
[0045] where is the area of the bounding box of the deformable object, is the weighting coefficient, is the object contour extracted from the deformable object image. The constraint conditions are that the first grasping point and the second grasping point are located on the object contour, and the first grasping point and the second grasping point are located on the left side of the folding symmetry axis; the first placement point is the symmetric point of the first grasping point with respect to the folding symmetry axis, and the second placement point is the symmetric point of the second grasping point with respect to the folding symmetry axis.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] 1) The present invention has a powerful ability to fuse multi-modal target information, and can comprehensively utilize visual targets and language targets to calculate deformable object shape control actions. It combines the guidance at the microscopic pixel level in visual targets and uses the information in language targets for macro-semantic level task planning and logical reasoning. Finally, a vision-language model is used to fuse the target information of the two modalities, enabling the visual target and the language target to promote and cooperate with each other, ensuring the precise planning of the deformable object shape control task.
[0048] 2) The present invention proposes a method for representing deformable objects based on a parameterized polygon model, and uses a black-box optimization algorithm for online parameter estimation to achieve real-time tracking and updating of the object deformation state. This geometric representation of the parameterized polygon model compresses high-dimensional visual observation data into a low-dimensional parameter space, significantly reducing the dimension of the deformable object state observation and only retaining the key state information relevant to the task. At the same time, the polygon model is easily converted into a text format prompt for input into the vision-language model, facilitating the large language model to understand, analyze, and calculate the deformation state of the deformable object.
[0049] 3) The present invention proposes a method for constructing a classified and stored historical database based on historical execution data. By preprocessing the historical execution data, extracting the useful information therein, and combining with the vision-language model to evaluate the data quality, and finally classifying and storing according to the type of the target. This method can not only effectively integrate and store a large amount of historical execution data, but also greatly improve the data retrieval efficiency, making the retrieval process more efficient and convenient. In addition, the database has the ability to expand online, can dynamically adapt to different types of data and requirements, and ensure that the database always maintains high efficiency and flexibility during use.
[0050] 4) The present invention proposes a two-layer retrieval and matching method of task matching - object matching, which can efficiently and accurately retrieve the historical data most relevant to the current task from the historical database. The task matching layer screens out the historical data most relevant to the current task level from the historical database to ensure a high degree of task relevance; the object matching layer further selects the historical data closest to the current deformable object shape from the screened data, thus ensuring the applicability of the historical data to the current object. This retrieval and matching method is not only simple in calculation and efficient in implementation, but also can screen out high-quality historical data from a huge historical dataset with extremely high efficiency, providing precise support for practical applications.
[0051] 5) The present invention can learn from historical data. By comparing the execution processes of historical data related to the current task, the vision-language model can extract valuable information from the context information, so as to accurately calculate the action representation required to complete the current deformable object shape control task. This method of learning based on historical data can continuously optimize the model performance with the accumulation of data volume and improve the accuracy of task execution. In addition, this method can better adapt to different task scenarios, and can dynamically adjust according to the changes in historical data, so as to ensure continuous and stable performance.
[0052] 6) By introducing a vision-language model and leveraging its powerful ability in common sense understanding, the present invention significantly improves the performance on unseen tasks or unseen objects, ensuring the efficiency and reliability of the system in a changing environment.
[0053] 7) The present invention proposes an operability optimization algorithm, which can effectively avoid action conflicts that may occur during the execution of deformable object shape control actions, ensuring the coordination and stability of actions. At the same time, the optimization algorithm can also minimize the internal deformation that occurs during the shape control process, ensuring that the actual obtained object shape is highly consistent with the target shape. This not only improves the operation accuracy of the system, but also enhances the stability of the execution process, enabling more accurate and reliable shape control in complex tasks and dynamic environments.
[0054] 8) The present invention performs excellently in complex deformable object shape control tasks and can handle various shape control tasks of various types of deformable objects. The present invention can flexibly adjust the shape control actions according to specific task requirements, ensuring the accuracy and efficiency of the shape control process. At the same time, the multi-task processing ability of the present invention enables it to adapt to various scenarios and requirements, providing targeted operation solutions for different types of deformable objects, and greatly enhancing the applicability and flexibility of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is the algorithm flowchart of a deformable object shape control method based on a vision-language model and historical data learning according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The present invention discloses a deformable object shape control method based on a vision-language model and historical data learning, Figure 1 is the algorithm flowchart of a deformable object shape control method based on a vision-language model and historical data learning according to the present invention. The method includes:
[0057] Obtain the historical execution data of the deformable object shape control, preprocess the data, and construct a historical database for classified storage; the historical execution data of the deformable object shape control is divided into historical data based on language goals and historical data based on visual goals according to the target types;
[0058] For each piece of historical data based on language goals, the preprocessing process includes: using a pre-trained text embedding encoder for the user language instructions and language sub-goals in the historical data based on language goals, mapping them to feature vectors with a dimension of , and at the same time fitting the corresponding parameterized polygon model according to the images of the deformable object before and after deformation in the historical data, scoring the execution effect of the historical data through a vision-language model, and the evaluation score takes values in a discrete range of three values: 1 represents that the execution effect is very good, 0 represents that the execution effect is average, and -1 represents that the execution effect is poor. The inputs of the vision-language model are the language sub-goals , the images of the deformable object before and after shape change and , the parameterized polygon models of the deformable object before and after shape change and , the folding symmetry axis , the two-arm grasping and placing actions , calculate the evaluation score through the vision-language model , merge the historical data, the obtained feature vectors, parameterized polygon models, and evaluation scores, and construct a historical database based on language goals;
[0059] For each piece of historical data based on visual goals, fit the parameterized polygon models of the deformable object before and after shape change in the historical data and the parameterized polygon models of the visual demonstration sub-goal sequences in the historical data. In addition, the process of calculating the evaluation score of the historical data based on language goals is the same. Also use the vision-language model to score the execution effect of the historical data, and merge the historical data, the obtained parameterized polygon models, and evaluation scores to construct a historical database based on visual goals;
[0060] Integrate the historical databases based on language goals and visual goals to construct a historical database for classified storage.
[0061] Obtain the current task user language goal instruction and the visual demonstration sub-goal sequence;
[0062] The current task user language target instruction is processed by a vision - language model, and decomposed into a language sub - target sequence that corresponds one - to - one with the vision demonstration sub - target sequence; the user language target instruction and the vision demonstration sub - target sequence are input into the vision - language model, and the vision - language model conducts a task - level analysis of the user language target instruction, decomposing the user language instruction into a series of language sub - targets. Each language sub - target represents an independent single - step sub - task and corresponds one - to - one with the vision demonstration sub - target sequence.
[0063] For each language sub - target in the language sub - target sequence and its corresponding vision demonstration sub - target, the following steps are executed in a loop until all sub - targets in the sequence are completed:
[0064] Obtain the image of the current deformable object;
[0065] Fit the parameterized polygon model of the current deformable object according to the image of the current deformable object;
[0066] Fit the parameterized polygon model of the vision demonstration sub - target according to the current vision demonstration sub - target;
[0067] Calculate the task matching score between each data in the historical database and the current task;
[0068] For the historical data based on the language target in the historical database, if the user language target instruction of a set of data is , the language sub - target is , and the corresponding feature vectors are and respectively; for the user language target instruction and the current language sub - target in the current task, as well as the corresponding feature vectors and , calculate the task matching score , and the formula is as follows:
[0069] ;
[0070] where represents the cosine similarity between two feature vectors, is the evaluation score calculated by the vision - language model, is the weighting coefficient;
[0071] For the historical data based on the vision target in the historical database, if the parameterized polygon models corresponding to the vision demonstration sub - targets of two consecutive frames in the current step of a set of data are and , and the parameterized polygon models corresponding to the vision demonstration sub - targets of two consecutive frames in the current step of the current task are and Normalize the coordinates of all polygon models to the range [0, 1] and calculate the task matching score :
[0072] ;
[0073] where is the evaluation score calculated by the vision-language model, is the weighting coefficient, is the Euclidean distance between polygon models, and the calculation method is as follows: First, determine whether their levels are the same, and verify the number of nodes and node labels in each level. If there are inconsistent situations, then , indicating that the two polygon models cannot be matched; for the case where they can be matched, the Euclidean distance of the polygon models is defined as the sum of the Euclidean distances of each layer, where the Euclidean distance of each layer is defined as the sum of the distances between all nodes with the same label in that layer.
[0074] Calculate the object matching score for each data in the historical database and the current deformable object;
[0075] If the image before the shape change of the deformable object in a set of data in the historical data is , and the corresponding parameterized polygon model is ; if the parameterized polygon model of the current deformable object is ; Rotate the parameterized polygon model and the parameterized polygon model to the standard orthogonal position and calculate the Euclidean distance between the polygon models as the object matching score , which can be expressed by the formula as follows:
[0076] .
[0077] Filter out a specific number of historical data from the historical database according to the task matching score and the object matching score;
[0078] Based on the current language sub-goal, the visual demonstration sub-goal and the parameterized polygon model of the visual demonstration sub-goal, the selected historical data, the current deformable object image and the parameterized polygon model of the current deformable object, calculate the folding symmetry axis required for the shape control action of the deformable object through the vision-language model; if the current language sub-goal is , the visual demonstration sub-goal is , the parameterized polygon model corresponding to the visual demonstration sub-goal is , and the current image is , the corresponding parametric polygon model is , obtained by screening from the historical database groups of historical data based on the language objective are , groups of historical data based on the visual objective are , taking this as the multi-modal input, through the vision-language model to perform inference calculation on the current task, generating the folding symmetry axis required for the deformable object shape control action , which is expressed by the formula as follows:
[0079] .
[0080] According to the folding symmetry axis and the image of the current deformable object, calculate the first grasping point, the first placement point, the second grasping point and the second placement point of the shape control action through the manipulability optimization algorithm; the manipulability optimization algorithm determines the first grasping point and the second grasping point by maximizing the weighted sum of the quadrilateral area and the distance from the point to the folding symmetry axis. The specific steps are as follows: if the first grasping point at the current step , the second grasping point , the starting point of the folding symmetry axis is , and the ending point is ; calculate , , and to enclose the quadrilateral area ; calculate the distance from the first grasping point to the folding symmetry axis ; calculate the distance from the second grasping point to the folding symmetry axis , and the specific formula is as follows:
[0081]
[0082] ;
[0083] where is the area of the bounding box of the deformable object, is the weighting coefficient, is the object contour extracted from the deformable object image. The constraint conditions are that the first grasping point and the second grasping point are located on the object contour, and the first grasping point and the second grasping point are located on the left side of the folding symmetry axis; the first placement point is the symmetric point of the first grasping point with respect to the folding symmetry axis, and the second placement point is the symmetric point of the second grasping point with respect to the folding symmetry axis.
[0084] Execute the shape control action of the deformable object according to the first grasping point, the first placement point, the second grasping point, and the second placement point.
[0085] The specific implementation method of the present invention is as follows:
[0086] Step 1: Obtain the historical execution data of the shape control of the deformable object. The historical data may include the execution data of various shape control tasks of various deformable objects such as square cloth, shirt, and trousers. For example, the data of tasks such as folding the square cloth in half twice along the opposite sides, folding the shirt in half, and folding the trouser legs of the trousers upwards. Preprocess the data and construct a historical database stored by classification, and divide it into a historical database based on language targets and a database based on visual targets according to the types of targets.
[0087] Step 2: Obtain the current task user language target instruction and the visual demonstration sub-target sequence. For example, the user language target instruction is "fold the trousers into a rectangular square", and the visual demonstration sub-target sequence includes images of the trousers in the flat state, the trousers folded from left to right, and the trousers folded from bottom to top.
[0088] Step 3: Process the current task user language target instruction through a visual language model, and decompose it into a language sub-target sequence corresponding one by one to the visual demonstration sub-target sequence. The decomposed language sub-target sequence is:
[0089] Sub-target 1: Fold the trousers in half from left to right;
[0090] Sub-target 2: Fold the folded trousers in half from bottom to top.
[0091] For each language sub-target and its corresponding visual demonstration sub-target, loop through the following steps until all sub-targets in the sequence are completed. For example, for sub-target 1 "Fold the trousers in half from left to right" and its corresponding visual demonstration sub-target, execute all the following steps:
[0092] Step 4: Obtain the image of the current deformable object through the camera. In this embodiment, an Intel RealSense D435i camera is selected to obtain an RGB image, and the resolution of the image is 400*400.
[0093] Step 5: Fit the corresponding parameterized polygon model according to the image of the current deformable object. Perform edge detection on the RGB image obtained in Step 5 to extract the contour of the deformable object, which is the trousers Extract the contour of the initial state parameterized polygon model ; Uniformly sample 10 points on each side of the object contour and the initial parameterized polygon model contour to obtain the deformable object contour sampling point set respectively and the polygon model contour sampling point set ; Calculate the deformable object contour sampling points and the polygon model contour sampling points between distance , the formula is:
[0094]
[0095]
[0096] Continuously adjust the parameters of the polygon model using black-box optimization so that is minimized, and the parameterized polygon model obtained by fitting is the optimal model that best conforms to the current RGB image.
[0097] Step 6: Fit the parameterized polygon model corresponding to the current visual teaching sub-goal according to the current visual teaching sub-goal, and the fitting process is the same as that in Step 5.
[0098] Step 7: Calculate the feature vectors of the user language target instruction "Fold the pants into a rectangular square" and the current language sub-goal "Fold the pants in half from left to right" respectively through the pre-trained text embedding encoder. The text embedding encoder is selected as the pre-trained OpenAI's text-embedding-3-large model, and the generated feature vector dimension is .
[0099] Step 8: Calculate the task matching score in the historical database based on the language goal according to the feature vectors corresponding to the user language target instruction "Fold the pants into a rectangular square" and the current language sub-goal "Fold the pants in half from left to right". Suppose a set of data in the historical database based on the language goal has a user language target instruction of , and the language sub-goal is , and their corresponding feature vectors are and . For the user language target instruction and the current language sub-goal in the current task, as well as the corresponding feature vectors and , calculate the task matching score , the formula is as follows:
[0100]
[0101] where represents the cosine similarity between the two feature vectors. The higher the value, the higher the similarity between the two vectors, and thus the higher the semantic similarity. is the evaluation score of the execution effect calculated by the visual language model, is the weighting coefficient.
[0102] Suppose a set of data in the language-based historical database, the user language target instruction is "Fold the shirt into a rectangular square", and the language sub-goal is "Fold the right sleeve of the shirt from right to left and inwards". At the same time, the execution process of this historical data is successful, then .
[0103] Suppose another set of data in the language-based historical database, the user language target instruction is "Fold the trousers into a square", and the language sub-goal is "Fold the trousers from left to right". At the same time, the execution process of this historical data is successful, then . Therefore, compared with the previous set of historical data, the task matching score of this set of data is higher and will be selected preferentially.
[0104] Calculate the task matching scores for all historical data in the language target-based historical database, and sort the historical data in descending order according to the task matching scores.
[0105] Step Nine: Screen out 20 pieces of language target-based historical data from the language target-based historical database in descending order of the task matching score.
[0106] Step Ten: Calculate the object matching score with the current deformable object among the 20 pieces of language target-based historical data screened out according to the parametric polygon model of the current deformable object, which is the trousers. Suppose the image before the shape change of the deformable object in a set of the screened language target-based historical data is , and the corresponding parametric polygon model is ; Suppose the parametric polygon model of the current deformable object is . Rotate the parametric polygon model and the parametric polygon model to the standard orthonormal position, and calculate the Euclidean distance between the polygon models as the object matching score.
[0107] The Euclidean distance between the polygon models is calculated as follows: For the given polygon models and , first judge whether their levels are the same, and verify whether the number of nodes and the node labels in each level are the same. If there are inconsistent situations, then , indicating that the two polygon models cannot be matched; for the case where they can be matched, the Euclidean distance of the polygon models is defined as the cumulative sum of the Euclidean distances of each layer, where the Euclidean distance of each layer is defined as the cumulative sum of the distances between all nodes with the same label in that layer, and can be expressed by the formula:
[0108]
[0109] where is the polygon model and is the number of layers, is the number of nodes in the -th layer of the polygon model, and represent the and -th -th node in the -th layer, and is the distance between the two nodes. Therefore, the object matching score
[0110]
[0111] For the parameterized polygon model corresponding to the current pair of pants, the number of layers , the number of nodes in the first layer is , so the calculation formula of the object matching score can be specified as: .
[0112] Calculate the object matching scores for all 20 pieces of historical data filtered based on the language goal, and sort the historical data in descending order according to the object matching scores.
[0113] Step Eleven: Further screen out 4 pieces of historical data based on the language goal according to the object matching score.
[0114] Step Twelve: Calculate the task matching score in the historical database based on the visual goal according to the parameterized polygon model of the current visual sub-goal. Assume that the parameterized polygon models corresponding to the visual sub-goals of the two frames before and after the current step in a set of data in the historical database based on the visual goal are and , and the parameterized polygon models corresponding to the visual sub-goals of the two frames before and after the current step of the current task are and . Normalize the coordinates of all polygon models to the range [0, 1], and calculate the task matching score :
[0115]
[0116] where is the Euclidean distance between polygon models, is the evaluation score calculated by the vision-language model, is the weighting coefficient. The task matching scores are calculated for all historical data in the historical database based on visual targets, and the historical data are sorted in descending order according to the task matching scores.
[0117] Step Thirteen: Screen out 20 pieces of historical data based on visual targets from the historical database based on visual targets in descending order of the task matching scores.
[0118] Step Fourteen: Calculate the object matching scores between the screened historical data based on visual targets and the current deformable object according to the parameterized polygon model of the current deformable object, and sort them in descending order of the object matching scores. The calculation process of the object matching scores is exactly the same as that in Step Ten.
[0119] Step Fifteen: Further screen out 4 pieces of historical data based on visual targets according to the object matching scores.
[0120] Step Sixteen: Based on the current language sub-goal, visual demonstration sub-goal, and the corresponding parameterized polygon model, the screened historical data based on language and visual targets, the current deformable object image, and the corresponding parameterized polygon model, calculate the folding symmetry axis required for the deformable object shape control action through the vision-language model. Let the current language sub-goal be , the visual demonstration sub-goal be , the parameterized polygon model corresponding to the visual sub-goal be , the current image be , the corresponding parameterized polygon model be , and the 4 groups of historical data based on language targets screened from the historical database be , and the 4 groups of historical data based on visual targets be . Using these as multi-modal inputs, through the vision-language model perform inference calculations on the current task to generate the folding symmetry axis required for the deformable object shape control action. It is expressed by the formula as follows:
[0121]
[0122] The folding symmetry axis calculated in the current step is the vector from the midpoint of the lower left vertex and the lower right vertex of the pants to the midpoint of the upper left vertex and the upper right vertex.
[0123] Step Seventeen: According to the folding symmetry axis and the current RGB image, calculate the first grasping point, the first placing point, the second grasping point, and the second placing point of the shape control action through the manipulability optimization algorithm. The manipulability optimization algorithm determines the first grasping point and the second grasping point by maximizing the weighted sum of the quadrilateral area and the distance from the point to the folding symmetry axis. The specific steps are as follows: Let the first grasping point and the second grasping point at the current step, and the starting point of the folding symmetry axis be , and the ending point be ; calculate the quadrilateral area , , , and enclosed by ; calculate the distance from the first grasping point to the folding symmetry axis ; calculate the distance from the second grasping point to the folding symmetry axis . The specific formulas are as follows:
[0124]
[0125]
[0126] where is the area of the bounding box of the deformable object, is the weighting coefficient, is the object contour extracted from the deformable object image. The constraint conditions are that the first grasping point and the second grasping point are located on the object contour, and the first grasping point and the second grasping point are located on the left side of the folding symmetry axis. The first placing point is the symmetric point of the first grasping point with respect to the folding symmetry axis, and the second placing point is the symmetric point of the second grasping point with respect to the folding symmetry axis.
[0127] Step Eighteen: Execute the shape control action of the deformable object according to the first grasping point, the first placing point, the second grasping point, and the second placing point. The specific execution process is to simultaneously grasp the deformable object at the positions of the first grasping point and the second grasping point, move them to the positions of the first placing point and the second placing point respectively and release them. The motion trajectory is generated by the cubic spline interpolation algorithm to ensure the smoothness and accuracy of the process.
[0128] After that, enter the next sub-goal, that is, sub-goal 2 "Fold the folded pants in half from bottom to top", and loop through the process of steps four to eighteen.
[0129] The above is not a limitation of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the essence of the present invention, several changes, modifications, additions or substitutions can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A deformable object shape control method based on visual language model and historical data learning, characterized in that: include: Obtain the historical execution data of the shape control of the deformable object, pre-process the data and build a historical database for classified storage; Obtain the user's language target instructions and visual teaching sub-target sequence for the current task; The visual language model is used to process the user's language target instructions for the current task and decompose them into language sub-target sequences that correspond one-to-one to the visual teaching sub-target sequences; For each language sub-goal and its corresponding visual teaching sub-goal in the language sub-goal sequence, loop through the following steps until all sub-goals in the sequence are completed: Get the image of the current deformable object; Fitting a parameterized polygonal model of the current deformable object according to the image of the current deformable object; Fitting a parameterized polygonal model of a visual teaching sub-goal according to the current visual teaching sub-goal; Calculate the task matching score between each data in the history database and the current task; Calculate the object matching score between each data in the history database and the current deformable object; Filtering a specific amount of historical data from the historical database according to the task matching score and the object matching score; Based on the current language sub-goal, the visual teaching sub-goal and the parametric polygonal model of the visual teaching sub-goal, the filtered historical data, the current deformable object image and the parametric polygonal model of the current deformable object, the folding symmetry axis required for the shape control action of the deformable object is calculated through the visual language model; According to the folding symmetry axis and the image of the current deformable object, a first grasping point, a first placement point, a second grasping point and a second placement point of the shape control action are calculated by an operability optimization algorithm; performing a shape control action of the deformable object according to the first grasping point, the first placement point, the second grasping point, and the second placement point; The preprocessing of data and construction of a historical database for classified storage specifically includes: The deformable object shape control historical execution data is divided into historical data based on language targets and historical data based on visual targets according to the target type; For each piece of historical data based on a language target, the preprocessing process includes: using a pre-trained text embedding encoder to map the user language instructions and language sub-targets in the historical data based on the language target into a feature vector with a dimension of H, and fitting the corresponding parameterized polygonal model according to the images of the deformable objects before and after deformation in the historical data, and scoring the execution effect of the historical data through the visual language model. The evaluation score ranges from three discrete values: 1 represents a very good execution effect, 0 represents a general execution effect, and -1 represents a poor execution effect. The input of the visual language model is the language sub-target L t , images of deformable objects before and after shape change and Parametric polygonal model of a deformable object before and after shape change and Folding symmetry axis F t , double-arm grab-place action a t , calculating the evaluation score score through the visual language model, merging the historical data and the preprocessed feature vector, parameterized polygon model and the evaluation score to construct a historical database based on the language target; For each piece of historical data based on visual targets, the parameterized polygonal model of the deformable object before and after the shape change in the historical data and the parameterized polygonal model of the visual teaching sub-target sequence in the historical data are fitted. In addition, the evaluation score calculation process of the historical data based on language targets is the same, and the visual language model is also used to score the execution effect of the historical data. The historical data and the parameterized polygonal model obtained by preprocessing and the evaluation score are merged to construct a historical database based on visual targets; Combining the historical database based on language goals and visual goals, constructing the classified storage historical database; The method of calculating the folding symmetry axis required for the shape control action of the deformable object through the visual language model specifically includes: If the current language sub-goal is L t , the visual teaching sub-goal is The parameterized polygonal model corresponding to the visual teaching sub-goal is: The current image is The corresponding parameterized polygonal model is The K groups of historical data based on language targets are obtained by screening in the historical database: K groups of historical data based on visual targets are Using this as the multimodal input, the visual language model VLM is used to perform reasoning calculations on the current task to generate the folding symmetry axis F required for the shape control action of the deformable object. t , which can be expressed as follows: The operability optimization algorithm specifically includes: The operability optimization algorithm determines the first grasping point and the second grasping point by maximizing the weighted sum of the area of the quadrilateral and the distance from the point to the folding symmetry axis. The specific steps are as follows: if the first grasping point Second grip point The folding symmetry axis F t The starting point is f t1 , the end point is f t2 ;calculate f t1 and f t2 The area of the quadrilateral S is enclosed; Calculate the first grasping point To the folding symmetry axis F t Distance l1; Calculate the second grasping point To the folding symmetry axis F t The distance l2 is as follows: Where S max is the area of the bounding box of the deformable object, γ is the weighting coefficient, C img It is an object contour extracted from the deformable object image, and the constraint conditions are that the first grasping point and the second grasping point are located on the object contour, and the first grasping point and the second grasping point are located on the left side of the folding symmetry axis; the first placement point is a symmetrical point of the first grasping point about the folding symmetry axis, and the second placement point is a symmetrical point of the second grasping point about the folding symmetry axis.
2. The method for controlling the shape of a deformable object based on a visual language model and historical data learning according to claim 1, characterized in that: The user's language target instructions are processed by the visual language model, and decomposed into language sub-target sequences corresponding to the visual teaching sub-target sequences, specifically including: The user language target instructions and the visual teaching sub-target sequence are input into the visual language model, and the user language target instructions are analyzed at the task level through the visual language model to decompose the user language instructions into a series of language sub-targets, each language sub-target represents an independent single-step sub-task and corresponds one-to-one to the visual teaching sub-target sequence.
3. The deformable object shape control method based on visual language model and historical data learning according to claim 1, characterized in that: The calculation of the task matching score between each data in the history database and the current task specifically includes: For the historical data based on language targets in the historical database, if the user language target instruction of one set of data is The language sub-goal is The corresponding eigenvectors are and For the user language target instruction in the current task and the current language sub-goal And the corresponding feature vector and Calculate the task matching score δ tm , the formula is as follows: in represents the cosine similarity between two feature vectors, score is the evaluation score calculated by the visual language model, and α is the weighting coefficient; For the historical data based on visual targets in the historical database, if the parameterized polygonal model corresponding to the visual teaching sub-targets of the two frames before and after the current step in a set of data is and The parameterized polygonal model corresponding to the visual teaching sub-goal of the current step of the current task in the two frames before and after is and Normalize the coordinates of all polygonal models to the range [0,1] and calculate the task matching score δ tm : Wherein score is the evaluation score calculated by the visual language model, β is the weighting coefficient, and d is the Euclidean distance between polygonal models, which is calculated as follows: first determine whether their number of levels is consistent, and verify whether the number of nodes in each level and the node labels are the same. If there is an inconsistency, d=∞, indicating that the two polygonal models cannot match; for the case where they can be matched, the Euclidean distance d of the polygonal model is defined as the cumulative sum of the Euclidean distances of each layer, where the Euclidean distance of each layer is defined as the cumulative sum of the distances between all nodes with the same label in that layer.
4. The method for controlling the shape of a deformable object based on a visual language model and historical data learning according to claim 1, characterized in that: The object matching score between each data in the historical database and the current deformable object is calculated, specifically including: If the image of the deformable object before the shape change in a set of data in the historical data is X db , the corresponding parameterized polygonal model is P db If the parameterized polygonal model of the current deformable object is p now ; The parameterized polygonal model is P db and the parameterized polygonal model is P now Rotate to the standard orthogonal position and calculate the Euclidean distance d between the polygonal models as the object matching score δ om , which can be expressed by the formula as follows: d om =-d(P db ,P now )。
Citation Information
Patent Citations
Demonstration enhancement depth deterministic strategy gradient-based deformable object robot operation method
CN118181285A
Deformable object interactive operation control method based on visual touch-language-action multi-mode model
CN119526422A