Industrial robot double-point incremental forming method and system based on multi-modal large model

Through the multimodal large model combined with robot multimodal data and deep learning, the problem of insufficient forming accuracy in the robot's double-point progressive forming process is solved, and higher compensation accuracy and forming path optimization are achieved.

CN120503178AActive Publication Date: 2025-08-19SHANDONG UNIV

Patent Information

Application Number
CN202510999090.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-08-19
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

In the prior art, the robot's double-point progressive forming process has insufficient forming accuracy, especially due to factors such as insufficient stiffness of the robot body and rebound of sheets, the existing deep learning methods have failed to effectively solve the robot's forming accuracy problem.

Method used

A multimodal large model is introduced. By acquiring robot multimodal data, including robot basic structure data, six-dimensional force value data, point cloud data and text description data, the deep learning model is used to perform joint stiffness coefficient and point cloud registration, combining visual and force awareness data to predict compensation values, and optimize the forming path.

Benefits of technology

The compensation accuracy of the robot's double-point progressive forming is improved, forming data is obtained in real time, errors caused by the robot's body attributes are eliminated, and the formation accuracy and path optimization are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120503178A_ABST
    Figure CN120503178A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of incremental forming, in particular to an industrial robot double-point incremental forming method and system based on a multi-modal large model, and the method comprises the steps: obtaining multi-modal data of a main forming robot and an auxiliary forming robot in a double-point incremental forming process; based on the multi-modal data, a joint stiffness coefficient is obtained; obtaining registered point cloud data based on the source point cloud data and the target point cloud data; based on the six-dimensional forming force value data, the data source point cloud data and the text description data of the forming scene, obtaining a description text and a preliminary compensation value; obtaining a prediction compensation value based on the registered point cloud data, the joint stiffness coefficient, the description text and the initial compensation value; and according to the predicted compensation value, the forming process of the to-be-machined plate is completed. According to the method, multi-dimensional data input of vision and the like is introduced, and a multi-modal large model is introduced as a training model, so that the compensation precision of double-point incremental forming is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of incremental forming, and in particular to an industrial robot double-point incremental forming method and system based on a multi-modal large model. Background Art

[0002] Incremental forming is a dieless forming process that utilizes a specialized tool head that moves incrementally across the material surface, applying localized pressure to induce plastic deformation in adjacent areas. The process is primarily categorized into single-point and dual-point incremental forming. The dual-point incremental forming process utilizes the coordinated control of two robots, allowing the tool heads to move synchronously on either side of the sheet to form the part, improving forming accuracy. However, this accuracy remains limited due to factors such as the robot's insufficient rigidity and the springback of the sheet during the forming process.

[0003] With the development and widespread application of deep learning in recent years, the application of deep learning in the field of incremental forming has also gradually deepened. In the existing technology, the defect points in the forming process are repaired by combining the visual system with deep learning, thereby improving the stiffness and flatness of the incrementally formed parts, but the accuracy of robot incremental forming is not considered; the existing technology also optimizes the forming path by training the error value in the forming process as the state vector, and only uses the deviation between the actual state and the target state as training data, but the forming accuracy in the forming process is affected by multimodal factors. Summary of the Invention

[0004] In order to address the shortcomings of the existing technology, the present invention provides an industrial robot dual-point incremental forming method and system based on a multimodal large model; the present invention introduces multi-dimensional data input such as vision and force perception, and the interactive data is more specific and richer, and introduces a multimodal large model as a training model, thereby improving the incremental forming path deviation compensation accuracy, thereby improving the forming accuracy of parts.

[0005] On the one hand, a method for industrial robot double-point incremental forming based on a multimodal large model is provided, including: Acquire multimodal data of the main forming robot and the auxiliary forming robot during the dual-point incremental forming process; the multimodal data includes: basic structure data of the robot, six-dimensional force value data, source point cloud data of the plate to be processed, robot posture data, and tool head forming motion deviation value; The multimodal data is input into the trained first deep learning model to obtain the joint stiffness coefficient; The source point cloud data and target point cloud data of the plate to be processed are input into the trained second deep learning model to obtain the registered point cloud data; The six-dimensional force data collected by the sensor, the robot's posture data, the source point cloud data of the plate to be processed, the tool head forming motion deviation value and the text description data of the forming scene are input into the trained third deep learning model to obtain the description text and preliminary compensation value; The registered point cloud data, joint stiffness coefficient, description text and preliminary compensation value are input into the fourth deep learning model to obtain the predicted compensation value; According to the predicted compensation value, the forming process of the sheet material to be processed is completed.

[0006] It should be understood that the source point cloud data of the plate to be processed is the surface contour point cloud of the actual formed component, and the target point cloud data is the contour point cloud data of the designed component.

[0007] The robot posture data is the joint angle of each joint of the robot in a certain posture, which is derived from the robot control cabinet.

[0008] The joint stiffness coefficient is a property value of the robot body obtained through training and is a fixed value (the Cartesian stiffness coefficient at different robot joint angles derived from it is a variable); The tool head forming motion deviation value is the deviation between the actual position and the target position of the robot forming tool when processing a certain point.

[0009] On the other hand, an industrial robot double-point incremental forming system based on a multimodal large model is provided, including: A compensation information module is configured to obtain multimodal data of the main forming robot and the auxiliary forming robot during the dual-point incremental forming process; the multimodal data includes: basic structure data of the robot, six-dimensional force value data, source point cloud data of the plate to be processed, robot posture data, and tool head forming motion deviation value; A joint stiffness coefficient determination module is configured to: input the multimodal data into the trained first deep learning model to obtain the joint stiffness coefficient; A module for determining registered point cloud data is configured to: input the source point cloud data and the target point cloud data of the plate to be processed into the trained second deep learning model to obtain registered point cloud data; A preliminary compensation value determination module is configured to: input the six-dimensional force data collected by the sensor, the robot posture data, the source point cloud data of the plate to be processed, the tool head forming motion deviation value, and the text description data of the forming scene into the trained third deep learning model to obtain the description text and the preliminary compensation value; a predicted compensation value determination module, configured to: input the registered point cloud data, joint stiffness coefficient, description text, and preliminary compensation value into a fourth deep learning model to obtain a predicted compensation value; The output module is configured to complete the forming process of the plate to be processed according to the predicted compensation value.

[0010] The above technical solution has the following advantages or beneficial effects: 1. Build a multimodal perception system, add vision and force sensors to the robot's dual-point incremental forming system, and output joint angle information in the robot control cabinet to obtain comprehensive forming data during the forming process in real time, thereby improving the comprehensiveness of the training model.

[0011] 2. The robot joint stiffness properties are added to the mathematical model of the robot constructed in deep learning. Starting from the robot body properties, the LSTM model is used to train and learn the posture and conventional force feedback compensation model based on the joint stiffness properties.

[0012] 3. After eliminating the forming errors caused by the robot's own properties, a camera is used to obtain point cloud information on the surface of the formed part, and multimodal interaction is achieved with information on various process parameters such as forming force and wall angle. A large vision-language-action multimodal model is used to fuse the network to build a large visual language model (VLM). Through the transformer attention mechanism and cross-modal interaction module, the optimized robot forming posture and forming path are trained, and the point cloud registration method is used for error analysis.

[0013] 4. Convert the image-text pair into an appropriate instruction format and generate a series of questions to guide the assistant to describe the image content, thus realizing the reshaping of multimodal instruction data. The model adopts a visual instruction adjustment method and maps the image features to the language feature space through a linear projection layer to realize the fusion of image and text information. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0015] Figure 1 It is a schematic diagram of a double-point incremental forming device; Figure 2 This is the robot stiffness principle diagram; Figure 3 is a flow chart of the method of the present invention; Among them, 1. the plate to be processed; 2. the plate fixture; 3. the camera; 4. the first force sensor; 5. the second force sensor; 6. the main forming tool head; 7. the auxiliary forming tool head. DETAILED DESCRIPTION

[0016] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0017] Example 1, as Figure 3 As shown, the industrial robot double-point incremental forming method based on a multimodal large model includes: S101: Acquire multimodal data of the main forming robot and the auxiliary forming robot during the dual-point incremental forming process; the multimodal data includes: basic structure data of the robot, six-dimensional force value data, source point cloud data of the plate to be processed, robot posture data, and tool head forming motion deviation value; S102: Inputting the multimodal data into the trained first deep learning model to obtain joint stiffness coefficients; S103: Inputting the source point cloud data and the target point cloud data of the plate to be processed into the trained second deep learning model to obtain the registered point cloud data; S104: Inputting the six-dimensional force data collected by the sensor, the robot posture data, the source point cloud data of the plate to be processed, the tool head forming motion deviation value, and the text description data of the forming scene into the trained third deep learning model to obtain the description text and the preliminary compensation value; S105: Inputting the registered point cloud data, joint stiffness coefficient, description text and preliminary compensation value into the fourth deep learning model to obtain a predicted compensation value; S106: completing the forming process of the plate to be processed according to the predicted compensation value.

[0018] Furthermore, the method further comprises: S107: Calculate the Euclidean distance between each point in the target point cloud of the surface of the plate to be processed and the nearest point in the source point cloud data of the plate to be processed, calculate the average distance of all Euclidean distances, and output the average forming error, maximum error, and minimum error.

[0019] Further, S101: multimodal data, including: robot basic structure data, six-dimensional force value data, source point cloud data of the plate to be processed, robot posture data and tool head forming motion deviation value; The basic structural data of the robot include: connecting rod length, connecting rod offset, joint angle, and joint torsion angle; Based on the basic structure data of the robot, the transformation matrix between adjacent links is established; based on the transformation matrix between adjacent links, the total transformation matrix from the base coordinate system to the end effector coordinate system is obtained; The tool head forming motion deviation value refers to the position deviation of the main forming tool head at each forming point; The six-dimensional force value data refers to the six-dimensional force value of the main forming tool head at each forming point.

[0020] Furthermore, if Figure 1 As shown, the main forming robot and the auxiliary forming robot are both industrial robots, the main forming robot is provided with a first end effector at the end of its robotic arm, and the auxiliary forming robot is provided with a second end effector at the end of its robotic arm; the first end effector is provided with a main forming tool head 6, and the second end effector is provided with an auxiliary forming tool head 7, a first force sensor 4 is provided between the first end effector and the main forming tool head 6, and a second force sensor 5 is provided between the second end effector and the auxiliary forming tool head 7.

[0021] A plate 1 to be processed is arranged between the main forming tool head 6 and the auxiliary forming tool head 7, and the plate 1 to be processed is fixed by a plate clamp 2; the first end effector is also connected to a camera 3; the camera 3 is used to shoot the processing process of the surface of the plate to be processed.

[0022] Furthermore, the transformation matrix between adjacent links is established based on the basic structure data of the robot; based on the transformation matrix between adjacent links, a total transformation matrix from the base coordinate system to the end effector coordinate system is obtained, including: Using the mDH modeling method (Craig JJ. Introduction to Robotics: Mechanics and Control [M]. Pearson Education, Inc, 1986.), four parameters are defined for each link: link length , connecting rod offset , joint angle , joint torsion angle , the transformation matrix between adjacent links is established through four parameters , the transformation matrix Set the current coordinate system Transform to target coordinate system .

[0023]

[0024] By successively applying the transformation matrix , calculate the total transformation matrix from the base coordinate system to the end effector coordinate system ,in, Indicates the A transformation matrix.

[0025] It should be understood that if there is a robotic arm, it consists of joints, then define a DH parameter transformation matrix for each joint, the transformation matrix Describes the The coordinate system of the first joint to the The total transformation matrix is obtained by multiplying the transformation matrix in sequence. , represents the coordinate system from the base of the robot arm (or the first joint) to the end effector (or the The complete transformation of the coordinate system of each joint. Contains all joint angle information of the robotic arm and can be used to calculate the position and orientation of the end effector.

[0026] Furthermore, if Figure 2 As shown, the tool head forming motion deviation value refers to the position deviation of the main forming tool head or the auxiliary forming tool head at each forming point, wherein the position deviation of the main forming tool head or the auxiliary forming tool head at each forming point refers to the deviation between the actual position reached by the tool head and the target position. Specifically, it refers to using a camera to measure the position deviation of the main forming tool head or the auxiliary forming tool head at each forming point in real time. Based on the previous point cloud registration data, the spatial correspondence between the forming position point and the target mold point cloud data is obtained. Therefore, the spatial distance difference between the two in the X, Y, and Z directions is expressed as .

[0027] Furthermore, the six-dimensional force data refers to the six-dimensional force value of the main forming tool head or the auxiliary forming tool head at each forming point, and the six-dimensional force value of the corresponding point is recorded by the force sensor. .

[0028] The six-dimensional force value refers to the six independent force and torque components that the indenter can measure or apply at each forming point. Specifically, the six-dimensional force value includes three force components: and three torque components .

[0029] Furthermore, the step S102: inputting the multimodal data into a trained first deep learning model to obtain a joint stiffness coefficient, wherein the trained first deep learning model includes: The first LSTM layer, the first Dropout layer, the second LSTM layer, the second Dropout layer, the computation layer, and the first fully connected layer are connected in sequence.

[0030] Furthermore, the functions of the first LSTM layer and the second LSTM layer are consistent, and the first LSTM layer is used to extract features.

[0031] Furthermore, the working processes of the first Dropout layer and the second Dropout layer are consistent, and both are used to improve the generalization ability of the first deep learning model.

[0032] Furthermore, the computing layer is implemented by a multi-layer perceptron. hidden layer neurons, which are calculated as follows: ; in, represents the weight between the input layer and the first hidden layer, represents the bias term of the first hidden layer, Represents the activation function.

[0033] For the The output layer neurons are calculated as follows: ; in, represents the weight between the first hidden layer and the output layer, represents the bias term of the output layer, Represents the activation function.

[0034] Furthermore, the first fully connected layer is used as an output layer.

[0035] Furthermore, in S102: inputting the multimodal data into the trained first deep learning model to obtain the joint stiffness coefficient, wherein the training process of the trained first deep learning model includes: Construct a first dataset, which is multimodal data of known robot joint stiffness coefficients; divide the first dataset into training set, test set and validation set in a ratio of 6:2:2, and set the batch_size of the number of samples in one training to 50; output the joint stiffness coefficients through training .

[0036] The training set is input into the first deep learning model, and the first deep learning model is trained. When the loss function value of the first deep learning model no longer decreases, the training is stopped to obtain the trained first deep learning model.

[0037] Furthermore, the loss function of the first deep learning model is selected as the mean square error As the loss function, and the six-dimensional force value Substitute the joint stiffness coefficient into the robot control cabinet to determine the target motion position Motion position after stiffness compensation gap and conduct training.

[0038] The loss function of the first deep learning model is as follows: ; ; in, Indicates that the stiffness is set under the condition of Repeat the experiment to compensate. Indicates the joint stiffness coefficient of the current robotic arm, Indicates the six-dimensional force value of the current end of the robotic arm. It means that the six-dimensional force value is first decomposed into components along the manipulator coordinate system through coordinate transformation. , transform the force value from the world coordinate system to the robot coordinate system, and then convert the joint stiffness coefficient composition The diagonal matrix of , and use the diagonal matrix Calculate the stiffness coefficient of the forming tool head in the Cartesian coordinate system ,in, ; Calculate the deformation of the robot arm due to external forces: ; is the position change, the calculated position change Added to the current position of the end of the robot arm to obtain the compensated position; yes The matrix, yes The matrix of .

[0039] Furthermore, in S102: inputting the multimodal data into the trained first deep learning model to obtain the joint stiffness coefficient, wherein the trained first deep learning model has a working process including: First, two stacked LSTM layers are introduced, with the number of hidden units in each layer set to 64. A Dropout layer is followed by each LSTM layer, and the Dropout rate is set to 0.2 to improve the generalization ability of the model. A calculation layer is added after the two LSTM layers to embed the stiffness model. A fully connected layer is connected as the output layer, and the number of neurons in the fully connected layer is set to the data feature of the joint stiffness coefficient. .

[0040] Furthermore, in S103, the source point cloud data and the target point cloud data of the plate to be processed are input into the trained second deep learning model to obtain the registered point cloud data, wherein the trained second deep learning model includes: A point cloud feature extraction module and a point cloud feature fusion module connected in sequence; The point cloud feature extraction module includes: a first branch and a second branch arranged in parallel; The first branch includes: a first K-nearest neighbor algorithm layer, a first maximum pooling layer, a first multi-layer perceptron, a first pooling layer, a second multi-layer perceptron, and a second maximum pooling layer connected in sequence; the input value of the first K-nearest neighbor algorithm layer is the source point cloud data; The second branch includes: a second K-nearest neighbor algorithm layer, a third maximum pooling layer, a third multi-layer perceptron, a second pooling layer, a fourth multi-layer perceptron, and a fourth maximum pooling layer connected in sequence; the input value of the second K-nearest neighbor algorithm layer is the target point cloud data; The point cloud feature fusion module includes: a stitching operation layer and a second fully connected layer connected in sequence; the output end of the second fully connected layer outputs a posture transformation vector; Among them, the output end of the second maximum pooling layer and the output end of the fourth maximum pooling layer are both connected to the input end of the splicing operation layer.

[0041] Furthermore, the training process of the trained second deep learning model includes: Constructing a second data set, wherein the second data set is source point cloud data and target point cloud data with known posture transformation vectors; The second data set is input into the second deep learning model, and the second deep learning model is trained. When the loss function value of the second deep learning model no longer decreases, the training is stopped to obtain the trained second deep learning model.

[0042] Furthermore, the loss function of the second deep learning model is set to minimize the distance between corresponding points in the source point cloud and the target point cloud, which is achieved by calculating the distance between corresponding points in the source point cloud and the target point cloud after the predicted transformation. The loss function of the second deep learning model is set to: ; in, represents the source point cloud dataset, Represents the data obtained by mapping the source point cloud data to the target point cloud. represents the target point cloud dataset, X Represents a single sample in the source point cloud dataset, Represents the inverse of the number of dataset samples in the source point cloud.

[0043] Furthermore, the working process of the trained second deep learning model includes: The input of the trained second deep learning model is the source point cloud and target point cloud ; The working process of the second branch is the same as that of the first branch. The difference is that the input value of the first branch is the source point cloud, and the input value of the second branch is the target point cloud. Only the working process of the first branch is explained: Import point cloud data , For each point , find its nearest k points as the edge set , for each edge combination Calculate its edge convolution features: ; Represents the input point cloud location point, express The spatial distance between the edge set point and it can reflect the contour features of its surroundings, and a pooling operation is performed on it to extract the point cloud contour features.

[0044] The edge features of each point are calculated by the first multi-layer perceptron Tensor , where the multilayer perceptron contains four layers of neural networks connected in sequence; each layer of the neural network includes: an input layer, a first hidden layer, a second hidden layer, and an output layer connected in sequence.

[0045] Next, the first pooling layer, for each point of dimensional edge features are pooled to obtain Tensor ; Afterwards, The second multi-layer perceptron is used to increase the dimension and obtain a 1024-dimensional tensor. ; Finally, the second max pooling layer is Perform the maximum pooling operation to obtain the global feature vector .

[0046] Furthermore, the point cloud feature fusion module combines the source point cloud The global eigenvector of and target point cloud The global eigenvector of Perform data merging operation to obtain a 2048-dimensional feature vector .

[0047] Furthermore, the point cloud feature fusion module uses the second fully connected layer (FC) to regress the fused feature vector. The second fully connected layer has four hidden layers with sizes of 1024, 512, 512, and 256, respectively, and outputs a seven-dimensional estimated pose transformation vector , the posture transformation vector By translating the 3D vector and rotation quaternions composition.

[0048] Transform the pose vector The rotation quaternion in is converted to the rotation matrix R: , in, Indicates the rotation angle information, Represents the unit vector of the rotation axis Multiply by the sine of half the rotation angle, Represents the unit vector of the rotation axis Multiply by the sine of half the rotation angle, Represents the unit vector of the rotation axis Multiply by the sine of half the rotation angle.

[0049] Combined with the translation matrix Acting on the source point cloud, the point after the posture transformation Expressed as: ; in, represents the transformed point cloud, represents the input point cloud, represents the rotation matrix, represents the translation matrix; Complete the registration of the source point cloud and the target point cloud and get the deviation value: ; in, represents the target point cloud, Indicates the deviation value in the first direction, Indicates the deviation value in the second direction, Indicates the deviation value of the third direction, wherein the first direction may be the X direction of the coordinate axis, the second direction may be the Y direction of the coordinate axis, and the third direction may be the Z direction of the coordinate axis.

[0050] Furthermore, in step S104, the six-dimensional force data collected by the sensor, the robot posture data, the source point cloud data of the plate to be processed, the tool head forming motion deviation value, and the text description data of the forming scene are input into the trained third deep learning model to obtain the description text and the preliminary compensation value, wherein the trained third deep learning model includes: The third, fourth, and fifth branches are parallel; The third branch includes: a visual encoder, a first projection layer, and a fifth multi-layer perceptron connected in sequence; The fourth branch includes: a sensor encoder, a second projection layer, and a sixth multi-layer perceptron connected in sequence; The fifth branch includes: a language encoder; Among them, the output end of the fifth multi-layer perceptron, the output end of the sixth multi-layer perceptron and the output end of the language encoder are all connected to the input end of the splicing unit, the output end of the splicing unit is connected to the input end of the large language model, and the output end of the large language model is used to output the description text and the preliminary compensation value.

[0051] The visual encoder, sensor encoder and language encoder can all be implemented using a transformer model.

[0052] The target data is the joint angle data that controls the robot's movement, and the description text is a text description of the operation that the robot is about to perform, such as polishing the raised parts of the mold.

[0053] Furthermore, the first and second projection layers are used to map the encoding results of different modalities into a unified representation space, which is then input into the Transformer model. The tokens output by the model are then mapped back to the decoder of the original modality through the projection layer to generate the corresponding shaping instructions.

[0054] The Projection layer is just a matrix multiplication, a regular / dense / linear layer, without any nonlinear activation (sigmoid / tanh) at the end. For example, it projects a 100K-dimensional discrete vector into a 600-dimensional continuous vector (here the dimensions are randomly chosen). The exact matrix parameters are learned during training. The Projection layer introduces a learnable projection matrix To achieve compression, the form The multiplication is replaced by ,pass Will Projecting to a lower dimensional space usually requires less storage space to store the learnable parameters.

[0055] For model building, 595K high-quality image-text data pairs were selected from the multimodal forming dataset, i.e., the corresponding forming scene information and point cloud deviation value information. , forming sensor information, etc.; input language commands in the incremental forming scenario, using real-time image information and sensor data information during the forming process, within the existing target deviation value The forming accuracy compensation is carried out based on the Clip-Sensor model, which is mainly composed of a visual encoder, a text encoder, and a sensor encoder.

[0056] Furthermore, in step S104, the six-dimensional force data, robot posture data, source point cloud data of the plate to be processed, tool head forming motion deviation value, and text description data of the forming scene collected by the sensor are input into the trained third deep learning model to obtain description text and preliminary compensation value. The working process of the visual encoder of the trained third deep learning model includes: In the forming scenario, the dimension extracted is Original image ; For the original image conduct The segmentation is obtained Initial image of dimension ; Initial Image go through The flattening operation and the linear layer operation obtain a 512-dimensional embedding vector : ; Among them, the weight vector The dimension is ; For the embedding vector Perform absolute position encoding; flattening operation is about to Convert to ; The feature extraction and spatial mapping of the graph are achieved by using 8 layers of Transformer encoding layers (including feedforward networks) connected in series. The Transformer encoding layer uses 8 attention heads, and the number of nodes in the hidden layer of the Transformer encoding layer is (512, 1024, 1024, 512, 1024, 1024, 2048, 2048) respectively. The final output visual matching embedding vector: , The output dimension is .

[0057] It should be understood that the Transformer encoding layer, see Attention Is All You Need, Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, LlionJones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin.

[0058] It should be understood that absolute position encoding refers to adding a position vector to each element in the input sequence to indicate its specific position in the sequence. This position vector is typically generated by a custom CLIP function and is independent of the input data, capturing relative position information within the sequence.

[0059] Furthermore, the working process of the sensor encoder of the trained third deep learning model includes: First, the forming wall angle extracted in the forming scene and three-dimensional force value Expand the dimension to obtain augmented data ; Perform feature concatenation on the generated data and merge them into a new feature vector: , Represents the data of the source sensor, with a dimension of (n, 256); A four-layer Transformer encoding layer is used in series. The Transformer encoding layer uses four attention heads. The number of nodes in the four hidden layers of the Transformer encoding layer are (512, 512, 1024, 1024) respectively, and each hidden layer has four attention heads. The final output sensor matching embedding vector .

[0060] It should be understood that the means of dimensional expansion are to calculate the statistical characteristics of the data, the normalized characteristics of the data and the time series characteristics of the data. The statistical characteristics include: mean value, standard deviation and variance; the time series characteristics are the sensor data over a period of time; the normalized characteristics of the data are the results of the normalized data.

[0061] It should be understood that because incremental forming involves the tool head pressing down on the sheet material layer by layer, the forming wall angle is the angle between the line connecting the instantaneous forming position of the workpiece and the corresponding position of the previous layer, and the plane of the unformed sheet material. This angle is a key parameter affecting forming quality and a key indicator for evaluating the application of forming technology for complex surfaces. The size of the forming wall angle directly affects the geometric accuracy of the formed part and the feasibility of the incremental forming process.

[0062] Furthermore, the language encoder of the trained third deep learning model works as follows: First, enter the text Perform the cutting operation, train the word segmentation data set for WordPiece, and use WordPiece to preprocess the text information to obtain: ; in, Represents a tokenizer in a large language model.

[0063] For the text embedding operation, the lookup table is first randomly initialized, the index of the text information corresponds to the row sequence of the lookup table, and the text information is (excluding ) is mapped to a 2048-dimensional vector space to obtain a preliminary text embedding vector , the dimension is ; At the same time, relative position encoding is used for the preliminary embedding vector, and finally the text information is encoded into a text latent vector By sequentially connecting 12 layers of Transformer, stacking the self-attention mechanism and feedforward neural network, and setting the final number of hidden layers to 4096, the language matching embedding vector is obtained. , the output dimension is The language encoder is implemented using the large language model GPT-4.

[0064] Furthermore, the working process of the trained third deep learning model also includes: A projection layer is set after the visual encoder and the sensor encoder. In order to align the visual features, sensor features and text features in dimension, a multi-layer perceptron is used to embed the sensor matching vector. Matching embedding vectors with visual For processing, in addition to the linear transformation of the input features, the multi-layer perceptron also uses the RELU function as the activation function to increase the nonlinear ability of the network, and the matching vector ; ; Visual vector , sensor vector Vector with text Perform serial splicing to obtain feature fusion vector ( ); Call the large language model Vicuna LLM to fusion vector feature Perform reasoning to obtain description text and preliminary compensation value.

[0065] It should be understood that a large industrial model combining language, vision, and sensor modalities is constructed. The model uses cameras and force sensors to achieve the ability to interact with the forming environment. Based on the existing language model GPT-4, it builds a multimodal instruction forming dataset LLM-PCR-PS with language understanding and generation capabilities.

[0066] GPT-4 has multimodal input capabilities, accepting both images and text as input and generating text output, making it suitable for multimodal tasks. The multimodal instruction dataset, LLM-PCR-PS, combines GPT-4's language understanding and generation capabilities to generate motion plans and text descriptions suitable for forming scenarios. By combining visual information (such as images and point clouds) with language instructions and angle instructions, the model can understand and generate specific instructions for forming tasks.

[0067] LLM stands for Large Language Model. Pre-trained Cross-Modal Representations: This refers to a dataset of multimodal representations generated using GPT4. Processing and Synthesis: This may refer to the dataset or model's ability to process multimodal input (such as images, text, and sensors) and synthesize new outputs (such as shaped instructions).

[0068] A method for constructing an industrial large model that calls a large language model (LLM) for industrial scenarios. The industrial large model combines language, vision, and sensor modalities. It realizes the ability to interact with the forming environment through cameras and force sensors. Based on the existing language model GPT-4 with multimodal input capabilities (which can accept images and text as input and generate text output to process multimodal tasks), a multimodal instruction forming dataset LLM-PCR-PS with language understanding and generation capabilities is constructed. This multimodal instruction dataset combines the language understanding and generation capabilities of GPT-4 to generate motion planning and text descriptions suitable for forming scenarios. Specifically, it combines visual information (such as images and point clouds) with language instructions and angle instructions to enable the model to understand and generate specific instructions in forming tasks.

[0069] It should be understood that the large language model Vicuna LLM is an existing large language model. What is called here is the model backbone, and the language encoder therein is from Google's LLaMA project.

[0070] Furthermore, the training process of the trained third deep learning model in S104 includes: Constructing a third data set, wherein the third data set includes six-dimensional force data with known description text and preliminary compensation values, robot posture data, source point cloud data of the plate to be processed, tool head forming motion deviation values, and text description data of the forming scene; The third data set is input into the third deep learning model, and the third deep learning model is trained. When the loss function value of the third deep learning model no longer decreases, the training is stopped to obtain the trained third deep learning model.

[0071] The text description data of the forming scene is an objective description and target of the current forming task of the incremental forming. For example, a robot is completing the grinding work of a cylinder, and the robot's task goal is to correct and improve the original irregularities.

[0072] The third deep learning model is pre-trained. Because the features extracted from vision and sensors are not in the same semantic space as text features, pre-training is required to align the image and sensor features. During this stage, the weight parameters of the visual encoder, sensor encoder, language encoder, and large language model are frozen, and only the weights of the interpolation layer (projection) are trained.

[0073] Fusion is performed at different depths in the Projection layer, allowing the model to learn interactions between features at different levels of abstraction. Training is performed with 1 epoch, a learning rate of 2e-3, and a batch size of 128. The weights of the Projection layer are optimized by minimizing an alignment loss (such as cosine similarity loss).

[0074] For example, the weight of the Projection layer is , then the image feature after Projection layer processing is represented as , then the loss value is expressed as , Represents the image features before Projection layer processing.

[0075] To optimize the model's performance on specific tasks, end-to-end training was performed. During training, the interpolation layer (Projection) and the Transformer block of the Vicuna LLM were updated simultaneously. Weights of these layers were allowed to update during training. Automatic Mixed Precision (AMP) was enabled to handle multi-type data. The number of training epochs was set to 3, which could be adjusted based on the model's performance on the validation set. The learning rate was reduced to 2e-5, the batch size was adjusted to 32, and the AdamW optimizer was used. When using AMP, careful attention should be paid to the optimizer settings to ensure compatibility with mixed precision. The parameters of the Projection layer and LLM were optimized by minimizing the cross-entropy loss of the generated text.

[0076] Based on the above model, it is possible to realize the registration information of part surface point cloud and joint stiffness coefficient Combined with the three-dimensional force value F, robot joint angle, feed rate, cutting distance, temperature, and speed, the rebound stiffness compensation instruction is input to the third deep learning model. , then starting from the target language information and combining multimodal information, the language model is used to generate a preliminary compensation value ,in, represents the third deep learning model obtained through training, represents the joint stiffness coefficient, Indicates the current robot joint angle information, Indicates the current force status of the robot's end sensor.

[0077] Furthermore, the step S105: inputting the registered point cloud data, joint stiffness coefficient, description text and preliminary compensation value into the fourth deep learning model to obtain a predicted compensation value includes: The fourth deep learning model after training is implemented using the transformer model; The training process includes: constructing a training set, wherein the training set is the registered point cloud data, joint stiffness coefficient, description text and preliminary compensation value with known preliminary compensation value; The fourth deep learning model is trained using the training set to obtain a trained fourth deep learning model.

[0078] Furthermore, S107: calculating the Euclidean distance between each point in the target point cloud of the surface of the plate to be processed and the nearest point in the source point cloud data of the plate to be processed, and calculating the average distance of all Euclidean distances, and outputting the average forming error, maximum error and minimum error, specifically including: After the forming is completed, based on the second deep learning model, the formed mold is subjected to a second point cloud registration, the formed mold and the target mold are point cloud registered, and the average Euclidean distance is calculated. The Euclidean distance of each point in the actual part surface point cloud relative to the nearest point in the theoretical model surface point cloud is calculated, and the average distance is calculated to output the average forming error and the maximum and minimum errors to verify the working effect of the model.

[0079] The error function is: ; in, Represents the point cloud dataset after forming, Represents the data obtained by mapping the formed point cloud data to the target point cloud. Represents the target point cloud dataset, X represents a single sample in the formed point cloud dataset, Represents the inverse of the number of data set samples in the formed point cloud; When the calculated average Euclidean distance is the smallest, the best forming effect is achieved.

[0080] It should be understood that the robot body attribute index is collected, and the robot joint stiffness attribute, robot posture dexterity and robot stiffness index are measured by a camera and a force sensor.

[0081] The collected multimodal data is processed and the joint stiffness coefficient of the robot is calculated through the LSTM algorithm. Make predictions.

[0082] The surface point cloud information of the formed parts is obtained in real time through a visual camera. Based on the text describing the point cloud scene, a point cloud encoder is used to obtain point cloud embedding features. In the multimodal large model, a hidden layer is set at each interval to introduce fusion features for multimodal fusion, obtain panoramic features, and realize point cloud registration.

[0083] Realize the registration of point cloud information and joint stiffness coefficient , and multimodal interaction with visual extraction information, forming arm angle sensor information, and three-dimensional force values, using multimodal feature extraction, and fusion network to build a large language vision model (VLM), through the transformer attention mechanism and cross-modal interaction module, to achieve accurate prediction of the compensation amount of the robotic arm based on different feature quantities, and to use language information to make forming adjustments based on real-time visual information.

[0084] Integrate multimodal large models, select models that exceed the set threshold after training, and encapsulate and deploy them to compensate for the incremental forming process.

[0085] Perform mathematical modeling of the robot and introduce the joint stiffness model. Establish the mDH model based on the actual structural parameters of the robot, and then add stiffness attributes to each joint.

[0086] Different forming objectives can also be set, such as using language information to provide real-time feedback on forming progress, improving responsiveness to unexpected forming scenarios. This model can also be applied to automated production and intelligent manufacturing in diverse fields (such as automotive, aerospace, and precision machinery). For example, by leveraging the model's language generation capabilities, combined with sensor data and visual information, it can enable remote monitoring of production line status and automatic diagnosis of potential faults. Furthermore, this multimodal large model optimizes computational efficiency and response time for real-time forming scenarios, promoting efficient intelligent manufacturing.

[0087] Example 2 This embodiment provides an industrial robot dual-point incremental forming system based on a multimodal large model, including: A compensation information module is configured to obtain multimodal data of the main forming robot and the auxiliary forming robot during the dual-point incremental forming process; the multimodal data includes: basic structure data of the robot, six-dimensional force value data, source point cloud data of the plate to be processed, robot posture data, and tool head forming motion deviation value; A joint stiffness coefficient determination module is configured to: input the multimodal data into the trained first deep learning model to obtain the joint stiffness coefficient; A module for determining registered point cloud data is configured to: input the source point cloud data and the target point cloud data of the plate to be processed into the trained second deep learning model to obtain registered point cloud data; A preliminary compensation value determination module is configured to: input the six-dimensional force data collected by the sensor, the robot posture data, the source point cloud data of the plate to be processed, the tool head forming motion deviation value, and the text description data of the forming scene into the trained third deep learning model to obtain the description text and the preliminary compensation value; a predicted compensation value determination module, configured to: input the registered point cloud data, joint stiffness coefficient, description text, and preliminary compensation value into a fourth deep learning model to obtain a predicted compensation value; The output module is configured to complete the forming process of the plate to be processed according to the predicted compensation value.

[0088] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. The industrial robot double-point incremental forming method based on a multi-modal large model is characterized by: include: Acquire multimodal data of the main forming robot and the auxiliary forming robot during the dual-point incremental forming process; The multimodal data includes: basic structure data of the robot, six-dimensional force value data, source point cloud data of the plate to be processed, robot posture data and tool head forming motion deviation value; The multimodal data is input into the trained first deep learning model to obtain the joint stiffness coefficient; The source point cloud data and target point cloud data of the plate to be processed are input into the trained second deep learning model to obtain the registered point cloud data; The six-dimensional force data collected by the sensor, the robot's posture data, the source point cloud data of the plate to be processed, the tool head forming motion deviation value and the text description data of the forming scene are input into the trained third deep learning model to obtain the description text and preliminary compensation value; The registered point cloud data, joint stiffness coefficient, description text and preliminary compensation value are input into the fourth deep learning model to obtain the predicted compensation value; According to the predicted compensation value, the forming process of the sheet material to be processed is completed.

2. The industrial robot double-point incremental forming method based on a multimodal large model as claimed in claim 1 is characterized in that: The multimodal data is input into the trained first deep learning model to obtain the joint stiffness coefficient, wherein the trained first deep learning model includes: The first LSTM layer, the first Dropout layer, the second LSTM layer, the second Dropout layer, the computation layer, and the first fully connected layer are connected in sequence; The loss function of the first deep learning model is the mean square error. As the loss function, and the six-dimensional force value Substitute the joint stiffness coefficient into the robot control cabinet to determine the target motion position Motion position after stiffness compensation gaps and conduct training; The loss function of the first deep learning model is as follows: ; ; in, Indicates that the stiffness is set under the condition of Repeat the experiment to compensate. Indicates the joint stiffness coefficient of the current robotic arm, Indicates the six-dimensional force value of the current end of the robotic arm. It means that the six-dimensional force value is first decomposed into components along the manipulator coordinate system through coordinate transformation. , transform the force value from the world coordinate system to the robot coordinate system, and then convert the joint stiffness coefficient composition The diagonal matrix of , and use the diagonal matrix Calculate the stiffness coefficient of the forming tool head in the Cartesian coordinate system ,in, ; Calculate the deformation of the robot arm due to external forces: ; is the position change, the calculated position change Added to the current position of the end of the robot arm to obtain the compensated position; yes The matrix, yes Matrix of The first deep learning model after training includes the following steps: first, introducing two stacked LSTM layers, followed by a Dropout layer after each LSTM layer to improve the generalization ability of the model; adding a computational layer after the two LSTM layers to embed the stiffness model; connecting a fully connected layer as the output layer, and setting the number of neurons in the fully connected layer to the data characteristics of the joint stiffness coefficient. .

3. The industrial robot double-point incremental forming method based on a multimodal large model as claimed in claim 1 is characterized in that: The source point cloud data and the target point cloud data of the plate to be processed are input into the trained second deep learning model to obtain the registered point cloud data, wherein the trained second deep learning model includes: A point cloud feature extraction module and a point cloud feature fusion module connected in sequence; The point cloud feature extraction module includes: a first branch and a second branch arranged in parallel; The first branch includes: a first K-nearest neighbor algorithm layer, a first maximum pooling layer, a first multi-layer perceptron, a first pooling layer, a second multi-layer perceptron, and a second maximum pooling layer connected in sequence; the input value of the first K-nearest neighbor algorithm layer is the source point cloud data; The second branch includes: a second K-nearest neighbor algorithm layer, a third maximum pooling layer, a third multi-layer perceptron, a second pooling layer, a fourth multi-layer perceptron, and a fourth maximum pooling layer connected in sequence; the input value of the second K-nearest neighbor algorithm layer is the target point cloud data; The point cloud feature fusion module includes: a stitching operation layer and a second fully connected layer connected in sequence; the output end of the second fully connected layer outputs a posture transformation vector; Among them, the output end of the second maximum pooling layer and the output end of the fourth maximum pooling layer are both connected to the input end of the splicing operation layer.

4. The industrial robot double-point incremental forming method based on a multimodal large model as claimed in claim 1 is characterized in that: The loss function of the second deep learning model is set to minimize the distance between the corresponding points in the source point cloud and the target point cloud. This is achieved by calculating the distance between the corresponding points in the source point cloud and the target point cloud after the predicted transformation. The loss function of the second deep learning model is set as: ; in, represents the source point cloud dataset, Represents the data obtained by mapping the source point cloud data to the target point cloud. represents the target point cloud dataset, X Represents a single sample in the source point cloud dataset, Represents the inverse of the number of dataset samples in the source point cloud; The working process of the trained second deep learning model includes: The input of the trained second deep learning model is the source point cloud and target point cloud ; The working process of the second branch is the same as that of the first branch. The difference is that the input value of the first branch is the source point cloud, and the input value of the second branch is the target point cloud. Only the working process of the first branch is explained: Import point cloud data , For each point , find its nearest k points as the edge set , for each edge combination Calculate its edge convolution features: ; Represents the input point cloud location point, express The spatial distance between the edge set point and the edge set point is used to reflect the contour features of the surrounding area, and a pooling operation is performed on it to extract the point cloud contour features; The edge features of each point are calculated by the first multi-layer perceptron Tensor , wherein the multilayer perceptron includes four layers of neural networks connected in sequence; each layer of the neural network includes: an input layer, a first hidden layer, a second hidden layer, and an output layer connected in sequence; Next, the first pooling layer, for each point of dimensional edge features are pooled to obtain Tensor ; Afterwards, The second multi-layer perceptron is used to increase the dimension and obtain a 1024-dimensional tensor. ; Finally, the second max pooling layer is Perform the maximum pooling operation to obtain the global feature vector .

5. The industrial robot double-point incremental forming method based on a multimodal large model as claimed in claim 3 is characterized in that: The point cloud feature fusion module combines the source point cloud The global eigenvector of and target point cloud The global eigenvector of Perform data merging operation to obtain a 2048-dimensional feature vector ; The point cloud feature fusion module uses the second fully connected layer to regress the fused feature vector. The second fully connected layer has four hidden layers with sizes of 1024, 512, 512, and 256 respectively, and outputs a seven-dimensional estimated posture transformation vector , the posture transformation vector By translating the 3D vector and rotation quaternions composition; Transform the pose vector The rotation quaternion in is converted to the rotation matrix R: , in, Indicates the rotation angle information, Represents the unit vector of the rotation axis Multiply by the sine of half the rotation angle, Represents the unit vector of the rotation axis Multiply by the sine of half the rotation angle, Represents the unit vector of the rotation axis Multiply by the sine of half the rotation angle; Combined with the translation matrix Acting on the source point cloud, the point after the posture transformation Expressed as: ; in, represents the transformed point cloud, represents the input point cloud, represents the rotation matrix, represents the translation matrix; Complete the registration of the source point cloud and the target point cloud and get the deviation value: ; in, represents the target point cloud, Indicates the deviation value in the first direction, Indicates the deviation value in the second direction, Indicates the deviation value in the third direction.

6. The industrial robot double-point incremental forming method based on a multimodal large model as claimed in claim 1 is characterized in that: The trained third deep learning model includes: The third, fourth, and fifth branches are parallel; The third branch includes: a visual encoder, a first projection layer, and a fifth multi-layer perceptron connected in sequence; The fourth branch includes: a sensor encoder, a second projection layer, and a sixth multi-layer perceptron connected in sequence; The fifth branch includes: a language encoder; Among them, the output end of the fifth multi-layer perceptron, the output end of the sixth multi-layer perceptron and the output end of the language encoder are all connected to the input end of the splicing unit, the output end of the splicing unit is connected to the input end of the large language model, and the output end of the large language model is used to output the description text and the preliminary compensation value.

7. The industrial robot double-point incremental forming method based on a multimodal large model as claimed in claim 6 is characterized in that: The working process of the visual encoder of the trained third deep learning model includes: In the forming scenario, the dimension extracted is Original image ; For the original image conduct The segmentation is obtained Initial image of dimension ; Initial Image go through The flattening operation and the linear layer operation obtain a 512-dimensional embedding vector : ; Among them, the weight vector The dimension is ; For the embedding vector Perform absolute position encoding; flattening operation is about to Convert to ; The feature extraction and spatial mapping of the graph are achieved by using 8 layers of Transformer encoding layers connected in series. The Transformer encoding layer uses 8 attention heads, and the number of nodes in the hidden layer of the Transformer encoding layer is (512, 1024, 1024, 512, 1024, 1024, 2048, 2048) respectively. The final output visual matching embedding vector: , The output dimension is .

8. The industrial robot double-point incremental forming method based on a multimodal large model as claimed in claim 6 is characterized in that: The working process of the sensor encoder of the trained third deep learning model includes: First, the forming wall angle extracted in the forming scene and three-dimensional force value Expand the dimension to obtain augmented data ; Perform feature concatenation on the generated data and merge them into a new feature vector: , Represents the data of the source sensor, with a dimension of (n, 256); A four-layer Transformer encoding layer is used in series. The Transformer encoding layer uses four attention heads. The number of nodes in the four hidden layers of the Transformer encoding layer are (512, 512, 1024, 1024) respectively, and each hidden layer has four attention heads. The final output sensor matching embedding vector .

9. The industrial robot double-point incremental forming method based on a multimodal large model as claimed in claim 6, characterized in that: The language encoder of the trained third deep learning model works as follows: First, enter the text Perform the cutting operation, train the word segmentation data set for WordPiece, and use WordPiece to preprocess the text information to obtain: ; in, Represents the word segmenter in a large language model; For the text embedding operation, the lookup table is first randomly initialized, the index of the text information corresponds to the row sequence of the lookup table, and the text information is Mapped to a 2048-dimensional vector space to obtain a preliminary text embedding vector , the dimension is ; At the same time, relative position encoding is used for the preliminary embedding vector, and finally the text information is encoded into a text latent vector By sequentially connecting 12 layers of Transformer, stacking the self-attention mechanism and feedforward neural network, and setting the final number of hidden layers to 4096, the language matching embedding vector is obtained. , the output dimension is ; The third deep learning model after training also includes the following working process: setting up a projection layer after the visual encoder and sensor encoder, in order to realize the alignment of visual features, sensor features and text features in dimension, and using a multi-layer perceptron to embed the sensor matching vector Matching embedding vectors with visual For processing, in addition to the linear transformation of the input features, the multi-layer perceptron also uses the RELU function as the activation function to increase the nonlinear ability of the network, and the matching vector ; ; Visual vector , sensor vector Vector with text Perform serial splicing to obtain feature fusion vector ( ); Call the large language model Vicuna LLM to fusion vector feature Perform reasoning to obtain description text and preliminary compensation value.

10. The industrial robot double-point incremental forming system based on a multimodal large model is characterized by: include: A compensation information module is configured to: obtain multimodal data of a main forming robot and an auxiliary forming robot during a dual-point incremental forming process; The multimodal data includes: basic structure data of the robot, six-dimensional force value data, source point cloud data of the plate to be processed, robot posture data and tool head forming motion deviation value; A joint stiffness coefficient determination module is configured to: input the multimodal data into the trained first deep learning model to obtain the joint stiffness coefficient; A module for determining registered point cloud data is configured to: input the source point cloud data and the target point cloud data of the plate to be processed into the trained second deep learning model to obtain registered point cloud data; A preliminary compensation value determination module is configured to: input the six-dimensional force data collected by the sensor, the robot posture data, the source point cloud data of the plate to be processed, the tool head forming motion deviation value, and the text description data of the forming scene into the trained third deep learning model to obtain the description text and the preliminary compensation value; a predicted compensation value determination module, configured to: input the registered point cloud data, joint stiffness coefficient, description text, and preliminary compensation value into a fourth deep learning model to obtain a predicted compensation value; The output module is configured to complete the forming process of the plate to be processed according to the predicted compensation value.

Citation Information

Patent Citations

  • Industrial part rapid pose estimation method based on deep learning and point cloud

    CN116580084A

  • Industrial robot motion planning method based on diffusion model

    CN119217373A

  • Double-point incremental forming manufacturing method and apparatus based on deep reinforcement learning

    WO2023202312A1

Cited By

  • Aluminum die-casting spraying steady-state compensation control method based on multi-dimensional perception and production line

    CN121928570A

  • Intelligent tool path compensation method for robot incremental forming

    CN122219305A