A Two-Point Incremental Forming Method and System for Industrial Robots Based on Multimodal Large Model
By fusing visual and force data through a multimodal large model, the dual-point progressive forming process of industrial robots was optimized, solving the problem of insufficient forming accuracy and achieving higher forming accuracy and path optimization.
Patent Information
- Application Number
- CN202510999090.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-07-21
AI Technical Summary
In existing technologies, the two-point incremental forming process of industrial robots suffers from insufficient forming accuracy. In particular, due to factors such as insufficient rigidity of the robot body and springback of the sheet metal, the forming accuracy is not high. Furthermore, existing deep learning methods have failed to effectively solve the error problem in the forming process.
A multimodal large model is adopted, integrating visual and force data. Multimodal data fusion is performed through a deep learning model, including the robot's basic structure, six-dimensional force values, point cloud data, and tool head motion deviation, to build a multimodal perception system and optimize the forming path to improve accuracy.
This technology improves the forming accuracy of robots. By using multimodal data interaction and deep learning models, it acquires forming process data in real time, optimizes the forming path, eliminates errors caused by robot body properties, and improves the forming accuracy of parts.
Smart Images

Figure CN120503178B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of incremental forming technology, and in particular to a two-point incremental forming method and system for industrial robots based on a multimodal large model. Background Technology
[0002] Incremental forming is a moldless forming process that uses a specialized tool head to gradually move across the material surface, applying continuous local pressure to induce plastic deformation in adjacent areas. It is mainly divided into single-point incremental forming and two-point incremental forming. Two-point incremental forming improves forming accuracy by using the coordinated control of two robots to move two tool heads synchronously on both sides of the sheet metal. However, due to insufficient rigidity of the robot body and factors such as springback of the sheet metal during forming, the forming accuracy still has certain limitations.
[0003] With the development and widespread application of deep learning in recent years, its application in the field of incremental forming has also gradually deepened. Existing technologies combine vision systems with deep learning to repair defects in the forming process, thereby improving the stiffness and flatness of incrementally formed parts, but do not consider the accuracy of robot incremental forming. Existing technologies also use the error values in the forming process as state vectors for training to optimize the forming path, using only the deviation between the actual state and the target state as training data. However, the forming accuracy in the forming process is affected by multimodal factors. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a two-point progressive forming method and system for industrial robots based on a multimodal large model. This invention introduces more multidimensional data inputs, such as vision and force perception, resulting in more specific and richer interactive data. Furthermore, it incorporates a multimodal large model as a training model, thereby improving the accuracy of progressive forming path deviation compensation and enhancing the forming accuracy of parts.
[0005] On the one hand, a two-point incremental forming method for industrial robots based on a multimodal large model is provided, including:
[0006] Acquire multimodal data of the main forming robot and the auxiliary forming robot during the two-point progressive forming process; the multimodal data includes: basic robot structure data, six-dimensional force value data, source point cloud data of the sheet material to be processed, robot pose data, and tool head forming motion deviation value;
[0007] Multimodal data is input into the first deep learning model after training to obtain joint stiffness coefficients;
[0008] The source point cloud data and target point cloud data of the board to be processed are input together into the trained second deep learning model to obtain the registered point cloud data.
[0009] The six-dimensional force data collected by the sensor, the robot pose data, the source point cloud data of the sheet material to be processed, the tool head forming motion deviation value, and the text description data of the forming scene are all input into the trained third deep learning model to obtain the description text and the preliminary compensation value.
[0010] The registered point cloud data, joint stiffness coefficients, descriptive text, and preliminary compensation values are input into the fourth deep learning model to obtain the predicted compensation values.
[0011] Based on the predicted compensation value, the forming process of the sheet material to be processed is completed.
[0012] It should be understood that the source point cloud data of the sheet material to be processed is the surface contour point cloud of the actual formed part. The target point cloud data is the contour point cloud data of the part design.
[0013] The robot pose data consists of the joint angles of each joint of the robot in a certain posture, and is exported from the robot control cabinet.
[0014] The joint stiffness coefficient is a constant value, which is the robot body attribute value obtained through training (the Cartesian stiffness coefficients at different robot joint angles derived from this are variables).
[0015] The tool head forming motion deviation value is the deviation between the actual position and the target position of the robot forming tool when processing a certain point.
[0016] On the other hand, a two-point progressive forming system for industrial robots based on a multimodal large model is provided, including:
[0017] The compensation information module is configured to: acquire multimodal data of the main forming robot and the auxiliary forming robot during the two-point progressive forming process; the multimodal data includes: robot basic structure data, six-dimensional force value data, source point cloud data of the sheet material to be processed, robot pose data, and tool head forming motion deviation value;
[0018] The joint stiffness coefficient determination module is configured to: input multimodal data into the trained first deep learning model to obtain the joint stiffness coefficients;
[0019] The registered point cloud data determination module is configured to input the source point cloud data and target point cloud data of the board to be processed into the trained second deep learning model to obtain the registered point cloud data.
[0020] The preliminary compensation value determination module is configured to input the six-dimensional force data collected by the sensor, the robot pose data, the source point cloud data of the sheet material to be processed, the tool head forming motion deviation value, and the text description data of the forming scene into the trained third deep learning model to obtain the description text and the preliminary compensation value.
[0021] The prediction compensation value determination module is configured to input the registered point cloud data, joint stiffness coefficients, descriptive text, and preliminary compensation value into the fourth deep learning model to obtain the prediction compensation value.
[0022] The output module is configured to complete the forming process of the sheet material to be processed based on the predicted compensation value.
[0023] The above technical solution has the following advantages or beneficial effects:
[0024] 1. Build a multimodal perception system, add vision and force sensors to the robot's two-point progressive forming system, and output joint angle information in the robot control cabinet to obtain comprehensive forming data in real time during the forming process, thereby improving the comprehensiveness of the training model.
[0025] 2. Incorporate robot joint stiffness attributes into the mathematical model of the robot constructed in deep learning. Starting from the robot's body attributes, use an LSTM model to train and learn the forming pose and the conventional force feedback compensation model based on joint stiffness attributes.
[0026] 3. After eliminating forming errors caused by robot body attributes, the point cloud information of the surface of the part to be formed is acquired by a camera. This information is then used to achieve multimodal interaction with various process parameters such as forming force and wall angle. A large-scale visual-language-motion multimodal model is constructed by fusing networks using a vision-language-action multimodal model. Through the transformer attention mechanism and cross-modal interaction module, the optimized robot forming pose and forming path are trained, and error analysis is performed using point cloud registration methods.
[0027] 4. The image-text pairs are converted into appropriate instruction formats. A series of questions are generated to guide the assistant in describing the image content, thus realizing the reshaping of multimodal instruction data. The model adopts a visual instruction adjustment method, which maps image features to language feature space through a linear projection layer to achieve the fusion of image and text information. Attached Figure Description
[0028] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0029] Figure 1 This is a schematic diagram of a two-point progressive forming device;
[0030] Figure 2 This is a schematic diagram of robot stiffness.
[0031] Figure 3 This is a flowchart of the method of the present invention;
[0032] The components include: 1. the sheet material to be processed; 2. the sheet material clamp; 3. the camera; 4. the first force sensor; 5. the second force sensor; 6. the main forming tool head; and 7. the auxiliary forming tool head. Detailed Implementation
[0033] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0034] Example 1, as Figure 3 As shown, the two-point progressive forming method for industrial robots based on a multimodal large model includes:
[0035] S101: Acquire multimodal data of the main forming robot and the auxiliary forming robot during the dual-point progressive forming process; the multimodal data includes: robot basic structure data, six-dimensional force value data, source point cloud data of the sheet material to be processed, robot pose data, and tool head forming motion deviation value;
[0036] S102: Input the multimodal data into the trained first deep learning model to obtain the joint stiffness coefficients;
[0037] S103: Input the source point cloud data and target point cloud data of the board to be processed into the trained second deep learning model to obtain the registered point cloud data;
[0038] S104: Input the six-dimensional force data collected by the sensor, the robot pose data, the source point cloud data of the sheet material to be processed, the tool head forming motion deviation value, and the text description data of the forming scene into the trained third deep learning model to obtain the description text and the preliminary compensation value.
[0039] S105: Input the registered point cloud data, joint stiffness coefficients, descriptive text, and preliminary compensation values into the fourth deep learning model to obtain the predicted compensation values;
[0040] S106: Based on the predicted compensation value, complete the forming process of the sheet material to be processed.
[0041] Furthermore, the method also includes:
[0042] S107: Calculate the Euclidean distance between each point in the target point cloud on the surface of the material to be processed and the nearest point in the source point cloud data of the material to be processed, and calculate the average distance of all Euclidean distances, and output the average forming error, maximum error and minimum error.
[0043] Furthermore, S101: Multimodal data, including: robot basic structure data, six-dimensional force value data, source point cloud data of the sheet material to be processed, robot pose data, and tool head forming motion deviation value;
[0044] The basic structural data of the robot includes: link length, link offset, joint angle, and joint torsion angle;
[0045] Based on the robot's basic structural data, a transformation matrix between adjacent links is established; based on the transformation matrix between adjacent links, the total transformation matrix from the base coordinate system to the end effector coordinate system is obtained.
[0046] The tool head forming motion deviation value refers to the positional deviation of the main forming tool head at each forming point;
[0047] The six-dimensional force data refers to the six-dimensional force value of the main forming tool head at each forming point.
[0048] Furthermore, such as Figure 1 As shown, both the main forming robot and the auxiliary forming robot are industrial robots. The main forming robot has a first end effector installed at the end of its robotic arm, and the auxiliary forming robot has a second end effector installed at the end of its robotic arm. The first end effector is equipped with a main forming tool head 6, and the second end effector is equipped with an auxiliary forming tool head 7. A first force sensor 4 is provided between the first end effector and the main forming tool head 6, and a second force sensor 5 is provided between the second end effector and the auxiliary forming tool head 7.
[0049] The main forming tool head 6 and the auxiliary forming tool head 7 are positioned between a plate to be processed 1, which is fixed by a plate clamp 2; the first end effector is also connected to a camera 3; the camera 3 is used to capture images of the processing of the surface of the plate to be processed.
[0050] Furthermore, based on the robot's basic structural data, a transformation matrix is established between adjacent links; based on the transformation matrix between adjacent links, a total transformation matrix is obtained from the base coordinate system to the end effector coordinate system, including:
[0051] Using the mDH modeling method (Craig JJ. Introduction to Robotics: Mechanics and Control [M]. Pearson Education, Inc., 1986.), four parameters are defined for each link: link length. Linkage offset Joint angle Joint torsion angle The transformation matrix between adjacent links is established using four parameters. Transformation matrix Set the current coordinate system Transform to target coordinate system .
[0052]
[0053] By applying the transformation matrix continuously Calculate the total transformation matrix from the base coordinate system to the end effector coordinate system. ,in, Indicates the first Transformation matrices.
[0054] It should be understood that if there is a robotic arm, it is... It consists of several joints, so a DH parameter transformation matrix is defined for each joint. Described from the first The coordinate system of the first joint is transferred to the first... The transformation of the coordinate systems of each joint. The total transformation matrix is obtained by successively multiplying the transformation matrices. This represents the coordinate system from the base of the robotic arm (or the coordinate system of the first joint) to the end effector (or the coordinate system of the second joint). Complete transformation of the coordinate systems of each joint. Total transformation matrix. It contains all the joint angle information of the robotic arm and can be used to calculate the position and orientation of the end effector.
[0055] Furthermore, such as Figure 2 As shown, the tool head forming motion deviation value refers to the positional deviation of the main forming tool head or auxiliary forming tool head at each forming point. Specifically, the positional deviation of the main forming tool head or auxiliary forming tool head at each forming point refers to the deviation between the actual position reached by the tool head and the target position. This means using a camera to measure the positional deviation of the main forming tool head or auxiliary forming tool head at each forming point in real time. Based on the previous point cloud registration data, the spatial correspondence between the forming position points and the target mold point cloud data is obtained. Therefore, the spatial distance difference between the two in the X, Y, and Z directions is expressed as... .
[0056] Furthermore, the six-dimensional force data refers to the six-dimensional force value of the main forming tool head or auxiliary forming tool head at each forming point, which is recorded by a force sensor at the corresponding point. .
[0057] The six-dimensional force value refers to the six independent force and torque components that this indenter can measure or apply at each forming point. Specifically, the six-dimensional force value includes three force components. and three torque components .
[0058] Further, in S102: the multimodal data is input into the trained first deep learning model to obtain joint stiffness coefficients, wherein the trained first deep learning model includes:
[0059] The first LSTM layer, the first Dropout layer, the second LSTM layer, the second Dropout layer, the computation layer, and the first fully connected layer are connected in sequence.
[0060] Furthermore, the first LSTM layer and the second LSTM layer have the same function; the first LSTM layer is used to extract features.
[0061] Furthermore, the first Dropout layer and the second Dropout layer operate in the same way, both serving to improve the generalization ability of the first deep learning model.
[0062] Furthermore, the computational layer is implemented using a multilayer perceptron, for the first... The hidden layer neurons are calculated as follows:
[0063] ;
[0064] in, This represents the weights between the input layer and the first hidden layer. This represents the bias term of the first hidden layer. This represents the activation function.
[0065] For the The calculation method for each output layer neuron is as follows:
[0066] ;
[0067] in, This represents the weights between the first hidden layer and the output layer. This represents the bias term of the output layer. This represents the activation function.
[0068] Furthermore, the first fully connected layer is used as an output layer.
[0069] Further, in S102: the multimodal data is input into the trained first deep learning model to obtain joint stiffness coefficients, wherein the training process of the trained first deep learning model includes:
[0070] Construct a first dataset, which consists of multimodal data with known robot joint stiffness coefficients. Divide the first dataset into a training set, a test set, and a validation set in a 6:2:2 ratio, and set the batch size (batch_size) to 50 samples per training iteration. Output the joint stiffness coefficients through training. .
[0071] The training set is input into the first deep learning model to train it. When the loss function value of the first deep learning model no longer decreases, the training is stopped, and the trained first deep learning model is obtained.
[0072] Furthermore, the loss function for the first deep learning model is chosen to be the mean squared error. As a loss function, and using the six-dimensional force value The joint stiffness coefficients are substituted into the robotic arm control cabinet to determine the target's motion position. Motion position after stiffness compensation To bridge the gap and provide training.
[0073] The loss function of the first deep learning model is as follows:
[0074] ;
[0075] ;
[0076] in, This indicates that under the given stiffness conditions, the following is performed. Compensation was performed through repeated experiments. This represents the joint stiffness coefficient of the current robotic arm. This represents the current six-dimensional force value at the end effector of the robotic arm. This means that the six-dimensional force value is first decomposed into components along the robot arm's coordinate system through coordinate transformation. The force values are transformed from the world coordinate system to the robot arm coordinate system, and then the joint stiffness coefficients are... composition diagonal matrix and using a diagonal matrix Calculate the stiffness coefficient of the forming tool head in Cartesian coordinates. ,in, ;
[0077] Calculate the deformation of the robotic arm caused by external forces: ; It is a change in position, and the calculated change in position will be used to further explain this. The current position is added to the end effector of the robotic arm to obtain the compensated position; where yes The matrix, yes The matrix.
[0078] Further, in step S102: the multimodal data is input into the trained first deep learning model to obtain joint stiffness coefficients, wherein the trained first deep learning model includes the following working process:
[0079] First, two stacked LSTM layers are introduced, with 64 hidden units in each layer. A Dropout layer is immediately following each LSTM layer, with a Dropout rate of 0.2 to improve the model's generalization ability. A computational layer is added after the two LSTM layers to embed the stiffness model. Finally, a fully connected layer is added as the output layer, with the number of neurons in the fully connected layer set to the number of joint stiffness coefficients as data features. .
[0080] Further, in step S103: the source point cloud data and target point cloud data of the board to be processed are jointly input into the trained second deep learning model to obtain the registered point cloud data, wherein the trained second deep learning model includes:
[0081] The point cloud feature extraction module and the point cloud feature fusion module are connected in sequence;
[0082] The point cloud feature extraction module includes: a first branch and a second branch in parallel;
[0083] The first branch includes: a first K-nearest neighbor algorithm layer, a first max pooling layer, a first multilayer perceptron, a first pooling layer, a second multilayer perceptron, and a second max pooling layer connected in sequence; the input value of the first K-nearest neighbor algorithm layer is source point cloud data;
[0084] The second branch includes: a second K-nearest neighbor algorithm layer, a third max pooling layer, a third multilayer perceptron, a second pooling layer, a fourth multilayer perceptron, and a fourth max pooling layer connected in sequence; the input value of the second K-nearest neighbor algorithm layer is the target point cloud data;
[0085] The point cloud feature fusion module includes: a stitching operation layer and a second fully connected layer connected in sequence; the output of the second fully connected layer outputs a pose transformation vector.
[0086] The outputs of the second and fourth max pooling layers are both connected to the input of the splicing operation layer.
[0087] Furthermore, the training process of the trained second deep learning model includes:
[0088] Construct a second dataset, which consists of source point cloud data and target point cloud data with known pose transformation vectors;
[0089] The second dataset is input into the second deep learning model to train it. Training is stopped when the loss function value of the second deep learning model no longer decreases, and the trained second deep learning model is obtained.
[0090] Furthermore, the loss function of the second deep learning model is set to minimize the distance between corresponding points in the source point cloud and the target point cloud. This is achieved by calculating the distance between corresponding points in the predicted transformed source point cloud and the target point cloud. The loss function of the second deep learning model is set as follows:
[0091] ;
[0092] in, Represents the source point cloud dataset. This represents the data obtained by mapping source point cloud data to target point cloud. Represents the target point cloud dataset. X This represents a single sample in the source point cloud dataset. This represents the reciprocal of the number of samples in the dataset in the source point cloud.
[0093] Furthermore, the working process of the trained second deep learning model includes:
[0094] The input to the trained second deep learning model is the source point cloud. and target point cloud ;
[0095] The second branch works in the same way as the first branch, except that the input value of the first branch is the source point cloud, while the input value of the second branch is the target point cloud. Only the working process of the first branch will be explained:
[0096] Input point cloud data , For each point Find its k nearest points as the edge set. For each edge combination Calculate its edge convolution features:
[0097] ;
[0098] Represents the location points of the input point cloud. express The spatial distance between the edge aggregation points and their surrounding contour features can be used to extract point cloud contour features through pooling.
[0099] The edge features of each point are calculated using the first multilayer perceptron. tensor The multilayer perceptron comprises four interconnected neural network layers; each neural network layer includes an input layer, a first hidden layer, a second hidden layer, and an output layer, which are connected in sequence.
[0100] Next, in the first pooling layer, for each point... of Pooling operation is performed on the edge features to obtain tensor ;
[0101] After that, The dimensionality is increased by using a second multilayer perceptron, resulting in a 1024-dimensional tensor. ;
[0102] Finally, the second max pooling layer... Perform max pooling to obtain the global feature vector. .
[0103] Furthermore, the point cloud feature fusion module will integrate the source point cloud... global feature vector and target point cloud global feature vector Perform a data merging operation to obtain a 2048-dimensional feature vector. .
[0104] Furthermore, the point cloud feature fusion module uses a second fully connected layer (FC) to regress the fused feature vector. The second fully connected layer has four hidden layers of sizes 1024, 512, 512, and 256, respectively, and outputs a seven-dimensional estimated pose transformation vector. Attitude transformation vector From the translation of three-dimensional vectors and rotation quaternions composition.
[0105] attitude transformation vector The rotation quaternion in the equation is converted to a rotation matrix R:
[0106] ,
[0107] in, Indicates rotation angle information. Represents the unit vector of the rotation axis Multiply by the sine of half the rotation angle, Represents the unit vector of the rotation axis Multiply by the sine of half the rotation angle, Represents the unit vector of the rotation axis Multiply by the sine of half the rotation angle.
[0108] And combined with translation matrix Applying to the source point cloud, the points after attitude transformation Represented as:
[0109] ;
[0110] in, This represents the transformed point cloud. Indicates the input point cloud. Represents the rotation matrix. Represents the translation matrix;
[0111] Complete the registration of the source point cloud and the target point cloud, and obtain the deviation value:
[0112] ;
[0113] in, Represents the target point cloud, This represents the deviation value in the first direction. This represents the deviation value in the second direction. This represents the deviation value in the third direction, where the first direction can be the X direction of the coordinate axis, the second direction can be the Y direction of the coordinate axis, and the third direction can be the Z direction of the coordinate axis.
[0114] Further, in step S104: the six-dimensional force data collected by the sensors, the robot pose data, the source point cloud data of the sheet material to be processed, the tool head forming motion deviation value, and the text description data of the forming scene are jointly input into the trained third deep learning model to obtain the description text and preliminary compensation value. The trained third deep learning model includes:
[0115] The third, fourth, and fifth branches are parallel;
[0116] The third branch includes: a visual encoder, a first projection layer, and a fifth multilayer perceptron connected in sequence;
[0117] The fourth branch includes: a sensor encoder, a second projection layer, and a sixth multilayer sensor connected in sequence;
[0118] The fifth branch includes: a language encoder;
[0119] The outputs of the fifth and sixth multilayer perceptrons and the language encoder are all connected to the input of the splicing unit. The output of the splicing unit is connected to the input of the large language model. The output of the large language model is used to output the descriptive text and the preliminary compensation value.
[0120] The visual encoder, sensor encoder, and language encoder can all be implemented using the transformer model.
[0121] The target data is the joint angle data that controls the robot's movement, and the descriptive text is a textual description of the operation that the robot is about to perform, such as: polishing the protrusions of the mold.
[0122] Furthermore, the first and second projection layers refer to projection layers, which are used to map the encoding results of different modalities to a unified representation space, and then input them into the Transformer model. The tokens output by the model are then mapped back to the decoder of the original modality through the projection layers to generate corresponding shaping instructions.
[0123] The projection layer is simply a matrix multiplication, a regular / dense / linear layer, without any non-linear activation (sigmoid / tanh). For example, it projects a 100K-dimensional discrete vector onto a 600-dimensional continuous vector (the dimensions are randomly chosen). The exact matrix parameters are learned through the training process. The projection layer introduces a learnable projection matrix. To achieve compression, the format is... Replace multiplication with ,pass Will Projecting to a lower-dimensional space typically requires less storage space to store learnable parameters.
[0124] For model building, 595K high-quality image-text data pairs were selected from the multimodal forming dataset, namely the corresponding forming scene information and point cloud deviation value information. Information such as forming sensor data; inputting language commands in progressive forming scenarios, utilizing real-time image information and sensor data during the forming process, within the existing target deviation value. Based on this, forming accuracy compensation was performed. A Clip-Sensor model was built, whose main structure consists of a visual encoder, a text encoder, and a sensor encoder.
[0125] Further, in step S104: the six-dimensional force data collected by the sensor, the robot pose data, the source point cloud data of the sheet material to be processed, the tool head forming motion deviation value, and the text description data of the forming scene are input together into the trained third deep learning model to obtain the description text and preliminary compensation value. The working process of the visual encoder of the trained third deep learning model includes:
[0126] Extracting dimensions of [dimensional value] from a defined scene The original image ;
[0127] For the original image conduct The segmentation is obtained Initial image of dimension ;
[0128] Initial image go through The flattening operation and the linear layer operation yield a 512-dimensional embedding vector. :
[0129] ;
[0130] Among them, the weight vector The dimension is For embedded vectors Perform absolute position encoding; flattening operation is about to begin. Turn to ;
[0131] Feature extraction and spatial mapping of graphics are achieved by using eight sequentially connected Transformer coding layers (including feedforward network). The Transformer coding layer uses eight attention heads, and the number of nodes in the hidden layers of the Transformer coding layer are (512, 1024, 1024, 512, 1024, 1024, 2048, 2048).
[0132] The final output is the visual matching embedding vector:
[0133] ,
[0134] Output dimension is .
[0135] It should be understood that the Transformer encoding layer, see Attention Is All You Need, Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin.
[0136] It should be understood that the absolute position encoding refers to adding a position vector to each element of the input sequence to represent the specific position of that element in the sequence. This position vector is typically generated by a CLIP-defined function, is independent of the input data, and can capture relative position information in the sequence.
[0137] Furthermore, the working process of the sensor encoder of the trained third deep learning model includes:
[0138] First, the forming wall corners extracted from the forming scene are... With three-dimensional force value Expand the dimensions to obtain augmented data. ;
[0139] The generated data is then concatenated to form a new feature vector:
[0140] ,
[0141] The data represents the source sensor data, with dimensions (n, 256).
[0142] The Transformer coding layer is constructed by cascading four layers in sequence. Each Transformer coding layer uses four attention heads. The number of nodes in the four hidden layers of the Transformer coding layer are (512, 512, 1024, 1024) respectively, and each hidden layer has four attention heads.
[0143] The final output sensor matching embedding vector .
[0144] It should be understood that dimensional expansion involves calculating the statistical characteristics, normalization characteristics, and time series characteristics of the data. Statistical characteristics include: mean, standard deviation, and variance; time series characteristics are the sensor data that change over a period of time; and normalization characteristics are the results after the data has been normalized.
[0145] It should be understood that, since incremental forming involves the tool head pressing down on the sheet material layer by layer, the forming angle is the angle between the line connecting the instantaneously formed position of the workpiece and the corresponding position of the previous layer, and the plane of the sheet material before forming. This angle is one of the key parameters affecting the forming quality and a crucial evaluation indicator for assessing the application of forming technology on complex surfaces. The size of the forming angle directly affects the geometric accuracy of the formed part and the feasibility of the incremental forming process.
[0146] Furthermore, the language encoder of the trained third deep learning model works as follows:
[0147] First, input the text. To perform the segmentation operation, we trained on a WordPiece word segmentation dataset, using WordPiece preprocessed text information to obtain:
[0148] ;
[0149] in, This refers to the word segmenter in a large language model.
[0150] For text embedding operations, the lookup table is first randomly initialized, and the index of the text information corresponds to the row sequence of the lookup table. The text information is then... (excluding) Mapping to a 2048-dimensional vector space yields preliminary text embedding vectors. , dimension ;
[0151] Simultaneously, relative position encoding is applied to the initial embedding vector, ultimately encoding the text information into a text latent vector. By sequentially cascading 12 Transformer layers, stacking self-attention mechanisms and feedforward neural networks, and setting the final number of hidden layers to 4096, language matching embedding vectors are obtained. The output dimension is The language encoder is implemented using the large language model GPT-4.
[0152] Furthermore, the working process of the third deep learning model after training also includes:
[0153] A projection layer is set after the visual encoder and the sensor encoder. To achieve visual features, sensor features, and text features are aligned in dimensions, and a multilayer perceptron is used to match the sensor embedding vector. Matching embedding vectors with vision In addition to the linear transformation of the input features, the multilayer perceptron also uses the ReLU function as an activation function to increase the nonlinearity of the network, resulting in a matching vector. ; ;
[0154] For visual vectors Sensor vector With text vectors By concatenating and splicing the data, a feature fusion vector is obtained. ( );
[0155] Calling the large language model Vicuna LLM to fuse feature vectors Reasoning is performed to obtain the descriptive text and preliminary compensation values.
[0156] Understandably, a large-scale industrial model combining language, vision, and sensor modalities is constructed. The model achieves interaction with the forming environment through cameras and force sensors; based on the existing language model GPT-4, a forming dataset LLM-PCR-PS with language understanding and multimodal instruction generation is constructed.
[0157] GPT-4 possesses multimodal input capabilities, accepting images and text as input and generating text output, enabling it to handle multimodal tasks. The multimodal instruction dataset LLM-PCR-PS, combined with GPT-4's language understanding and generation capabilities, generates motion planning and text descriptions suitable for shaping scenarios. Visual information (such as images and point clouds) is combined with language and angular instructions so that the model can understand and generate specific instructions in shaping tasks.
[0158] LLM stands for Large Language Model. Pre-trained Cross-Modal Representations: This refers to pre-trained cross-modal representations, using GPT4 to obtain a multimodal representation dataset. Processing and Synthesis: This likely refers to the functionality of the dataset or model, namely, processing the multimodal information input (such as images and text, sensor data) and synthesizing new outputs (such as shaped instructions).
[0159] A method for constructing a large industrial model using a large language model (LLM) for industrial scenarios is presented. This large industrial model combines language, vision, and sensor modalities, and enables interaction with the forming environment through cameras and force sensors. Based on the existing language model GPT-4, which has multimodal input capabilities (accepting images and text as input and generating text output to process multimodal tasks), a multimodal instruction forming dataset LLM-PCR-PS with language understanding and generation capabilities is constructed. This multimodal instruction dataset combines the language understanding and generation capabilities of GPT-4 to generate motion planning and text descriptions suitable for forming scenarios. Specifically, it combines visual information (such as images and point clouds) with language instructions and angle instructions, enabling the model to understand and generate specific instructions in the forming task.
[0160] It should be understood that the Vicuna LLM is an existing large language model, and what is being called here is the model backbone, which utilizes the language encoder from Google's LLaMA project.
[0161] Furthermore, S104: the trained third deep learning model, the specific training process includes:
[0162] Construct a third dataset, which consists of six-dimensional force data with known descriptive text and preliminary compensation values, robot pose data, source point cloud data of the sheet material to be processed, tool head forming motion deviation values, and textual description data of the forming scene;
[0163] The third dataset is input into the third deep learning model to train it. Training is stopped when the loss function value of the third deep learning model no longer decreases, and the trained third deep learning model is obtained.
[0164] The textual description data of the forming scene is an objective description of the current forming task and target situation in progressive forming. For example: The robot is completing the polishing work of a cylinder, and the robot's task objective is to correct and improve the original irregularities.
[0165] For training the third deep learning model, pre-training is first performed. Since the features extracted from vision and sensors are not in the same semantic representation space as text features, pre-training is required to align the image and sensor features. During this stage, the weight parameters of the visual encoder, sensor encoder, language encoder, and large language model are frozen, and only the weights of the projection interpolation layer are trained.
[0166] Fusion is performed at different depths of the Projection layer, allowing the model to learn the interactions between features at different levels of abstraction. For training, the number of training epochs is set to 1, the learning rate to 2e-3, and the batch size to 128. The weights of the Projection layer are optimized by minimizing a certain alignment loss (such as cosine similarity loss).
[0167] For example, the weight of the Projection layer is... The image features after the Projection layer are represented as follows: The loss value is then expressed as , This represents the image features before the Projection layer processing.
[0168] To optimize model performance on specific tasks, end-to-end training is performed; the Transformer blocks of the Projection interpolation layer and the Vicuna LLM are updated simultaneously during training. These layers are allowed to update their weights during training; Automatic Mixed Precision (AMP) is enabled to handle multiple data types; the training epochs are set to 3, which can be adjusted appropriately based on the model's performance on the validation set. The learning rate is reduced to 2e-5, the batch size is adjusted to 32, and the AdamW optimizer is used. When using AMP, careful optimization settings are necessary to ensure compatibility with mixed precision. The parameters of the Projection layer and LLM are optimized by minimizing the cross-entropy loss of the generated text.
[0169] Based on the above model, it is possible to realize the registration information of point cloud on the surface of the part and the joint stiffness coefficient. Combined with the three-dimensional force value F, robot joint angle, feed rate, depth of cut, temperature, and rotational speed, the springback stiffness compensation command is input into the third deep learning model. Then, starting from the target language information and combining it with multimodal information, a preliminary compensation value is generated using a language model. ,in, This represents the third deep learning model obtained through training. Indicates the joint stiffness coefficient. This indicates the current joint angle information of the robot. This indicates the current force conditions applied to the robot's end effector sensors.
[0170] Further, in step S105: the registered point cloud data, joint stiffness coefficients, descriptive text, and preliminary compensation values are input into the fourth deep learning model to obtain predicted compensation values, including:
[0171] The fourth deep learning model after training is implemented using the transformer model;
[0172] The training process includes: constructing a training set, which consists of registered point cloud data with known initial compensation values, joint stiffness coefficients, descriptive text, and initial compensation values;
[0173] The fourth deep learning model is trained using the training set to obtain the trained fourth deep learning model.
[0174] Further, S107: Calculate the Euclidean distance between each point in the target point cloud on the surface of the material to be processed and the nearest point in the source point cloud data of the material to be processed, and calculate the average distance of all Euclidean distances, outputting the average forming error, maximum error, and minimum error, specifically including:
[0175] After the forming is completed, based on the second deep learning model, the mold obtained by forming is registered for the second time. The mold after forming is registered with the target mold, and the average Euclidean distance is calculated. The Euclidean distance of each point in the actual part surface point cloud relative to the nearest point in the theoretical model surface point cloud is calculated, and the average distance is calculated to output the average forming error and the maximum and minimum errors to verify the working effect of the model.
[0176] The error function is:
[0177] ;
[0178] in, This represents the formed point cloud dataset. This represents the data obtained by mapping the formed point cloud data to the target point cloud. Let X represent the target point cloud dataset, and let X represent a single sample in the shaped point cloud dataset. This represents the reciprocal of the number of samples in the dataset after the point cloud is formed;
[0179] The optimal forming effect is achieved when the calculated average Euclidean distance is minimized.
[0180] It should be understood that robot body property indices are collected, and robot joint stiffness properties, robot pose dexterity, and robot stiffness indices are measured using cameras and force sensors.
[0181] The collected multimodal data is processed, and the joint stiffness coefficients of the robot are calculated using the LSTM algorithm. Make predictions.
[0182] The surface point cloud information of the part being formed is acquired in real time by a vision camera. Based on the text describing the point cloud scene, a point cloud encoder is used to obtain point cloud embedding features. In the multimodal large model, a hidden layer is set at each interval to introduce fusion features for multimodal fusion, thereby obtaining panoramic features and realizing point cloud registration.
[0183] Achieve registration of point cloud information and joint stiffness coefficients The system utilizes multimodal interaction with visual information extraction, forming arm angle sensor information, and three-dimensional force values. By extracting multimodal features and fusing them into a network, a large-scale language vision model (VLM) is constructed. Through the transformer attention mechanism and cross-modal interaction module, it achieves accurate prediction of the compensation amount of different feature quantities for the robotic arm and makes forming adjustments based on real-time visual information using language information.
[0184] By integrating multimodal large models, after training, models exceeding a set threshold are selected for encapsulation and deployment to compensate for the incremental shaping process.
[0185] We perform mathematical modeling of the robot and introduce a joint stiffness model. We establish an mDH model based on the actual structural parameters of the robot, and then add stiffness attributes to each joint.
[0186] Simultaneously, different forming objectives can be set, such as using linguistic information to provide real-time feedback on forming progress, improving responsiveness to unexpected situations in the forming scenario. This model can also be applied to automated production and intelligent manufacturing in various fields (such as automotive manufacturing, aerospace, and precision machinery). For example, by utilizing the model's language generation capabilities, combined with sensor data and visual information, it can remotely monitor the production line's operating status and automatically diagnose potential faults. Furthermore, this multimodal large-scale model, designed for real-time forming scenarios, can optimize the model's computational efficiency and response time, promoting efficient and intelligent manufacturing.
[0187] Example 2
[0188] This embodiment provides a two-point progressive forming system for industrial robots based on a multimodal large model, including:
[0189] The compensation information module is configured to: acquire multimodal data of the main forming robot and the auxiliary forming robot during the two-point progressive forming process; the multimodal data includes: robot basic structure data, six-dimensional force value data, source point cloud data of the sheet material to be processed, robot pose data, and tool head forming motion deviation value;
[0190] The joint stiffness coefficient determination module is configured to: input multimodal data into the trained first deep learning model to obtain the joint stiffness coefficients;
[0191] The registered point cloud data determination module is configured to input the source point cloud data and target point cloud data of the board to be processed into the trained second deep learning model to obtain the registered point cloud data.
[0192] The preliminary compensation value determination module is configured to input the six-dimensional force data collected by the sensor, the robot pose data, the source point cloud data of the sheet material to be processed, the tool head forming motion deviation value, and the text description data of the forming scene into the trained third deep learning model to obtain the description text and the preliminary compensation value.
[0193] The prediction compensation value determination module is configured to input the registered point cloud data, joint stiffness coefficients, descriptive text, and preliminary compensation value into the fourth deep learning model to obtain the prediction compensation value.
[0194] The output module is configured to complete the forming process of the sheet material to be processed based on the predicted compensation value.
[0195] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A two-point incremental forming method for industrial robots based on a multimodal large model, characterized by: include: Acquire multimodal data of the main forming robot and the auxiliary forming robot during the two-point progressive forming process; The multimodal data includes: basic robot structure data, six-dimensional force data, source point cloud data of the sheet material to be processed, robot pose data, and tool head forming motion deviation values; Multimodal data is input into the first deep learning model after training to obtain joint stiffness coefficients; The source point cloud data and target point cloud data of the board to be processed are input together into the trained second deep learning model to obtain the registered point cloud data. The six-dimensional force data collected by the sensor, the robot pose data, the source point cloud data of the sheet material to be processed, the tool head forming motion deviation value, and the text description data of the forming scene are all input into the trained third deep learning model to obtain the description text and the preliminary compensation value. The registered point cloud data, joint stiffness coefficients, descriptive text, and preliminary compensation values are input into the fourth deep learning model to obtain the predicted compensation values. Based on the predicted compensation value, the forming process of the sheet material to be processed is completed.
2. The two-point progressive forming method for industrial robots based on a multimodal large model as described in claim 1, characterized in that, Multimodal data is input into the trained first deep learning model to obtain joint stiffness coefficients. The trained first deep learning model includes: The first LSTM layer, the first Dropout layer, the second LSTM layer, the second Dropout layer, the computation layer, and the first fully connected layer are connected in sequence. The loss function for the first deep learning model is chosen to be mean squared error. As a loss function, and using the six-dimensional force value The joint stiffness coefficients are substituted into the robotic arm control cabinet to determine the target's motion position. Motion position after stiffness compensation To bridge the gap and provide training; The loss function of the first deep learning model is as follows: ; ; in, This indicates that under the given stiffness conditions, [the following is done]. Compensation was performed through repeated experiments. This represents the joint stiffness coefficient of the current robotic arm. This represents the current six-dimensional force value at the end effector of the robotic arm. This means that the six-dimensional force value is first decomposed into components along the robot arm's coordinate system through coordinate transformation. The force values are transformed from the world coordinate system to the robot arm coordinate system, and then the joint stiffness coefficients are... composition diagonal matrix and using a diagonal matrix Calculate the stiffness coefficient of the forming tool head in Cartesian coordinates. ,in, ; Calculate the deformation of the robotic arm caused by external forces: ; It is a change in position, and the calculated change in position will be used to further explain this. The current position is added to the end effector of the robotic arm to obtain the compensated position; where yes The matrix, yes Matrix; The first deep learning model after training works as follows: First, two stacked LSTM layers are introduced, followed by a Dropout layer after each LSTM layer to improve the model's generalization ability; a computational layer is added after the two LSTM layers to embed the stiffness model; a fully connected layer is then connected as the output layer, with the number of neurons in the fully connected layer set as the data feature of the joint stiffness coefficients. .
3. The two-point progressive forming method for industrial robots based on a multimodal large model as described in claim 1, characterized in that, The source point cloud data and target point cloud data of the board to be processed are jointly input into the trained second deep learning model to obtain the registered point cloud data. The trained second deep learning model includes: The point cloud feature extraction module and the point cloud feature fusion module are connected in sequence; The point cloud feature extraction module includes: a first branch and a second branch in parallel; The first branch includes: a first K-nearest neighbor algorithm layer, a first max pooling layer, a first multilayer perceptron, a first pooling layer, a second multilayer perceptron, and a second max pooling layer connected in sequence; the input value of the first K-nearest neighbor algorithm layer is source point cloud data; The second branch includes: a second K-nearest neighbor algorithm layer, a third max pooling layer, a third multilayer perceptron, a second pooling layer, a fourth multilayer perceptron, and a fourth max pooling layer connected in sequence; the input value of the second K-nearest neighbor algorithm layer is the target point cloud data; The point cloud feature fusion module includes: a stitching operation layer and a second fully connected layer connected in sequence; the output of the second fully connected layer outputs a pose transformation vector. The outputs of the second and fourth max pooling layers are both connected to the input of the splicing operation layer.
4. The two-point progressive forming method for industrial robots based on a multimodal large model as described in claim 1, characterized in that, The loss function of the second deep learning model is set to minimize the distance between corresponding points in the source point cloud and the target point cloud. This is achieved by calculating the distance between corresponding points in the predicted transformed source point cloud and the target point cloud. The loss function of the second deep learning model is set as follows: ; in, Represents the source point cloud dataset. This represents the data obtained by mapping source point cloud data to target point cloud. Represents the target point cloud dataset. X This represents a single sample in the source point cloud dataset. This represents the reciprocal of the number of samples in the dataset from the source point cloud; The second deep learning model after training operates as follows: The input to the trained second deep learning model is the source point cloud. and target point cloud ; The second branch works in the same way as the first branch, except that the input value of the first branch is the source point cloud, while the input value of the second branch is the target point cloud. Only the working process of the first branch will be explained: Input point cloud data , For each point Find its k nearest points as the edge set. For each edge combination Calculate its edge convolution features: ; Represents the location points of the input point cloud. express The spatial distance between the edge aggregation points and their surrounding contour features is used to extract point cloud contour features by pooling the data. The edge features of each point are calculated using the first multilayer perceptron. tensor The multilayer perceptron comprises four interconnected neural network layers; each neural network layer includes an input layer, a first hidden layer, a second hidden layer, and an output layer connected in sequence. Next, in the first pooling layer, for each point... of Pooling operation is performed on the edge features to obtain tensor ; After that, The dimensionality is increased by using a second multilayer perceptron, resulting in a 1024-dimensional tensor. ; Finally, the second max pooling layer... Perform max pooling to obtain the global feature vector. .
5. The two-point progressive forming method for industrial robots based on a multimodal large model as described in claim 3, characterized in that, The point cloud feature fusion module will integrate the source point cloud... global feature vector and target point cloud global feature vector Perform a data merging operation to obtain a 2048-dimensional feature vector. ; The point cloud feature fusion module uses a second fully connected layer to regress the fused feature vector. This second fully connected layer has four hidden layers of sizes 1024, 512, 512, and 256, and outputs a seven-dimensional estimated pose transformation vector. Attitude transformation vector From the translation of three-dimensional vectors and rotation quaternions composition; attitude transformation vector The rotation quaternion in the equation is converted to a rotation matrix R: , in, Indicates rotation angle information. Represents the unit vector of the rotation axis Multiply by the sine of half the rotation angle, Represents the unit vector of the rotation axis Multiply by the sine of half the rotation angle, Represents the unit vector of the rotation axis Multiply by the sine of half the rotation angle; And combined with translation matrix Applying to the source point cloud, the points after attitude transformation Represented as: ; in, This represents the transformed point cloud. Indicates the input point cloud. Represents the rotation matrix. Represents the translation matrix; Complete the registration of the source point cloud and the target point cloud, and obtain the deviation value: ; in, Represents the target point cloud, This represents the deviation value in the first direction. This indicates the deviation value in the second direction. This represents the deviation value in the third direction.
6. The two-point progressive forming method for industrial robots based on a multimodal large model as described in claim 1, characterized in that, The third deep learning model after training includes: The third, fourth, and fifth branches are parallel; The third branch includes: a visual encoder, a first projection layer, and a fifth multilayer perceptron connected in sequence; The fourth branch includes: a sensor encoder, a second projection layer, and a sixth multilayer sensor connected in sequence; The fifth branch includes: a language encoder; The outputs of the fifth and sixth multilayer perceptrons and the language encoder are all connected to the input of the splicing unit. The output of the splicing unit is connected to the input of the large language model. The output of the large language model is used to output the descriptive text and the preliminary compensation value.
7. The two-point progressive forming method for industrial robots based on a multimodal large model as described in claim 6, characterized in that, The visual encoder of the trained third deep learning model works as follows: Extracting dimensions of [dimensional value] from a defined scene The original image ; For the original image conduct The segmentation is obtained Initial image of dimension ; Initial image go through The flattening operation and the linear layer operation yield a 512-dimensional embedding vector. : ; Among them, the weight vector The dimension is For embedding vectors Perform absolute position encoding; flattening operation is about to begin. Turn to ; Feature extraction and spatial mapping of graphics are achieved by using eight Transformer coding layers connected in series. The Transformer coding layer uses eight attention heads, and the number of nodes in the hidden layers of the Transformer coding layer are (512, 1024, 1024, 512, 1024, 1024, 2048, 2048). The final output is the visual matching embedding vector: , Output dimension is .
8. The two-point progressive forming method for industrial robots based on a multimodal large model as described in claim 6, characterized in that, The working process of the sensor encoder in the trained third deep learning model includes: First, the shaped corners extracted from the shaped scene are... With three-dimensional force value Expand the dimensions to obtain augmented data. ; The generated data is then concatenated to form a new feature vector: , The data represents the source sensor data, with dimensions (n, 256). The Transformer coding layer is constructed by cascading four layers in sequence. Each Transformer coding layer uses four attention heads. The number of nodes in the four hidden layers of the Transformer coding layer are (512, 512, 1024, 1024) respectively, and each hidden layer has four attention heads. The final output sensor matching embedding vector .
9. The two-point progressive forming method for industrial robots based on a multimodal large model as described in claim 6, characterized in that, The language encoder of the third deep learning model after training includes the following processes: First, input the text. To perform the segmentation operation, we trained on a WordPiece word segmentation dataset, using WordPiece preprocessed text information to obtain: ; in, This refers to the word segmenter in a large language model. For text embedding operations, the lookup table is first randomly initialized, and the index of the text information corresponds to the row sequence of the lookup table. The text information is then... Mapping to a 2048-dimensional vector space yields preliminary text embedding vectors. , dimension ; Simultaneously, relative position encoding is applied to the initial embedding vector, ultimately encoding the text information into a text latent vector. By sequentially cascading 12 Transformer layers, stacking self-attention mechanisms and feedforward neural networks, and setting the final number of hidden layers to 4096, language matching embedding vectors are obtained. The output dimension is ; The third deep learning model after training also includes the following steps: a projection layer is set after the visual encoder and the sensor encoder; to achieve dimensional alignment of visual features, sensor features, and text features, a multilayer perceptron is used to match the sensor embedding vectors. Matching embedding vectors with vision In addition to the linear transformation of the input features, the multilayer perceptron also uses the ReLU function as an activation function to increase the nonlinearity of the network, resulting in a matching vector. ; ; For visual vectors Sensor vector With text vectors By concatenating and splicing the data, a feature fusion vector is obtained. ( ); Calling the large language model Vicuna LLM to fuse feature vectors Reasoning is performed to obtain the descriptive text and preliminary compensation values.
10. A two-point progressive forming system for industrial robots based on a multimodal large model, characterized in that: include: The compensation information module is configured to acquire multimodal data of the main forming robot and the auxiliary forming robot during the two-point progressive forming process. The multimodal data includes: basic robot structure data, six-dimensional force data, source point cloud data of the sheet material to be processed, robot pose data, and tool head forming motion deviation values; The joint stiffness coefficient determination module is configured to: input multimodal data into the trained first deep learning model to obtain the joint stiffness coefficients; The registered point cloud data determination module is configured to input the source point cloud data and target point cloud data of the board to be processed into the trained second deep learning model to obtain the registered point cloud data. The preliminary compensation value determination module is configured to input the six-dimensional force data collected by the sensor, the robot pose data, the source point cloud data of the sheet material to be processed, the tool head forming motion deviation value, and the text description data of the forming scene into the trained third deep learning model to obtain the description text and the preliminary compensation value. The prediction compensation value determination module is configured to input the registered point cloud data, joint stiffness coefficients, descriptive text, and preliminary compensation value into the fourth deep learning model to obtain the prediction compensation value. The output module is configured to complete the forming process of the sheet material to be processed based on the predicted compensation value.
Citation Information
Patent Citations
Industrial part rapid pose estimation method based on deep learning and point cloud
CN116580084A
Industrial robot motion planning method based on diffusion model
CN119217373A