Rice machine steel frame automatic welding method based on machine vision
By using a multimodal pruned attention network based on machine vision, combined with 3D CAD and finite element simulation data, the problem of insufficient weld recognition and trajectory generation between complex spatial components was solved, and high-precision and stable welding path planning and control were achieved.
Patent Information
- Application Number
- CN202511017332.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing welding robot systems lack the ability to identify weld structures and generate trajectories between complex spatial components, resulting in low welding efficiency and low path accuracy. They are also highly dependent on operating equipment and manual experience, making it difficult to meet complex and diverse welding needs.
A multimodal pruned attention network based on machine vision is used, combined with 3D CAD assembly models and finite element simulation data, to construct a weld trajectory generation model, achieving deep collaboration and path planning of multimodal information. This includes the fusion of weld image information, structural connection relationships, and stress distribution characteristics, generating a continuous 3D welding trajectory, and performing dynamic feasibility verification and welding gun posture planning.
It improves the accuracy of weld seam recognition and the physical feasibility of trajectory generation, enhances the closed-loop stability of control execution of the welding process, and improves the feasibility of the welding path and the intelligent adaptive capability of the system.
Smart Images

Figure CN120734573A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent manufacturing and welding automation, and in particular to an automatic welding method for a rice machine steel frame based on machine vision. Background Art
[0002] Currently, steel frames are widely used in heavy-load, high-precision industrial applications. Their complex components and rigorous structure place higher demands on the accuracy and strength of welding operations. Traditional welding methods rely primarily on manual or semi-automatic welding equipment. Although some high-end production lines have begun to use robotic automatic welding systems, in most cases they still lack the ability to intelligently identify and generate trajectories for weld structures between complex spatial components. This results in low welding efficiency, insufficient path accuracy, and a strong dependence on equipment. Furthermore, they rely heavily on operating equipment and manual experience, making it difficult to meet the highly complex and diverse welding needs of steel frame structures.
[0003] Existing welding robot systems typically perform path planning based on two-dimensional images or manually defined geometric trajectories, lacking a deep understanding of the spatial characteristics of welds. In terms of image processing, traditional methods often rely on single visual features such as edge detection and grayscale distribution to locate welds. This makes it difficult to cope with actual working conditions such as surface reflections, occlusions, or complex contact relationships between components, resulting in insufficient weld recognition accuracy. Furthermore, current path planning algorithms generally fail to effectively integrate structural correlation information and stress distribution characteristics between components, making it impossible to optimize trajectories based on the actual connection topology and load conditions between components. This limits the feasibility of welding paths and the stability of the welding process.
[0004] Although some studies have incorporated deep learning methods into weld inspection, these approaches have primarily focused on image classification or region recognition tasks. A continuous three-dimensional trajectory output mechanism for welding control has yet to be established. Furthermore, the fusion of multimodal information is relatively crude, making it difficult to fully exploit the synergistic relationship between image features, structural information, and stress distribution. Consequently, the generated weld trajectories lack the accuracy, stability, and physical feasibility required to meet practical requirements. Furthermore, the general lack of dynamic trajectory feasibility verification and welding gun posture planning mechanisms during welding further limits the direct deployment and application of these generated trajectories in high-precision, multi-degree-of-freedom industrial welding arms under complex working conditions.
[0005] Therefore, how to provide a method for automatic welding of rice machine steel frames based on machine vision is a problem that those skilled in the art urgently need to solve. Summary of the Invention
[0006] One purpose of the present invention is to propose an automatic welding method for a rice machine steel frame based on machine vision. The present invention fully integrates weld image information, structural connection relationship and stress distribution characteristics, constructs a weld trajectory generation model based on a multimodal pruning attention network, realizes deep collaboration and path planning of multimodal information, and has the advantages of high weld recognition accuracy, physically feasible trajectory generation, and stable control execution closed loop.
[0007] According to an embodiment of the present invention, a method for automatic welding of a steel frame of a rice machine based on machine vision includes the following steps:
[0008] S1. Capture the image of the area to be welded by an industrial camera to obtain the original image data;
[0009] S2, performing denoising and grayscale normalization on the original image data to generate an enhanced image;
[0010] S3. Based on the 3D CAD assembly model and the enhanced image, construct the weld candidate area image, weld structure description vector and stress weighted area image;
[0011] S4. A weld trajectory generation model is constructed through a multimodal pruning attention network. Multimodal fusion modeling is performed on the weld candidate region image, weld structure description vector, and stress-weighted region image to generate a continuous three-dimensional welding trajectory.
[0012] S5. Perform dynamic feasibility verification and welding gun posture planning on the continuous three-dimensional welding trajectory;
[0013] S6. Discretize the verified three-dimensional welding trajectory according to the set step size to generate a discrete trajectory point sequence;
[0014] S7, binding the matching welding parameters in the process parameter library to the discrete trajectory point sequence to generate a linkage execution instruction set;
[0015] S8. Send the linkage execution instruction set to the welding control unit, and collect real-time welding status data through the control feedback interface to complete the closed-loop control of the welding process.
[0016] Optionally, step S1 further includes:
[0017] S11. Install the industrial camera on the fixed bracket or motion platform of the welding equipment, and adjust the shooting angle and focal length of the industrial camera;
[0018] S12. Acquire high-resolution images of the weld area under standard lighting conditions to form an original image sequence, wherein the high-resolution images include the welding start point, end point, and boundaries of adjacent components;
[0019] S13. Perform time synchronization and coordinate unification processing on the original image sequence to generate original image data.
[0020] Optionally, step S2 further includes: using median filtering to remove isolated bright spots or granular noise in the original image sequence, and performing grayscale normalization processing on the denoised image using a contrast-limited adaptive histogram equalization method to generate an enhanced image.
[0021] Optionally, step S3 further includes:
[0022] S31. Based on the three-dimensional CAD assembly model file generated in the design phase of the rice mill steel frame, the assembly geometry data of the rice mill steel frame structure is imported, the component structure hierarchy and assembly relationship are analyzed, the direct contact area between the components is identified, the welding connection information between the components is extracted, and a component connection relationship matrix is established; the components include steel columns, beams, supports, node plates and stiffening ribs, which are all connecting parts that constitute the spatial structural framework of the rice mill steel frame. In the component connection relationship matrix, if there is a welding connection between two components, the corresponding matrix position is assigned a value of 1, and if there is no welding connection, the corresponding matrix position is assigned a value of 0;
[0023] S32. Based on the component connection relationship matrix, assign a unique identifier to each component, establish a component connection diagram, and perform interconnected component decomposition on the component connection diagram to identify multiple independent connection sub-diagrams, each connection sub-diagram corresponding to a welded structural unit;
[0024] In each connection subgraph, a depth-first traversal algorithm is used to extract the welding connection path, record the topological order of adjacent components and the location information of the connection surface, and combine the contact type, tolerance range and structural connection specifications specified in the welding process to screen the component pairs that meet the conditions from the traversal results to obtain the component space boundary. The component space boundary is then transformed into the image coordinate system of the enhanced image through the calibration matrix, and the image of the weld candidate area is cropped to generate.
[0025] S33. Construct a finite element simulation model based on the three-dimensional assembly model, apply boundary constraints and welding load conditions to each connection subgraph in the component connection diagram, calculate the equivalent stress distribution of the component connection nodes, and obtain the local stress response value of each component to the connection area, wherein the stress response value at each connection node is defined as the maximum equivalent stress in the component connection area. The results are used to construct a stress distribution data set in the form of correspondence between node numbers and stress amplitudes;
[0026] S34. Based on the weld candidate region image, mapping the local stress response value of the component to the connection region to the image coordinate domain to generate a stress-weighted regional image, wherein the stress-weighted regional image represents the distribution characteristics of the component connection strength in the weld candidate region;
[0027] S35. Based on the stress-weighted regional image, a regional guided feature extraction strategy is constructed. The candidate regions are spatially weighted according to the connection strength distribution. The weld edge structural features, grayscale gradient distribution features and heat-affected zone color features are extracted. The pixel-level weighted fusion is performed through the spatial guidance matrix to form a multi-layer feature map of the weld.
[0028] S36, fusing the weld multi-layer feature map with the stress-weighted regional image to generate a weld space priori map, and based on the weld space priori map, extracting the weld centerline, boundary contour, and start and end points by using spatial continuity modeling and geometric pattern fitting to generate a weld structure description vector, wherein the structure description vector includes weld morphology, length, curvature change, and structural distribution characteristics;
[0029] S37. Generate an initial three-dimensional welding trajectory based on the weld structure description vector. The trajectory is smoothed according to the curvature continuity principle, and a node compression strategy is used to reduce the complexity of motion control instructions, and a continuous executable trajectory sequence is output.
[0030] Optionally, the region-guided feature extraction strategy in step S35 is specifically:
[0031] S351, using the stress-weighted region image as a weight map, constructing a spatial guidance matrix, and assigning a spatial response weight to each pixel in the weld candidate region image;
[0032] S352. Use square sliding windows of 3×3, 5×5, and 7×7 pixels to perform sliding scanning on the weld candidate area image in the enhanced image. Perform the following feature extraction operations in each scale window: perform edge detection using the Sobel operator to obtain weld edge structural features, calculate the local gradient pattern using the directional gradient histogram to obtain grayscale gradient distribution features, convert the weld candidate area image to the Lab color space, extract the chromaticity and brightness differences, and obtain the color features of the heat-affected zone;
[0033] S353. The weld edge structural features, grayscale gradient distribution features and color features of the heat-affected zone are weightedly fused at the pixel level through a spatial guidance matrix to form a multi-layer feature map of the weld.
[0034] Optionally, step S4 includes:
[0035] S41. Constructing a weld trajectory generation model through a multimodal pruned attention network, wherein the multimodal pruned attention network includes a multimodal feature fusion encoding layer, a modality collaborative gating layer, a pruned attention layer, and a trajectory regression output layer;
[0036] S42, inputting the weld candidate region image, the weld structure description vector, and the stress-weighted region image into a multimodal feature fusion encoding layer, extracting image texture features, geometric structure features, and connection strength features using independent linear transformations, and mapping the image texture features, geometric structure features, and connection strength features to the same dimensional space through linear projection, and splicing them to form a multimodal joint feature;
[0037] S43. Based on the component connection graph, the multimodal feature fusion encoding layer is introduced into the structure-aware position bias. The structure-aware bias is calculated according to the component connection relationship matrix and the topological distance between component pairs in the component connection graph. The structure-aware bias is then embedded into the attention calculation to form the structure-aware attention weight.
[0038] S44. Fusing the three modal features of the weld candidate region image, the weld structure description vector, and the stress-weighted region image through a modal collaborative gating layer, setting the scalar score value of the weld candidate region image to 1, the scalar score value of the weld structure description vector to 2, and the scalar score value of the stress-weighted region image to 3, assigning a learnable gating weight to each modality, normalizing them so that the sum of the total weights is 1, and fusing them to generate a semantic joint feature;
[0039] S45. Based on the weight value of each pixel position in the weld space prior map, a pruning mask map is constructed through the pruning attention layer, and the pruning mask map is applied to the attention calculation process of the semantic joint feature to crop or locally restrict the attention path of the low-weight area to generate a pruning mask feature: when the weight value of a pixel position is greater than a preset threshold, the corresponding position in the mask map is assigned a value of 1, indicating that the full attention is retained; when the weight value is less than or equal to the threshold, the corresponding position in the mask map is assigned a value of 0, indicating that only local attention is retained or an attention pruning operation is performed;
[0040] S46. Inputting the pruned mask features into the trajectory regression output layer, and based on the principle of weld space topology and geometric continuity, using a multi-layer perceptron to perform point-by-point regression on the weld centerline to generate a regression trajectory point sequence. In the point-by-point regression process, the initial three-dimensional welding trajectory is introduced as a structural guidance path;
[0041] The initial three-dimensional welding trajectory is geometrically consistent with the regression trajectory point sequence, and the node compression algorithm is used to optimize the curvature continuity constraint, eliminating redundant relay points in the path. Only the position points where the curvature change exceeds the set threshold are retained to output a continuous three-dimensional welding trajectory.
[0042] Optionally, step S5 further includes:
[0043] S51, performing spatial accessibility verification on the continuous three-dimensional welding trajectory, and performing spatial interference detection on each trajectory point in the welding path according to the geometric boundary model of the steel frame component;
[0044] S52. Calculate feasible posture solutions corresponding to each trajectory point based on welding gun structural parameters and robot arm kinematic constraints, and perform posture optimization on redundant degrees of freedom to form a continuous posture sequence;
[0045] S53. Combine the trajectory points and posture sequence to determine whether the joint limit, trajectory smoothness and dynamic acceleration constraint requirements are met. If there is an unexecutable trajectory segment, adjust the multimodal pruning attention network parameters and regenerate the three-dimensional welding trajectory; the multimodal pruning attention network parameters include modal encoding parameters, structure-aware position bias parameters, modal collaborative gating parameters, pruning attention mechanism parameters and pruning attention mechanism parameters.
[0046] Optionally, step S6 further includes:
[0047] S61. Evaluate the trajectory length of the three-dimensional welding trajectory that has passed the verification, perform trajectory point interpolation sampling according to the set spatial step size or time interval, generate a trajectory point sequence with equal spacing, and assign a time sequence number to each trajectory point sequence according to the index order;
[0048] S62. Calculate the spatial position and posture Euler angles of each trajectory point sequence in the welding reference system, and solve the expected linear velocity and angular velocity per unit time;
[0049] S63, performing curvature analysis on the trajectory point sequence, dynamically adjusting the sampling density according to the local curvature change, increasing the sampling density in areas where the curvature change exceeds a set threshold, and reducing the sampling density in areas where the curvature change is below the threshold;
[0050] S64: Output a discrete trajectory point sequence, where the discrete trajectory point sequence includes spatial position, attitude Euler angles, and velocity information.
[0051] Optionally, step S7 further includes:
[0052] S71. Based on the spatial position, posture angle, and speed information of each trajectory point in the discrete welding trajectory point sequence, query a preset process parameter library and match optimal welding parameters, wherein the welding parameters include welding current, voltage, welding speed, wire feed rate, and shielding gas flow rate;
[0053] S72, binding the matched welding parameters to the discrete welding trajectory points, and constructing a linkage control structure including the timing number, spatial position, posture angle, and welding parameters;
[0054] S73. According to the welding equipment control interface protocol, the linkage control structure is translated into an executable instruction set to generate a linkage execution instruction set adapted to the welding control unit.
[0055] Optionally, the welding status data includes welding current, welding voltage, welding speed, molten pool temperature, penetration information and weld formation image, and the deviation between the collected welding status data and the expected trajectory is feedback-adjusted to dynamically adjust the welding parameters and trajectory execution rate to achieve continuous monitoring and control closed-loop feedback of the welding process.
[0056] The beneficial effects of the present invention are:
[0057] First, by introducing a 3D CAD assembly model and finite element simulation data, this invention constructs a spatial priori map of the weld seam before weld identification. This enables spatial guidance and stress-weighted control of the weld candidate region, significantly improving the accuracy and robustness of weld region extraction. A structural topology model established through a component connection relationship matrix enables weld location to be based not only on the visual features of the image but also on the actual connection relationships and stress distribution information between components. This overcomes the technical bottleneck of traditional methods in dealing with complex working conditions such as component occlusion, surface reflections, and blurred boundaries.
[0058] Secondly, the present invention constructs a multimodal pruning attention network, effectively fusing three types of modal data: weld images, structural description vectors, and stress-weighted images. It designs a multimodal feature fusion encoding layer, a modal collaborative gating mechanism, and a structure-aware pruning attention mechanism, achieving deep coupling of multi-source information and selective reduction of redundant paths. By introducing a structure-aware position bias and combining it with a weld spatial prior map to generate a pruning mask, the attention mechanism enhances its ability to model spatial connectivity and weld regions, significantly improving feature representation and trajectory prediction accuracy while reducing redundant computation and improving the overall efficiency of the model.
[0059] Furthermore, during the trajectory generation phase, the present invention introduces an initial welding trajectory as a structural guidance path, and combines it with the output of pruning attention for geometric consistency fusion. This approach improves the safety and rationality of the welding path during spatial execution through dynamic feasibility verification and posture planning. The generated welding trajectory is discretized and then linked to a preset process parameter library to form a control instruction set encompassing position, posture, speed, and welding parameters. This implements closed-loop feedback for welding control, effectively enhancing the system's intelligent adaptability and practical feasibility.
[0060] In summary, the present invention has the advantages of high weld seam recognition accuracy, physically feasible trajectory generation, and stable closed-loop control execution. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0062] Figure 1 Schematic diagram of a method for automatic welding of a steel frame of a rice machine based on machine vision proposed by the present invention;
[0063] Figure 2 It is a flow chart of the construction of weld prior information and the generation of initial trajectory in the present invention;
[0064] Figure 3 It is a schematic diagram of the weld trajectory generation model based on the multimodal pruned attention network in the present invention. DETAILED DESCRIPTION
[0065] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0066] refer to Figure 1-3 , a rice machine steel frame automatic welding method based on machine vision, comprising the following steps:
[0067] S1. Capture the image of the area to be welded by an industrial camera to obtain the original image data;
[0068] S2, performing denoising and grayscale normalization on the original image data to generate an enhanced image;
[0069] S3. Based on the 3D CAD assembly model and the enhanced image, construct the weld candidate area image, weld structure description vector and stress weighted area image;
[0070] S4. A weld trajectory generation model is constructed through a multimodal pruning attention network. Multimodal fusion modeling is performed on the weld candidate region image, weld structure description vector, and stress-weighted region image to generate a continuous three-dimensional welding trajectory.
[0071] S5. Perform dynamic feasibility verification and welding gun posture planning on the continuous three-dimensional welding trajectory;
[0072] S6. Discretize the verified three-dimensional welding trajectory according to the set step size to generate a discrete trajectory point sequence;
[0073] S7, binding the matching welding parameters in the process parameter library to the discrete trajectory point sequence to generate a linkage execution instruction set;
[0074] S8. Send the linkage execution instruction set to the welding control unit, and collect real-time welding status data through the control feedback interface to complete the closed-loop control of the welding process.
[0075] In this embodiment, the step S1 further includes:
[0076] S11. Install the industrial camera on the fixed bracket or motion platform of the welding equipment, and adjust the shooting angle and focal length of the industrial camera;
[0077] S12. Under standard lighting conditions, high-resolution images of the weld area are collected to form an original image sequence, where the high-resolution images include the welding start point, end point, and boundaries of adjacent components. The standard lighting conditions include uniform and stable white light illumination provided by a ring-shaped LED light source or a bar-shaped cold light source, with a color temperature between 4000K and 6000K, an illumination range of 3000lux to 8000lux, and a light direction coaxial with the optical axis of the industrial camera or incident at a low angle.
[0078] S13. Perform time synchronization and coordinate unification processing on the original image sequence to generate original image data.
[0079] In this embodiment, step S2 further includes: using median filtering to remove isolated bright spots or granular noise in the original image sequence, and using a contrast-limited adaptive histogram equalization method to perform grayscale normalization on the denoised image to generate an enhanced image; the contrast-limited adaptive histogram equalization method is: dividing the image into sub-blocks according to preset pixel width and height dimensions, calculating a histogram for each sub-block and setting a clipping threshold, clipping the grayscale levels whose frequency exceeds the threshold, normalizing the clipped histogram and mapping it to a new grayscale value, and performing boundary smoothing on adjacent areas through bilinear interpolation.
[0080] In this embodiment, step S3 further includes:
[0081] S31. Based on the three-dimensional CAD assembly model file generated in the design phase of the rice mill steel frame, the assembly geometry data of the rice mill steel frame structure is imported, the component structure hierarchy and assembly relationship are analyzed, the direct contact area between the components is identified, the welding connection information between the components is extracted, and a component connection relationship matrix is established; the components include steel columns, beams, supports, node plates and stiffening ribs, which are all connecting parts that constitute the spatial structural framework of the rice mill steel frame. In the component connection relationship matrix, if there is a welding connection between two components, the corresponding matrix position is assigned a value of 1, and if there is no welding connection, the corresponding matrix position is assigned a value of 0;
[0082] S32. Based on the component connection relationship matrix, assign a unique identifier to each component, establish a component connection graph, and perform connected component decomposition on the component connection graph to identify multiple independent connection subgraphs, each connection subgraph corresponding to a welded structural unit; the nodes of the component connection graph represent components, and the edges represent pairs of components with welded connections; within each connection subgraph, use a depth-first traversal algorithm to extract all weld connection paths, and record the topological order and connection surface location information of each pair of adjacent components;
[0083] In each connection subgraph, a depth-first traversal algorithm is used to extract the welding connection path, record the topological order of adjacent components and the location information of the connection surface, and combine the contact type, tolerance range and structural connection specifications specified in the welding process to screen the component pairs that meet the conditions from the traversal results to obtain the component space boundary. The component space boundary is then transformed into the image coordinate system of the enhanced image through the calibration matrix, and the image of the weld candidate area is cropped to generate.
[0084] S33. Construct a finite element simulation model based on the three-dimensional assembly model, apply boundary constraints and welding load conditions to each connection subgraph in the component connection diagram, calculate the equivalent stress distribution of the component connection nodes, and obtain the local stress response value of each component to the connection area, wherein the stress response value at each connection node is defined as the maximum equivalent stress in the component connection area. The results are used to construct a stress distribution data set in the form of correspondence between node numbers and stress amplitudes;
[0085] The boundary constraint is to set a zero displacement constraint condition at the fixed end of the component and apply a mirror boundary to the symmetrical structural component; the welding load condition is to apply a linear thermal load and a transient temperature rise boundary along the center line of the fitted weld in the connection area;
[0086] S34. Based on the weld candidate region image, mapping the local stress response value of the component to the connection region to the image coordinate domain to generate a stress-weighted regional image, wherein the stress-weighted regional image represents the distribution characteristics of the component connection strength in the weld candidate region;
[0087] S35. Based on the stress-weighted regional image, a regional guided feature extraction strategy is constructed. The candidate regions are spatially weighted according to the connection strength distribution. The weld edge structural features, grayscale gradient distribution features and heat-affected zone color features are extracted. The pixel-level weighted fusion is performed through the spatial guidance matrix to form a multi-layer feature map of the weld.
[0088] S36. Fusion modeling is performed on the weld multi-layer feature map and the stress-weighted regional image to generate a weld space priori map. The weight value P(x, y) of each pixel point in the weld space priori map is:
[0089] P(x,y)=α·F(x,y)+β·G(x,y);
[0090] Where F(x,y) represents the fusion feature intensity of the pixel point (x,y) in the weld multi-layer feature map, G(x,y) represents the intensity value of the pixel point (x,y) in the stress projection map, α and β are weighting coefficients, satisfying α+β=1;
[0091] Based on the weld space prior map, spatial continuity modeling and geometric pattern fitting are used to extract the weld centerline, boundary contour and start and end points, and generate a weld structure description vector. The structure description vector includes the weld shape, length, curvature change and structural distribution characteristics.
[0092] S37. Generate an initial three-dimensional welding trajectory based on the weld structure description vector. The trajectory is smoothed according to the curvature continuity principle, and a node compression strategy is used to reduce the complexity of motion control instructions, and a continuous executable trajectory sequence is output.
[0093] In this embodiment, the region-guided feature extraction strategy in step S35 is specifically as follows:
[0094] S351, using the stress-weighted region image as a weight map, constructing a spatial guidance matrix, and assigning a spatial response weight to each pixel in the weld candidate region image;
[0095] S352. Use square sliding windows of 3×3, 5×5, and 7×7 pixels to perform sliding scanning on the weld candidate area image in the enhanced image. Perform the following feature extraction operations in each scale window: perform edge detection using the Sobel operator to obtain weld edge structural features, calculate the local gradient pattern using the directional gradient histogram to obtain grayscale gradient distribution features, convert the weld candidate area image to the Lab color space, extract the chromaticity and brightness differences, and obtain the color features of the heat-affected zone;
[0096] S353. Perform pixel-level weighted fusion on the weld edge structure features, grayscale gradient distribution features, and color features of the heat-affected zone using a spatial guidance matrix to form a weld multi-layer feature map. The fusion feature intensity F(x, y) of each pixel in the weld multi-layer feature map satisfies:
[0097] F(x,y)=W(x,y)·φ(I(x,y));
[0098] Where I(x,y) represents the pixel value of the enhanced image, φ(·) represents the set of feature extraction operators, and W(x,y) is the spatial guidance weight value in the stress-weighted regional image.
[0099] In this embodiment, step S4 includes:
[0100] S41. Construct a weld trajectory generation model through a multimodal pruned attention network, wherein the multimodal pruned attention network includes a multimodal feature fusion encoding layer, a modal collaborative gating layer, a pruned attention layer, and a trajectory regression output layer; the multimodal feature fusion encoding layer receives weld candidate area images, weld structure description vectors, and stress-weighted area images, extracts image texture features, geometric structure features, and connection strength features, and generates a structure-aware position bias based on the component connection diagram. and embedded in the attention weight calculation; the modal collaborative gating layer dynamically adjusts the fusion strength of different modal features; the pruning attention layer generates a pruning mask based on the weld space prior map, retaining the global attention path in a specific area; the trajectory regression output layer regresses and outputs a continuous three-dimensional weld centerline trajectory point sequence;
[0101] S42, inputting the weld candidate region image, the weld structure description vector, and the stress-weighted region image into a multimodal feature fusion encoding layer, extracting image texture features, geometric structure features, and connection strength features using independent linear transformations, and mapping the image texture features, geometric structure features, and connection strength features to the same dimensional space through linear projection, and splicing them to form a multimodal joint feature;
[0102] S43. Based on the component connection graph, the multimodal feature fusion coding layer is introduced into the structure-aware position bias. The structure-aware bias Δ is calculated according to the component connection relationship matrix and the topological distance between component pairs in the component connection graph. struct , structure-aware bias Δ struct Each element in Represents the connection relationship weight between the components to which the features at positions i and j belong in the input multimodal joint feature: If the features corresponding to the two positions i and j are derived from a pair of components connected in the structure, then let If connected indirectly, set Satisfying δ1>δ2>0, if there is no structural association, then set And embed the structure-aware bias into the attention calculation to form the structure-aware attention weight:
[0103]
[0104] Among them, q i 、k j are the query and key vectors of the i-th and j-th positions in the multimodal joint feature, represents the structural perception bias between components, d k represents the dimension of the key vector;
[0105] The present invention introduces a structure-aware position bias term This can enhance the attention mechanism's ability to model structural connections, reflecting the strength of the structural association between components corresponding to different positions in the multimodal joint feature. For example, assume that the 10th position in the input feature belongs to component A, the 24th position belongs to component B, and the 33rd position belongs to component C. If components A and B have a direct weld connection, the corresponding position in the bias matrix is assigned a value of δ1 = 1.0. If components A and C are not directly welded but are indirectly connected through component B, the corresponding bias is δ2 = 0.5, satisfying δ1 > δ2 > 0. If a position belongs to component D and has no connection to A, B, or C, the bias is set to 0. Embedding a structurally aware bias in the attention weight calculation process can enhance the model's spatial perception and component association modeling capabilities in complex steel structure weld identification and trajectory regression tasks.
[0106] S44. The three modal features of the weld candidate region image, weld structure description vector and stress weighted region image are fused through the modal collaborative gating layer. The scalar score value of the weld candidate region image is set to 1, the scalar score value of the weld structure description vector is set to 2, and the scalar score value of the stress weighted region image is set to 3. A learnable gating weight γ is assigned to each modality. m , and normalize them so that the sum of the total weights is 1:
[0107]
[0108] Among them, g m scalar score for each modality and fusion to generate semantic joint feature X fusion :
[0109]
[0110] S45. Based on the weight value of each pixel position in the weld space prior map, a pruning mask map is constructed through the pruning attention layer, and the pruning mask map is applied to the attention calculation process of the semantic joint feature to crop or locally restrict the attention path of the low-weight area to generate a pruning mask feature: when the weight value of a pixel position is greater than a preset threshold, the corresponding position in the mask map is assigned a value of 1, indicating that the full attention is retained; when the weight value is less than or equal to the threshold, the corresponding position in the mask map is assigned a value of 0, indicating that only local attention is retained or an attention pruning operation is performed;
[0111] The pruning mask can perform spatial selective control of the attention path in the multimodal pruned attention network. According to the weight distribution of each pixel position in the weld space prior map, it dynamically determines which areas retain global attention calculations and which areas perform local attention or path pruning, significantly reducing redundant calculations and effectively improving the model's reasoning efficiency and the accuracy of weld trajectory generation.
[0112] S46, inputting the pruning mask features into the trajectory regression output layer, and based on the principle of weld space topology and geometric continuity, using a multi-layer perceptron to perform point-by-point regression on the weld centerline to generate a sequence of regression trajectory points. In the point-by-point regression process, the initial three-dimensional welding trajectory is introduced as a structural guidance path.
[0113] The initial three-dimensional welding trajectory is geometrically consistent with the regression trajectory point sequence, and the node compression algorithm is used to optimize the curvature continuity constraint, eliminating redundant relay points in the path. Only the position points where the curvature change exceeds the set threshold are retained to output a continuous three-dimensional welding trajectory.
[0114] In this embodiment, step S5 further includes:
[0115] S51, performing spatial accessibility verification on the continuous three-dimensional welding trajectory, and performing spatial interference detection on each trajectory point in the welding path according to the geometric boundary model of the steel frame component;
[0116] S52. Calculate feasible posture solutions corresponding to each trajectory point based on welding gun structural parameters and robot arm kinematic constraints, and perform posture optimization on redundant degrees of freedom to form a continuous posture sequence;
[0117] S53. Combine the trajectory points and posture sequence to determine whether the joint limit, trajectory smoothness and dynamic acceleration constraint requirements are met. If there is an unexecutable trajectory segment, adjust the multimodal pruning attention network parameters and regenerate the three-dimensional welding trajectory; the multimodal pruning attention network parameters include modal encoding parameters, structure-aware position bias parameters, modal collaborative gating parameters, pruning attention mechanism parameters and pruning attention mechanism parameters.
[0118] In this embodiment, step S6 further includes:
[0119] S61. Evaluate the trajectory length of the three-dimensional welding trajectory that has passed the verification, perform trajectory point interpolation sampling according to the set spatial step size or time interval, generate a trajectory point sequence with equal spacing, and assign a time sequence number to each trajectory point sequence according to the index order;
[0120] S62. Calculate the spatial position and posture Euler angles of each trajectory point sequence in the welding reference system, and solve the expected linear velocity and angular velocity per unit time;
[0121] S63, performing curvature analysis on the trajectory point sequence, dynamically adjusting the sampling density according to the local curvature change, increasing the sampling density in areas where the curvature change exceeds a set threshold, and reducing the sampling density in areas where the curvature change is below the threshold;
[0122] S64: Output a discrete trajectory point sequence, where the discrete trajectory point sequence includes spatial position, attitude Euler angles, and velocity information.
[0123] In this embodiment, step S7 further includes:
[0124] S71. Based on the spatial position, posture angle, and speed information of each trajectory point in the discrete welding trajectory point sequence, query a preset process parameter library and match optimal welding parameters, wherein the welding parameters include welding current, voltage, welding speed, wire feed rate, and shielding gas flow rate;
[0125] S72, binding the matched welding parameters to the discrete welding trajectory points, and constructing a linkage control structure including the timing number, spatial position, posture angle, and welding parameters;
[0126] S73. According to the welding equipment control interface protocol, the linkage control structure is translated into an executable instruction set to generate a linkage execution instruction set adapted to the welding control unit.
[0127] In this embodiment, the welding status data includes welding current, welding voltage, welding speed, molten pool temperature, penetration information and weld formation image, and the deviation between the collected welding status data and the expected trajectory is feedback-adjusted to dynamically adjust the welding parameters and trajectory execution rate to achieve continuous monitoring and control closed-loop feedback of the welding process.
[0128] Example 1:
[0129] To verify the practical application of the present invention, it was applied to a high-speed railway continuous beam steel structure welding project at a bridge factory. The steel frame used in this project primarily serves as a connecting member between the load-bearing main beam and vertical supports. Its complex structure, diverse component types, and high welding precision requirements, coupled with a tight construction schedule, placed even higher demands on welding automation.
[0130] In this scenario, the steel frame components include steel columns, beams, node plates and stiffeners, and the welds are distributed in the multi-angular space where the components meet. During the execution process, traditional manual welding and semi-automatic trajectory teaching are often affected by factors such as large changes in component posture, poor weld visibility, and limited space. Problems such as path setting difficulties, weld leaks, and welding gun interference are prone to occur, seriously affecting welding efficiency and weld quality. To this end, the automatic welding method based on machine vision proposed in this invention is adopted. The six-axis robot ABB IRB2600, the industrial camera Basler acA2500-60gm and the i9-13900K+RTX4090 computing platform are deployed at the experimental station to complete image processing, structural modeling, stress simulation and multimodal trajectory regression in sequence, realizing intelligent recognition of complex steel frame welds and feasible trajectory output.
[0131] The present invention collects images of the welding area of steel frame components, performs enhanced preprocessing on the images, and analyzes the component connection topology structure in combination with the three-dimensional CAD model, and uses finite element simulation to generate stress-weighted regional images. Based on the structural connection diagram and the weld space prior diagram, a pruned attention network is constructed to achieve multimodal feature fusion and weld trajectory generation. The welding trajectory is dynamically feasibility checked before output, and then discretized and bound to the corresponding welding parameters according to the set rules to form a control instruction set for driving the welding equipment to complete high-precision trajectory execution. The welding trajectory is dynamically feasibility checked before output, and then discretized and bound to the corresponding welding parameters according to the set rules to generate a control instruction set for driving the welding equipment.
[0132] To evaluate the performance of the proposed method in a welding scenario involving steel frame components of a rice mill, three comparison schemes were selected: (1) manual path setting, which relies on the operator's experience to manually set the welding path; (2) point cloud fitting, which reconstructs the component geometry from a 3D point cloud and infers the weld path based on the connection relationship; and (3) image recognition, which extracts a 2D image of the weld area based on a deep learning segmentation model. The test results of each scheme in terms of weld recognition accuracy, trajectory rationality, processing efficiency, and welding quality are shown in Table 1.
[0133] Table 1 Comparison of the effects of different welding path generation methods in the rice machine steel frame welding project
[0134]
[0135]
[0136] It can be seen from the data in Table 1 that the method of the present invention is significantly superior to the traditional manual path setting method, point cloud fitting method and image recognition method in all key indicators. In terms of trajectory generation efficiency, the average trajectory generation time of the present invention is 28.5 seconds, which is about 81% less than the manual path setting method, about 68% less than the point cloud fitting method, and about 57% less than the image recognition method, showing a significant planning speed advantage. In terms of welding accuracy, the average welding deviation is 0.9 mm, which is about 2.2 mm lower than the manual path setting method, 1.5 mm lower than the point cloud fitting method, and 0.9 mm lower than the image recognition method, which significantly improves the weld formation quality. In terms of weld integrity, the present invention reaches 99.2%, which is higher than other methods, indicating that the present invention has stronger robustness and coverage capabilities in the continuous identification and trajectory regression of welds of complex components.
[0137] At the execution level, the accessibility anomaly rate of the present invention is only 1.3%, which is significantly lower than the manual path setting method of 14.5%, the point cloud fitting method of 9.8%, and the image recognition method of 6.7%. This shows that the path planning of the present invention is more in line with the robot's kinematic constraints and can effectively avoid welding gun interference and operational anomalies. The welding cycle is improved by 73.4%, far exceeding the point cloud fitting method of 31.7% and the image recognition method of 41.7%, reflecting a higher trajectory execution efficiency. The welding defect rate is controlled at 0.4 per 10 meters, while the image recognition method is 1.2 per 10 meters, the point cloud fitting method is 1.9 per 10 meters, and the manual path setting method reaches 2.7 per 10 meters, verifying the comprehensive advantages of the present invention in weld recognition accuracy and trajectory feasibility.
[0138] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A method for automatic welding of rice machine steel frame based on machine vision, characterized in that: The steps include: S1. Capture the image of the area to be welded by an industrial camera to obtain the original image data; S2, performing denoising and grayscale normalization on the original image data to generate an enhanced image; S3. Based on the 3D CAD assembly model and the enhanced image, construct the weld candidate area image, weld structure description vector and stress weighted area image; S4. A weld trajectory generation model is constructed through a multimodal pruning attention network. Multimodal fusion modeling is performed on the weld candidate region image, weld structure description vector, and stress-weighted region image to generate a continuous three-dimensional welding trajectory. S5. Perform dynamic feasibility verification and welding gun posture planning on the continuous three-dimensional welding trajectory; S6. Discretize the verified three-dimensional welding trajectory according to the set step size to generate a discrete trajectory point sequence; S7, binding the matching welding parameters in the process parameter library to the discrete trajectory point sequence to generate a linkage execution instruction set; S8. Send the linkage execution instruction set to the welding control unit, and collect real-time welding status data through the control feedback interface to complete the closed-loop control of the welding process.
2. The automatic welding method for a rice machine steel frame based on machine vision according to claim 1, characterized in that: The step S1 further comprises: S11. Install the industrial camera on the fixed bracket or motion platform of the welding equipment, and adjust the shooting angle and focal length of the industrial camera; S12. Acquire high-resolution images of the weld area under standard lighting conditions to form an original image sequence, wherein the high-resolution images include the welding start point, end point, and boundaries of adjacent components; S13. Perform time synchronization and coordinate unification processing on the original image sequence to generate original image data.
3. The automatic welding method for a rice mill steel frame based on machine vision according to claim 1, characterized in that: The step S2 further includes: using median filtering to remove isolated bright spots or granular noise in the original image sequence, and performing grayscale normalization processing on the denoised image using a contrast-limited adaptive histogram equalization method to generate an enhanced image.
4. The automatic welding method for a rice mill steel frame based on machine vision according to claim 1, characterized in that: The step S3 further comprises: S31. Based on the 3D CAD assembly model file generated during the rice mill steel frame design phase, import the assembly geometry data of the rice mill steel frame structure, analyze the structural hierarchy and assembly relationship of the components, identify the direct contact areas between the components, extract the welding connection information between the components, and establish a component connection relationship matrix; S32. Based on the component connection relationship matrix, assign a unique identifier to each component, establish a component connection diagram, and perform interconnected component decomposition on the component connection diagram to identify multiple independent connection sub-diagrams, each connection sub-diagram corresponding to a welded structural unit; In each connection subgraph, a depth-first traversal algorithm is used to extract the welding connection path, record the topological order of adjacent components and the location information of the connection surface, and combine the contact type, tolerance range and structural connection specifications specified in the welding process to screen the component pairs that meet the conditions from the traversal results to obtain the component space boundary. The component space boundary is then transformed into the image coordinate system of the enhanced image through the calibration matrix, and the image of the weld candidate area is cropped to generate. S33. Construct a finite element simulation model based on the three-dimensional assembly model, apply boundary constraints and welding load conditions to each connection subgraph in the component connection diagram, calculate the equivalent stress distribution of the component connection nodes, and obtain the local stress response value of each component to the connection area, wherein the stress response value at each connection node is defined as the maximum equivalent stress in the component connection area. The results are used to construct a stress distribution data set in the form of correspondence between node numbers and stress amplitudes; S34. Based on the weld candidate region image, mapping the local stress response value of the component to the connection region to the image coordinate domain to generate a stress-weighted region image; S35. Based on the stress-weighted regional image, a regional guided feature extraction strategy is constructed. The candidate regions are spatially weighted according to the connection strength distribution. The weld edge structural features, grayscale gradient distribution features and heat-affected zone color features are extracted. The pixel-level weighted fusion is performed through the spatial guidance matrix to form a multi-layer feature map of the weld. S36, fusing the weld multi-layer feature map with the stress-weighted regional image to generate a weld spatial priori map, and based on the weld spatial priori map, extracting the weld centerline, boundary contour, and start and end points by using spatial continuity modeling and geometric pattern fitting to generate a weld structure description vector; S37. Generate an initial three-dimensional welding trajectory based on the weld structure description vector. The trajectory is smoothed according to the curvature continuity principle, and a node compression strategy is used to reduce the complexity of motion control instructions, and a continuous executable trajectory sequence is output.
5. The automatic welding method for a rice mill steel frame based on machine vision according to claim 1, characterized in that: The region-guided feature extraction strategy in step S35 is specifically as follows: S351, using the stress-weighted region image as a weight map, constructing a spatial guidance matrix, and assigning a spatial response weight to each pixel in the weld candidate region image; S352. Use square sliding windows of 3×3, 5×5, and 7×7 pixels to perform sliding scanning on the weld candidate area image in the enhanced image. Perform the following feature extraction operations in each scale window: perform edge detection using the Sobel operator to obtain weld edge structural features, calculate the local gradient pattern using the directional gradient histogram to obtain grayscale gradient distribution features, convert the weld candidate area image to the Lab color space, extract the chromaticity and brightness differences, and obtain the color features of the heat-affected zone; S353. The weld edge structural features, grayscale gradient distribution features and color features of the heat-affected zone are weightedly fused at the pixel level through a spatial guidance matrix to form a multi-layer feature map of the weld.
6. The automatic welding method for rice mill steel frame based on machine vision according to claim 1, characterized in that: The step S4 comprises: S41. Constructing a weld trajectory generation model through a multimodal pruned attention network, wherein the multimodal pruned attention network includes a multimodal feature fusion encoding layer, a modality collaborative gating layer, a pruned attention layer, and a trajectory regression output layer; S42, inputting the weld candidate region image, the weld structure description vector, and the stress-weighted region image into a multimodal feature fusion encoding layer, extracting image texture features, geometric structure features, and connection strength features using independent linear transformations, and mapping the image texture features, geometric structure features, and connection strength features to the same dimensional space through linear projection, and splicing them to form a multimodal joint feature; S43. Based on the component connection graph, the multimodal feature fusion encoding layer is introduced into the structure-aware position bias. The structure-aware bias is calculated according to the component connection relationship matrix and the topological distance between component pairs in the component connection graph. The structure-aware bias is then embedded into the attention calculation to form the structure-aware attention weight. S44, fusing the three modal features of the weld candidate region image, the weld structure description vector, and the stress weighted region image through a modal collaborative gating layer to generate a semantic joint feature; S45. Based on the weight value of each pixel position in the weld space prior map, a pruning mask map is constructed through the pruning attention layer, and the pruning mask map is applied to the attention calculation process of the semantic joint feature to crop or locally restrict the attention path of the low-weight area to generate a pruning mask feature: when the weight value of a pixel position is greater than a preset threshold, the corresponding position in the mask map is assigned a value of 1, indicating that the full attention is retained; when the weight value is less than or equal to the threshold, the corresponding position in the mask map is assigned a value of 0, indicating that only local attention is retained or an attention pruning operation is performed; S46. Input the pruned mask features into the trajectory regression output layer, and based on the weld space topology and geometric continuity principle, use a multi-layer perceptron to perform point-by-point regression on the weld centerline to generate a regression trajectory point sequence. In the point-by-point regression process, the initial three-dimensional welding trajectory is introduced as a structural guidance path, and the initial three-dimensional welding trajectory and the regression trajectory point sequence are geometrically fused. A node compression algorithm is used to optimize the curvature continuity constraint, eliminate redundant relay points in the path, and only retain position points where the curvature change exceeds the set threshold, and output a continuous three-dimensional welding trajectory.
7. The automatic welding method for a rice mill steel frame based on machine vision according to claim 1, characterized in that: The step S5 further comprises: S51, performing spatial accessibility verification on the continuous three-dimensional welding trajectory, and performing spatial interference detection on each trajectory point in the welding path according to the geometric boundary model of the steel frame component; S52. Calculate feasible posture solutions corresponding to each trajectory point based on welding gun structural parameters and robot arm kinematic constraints, and perform posture optimization on redundant degrees of freedom to form a continuous posture sequence; S53. Combine the trajectory points and posture sequence to determine whether the joint limit, trajectory smoothness and dynamic acceleration constraint requirements are met. If there is an unexecutable trajectory segment, adjust the multimodal pruning attention network parameters and regenerate the three-dimensional welding trajectory.
8. The automatic welding method for a rice mill steel frame based on machine vision according to claim 1, characterized in that: The step S6 further comprises: S61. Evaluate the trajectory length of the three-dimensional welding trajectory that has passed the verification, perform trajectory point interpolation sampling according to the set spatial step size or time interval, generate a trajectory point sequence with equal spacing, and assign a time sequence number to each trajectory point sequence according to the index order; S62. Calculate the spatial position and posture Euler angles of each trajectory point sequence in the welding reference system, and solve the expected linear velocity and angular velocity per unit time; S63, performing curvature analysis on the trajectory point sequence, dynamically adjusting the sampling density according to the local curvature change, increasing the sampling density in areas where the curvature change exceeds a set threshold, and reducing the sampling density in areas where the curvature change is below the threshold; S64: Output a discrete trajectory point sequence, where the discrete trajectory point sequence includes spatial position, attitude Euler angles, and velocity information.
9. The automatic welding method for a rice mill steel frame based on machine vision according to claim 1, characterized in that: The step S7 further comprises: S71. Based on the spatial position, posture angle, and speed information of each trajectory point in the discrete welding trajectory point sequence, query a preset process parameter library and match optimal welding parameters, wherein the welding parameters include welding current, voltage, welding speed, wire feed rate, and shielding gas flow rate; S72, binding the matched welding parameters to the discrete welding trajectory points, and constructing a linkage control structure including the timing number, spatial position, posture angle, and welding parameters; S73. According to the welding equipment control interface protocol, the linkage control structure is translated into an executable instruction set to generate a linkage execution instruction set adapted to the welding control unit.
10. The automatic welding method for a rice mill steel frame based on machine vision according to claim 1, characterized in that: The welding status data includes welding current, welding voltage, welding speed, molten pool temperature, penetration information and weld formation image, and the deviation between the collected welding status data and the expected trajectory is feedback-adjusted to dynamically adjust the welding parameters and trajectory execution rate to achieve continuous monitoring and control closed-loop feedback of the welding process.
Citation Information
Cited By
Robot self-adaptive impeller welding control method and system fusing multi-modal data
CN121223352A
Method for automatically generating welding track of robot
CN121235969A
Ball valve assembly welding method based on image recognition
CN121883960A
Steel structure construction welding quality evaluation method and system based on visual analysis
CN122065600A