Welding behavior capturing device and welding decision model establishing method

Through the welding behavior capture device and decision-making model, welder data is collected to establish the robot's autonomous decision-making capability, which solves the problem of insufficient real-time control capabilities of the welding robot and improves the intelligence and adaptability of the welding robot.

CN120688541APending Publication Date: 2025-09-23SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510810839.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing welding robots lack real-time control capabilities and are unable to adjust welding parameters according to the real-time conditions of the weld during the welding process. They are not adaptable enough and their intelligence level is not high.

Method used

A welding behavior capture device is designed, which combines an inertial measurement unit, a molten pool vision camera, a motion capture device and a screen eye tracker to collect the welder's posture, eye movement and hand movement data. A welding decision model is established through machine learning to enable the robot to make autonomous decisions.

Benefits of technology

The welding robot can adjust the welding process in real time, which broadens its application scope and flexibility, reduces trajectory programming work and improves its intelligence level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688541A_ABST
    Figure CN120688541A_ABST
Patent Text Reader

Abstract

The invention discloses a welding behavior capturing device and a welding decision model establishing method. The welding behavior capturing device comprises a welding gun; the sliding type operation box is used for containing a workpiece and a welding gun, can shield arc light and is provided with a channel hole through which the hand of a welder extends into the box body for welding operation; the inertial measurement unit is used for collecting a yaw angle, a pitch angle and a roll angle of the tail end of the welding gun, which are equal to hand postures of a welder; the molten pool visual camera is used for shooting real-time pictures of the molten pool; the motion capture instrument is used for acquiring and capturing the position and the motion speed of the tip of the welding gun in real time; the screen eye tracker is used for monitoring the position where the eyes of the welder watch the computer screen in real time; the computer is used for displaying a molten pool image in real time, observing by human eyes of a welder, training according to data acquired by the inertial measurement unit, the molten pool visual camera and the screen eye tracker and constructing an intelligent monitoring model and a behavior decision model, and the combination of the intelligent monitoring model and the behavior decision model is a welding decision model; according to the invention, the industrial robot forms an autonomous decision-making ability through machine learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent manufacturing and artificial intelligence, and in particular to a welding behavior capturing device and a welding decision model establishing method. Background Art

[0002] Welding, a crucial joining technology in modern industry, plays an irreplaceable role in the rapid development of modern manufacturing. It is widely used in a variety of fields, including automotive, aerospace, shipbuilding, and construction, earning it the nickname "industrial tailor." In recent years, with the continuous increase in labor costs in industrial production in my country, production models have shifted towards flexibility, refinement, and intelligence. Industrial robots, with their high efficiency and stable quality, have played a crucial role in the development of industrial automation. Currently, the majority of welding robots in use both domestically and internationally are teach-type robots, lacking the real-time control capabilities of skilled welders based on welding conditions. This limits their potential application. In recent years, some researchers have begun researching how robots can mimic human behavior, but this research is currently limited to controlling basic process parameters such as welding current and voltage. He et al. [He F, Yuan L, Mu H, et al. Integrating human expertise to optimize the fabrication of parts with complex geometries in WAAM [J]. Journal of Manufacturing Systems, 2024, 74: 858-868.] used ANFIS's forward-backward algorithm to learn from welders' judgment experience and construct a model of input parameters and weld bead shape. This enabled the robot to adaptively adjust welding speed and angle based on weld bead width, marking an important step toward intelligent humanization. However, their research focused solely on the single condition of weld bead width, and the mapping between width and parameters was limited to a single material. Consequently, their adaptability remains limited, and the robot's intelligence level is low, preventing them from adjusting parameters based on the real-time weld seam conditions during the welding process. Summary of the Invention

[0003] The purpose of the present invention is to provide a welding behavior capture device and a welding decision model establishment method to control and capture welding behavior and enable industrial robots to form autonomous decision-making capabilities through machine learning.

[0004] The technical solutions for achieving the purpose of the present invention are:

[0005] A welding behavior capturing device, comprising:

[0006] welding gun;

[0007] A sliding operating box is used to accommodate the workpiece and welding gun, can shield the arc light, and is provided with a passage hole for the welder's hands to extend into the box to perform welding operations;

[0008] An inertial measurement unit (IMU) is used to collect the yaw, pitch, and roll angles of the welding gun tip, which is equivalent to the welder's hand posture.

[0009] Molten pool vision camera, used to capture real-time images of the molten pool;

[0010] Motion capture device, used to capture the position and movement speed of the welding gun tip in real time;

[0011] A screen eye tracker is used to monitor in real time where the welder's eyes are looking at the computer screen;

[0012] The computer displays the molten pool image in real time for the welder's visual observation. The computer trains and builds an intelligent monitoring model and a behavioral decision model based on the data collected by the inertial measurement unit, the molten pool vision camera, and the screen eye tracker. The combination of the two is the welding decision model.

[0013] The intelligent monitoring model determines the key focus areas during the welding process based on the changes in the molten pool image and the eye movement signals; the decision model is used to establish a mapping relationship between the key focus area image and the welder's hand movement characteristics.

[0014] A method for establishing a welding decision model includes the following steps:

[0015] Step 1: Construct a paired sample library: Collect the yaw, pitch, and roll angles of the welding gun tip to equate to the welder's hand posture; capture the real-time image of the molten pool; capture the position and movement speed of the welding gun tip in real time; and monitor the position of the welder's eyes looking at the computer screen in real time.

[0016] Fully synchronize the eye movement, weld pool image and welding gun motion data timestamps;

[0017] Step 2: Build an intelligent monitoring model: Determine the key areas of focus during the welding process based on changes in the molten pool image and eye movement signals;

[0018] Step 3: Build a behavioral decision model: Establish a mapping relationship between the focus area image and the welder's hand movement features.

[0019] Compared with the prior art, the present invention has the following significant advantages:

[0020] The present invention does not rely on the conditions of the weldment itself or the presets of the welder. Starting from the phenomenon of the molten pool during the welding process, the robot learns from skilled workers, understands the welder's key focus on the molten pool during the welding process, and adjusts the welding gun posture in real time according to changes in the molten pool, thereby broadening the practicality and flexibility of the welding robot.

[0021] This invention first proposes a device for capturing welder welding behavior. It then uses this device to synchronously collect changes in the weld pool image, the welder's eye movements, and manual movements. This data is then used to train and establish an intelligent monitoring model and a behavioral decision-making model. The combination of these two models forms the welding decision-making model, which is then applied to the real-time operation of a welding robot, ultimately achieving intelligent welding robot control. This invention provides a new approach to the research of intelligent industrial robots, reducing the workload of trajectory programming while enabling the robot to automatically adjust the welding process. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a schematic diagram of the dimensions of the sliding operation box.

[0023] Figure 2 This is a schematic diagram of the sliding operation box structure.

[0024] Figure 3 Schematic diagram of welding gun assembly installation.

[0025] Figure 4 Schematic diagram of the data collection process.

[0026] Figure 5 Schematic diagram of the timestamp synchronization process.

[0027] Figure 6 Schematic diagram of the melt pool image partitioning method.

[0028] Figure 7 is the classification CNN encoder structure.

[0029] Figure 8 This is a flow chart of the welding decision model. DETAILED DESCRIPTION

[0030] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0031] Example 1

[0032] A device for capturing welding behavior of a welder in this embodiment includes a screen eye tracker 2, a screen eye tracker bracket 3, a sliding operation box 4, a workbench 5, a welding power supply, a welding gun 6, an inertial measurement unit 7, an inertial measurement unit bracket 8, a computer, a molten pool vision camera 9, a molten pool vision camera bracket 10, an image acquisition card, a molten pool vision real-time display system, a workpiece 11, three motion capture devices 12, three motion capture device tripod brackets 13, and a shielded signal line 14, wherein:

[0033] The sliding operation box 4 includes an operation box body (4-1), an operation channel hole 4-2, a retractable light-shielding rubber strip 4-3, a near-infrared light source 4-4, a material hole 4-5, a camera 4-6, and a sliding guide rail 4-7. The operation box body is introduced in the present invention to shield light and prevent the arc light from affecting the welder's eye movement signal; the operation box shell 4-1 is fixed to the workbench 5 by screws and nuts, and the box body is 1000mm long and 500mm high; the box body is tilted, the vertical part is 100mm high, and the horizontal length of the inclination angle is 80mm; the front of the operation box shell 4-1 is provided with an operation channel hole 4-2, and the diameter of the operation channel hole 4-2 is 200mm, which is convenient for the welder to extend his hand into the box for free arc welding. A rubber ring is added to the edge of the hole to prevent hand cuts; the operation channel hole 4-2 is provided on the sliding plate, and the sliding guide rails (4-7) are divided into horizontal guide rails and longitudinal guide rails. The sliding plate can slide left and right along the operating box shell 4-1 through the horizontal guide rails at the top and bottom. The six longitudinal guide rails provided on the left and right sides of the operating channel hole 4-2 are responsible for enabling the operating channel hole 4-2 to slide up and down along the slide plate; retractable light-shielding rubber strips 4-3 are provided between the left and right sides of the slide plate and the operating box shell 4-1. The retractable light-shielding rubber strips (4-3) are 80mm wide around the edges of the box body, which can block arc light while maximizing the range of hand movement.

[0034] A near-infrared light source 4-4 is installed at the top of the sliding operation box 4, replacing natural light and ordinary light sources to illuminate the weld, so that the molten pool visual camera with filters and dimming films can still clearly capture the weld morphology; there is a material hole 4-5 at the rear of the operation box, and an 80mm boundary is left between the material hole 4-5 and the edge of the box to facilitate the feeding or removal of workpieces. The purpose of designing the material hole larger is to facilitate the motion capture device 12 to collect the motion signal of the welding gun; the camera 4-6 is fixed to the inner wall of the operation box to capture the morphology of the workpiece 11 before welding. The basic dimensions of the operation box are shown in the attached figure. Figure 1 , the structural composition diagram is as attached Figure 2 .

[0035] Two mounting lugs are designed around the end of the welding gun 6 to facilitate the installation of the inertial measurement unit 7 and the weld pool vision camera 9. The computer has built-in inertial measurement unit, motion capture device, screen eye tracker, and supporting SDK and software for the weld pool vision camera.

[0036] The installation and debugging methods of the equipment are as follows:

[0037] The sliding operation box 4 is fixed to the workbench 5 by screws and nuts; an inertial measurement unit 7 is fixed to the hanging ear at the end of the welding gun through the inertial measurement unit bracket 8. The inertial measurement unit device used in this example is XsensDOT. The data collected by the supporting SDK program configured on the computer can be the yaw angle of the end of the welding gun. Pitch angle and roll angle The data is Equivalent to the welder's hand posture, t imu Record data timestamps for the inertial measurement unit;

[0038] The molten pool vision camera 9 is fixed to the welding gun through the molten pool vision camera bracket 10. The device used in this example is a Mecaweld molten pool vision camera, which captures the real-time image of the molten pool. A shielded signal line 14 is used to reduce interference. An image acquisition card is used to transmit the video image captured by the lens to the computer at high speed and accuracy. The molten pool image is displayed in real time using the supporting molten pool vision real-time display system. By developing the supporting SDK function, the molten pool image is output frame by frame, which is recorded as "image_{t image}.jpg";

[0039] The three motion capture devices 12 are fixed on three motion capture device tripods 13 respectively and placed on the back of the operating box to obtain the position and movement speed of the welding gun in real time. The device used in this example is the Opti-track motion capture device. The collected data are recorded as the welding gun tip position, welding gun tip speed, where t motive is the motion capture device timestamp, Represents the coordinate data of the welding gun tip in the X, Y, and Z directions in the world coordinate system of the motion capture device. It represents the velocity data of the welding gun tip in the X, Y, and Z directions in the world coordinate system of the motion capture device;

[0040] Fix the screen eye tracker 2 under the computer monitor through the screen eye tracker bracket 3. The device used in this example is the Tobii screen eye tracker, which is used to monitor the position of the human eye observing the screen in real time, denoted as The value is the proportion of the annotation gaze position in the entire screen. Indicates the relative position (ratio) of a point on the screen in the horizontal direction of the screen. It is the relative position (proportion) of a point on the screen in the vertical direction of the screen. The coordinates of the upper left corner of the screen are (0,0), and the coordinates of the lower right corner are (1,1). Before the screen eye tracker collects eye movement signals, data calibration should be performed in advance, the supporting software should be started, and the eye gaze direction should be calibrated by staring at the mark points displayed in the software. During the calibration process, the software will successively display marks at the four corners and the center of the screen. The operator stares at the mark points until the mark points break, and the calibration is completed. Due to physiological differences among operators, recalibration is required after changing operators. Through the development of the supporting SDK, the real-time transmission and recording of eye movement coordinates can be realized; the line of sight coordinates can be displayed on the screen using the GWindows GDI (Graphics Device Interface) function. This function provides multi-level collaborative office capabilities for screen images, and can be used according to eye movement data. Generates a specified shape marker point on the upper layer of the video stream layer to confirm the accuracy of the gaze point coordinates.

[0041] The method of using the equipment and the process of capturing welder behavior motion are as follows:

[0042] Before the experiment, turn on the computer, the weld pool vision camera 9, the screen eye tracker 2, the motion capture device 12, and the inertial measurement unit 7 equipment and corresponding software, install the equipment, and complete the calibration. After the preparation is completed, the arc welding skill master 1 inserts his hand into the operating box 4 through the operating channel hole 4-2 on the operating box 4, grasps the welding gun 6 and prepares to weld. The welding gun is fixed with the weld pool vision camera and the inertial measurement unit. The welding gun assembly is installed as shown in the attached figure. Figure 3 As shown. The arc welding skill master 1 determines the welding starting position by observing the image inside the operation box displayed on the screen, and turns on the fusion welding power supply to weld. During the welding process, the arc welding skill master observes the real-time molten pool image displayed on the screen. The position of his gaze is read by the screen eye tracker 2 fixed on the welding gun. When the arc welding skill master observes a certain phenomenon in a certain area, he forms a decision based on his experience and adjusts the speed, angle and height of the welding gun in time. The various values ​​are read by the inertial measurement unit 7 and the motion capture device 12 supporting program, and finally the welding is completed. The data collection process is shown in the attached figure. Figure 4 As shown, multiple sets of paired samples of “molten pool image-eye gaze data-welding gun movement data” are formed.

[0043] Example 2

[0044] This embodiment provides a method for establishing a welding decision model. This method, based on the aforementioned computer and the data acquisition device designed in Example 1, optimizes data acquisition and enables the robot to learn from arc welding masters to automatically adjust the welding process. The welding decision model consists of two parts: an intelligent monitoring model and a behavioral decision model. Its establishment method includes the following steps:

[0045] Step 1: Construction of paired sample library:

[0046] This step is based on the process of welder behavior motion capture described in Example 1, and its main purpose is to achieve complete synchronization of eye movement, molten pool image and welding gun action data timestamps. Development is based on the SDK of all devices, and the reading frequency of all devices is set to the same, which is set to 30Hz in this example. The datetime function of the Pandas database is used to unify the timestamp format of all devices to "%d-%H-%M-%S.%f". The OpenCV function package is used to realize the frame-by-frame output of the molten pool image, and the recorded timestamp is used as the image name. The timestamp of the molten pool image is counted as t image , which is convenient for the subsequent image extraction and matching, and the output image "image_{t image}.jpg" to the folder "image_{start-time}".

[0047] Using a multi-threaded approach, the first frame of the melt pool image is output as a signal to activate the desktop eye tracker, motion capture device, and inertial measurement unit program for synchronous reading. After the reading is completed, it enters a waiting state. When the melt pool image shows a second reading signal, the data of other devices is read again to achieve preliminary data synchronization. The inertial measurement unit data is written to the folder imu_data. The corresponding timestamp of the inertial measurement unit data is t imu , the output data is stored in "imu_{start-time}.csv", the data content is Calculate Indicates t imu Yaw angle at the moment, Indicates t imu The pitch angle at the moment, Indicates t imu The roll angle at the moment; the motion capture device 12 data is written to the folder motive_data, and the time of the motion capture device data is t motive , the output data is stored in "motive_{start-time}.csv", the data content is Calculate Respectively represent t motive The position of the welding gun tip on the X, Y, and Z axes in the world coordinate system of the motion capture device at any moment. Respectively represent t motive The speed of the welding gun tip on the X, Y, and Z axes in the world coordinate system of the motion capture device at this moment; the screen eye tracker 2 data is written to the folder gaze_data, and the timestamp corresponding to the eye tracker data is t gaze, the output data is stored in "gaze_{start-time}.csv", the data content is Calculate in Indicates that the welder's line of sight is at t gaze Always observe the relative position of the screen in the horizontal direction. Indicates that the welder's line of sight is at t gaze Always observe the relative position of the screen in the vertical direction.

[0048] Due to the influence of running errors between different programs, the timestamps of each device have microsecond differences, which requires further timestamp synchronization. image As a benchmark, the remaining data are interpolated to the melt pool image time to generate an approximate value, that is, to ensure that all devices have t image The interpolation approximation algorithm used for the data at each moment is quadratic spline interpolation. The principle is to define a quadratic polynomial between each adjacent known point and use the quadratic polynomial to interpolate in each space. The algorithm is as follows:

[0049] S i (m) = a i +b i (mm i )+c i (mm i ) 2 (1)

[0050] According to the sample point (m i ) i (m) Function and derivative S′ i (m) Continuous, natural spline boundary condition second-order derivative S″ i The coefficients of each position are solved under the condition that (m) is 0:

[0051]

[0052] Among them S i (m) is the i-th polynomial, a i 、b i 、c i is the coefficient to be determined for the i-th sample point, m i is the i-th known sample point, m is the interpolation variable, h i is the distance between the sample point and the interpolation point in mm i , n is the total number of samples. This algorithm can ensure continuity at known points and continuity of the first-order derivative in the interpolation part, providing better interpolation effect without significantly increasing the computational complexity.

[0053] In actual application, the collected data is substituted into the above interpolation algorithm formula (2):

[0054]

[0055] in: is the coefficient to be determined when interpolating the motion capture device data. is the coefficient to be determined when interpolating the inertial measurement unit data; is the coefficient to be determined when interpolating the eye tracker data; The inertial measurement unit, motion capture device, and eye tracker are interpolated to the image timestamp t image The principle of timestamp synchronization can be shown as follows: Figure 5 , after using synchronization at t image The data at time (a) forms paired data, where a is the interpolation point number. The interpolated data is output to "data_all.csv", the content is Among them, t image is the timestamp of the image, x gaze ,y gaze t image The welder should always pay attention to the horizontal and vertical relative positions of the screen. t image Yaw angle, pitch angle and roll angle of the welding gun tip at all times, x, y, z, v x , v y , v z t image The coordinates and speed of the welding gun tip on the X, Y, and Z axes in the world coordinate system of the motion capture device at each moment. All values ​​at each moment constitute a paired sample.

[0056] Step 2: Build an intelligent monitoring model:

[0057] The welder's judgment of the welding process is based on certain phenomena in a specific area of ​​the molten pool image, while other areas of the image are not focused on. The area where the welder's line of sight is assessed is the focus area. Step 2 is based on the paired sample library obtained in step 1, aiming to learn from the welder's experience and determine the focus area during the welding process based on the changes in the molten pool image and the characteristics of the eye movement signal. The data used is the molten pool image "image_{t image}.jpg" and the first three columns of data in "data_all.csv" (t image , x gaze ,y gazeThe Canny+ median filter algorithm is used to optimize the molten pool image, enhance its contour contrast, highlight the changes in the molten pool characteristics, clean the eye movement data, and filter out unconscious gaze data. Then, according to the calibration of the arc welding skill master on the welding molten pool image, the molten pool image is partitioned into three areas: nozzle area (a), arc column area (b) and molten pool area (c). The image partitioning example effect is shown in the attached figure. Figure 6 As shown in the figure, add area labels: nozzle area (a) → [1, 0, 0], arc column area (b) → [0, 1, 0], molten pool area (c) → [0, 0, 1]. image}.jpg”, eye movement signal (t image , x gaze ,y gaze ) and 70% of the samples paired with the regional division rule are used as the training set for model training. The model adopts the ResNet18 architecture. Since the welding image is a grayscale image, the ResNet18 model receives RGB three-channel input. In this example, single-channel input is forced to avoid feature information loss caused by RGB conversion. The loss function is weighted cross-entropy loss (Cross-Entropy Loss). Combined with the image features, the arc column area is small, and the arc column area (b) is weighted to increase the model's attention to the arc column area. The input of the model is a dynamic molten pool image sequence, such as Figure 6 The structure shown in the figure outputs three types of regional classification: probability distribution of electrode area (nozzle), arc column area (arc), and molten pool area (including boundary and weld bead), and local images of key focus areas cropped according to probability. This is used for model training to form a mapping relationship model between dynamic molten pool images and local images of key focus areas. This mapping relationship model is the intelligent monitoring model, which can ensure the automatic extraction of welder's key focus areas based on real-time molten pool images.

[0058] The remaining 30% of the data is used as a test set, and the regional overlap evaluation is introduced, as shown in the formula:

[0059]

[0060] Where Pred represents the predicted area, GT represents the true labeled area, Area(GT∩Pred) represents the overlapping area of ​​the true and predicted areas, and Area(GT∪Pred) represents the combined area of ​​the true and predicted areas. The results are scored from 0 to 1. If it is close to 0, it means that the prediction accuracy is low, and more training data is needed to improve the prediction ability of the model to ensure the accuracy of the model. The model outputs each timestamp t image The local image of the focus area is “focus_{t image}.jpg".

[0061] Step 3: Build a behavioral decision model:

[0062] This step is based on the local image of the focus area obtained in step 2 of Example 2. The main purpose is to establish a mapping relationship between the focus area image and the arc welding skill master's behavior. The data used is the focus area image "focus_{t image}.jpg" and the 1st and 3-12th columns of the hand behavior data "data_all.csv" obtained in step 1, the contents are

[0063] Building a behavioral decision-making model requires a multimodal learning system that extracts visual information from images and generates corresponding hand movement data based on this information. This example uses a deep learning framework to achieve this, combining a convolutional neural network (CNN) with a sequence generation model (LSTM). Using an encoder-decoder architecture, the CNN acts as the encoder to extract image features, and the LSTM acts as the decoder to generate hand movement sequences from images.

[0064] For different focus points of the focus area image, algorithms are designed to extract features: If the focus area is a, that is, the nozzle area, the electrode tip state needs to be monitored ( Figure 6 (a) Nozzle position), identify the shielding gas flow pattern ( Figure 6 (a) grayscale gradient area around the nozzle), select CoordConv+ dilated convolution structure to enhance the perception of the electrode spatial position offset and capture the large-scale characteristics of gas diffusion around the nozzle; if the focus area is b, that is, the arc column area, it is necessary to quantify the arc plasma morphology ( Figure 6 (b) Highlighted strip area) and arc stability, a high-frequency filter kernel is selected to enhance the ability to capture sudden changes in arc brightness in the image, and a spatiotemporal attention mechanism is introduced to extract the high-frequency oscillation of the arc, while processing spatial brightness and temporal flicker characteristics; if the focus area is c, that is, the molten pool area, the main analysis is the molten pool shape and the continuity of the molten droplet. U-Net is selected as the basic architecture, and the Canny algorithm is introduced in the input layer and optimized with a median filter to enhance edge information and facilitate accurate edge positioning. The classification CNN encoder structure is as follows Figure 7 Targeted algorithm design in this way helps improve the model's accuracy and robustness in predicting melt pool image features. Hand motion data is complex and contains physiological jitter noise. In this example, an exponential moving average filter is introduced to smooth the signal, and a long short-term memory network (LSTM) is selected to capture the temporal dependencies in the motion data and extract motion features.

[0065] The encoder and decoder are connected through features, and the mapping relationship between the weld pool image features and the hand movement features is established to form a behavior decision model. The entire welding decision model operation process is shown in the attached figure. Figure 8 shown.

Claims

1. A welding behavior capture device, characterized in that: include: welding gun; A sliding operating box is used to accommodate the workpiece and welding gun, can shield the arc light, and is provided with a passage hole for the welder's hands to extend into the box to perform welding operations; An inertial measurement unit (IMU) is used to collect the yaw, pitch, and roll angles of the welding gun tip, which is equivalent to the welder's hand posture. Molten pool vision camera, used to capture real-time images of the molten pool; Motion capture device, used to capture the position and movement speed of the welding gun tip in real time; A screen eye tracker is used to monitor in real time where the welder's eyes are looking at the computer screen; The computer displays the molten pool image in real time for the welder's visual observation. The computer trains and builds an intelligent monitoring model and a behavioral decision model based on the data collected by the inertial measurement unit, the molten pool vision camera, and the screen eye tracker. The combination of the two is the welding decision model. The intelligent monitoring model determines the key focus areas during the welding process based on the changes in the molten pool image and the eye movement signals; the decision model is used to establish a mapping relationship between the key focus area image and the welder's hand movement characteristics.

2. The welding behavior capturing device according to claim 1, characterized in that: The process of constructing the intelligent monitoring model by the computer is as follows: Fully synchronize the eye movement, weld pool image and welding gun motion data timestamps; The Canny+median filter algorithm is used to optimize the molten pool image, clean the eye movement data, and filter out unconscious gaze data; the molten pool image is partitioned according to the calibration of the welding molten pool image used by the welder, and the molten pool image is divided into three areas: nozzle area, arc column area, and molten pool area. The molten pool image, eye movement signal, and regional division rule are paired to form a training set and a test set; the training model adopts the ResNet18 architecture, and the loss function is the weighted cross entropy loss; the input of the training model is a dynamic molten pool image sequence, and the output is a regional classification, forming a mapping relationship model between the dynamic molten pool image at each timestamp and the local image of the key focus area.

3. The welding behavior capturing device according to claim 2, characterized in that: The synchronization process is: Set the reading frequencies of the inertial measurement unit, melt pool vision camera, motion capture device, and screen eye tracker to be the same, and output the melt pool image frame by frame; Using a multi-threaded approach, the first frame of the melt pool image is output as a signal to activate the screen eye tracker, motion capture device, and inertial measurement unit program for synchronous reading. Read the timestamp t in the melt pool image image As a benchmark, the remaining data are interpolated to the melt pool image time to generate approximate values ​​to ensure that all devices have t image Data at this moment: in: is the coefficient to be determined when interpolating the motion capture device data. is the coefficient to be determined when interpolating the inertial measurement unit data; is the coefficient to be determined when interpolating the eye tracker data; The inertial measurement unit, motion capture device, and screen eye tracker are interpolated to the image timestamp t image data; Save the interpolated data: (t image , x gaze ,y gaze , x, y, z, v x , v y , v z ) where t image is the timestamp of the image, x gaze ,y gaze t image The welder should always pay attention to the horizontal and vertical relative positions of the screen. t image Yaw angle, pitch angle and roll angle of the welding gun tip at all times, x, y, z, v x , v y , v z t image The coordinates and speed of the welding gun tip on the X, Y, and Z axes in the world coordinate system of the motion capture device at each moment. All values ​​at each moment constitute a paired sample.

4. The welding behavior capturing device according to claim 2, characterized in that: The test of the trained model introduces the regional overlap IoU evaluation: Where Pred represents the predicted area, GT represents the true labeled area, Area(GT∩Pred) represents the overlapping area of ​​the true and predicted areas, and Area(GT∪Pred) represents the joint area of ​​the true and predicted areas.

5. The welding behavior capturing device according to claim 2, characterized in that: The computer-built behavioral decision model is implemented through a deep learning framework: Combining convolutional neural networks and sequence generation models, using an encoder-decoder architecture, with the convolutional neural network as the encoder to extract image features and the sequence generation model as the decoder to generate hand movement sequences based on the image; Feature extraction is performed on different focus points of the key focus area image. If the key focus area is the nozzle area, the electrode tip status is monitored and the shielding gas flow pattern is identified. The CoordConv+dilated convolution structure is selected to enhance the perception of the electrode spatial position offset and capture the large-scale characteristics of the gas diffusion around the nozzle. If the focus area is the arc column, it is necessary to quantify the arc plasma morphology and arc stability. A high-frequency filter kernel is selected to enhance the ability to capture sudden changes in arc brightness in the image. A spatiotemporal attention mechanism is introduced to extract high-frequency oscillations of the arc and process both spatial brightness and temporal flicker characteristics. If the focus area is the molten pool area, it is necessary to analyze the molten pool shape and the continuity of the molten droplets. U-Net is selected as the basic architecture, and the Canny algorithm is introduced in the input layer and optimized with median filtering to enhance the edge information.

6. The welding behavior capturing device according to claim 1, wherein: The sliding operation box includes an operation box body, an operation channel hole, a retractable light-shielding rubber strip, a near-infrared light source, a material hole, a camera and a sliding guide rail; An operation channel hole is provided at the front of the operation box shell, and the operation channel hole is arranged on a sliding plate. The sliding plate can slide left and right along the operation box shell through the horizontal guide rails at the top and bottom; the operation channel hole can slide up and down along the slide plate through the longitudinal guide rails; retractable light-shielding strips are provided between the left and right sides of the slide plate and the operation box shell; there is a material hole at the rear of the operation box, and a near-infrared light source for photography at the top.

7. A method for establishing a welding decision model, characterized in that: The following steps are involved: Step 1: Construction of paired sample library: Collect the yaw angle, pitch angle and roll angle of the welding gun end to be equivalent to the welder's hand posture; Capture the real-time image of the molten pool; obtain and capture the position and movement speed of the welding gun tip in real time; and monitor the position of the welder's eyes looking at the computer screen in real time; Fully synchronize the eye movement, weld pool image and welding gun motion data timestamps; Step 2: Build an intelligent monitoring model: Determine the key areas of focus during the welding process based on changes in the molten pool image and eye movement signals; Step 3: Build a behavioral decision model: Establish a mapping relationship between the focus area image and the welder's hand movement features.

8. The method for establishing a welding decision model according to claim 7, characterized in that: The synchronization process is: The reading frequencies of the molten pool image, welder's hand posture data, the position data of the human eye observing the screen, and the welding gun tip data are set to be the same, and the molten pool image is output frame by frame; Using a multi-threaded approach, the first frame of the weld pool image is output as a signal to synchronously read the welder's hand posture data, the position of the human eye observing the screen, and the welding gun tip data; Read the timestamp t in the melt pool image image As a benchmark, the remaining data are interpolated to the melt pool image time to generate approximate values ​​to ensure that all devices have t image Data at this moment: in: is the coefficient to be determined when interpolating the welding gun tip data. is the coefficient to be determined when interpolating the welder's hand posture data; The coefficients to be determined when interpolating the position data of the human eye observing the screen; Interpolated to image timestamp t image The welder's hand posture and the position data of the welding gun tip and the human eye observing the screen; Save the interpolated data: (t image , x gaze ,y gaze , x, y, z, v x , v y , v z ) where t image is the timestamp of the image, x gaze , x gaze t image The welder should always pay attention to the horizontal and vertical relative positions of the screen. t image Yaw angle, pitch angle and roll angle of the welding gun tip at all times, x, y, z, v x , v y , v z t image The coordinates and speed of the welding gun tip on the X, Y, and Z axes in the world coordinate system of the motion capture device at each moment. All values ​​at each moment constitute a paired sample.

9. The method for establishing a welding decision model according to claim 7, characterized in that: The process of building an intelligent monitoring model is as follows: the Canny+median filter algorithm is used to optimize the molten pool image, clean the eye movement data, and filter out unconscious gaze data; the molten pool image is partitioned according to the calibration of the welding molten pool image used by the welder, and the molten pool image is divided into three areas: nozzle area, arc column area, and molten pool area. The molten pool image, eye movement signal, and regional division rule are paired to form a training set and a test volume set; the training model adopts the ResNet18 architecture, and the loss function is the weighted cross entropy loss; the input of the training model is a dynamic molten pool image sequence, and the output is regional classification, forming a mapping relationship model between the dynamic molten pool image at each timestamp and the local image of the key focus area.

10. The method for establishing a welding decision model according to claim 7, characterized in that: The behavioral decision model is implemented through a deep learning framework: Combining convolutional neural networks and sequence generation models, using an encoder-decoder architecture, with the convolutional neural network as the encoder to extract image features and the sequence generation model as the decoder to generate hand movement sequences based on the image; Feature extraction is performed on different focus points of the key focus area image. If the key focus area is the nozzle area, the electrode tip status is monitored and the shielding gas flow pattern is identified. The CoordConv+dilated convolution structure is selected to enhance the perception of the electrode spatial position offset and capture the large-scale characteristics of the gas diffusion around the nozzle. If the focus area is the arc column, it is necessary to quantify the arc plasma morphology and arc stability. A high-frequency filter kernel is selected to enhance the ability to capture sudden changes in arc brightness in the image. A spatiotemporal attention mechanism is introduced to extract high-frequency oscillations of the arc and process both spatial brightness and temporal flicker characteristics. If the focus area is the molten pool area, it is necessary to analyze the molten pool shape and the continuity of the molten droplets. U-Net is selected as the basic architecture, and the Canny algorithm is introduced in the input layer and optimized with median filtering to enhance the edge information.