Dual-camera keyhole TIG welding on-line vision detection device and method
By combining a dual-camera vision inspection device with an AI computing rod, the problem of unstable weld quality during the welding of long curved welds on medium-thick plates was solved, achieving efficient welding quality monitoring and trajectory correction, and reducing equipment complexity and cost.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-21
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies struggle to achieve real-time welding quality monitoring and trajectory correction during the welding of long, curved welds on medium-thick plates. In particular, deep penetration keyhole TIG welding suffers from unstable weld quality and inaccurate information extraction. Furthermore, existing equipment is either structurally complex or costly.
An online visual inspection device for keyhole TIG welding robots based on dual cameras is adopted. Welding images are acquired synchronously by dual CMOS cameras and high dynamic range image synthesis and stitching are performed. The embedded main control board and AI computing rod are combined to identify the penetration and correct the trajectory. Lightweight CNN segmentation network and RANSAC method are used to extract weld features.
It enables online welding quality inspection and trajectory correction for long curved welds in medium and thick plates, improving welding accuracy and efficiency, reducing equipment costs, and enhancing the visualization and automated control of the welding process.
Smart Images

Figure CN117102635B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of welding monitoring and image processing, in particular to a double-camera-based keyhole TIG welding robot online visual detection device and method. BACKGROUND
[0002] Plate welding structures are widely used in many industrial fields, such as shipbuilding, bridges, petrochemical industry, boilers and containers, and heavy machinery. However, large welding parts usually have a large number of large curved long welds, which leads to poor weld quality stability, low production efficiency and high labor intensity. In recent years, deep penetration keyhole TIG welding technology as a new welding method can achieve large penetration by using keyhole effect, one-time welding, single-sided welding and double-sided forming, good weld forming and high welding quality. Combined with robots, it can well replace manual welding to improve production efficiency, reduce labor intensity and improve weld quality stability. However, in the actual welding process, deep penetration keyhole TIG welding still has some problems, such as incomplete penetration, partial penetration and over penetration of the weld.
[0003] And the assembly, deformation and other factors of the curved long weld of the plate often lead to irregular and random state of the weld, and the robot welding needs to adjust the welding gun in real time for correction. At present, the mature industrial computer + industrial robot mode is heavy and cannot meet the real-time welding detection and control requirements of the curved long weld during movement.
[0004] In addition, the arc light interference is strong during welding, and it is difficult to extract visual information. The mainstream "single camera + combined filter" scheme cannot fully extract the welding process information, which makes it difficult to further improve the welding precision and also cannot realize the penetration recognition and trajectory correction functions at the same time. The existence of these problems seriously limits the online quality monitoring and control of the welding process.
[0005] Invention patent CN 109719368 B discloses a robot welding process multi-information acquisition and monitoring system and method, which uses a master control computer + fixed industrial robot to realize real-time monitoring of welding. This system and method are suitable for welding applications on the assembly line, but are difficult to apply to real-time identification and remote monitoring of welding quality on curved long welds that need to move or crawl. And the "single camera + combined filter" method used by it to obtain clear welding process information not only has a complex structure design, but also a single camera can only collect information on one side of the molten pool, the information extraction is not accurate enough, and the recognition accuracy is difficult to improve.
[0006] The patent CN202210550090 discloses a full-position robot deep penetration K-TIG welding system and control method, which uses an expensive HDR camera scheme. Not only is it expensive, but the information extraction of the single camera to the molten pool is not complete, and other information needs to be combined to make up for the accuracy of welding detection and control. SUMMARY
[0007] The purpose of the present application is to solve the above-mentioned defects in the prior art, and to provide a double-camera-based lock hole TIG welding robot online visual detection device and method, which is suitable for online welding quality detection and trajectory correction of lock hole TIG welding on medium-thick plate curved long welds.
[0008] The purpose of the present application can be achieved by adopting the following technical solutions:
[0009] A visual detection method of a double-camera-based lock hole TIG welding robot online visual detection device, wherein the visual detection device comprises: a first CMOS camera 1 connected to an embedded main control board 3, used for real-time acquisition of a welding scene image of a to-be-welded welding direction of a lock hole TIG welding, synchronous acquisition of multi-exposure images through a hard trigger signal and a second CMOS camera 2, and transmission of image data to the embedded main control board 3; the second CMOS camera 2 is connected to the embedded main control board 3, used for real-time acquisition of a welding scene image of a welded welding direction of a lock hole TIG welding, synchronous acquisition of multi-exposure images through a hard trigger signal and a first CMOS camera 1, and transmission of image data to the embedded main control board 3;
[0010] The embedded main control board 3 communicates data with other components of the visual detection device, and integrates a software package for lock hole TIG welding penetration recognition and trajectory correction;
[0011] The HDMI screen 4 is connected to the embedded main control board 3, used for display output of local real-time low dynamic welding images, segmentation and deviation identification images, penetration state and early warning information;
[0012] The mouse 5 is connected to the embedded main control board 3, used for software interface clicking and selection;
[0013] The keyboard 6 is connected to the embedded main control board 3, used for parameter setting input of the software interface;
[0014] The AI computing stick 7 is connected to the embedded main control board 3, receives high dynamic image data preprocessed by the embedded main control board 3, infers the segmentation model and returns the result data, and communicates with the embedded main control board 3 through USB3.0;
[0015] The motion controller 8 is connected with the embedded main control board 3, receives the motion instruction of the embedded main control board 3, controls the motion of the crawling robot, and feeds back the real-time pose information of the welding gun, and communicates with the embedded main control board 3 through a network port;
[0016] The teaching device 9 is connected with the motion controller 8, is used for manually controlling the pose of the crawling robot 10, teaching the starting point, the intermediate process point and the termination point of welding, completing the welding predetermined trajectory planning through the drag type programming, and communicating with the embedded main control board 3 through a network port;
[0017] The crawling robot 10 is connected with the motion controller 8, executes the control instruction sent by the motion controller 8, and feeds back the real-time axis state of the robot, and communicates with the motion controller 8 through an EtherCat bus protocol through a network port;
[0018] The WIFI module 11 is connected with the embedded main control board 3, is used for transmitting the real-time low dynamic welding image, the segmentation and deviation identification image, the penetration state and the early warning information to a remote server;
[0019] The signal distributor 12 is connected with the embedded main control board 3, the first CMOS camera 1 and the second CMOS camera 2 IO ports, is used for generating two-way output signals of the IO port output signals of the embedded main control board 3, and is used for realizing the synchronous triggering function of sending one signal output to two cameras at the same time.
[0020] The process of the visual detection method based on the online visual detection device is as follows:
[0021] Step S1: The first CMOS camera 1 and the second CMOS camera 2 are respectively installed at the front and rear positions of the welding gun, the included angle between the first CMOS camera 1 and the center line of the welding gun is 50 degrees, and the included angle between the second CMOS camera 2 and the center line of the welding gun is 60 degrees.
[0022] Step S2: The exposure time sequence is calibrated, the exposure time ranges of the arc, the keyhole, the molten pool and the to-be-welded / have-welded areas are artificially roughly selected, the equal-step exposure time sequences are generated in the respective ranges, the trial welding is respectively performed, the trial welding image results are observed, the clear exposure time of the arc, the keyhole, the molten pool and the to-be-welded area is selected by the first CMOS camera 1, the clear exposure time of the arc, the keyhole, the molten pool and the have-welded area is selected by the second CMOS camera 2, the square root of the product of the exposure time of the arc-arc, the keyhole-keyhole, the molten pool-molten pool and the to-be-welded area-have-welded area of the two cameras is taken as the final four exposure times, and the four exposure times are synchronously collected through the external triggering control of the two cameras;
[0023] Step S3: Calibrate the stitching parameters, design a 13x9 checkerboard with each square being 3mmx3mm, place it in the overlapping area of the two cameras, detect the corner points of the checkerboard, map the corner points of the two cameras to the same real-world plane, which is chosen to be the welding plate surface, the area is chosen to be part of the welded area, the molten pool area and part of the to-be-welded area, the pixel values of the mapped images and the real world have a corresponding proportional relationship, and the transformation matrix mapped to the plane is calculated respectively;
[0024] Step S4: stitching and fusion, use the transformation matrix to map the images of the two cameras to the same plane. By converting the pixel coordinates in each camera image to coordinates on the plane, the lockhole area is used as the overlapping area for linear transition fusion to ensure that there is no obvious joint at the stitching position. The clear images of the first CMOS camera 1 and the second CMOS camera 2 are stitched and fused respectively, arc-light-arc-light, lockhole-lockhole, molten pool-molten pool, and to-be-welded area-welded area;
[0025] Step S5: high dynamic synthesis, estimate the camera response curve function according to the four stitched and fused images with different exposure times, and synthesize the four stitched and fused images with different exposure times into a high dynamic image according to the mapping relationship between pixel value and scene illumination, and linearly map each channel of the high dynamic image to the 0-255 range low dynamic image for display;
[0026] Step S6: use the high dynamic image synthesized in step S5 as input, and use the CNN segmentation network to output a predicted feature map containing the lockhole, molten pool and to-be-welded weld;
[0027] Step S7: separate the lockhole image and the molten pool image according to the predicted feature map obtained in step S6, and extract 16 key points respectively;
[0028] Step S8: take the 32 key points of the lockhole and the molten pool as features, take the features of the current frame and the previous 3 frames as input, and identify the state of the penetration according to the constructed neural network;
[0029] Step S9: separate the molten pool image and the to-be-welded weld image according to the predicted feature map obtained in step S6, calculate the center of gravity of the molten pool, and fit a cubic curve to the center of the to-be-welded weld. Get the longitudinal deviation through the center of gravity of the molten pool and the center of the weld, and superimpose the longitudinal deviation feedback to the robot controller through the original path planning to control the welding deviation.
[0030] Further, the embedded main control board is small in size, can replace the heavy industrial robot, and can be flexibly installed on the crawling robot. The embedded main control board 3 is 80mmx55mm in size, adopts a 4-core Cortex-A9 processor of an NXP i.MX6 series, has a 1GHz main frequency, has a 1GB DDR3 memory, has a 4G EMMC storage, has interfaces including 1-way HDMI 2.0, 2-way gigabit Ethernet, 2-way USB2.0, 2-way USB3.0 ports, 1-way WiFi module interface, 5-way serial ports (including 1-way debugging serial port), 2-way TF cards, 3-way IIC, 1-way SPI, 1-way PCIE, 23-way GPIO, and 1-way PWM, and is provided with an independent hardware watchdog;
[0031] The first CMOS camera 1 and the second CMOS camera 2 are respectively installed at positions in front of and behind the welding gun. The first CMOS camera 1 has an angle of 50 degrees with the center line of the welding gun, and the second CMOS camera 2 has an angle of 60 degrees with the center line of the welding gun. Hard triggering is supported. One IO port of the embedded main control board 3 is connected to the signal distributor 12 to divide two IO ports connected to the two cameras. Different pulse signals are used to synchronously control the collection of different exposure times of the two cameras. The exposure time range should meet 10us-1s, the dynamic range is greater than 60dB, the resolution is higher than 1080P, and the frame rate is greater than 90 frames / s.
[0032] Further, the chip with AI hardware acceleration has few choices. The AI acceleration stick is used in a mode independent of the main control. The selected chip range is wide, and the computing power can be dynamically expanded to adapt to different application scenarios. The AI computing stick 7 transmits data through a USB3.0 port, is used for inference of an AI model, is virtually connected to a network card after connection, and user can complete input and output of data through socket programming.
[0033] The motion controller 8 integrates an EtherCAT master station library, supports 1ms cycle communication, contains motion functions such as point motion, continuous trajectory, straight line and arc interpolation, continuous interpolation, can freely set running speed, stopping speed, acceleration and deceleration time which can be independently set, S-shaped curve smoothing and other parameters, and can modify, add and delete motion online.
[0034] The teach pendant 9 can manually teach joints and a rectangular coordinate system through keys, complete teaching of starting points, intermediate process points and ending points, and an image module integrates motion functions. Path planning program scripts can be generated through drag-and-drop graphical programming. The program scripts are sent to the motion controller 8 through a network port during running. The motion controller 8 analyzes the scripts to control the motion of the robot.
[0035] The WIFI module 11 transmits real-time low dynamic welding images, segmentation and deviation identification images, penetration states and early warning information to a remote server.
[0036] Further, the conventional CMOS camera dynamic range does not exceed 80dB, and it is necessary to use different exposure times to collect clear images of different regions of interest under strong arc light, which requires a camera with a dynamic range exceeding 140dB. Without purchasing an expensive high dynamic camera, different exposure time ranges for arc light, keyhole, molten pool, and to-be-welded / welded regions are manually selected in step S2, and equal-step exposure time sequences are generated within the respective ranges to perform trial welding, and images of arc light, keyhole, molten pool, and to-be-welded / welded regions are respectively intercepted, and an image evaluation standard is established to automatically select the one with the highest exposure time score, wherein in the image evaluation standard, a percentage method is used to determine the average of the pixel values ranked in the top 5% of brightness and the average of the pixel values ranked in the bottom 5% of brightness, and then the difference between the two is calculated as the evaluation standard. The larger the difference, the higher the dynamic range of the image and the richer and clearer the content.
[0037] Further, the images collected by the two cameras are not at the same angle, but correspond to different planes, and perspective transformation is needed to transform the two cameras to the same plane for subsequent splicing. The transformed image pixel size and the physical world have a proportional relationship, which can be used for deviation calculation. The transformation matrix is: wherein are the 8 unknown parameters to be solved, and the following equation is used to calculate and solve the unknown parameters:
[0038]
[0039] wherein and are the pixel coordinates of the i=1, 2, …, t corner points before transformation, and are the pixel coordinates of the i=1, 2, …, t corner points after transformation;
[0040] The proportional relationship between the pixel size and the actual size of the transformed image is as follows:
[0041]
[0042] wherein , and , are the horizontal and vertical coordinates of the two points on the transformed image, , and , are the horizontal and vertical coordinates of the corresponding two points in the real world, and are the horizontal and vertical scaling coefficients, respectively.
[0043] Further, after transformation, seamless splicing needs to be performed for the overlapping area, and simple superposition and averaging operations will produce ghosting, and a linear gradient fusion method is adopted, and in the splicing process, a linear gradient weight is used to realize smooth transition, the weight is calculated according to the position of the pixel in the overlapping area, and interpolation is performed between the pixel value 1 and the pixel value 2, the distance of the pixel is normalized to the range [0, 1] to obtain the weight coefficient a, and the calculation formula of the weight coefficient a is as follows:
[0044]
[0045] Wherein, d represents the distance of the pixel position to the boundary of the overlapping area, is the diagonal distance of the overlapping area, the weight coefficient a is used for linear interpolation to obtain the fused pixel value.
[0046] Further, the clear images of different regions of interest corresponding to the four fused images with different exposure times, based on the light irradiance and image imaging principle, the four images are synthesized into one high dynamic image with clear regions of interest, and the dynamic range of the synthesized high dynamic image can be improved to more than 140dB, in order to facilitate visual observation, the high dynamic image is mapped to a low dynamic image for display, and in the high dynamic synthesis, the camera response curve function is estimated based on the four spliced and fused images with different exposure times, and the following formula is minimized:
[0047]
[0048] Wherein, is the target loss function, is the spatial index of the pixel, is the exposure time corresponding to the spliced image index position, represents the corresponding pixel value, g( ) is the irradiance recovery function, represents the irradiance of the index position m, and the statically indeterminate equation is established to solve the 256 and pixel spaces in the formula;
[0049] The four spliced images are mapped from pixel values to irradiance to synthesize high dynamic images , and the synthesis formula of the m point pixel is as follows:
[0050] In addition, the high dynamic image is linearly mapped to the 0-255 range to obtain the low dynamic image for display, and the mapping formula of the m point is as follows: *255;
[0051] wherein and are the maximum and minimum values in the high dynamic image respectively.
[0052] Further, more information is mined by replacing low dynamic RGB images with high dynamic images with more information, and the existing AI computing stick does not exceed 3Tops computing power. Due to the limitation of computing power, the segmentation network needs to be lightweight, and a convolution block with residual structure is designed to solve the problems of gradient disappearance and gradient explosion. An expansion convolution is introduced to expand the receptive field of the network. The input of the CNN segmentation network is a 512x512x3 high dynamic image, and the output is 512x512x4. The category of each spatial pixel point of 512x512 can be determined by the following method: comparing the size of the output 4 layers point by point, finding the maximum value index, when the index is 0, it corresponds to the background point, when the index is 1, it corresponds to the keyhole point, when the index is 2, it corresponds to the molten pool point, and when the index is 3, it corresponds to the to-be-welded weld point; The network structure is sequentially connected as follows: input layer: the input is 512x512x3;
[0053] First convolution layer: using 64 3x3 convolution kernels, the output is a 512x512x64 feature map;
[0054] First down-sampling layer: using a 2x2 pooling window for down-sampling, reducing the feature map size by half to 256x256x64;
[0055] The first convolution block with residual structure includes: a first main convolution layer, a first dilated convolution layer and a first residual connection, wherein the first main convolution layer: using 128 3x3 convolution kernels, generating a 256x256x128 feature map; the first dilated convolution layer: using 128 3x3 dilated convolution kernels, the expansion factor is 2, generating a 256x256x128 feature map; the first residual connection: adding the output of the first main convolution layer and the output of the first dilated convolution layer to obtain a 256x256x128 feature map;
[0056] Second down-sampling layer: using a 2x2 pooling window for down-sampling, reducing the feature map size by half to 128x128x128;
[0057] The second convolutional block with a residual structure comprises a second main convolutional layer, a second dilated convolutional layer and a second residual connection, wherein the second main convolutional layer generates a feature map of 128x128x128 using 128 3x3 convolutional kernels; the second dilated convolutional layer generates a feature map of 128x128x128 using 256 3x3 dilated convolutional kernels with a dilated factor of 2; and the second residual connection adds the output of the second main convolutional layer and the output of the second dilated convolutional layer to obtain a feature map of 128x128x128.
[0058] The first up-sampling layer restores the size of the feature map to 256x256x128 using 2x2 up-pooling.
[0059] The first concatenation layer concatenates the feature map of the first up-sampling layer and the feature map of the convolutional block with a residual structure to obtain a feature map of 256x256x256; the third convolutional block with a residual structure comprises a third main convolutional layer, a third dilated convolutional layer and a third residual connection, wherein the third main convolutional layer generates a feature map of 256x256x64 using 64 3x3 convolutional kernels; the third dilated convolutional layer generates a feature map of 256x256x64 using 64 3x3 dilated convolutional kernels with a dilated factor of 2; and the third residual connection adds the output of the third main convolutional layer and the output of the third dilated convolutional layer to obtain a feature map of 256x256x64.
[0060] The second up-sampling layer restores the size of the feature map to 512x512x64 using 2x2 up-pooling.
[0061] The third convolutional layer adopts 4 3x3 convolutional kernels and outputs a feature map of 512x512x4.
[0062] The output layer outputs a feature map of 512x512x4.
[0063] Further, the shapes of the keyhole and the molten pool change diversely during the welding process. The conventional assumption that the keyhole and the molten pool are ellipses only extracts the length and the width, which is not in line with the actual situation and will lose a lot of useful information. The boundary key point mode can more specifically describe the shape characteristics of the keyhole and the molten pool. In step S7, the keyhole graph is separated out, the keyhole region barycenter O is obtained according to the barycenter method, the uppermost point T, the lowermost point B, the leftmost point L and the rightmost point R of the keyhole region are calculated, the line connecting the points T, B, L and R with O as the center is divided into four regions, each region is equally divided by the central angle, there are 16 key points at the intersection of the division line and the boundary as key feature points, and the four regions contain the points T, B, L and R. The same operation is performed on the molten pool, and 16 key feature points of the molten pool are obtained.
[0064] Further, the BP network is used to automatically extract features instead of manually extracting the length and width and other shape features of the keyhole and the molten pool, the mining ability is deeper, and the key points of the current frame and the previous 3 frames are used as inputs at the same time, the welding is considered as a gradual process, and the historical frames have relevance, compared with only using the current frame, thereby improving the recognition accuracy, the penetration recognition network in the step S8 has 4 inputs and 1 output, the 4 inputs are respectively the current frame, and the 16 keyhole key points and the 16 molten pool key points of the previous 3 frames, each key point contains horizontal and vertical coordinates, and each input has 64 nodes; the output is 4 nodes, the maximum index of the 4 nodes is obtained, the index is 0, which represents under-penetration, the index is 1, which represents partial penetration, the index is 2, which represents good penetration, and the index is 3, which represents over-penetration; the network structure is sequentially connected as follows:
[0065] Input 1: 16 keyhole key points and 16 molten pool key points of the current frame, each key point contains horizontal and vertical coordinates, input 1 has 64 nodes;
[0066] Input 2: 16 keyhole key points and 16 molten pool key points of the first previous frame, each key point contains horizontal and vertical coordinates, input 2 has 64 nodes;
[0067] Input 3: 16 keyhole key points and 16 molten pool key points of the second previous frame, each key point contains horizontal and vertical coordinates, input 3 has 64 nodes;
[0068] Input 4: 16 keyhole key points and 16 molten pool key points of the third previous frame, each key point contains horizontal and vertical coordinates, input 4 has 64 nodes;
[0069] First BP network: the output of input 1 as input, output 64 nodes, the network structure contains 3 layers, the first hidden layer has 256 nodes, the second hidden layer has 1024 nodes, and the output layer has 64 nodes;
[0070] Second BP network: the output of input 2 as input, output 64 nodes, the network structure is the same as that of the first BP network;
[0071] Third BP network: the output of input 3 as input, output 64 nodes, the network structure is the same as that of the first BP network;
[0072] Fourth BP network: the output of input 4 as input, output 64 nodes, the network structure is the same as that of the first BP network;
[0073] Feature 1: the output of the first BP network as the feature, a total of 64 feature points;
[0074] Feature 2: the output of the second BP network as the feature, a total of 64 feature points;
[0075] Feature 3: the output of the third BP network as the feature, a total of 64 feature points;
[0076] Feature 4: the output of the fourth BP network as the feature, a total of 64 feature points;
[0077] First fully connected layer: 256 total feature points formed by connecting feature 1, feature 2, feature 3, and feature 4 in turn as input, and output 2048 nodes;
[0078] Second fully connected layer: connected to the first fully connected layer, and output 256 nodes;
[0079] Third fully connected layer: connected to the second fully connected layer, and output 4 nodes;
[0080] Output layer: output 4 nodes.
[0081] Further, the method of replacing the incomplete molten pool fitting and extracting the center of gravity with the complete molten pool extracted by the dual-camera synthesized image ensures more accurate center of gravity calculation; the method of replacing the Hough straight line transformation or direct fitting straight line method with the method of extracting the skeleton points of the weld to be welded and the RANSAC robust cubic polynomial method to fit the straight line of the weld to be welded ensures the fitting accuracy and stability of the weld center, and the welding deviation detection in step S9 is as follows:
[0082] S91, extract the 512x512 output layer, and separate the weld-to-be-welded segmentation image according to the category;
[0083] S92, perform binaryzation processing on the weld-to-be-welded segmentation image, and convert the weld area to white and the non-weld area to black;
[0084] S93, perform thinning on the binaryzation processed weld segmentation image to extract the skeleton of the weld;
[0085] S94, extract the weld skeleton points after thinning; wherein the thinning extraction process is as follows:
[0086] S941, initialize the input image and create an output image of the same size as the input image to store the thinning result;
[0087] S942, copy the input image to the output image;
[0088] S943, perform iterative processing, and the iteration process is as follows:
[0089] S9431, traverse each pixel in the output image, and skip the edge pixels;
[0090] S9432, for the current pixel P, obtain the values of its 8 adjacent pixels (up, down, left, right, top-left, top-right, bottom-left, bottom-right);
[0091] S9433, check the following conditions: the value of P is foreground (white), at least two of the 8 adjacent pixels of P are background (black), and the number of transitions of adjacent pixels in the 8 adjacent pixels of P is equal to 1; if all the above conditions are met, mark the current pixel P as a deletion candidate (background color);
[0092] S9434, update the output image according to the marked deletion candidate pixel, and delete the marked pixel;
[0093] S9435, repeat step S9433 and step S9434 until no pixel can be marked as a deletion candidate;
[0094] S9436, return the final refined result image.
[0095] S95, combine the cubic polynomial to fit the curve of the skeleton point, and the fitting process is as follows:
[0096] S951, randomly select a small part of points from the weld skeleton as samples;
[0097] S952, use these sample points to fit a cubic polynomial curve;
[0098] S953, calculate the distance of other non-sample points to the fitted curve, and divide them into inliers and outliers according to the set threshold;
[0099] S954, repeat the above steps, and select the model with the maximum number of inliers as the final fitting result;
[0100] S96, based on the horizontal coordinate of the current frame of the molten pool center of gravity as input, substitute the fitted cubic polynomial to calculate the vertical coordinate of the current frame of the weld curve, then the current vertical deviation is the vertical coordinate of the current penetration center of gravity minus the vertical coordinate of the weld curve;
[0101] S97, based on the proportional conversion relationship, get the actual vertical deviation, superimpose the original path planning between the current frame and the next frame with the vertical deviation path planning, and send the motion superposition instruction to the motion controller.
[0102] The present application has the following advantages and effects compared with the prior art:
[0103] 1) The present application is suitable for online welding quality detection and trajectory correction of lock hole TIG welding in medium thick plate curved long weld, the designed embedded device is small and flexible, can be installed on various crawling robots, and has wide application range;
[0104] 2) Dual camera uses hard trigger multi-exposure synchronous acquisition of welding image, and generates high dynamic image by software method, which is simpler than complex "single camera + filter combination" structure design, and the cost of the application is low compared with expensive high dynamic camera.
[0105] 3) Dual camera simultaneously collects in front and rear directions of the welding torch, and maps to the welding plate surface for splicing to obtain complete keyhole and molten pool image information, which is more complete than the single camera scheme in extracting keyhole and molten pool information, and the present scheme can simultaneously perform penetration detection and trajectory correction, while the single camera can only perform one of them. Compared with the double camera scheme of independent collection without splicing, it is more intuitive in vision, and the complete keyhole and molten pool image information can be directly observed on one image, which is more humanized in application to welding monitoring.
[0106] 4) The present application directly uses high dynamic image as the input of the segmentation network, which has higher dynamic range and richer input information than the conventional RGB image input. The designed residual structure convolution block not only solves the problems of gradient disappearance and gradient explosion, but also expands the receptive field of the network by introducing dilated convolution. The designed segmentation network is not only light, but also has high accuracy.
[0107] 5) The present application uses 32 key points of keyhole and molten pool to automatically extract features using BP network instead of manually extracting length, width and other morphological features of keyhole and molten pool, and the mining ability is deeper. The key points of the current frame and the previous 3 frames are used as input at the same time, which considers that welding is a gradual process and has relevance with historical frames, and has higher recognition accuracy compared with using only the current frame.
[0108] 6) The method of extracting the center of gravity of the complete molten pool segmented by the dual camera synthetic image is more accurate than the method of extracting the center of gravity of the molten pool fitted by the incomplete molten pool of the single camera. The method of fitting the straight line of the weld seam by extracting the skeleton points of the weld seam and using the RANSAC robust cubic polynomial method is more accurate and stable than the method of directly fitting the straight line. BRIEF DESCRIPTION OF DRAWINGS
[0109] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and serve to explain the principles of the application. In the drawings:
[0110] Figure 1 is a structure diagram of a keyhole TIG welding online visual detection device based on a dual camera according to an embodiment of the present application;
[0111] Figure 2is a double-camera-based lock hole TIG welding online visual inspection flowchart disclosed in the embodiments of the present application;
[0112] Figure 3 is a double-camera-based visual mounting schematic diagram disclosed in the embodiments of the present application;
[0113] Figure 4 is a neural network structure diagram for segmenting a lock hole, a molten pool and a weld to be welded, disclosed in the embodiments of the present application;
[0114] Figure 5 is a neural network structure diagram for identifying a penetration state, disclosed in the embodiments of the present application. DETAILED DESCRIPTION
[0115] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0116] EMBODIMENT
[0117] As shown in the figure, the double-camera-based lock hole TIG welding online visual inspection embedded device comprises: Figure 1
[0118] The first CMOS camera 1 and the second CMOS camera 2 are respectively connected with the embedded main control board 3, and are used to collect real-time welding scene images of the forward and reverse directions of the welding progress of the lock hole TIG welding. The two CMOS cameras synchronously collect multi-exposure images through a hard trigger signal, and transmit image data to the embedded main control board 3. One IO port of the embedded main control board 3 is connected with a signal distributor 12, which divides two IO ports connected with the two cameras, and synchronously controls the collection of different exposure times of the two cameras through different pulse signals. The two exposure time ranges of the two cameras should meet 10us-1s, the dynamic range should be greater than 60dB, the resolution should be higher than 1080P, and the frame rate should be greater than 90 frames / s. The two cameras collect the welding scene images of the forward and reverse directions of the welding progress, and the collected welding information is more comprehensive. The dynamic range of a conventional camera is only about 60dB, which is insufficient to observe the welding information, so the multi-exposure images are collected through a wide exposure time range to synthesize a high dynamic image, and welding observation is realized;
[0119] The embedded main control board 3 communicates data with other components of the visual detection device, and integrates a software package for keyhole TIG welding penetration identification and trajectory deviation correction; the embedded main control board 3 has a size of 80mmx55mm, adopts a 4-core Cortex-A9 processor of the NXP i.MX6 series, has a 1GHz main frequency, has a 1GB DDR3 memory, has a 4G EMMC storage, interfaces include: 1-way HDMI 2.0, 2-way gigabit Ethernet, 2-way USB2.0, 2-way USB3.0 ports, 1-way WiFi module interface, 5-way serial ports (including 1-way debugging serial port), 2-way TF cards, 3-way IIC, 1-way SPI, 1-way PCIE, 23-way GPIO and 1-way PWM, and is provided with an independent hardware watchdog; the embedded main control board 3 is small in size and rich in interfaces, and can be conveniently installed on the crawling robot.
[0120] The HDMI screen 4 is connected with the embedded main control board 3, and is used for display output of local real-time low dynamic welding images, segmentation and deviation identification images, penetration states and early warning information.
[0121] The mouse 5 is connected with the embedded main control board 3, and is used for software interface clicking and selection.
[0122] The keyboard 6 is connected with the embedded main control board 3, and is used for parameter setting input of a software interface.
[0123] The AI computing stick 7 communicates with the embedded main control board 3, transmits data through a USB3.0 port, is virtually connected with a network card after connection, a user can complete input and output of data only by socket programming, receives high dynamic image data preprocessed by the embedded main control board 3, infers a segmentation model and returns result data; the AI computing stick is more flexible, there are much more chips without AI hardware acceleration than chips with AI hardware acceleration, and the function and performance are powerful, the mode is selective, and when the single AI computing stick is insufficient in computing power, multiple AI computing sticks can be used, have the characteristics of computing power superposition, flexibility and convenience, and are suitable for intelligent needs of different keyhole TIG welding crawling.
[0124] The motion controller 8 is connected with the embedded main control board 3, receives a motion instruction of the embedded main control board 3 to control motion of the crawling robot, and feeds back real-time pose information of a welding torch, communicates with the embedded main control board 3 through a network port; the motion controller 8 integrates an EtherCAT master station library, supports 1ms cycle communication, can modify and increase motion online, plans a trajectory for a deviation after obtaining a welding deviation value, superimposes the trajectory on an original trajectory, and realizes a deviation correction function.
[0125] The demonstrator 9, connected with the motion controller 8, can manually teach the joint shaft and the rectangular coordinate system by key, complete the teaching of the starting point, the intermediate process point and the termination point, the image module integrates the motion function, the path planning program script can be generated by dragging the graphical programming, the program script is sent to the motion controller through the network port during running, the script is parsed by the motion controller to control the robot to execute the motion, and the teaching and reproduction of the predetermined trajectory are mainly completed;
[0126] The crawling robot 10 is connected with the motion controller 8, executes the control instructions sent by the motion controller 8, and feeds back the real-time shaft state of the robot, and communicates with the motion controller 8 through the network port;
[0127] The WIFI module 11 is connected with the embedded main control board 3, and is used for transmitting real-time low dynamic welding images, segmentation and deviation identification images, penetration states and early warning information to a remote server;
[0128] The signal distributor 12 is connected with the embedded main control board 3, the first CMOS camera 1 and the second CMOS camera 2 IO ports, is used for generating two-way output signals from the IO port output signals of the embedded main control board 3, is used for synchronously triggering two cameras, mainly realizes the synchronous collection of two cameras, and is used for splicing the two cameras.
[0129] The process based on the online visual detection method of the embedded device is described as follows:
[0130] Due to strong arc light interference in the welding process, it is difficult to extract visual information. The mainstream "single vision + filter combination" scheme cannot comprehensively extract welding process information, so that the welding precision is difficult to further improve, and meanwhile, the penetration recognition and trajectory correction functions cannot be realized at the same time. However, the double camera and the high dynamic synthesis + splicing fusion scheme used in the application can clearly extract complete molten pool and keyhole information, further improve the recognition precision, and meet the requirements of realizing penetration recognition and trajectory correction. As shown in Figure 2 The steps are as follows:
[0131] As shown in Figure 3 The first CMOS camera 1 and the second CMOS camera 2 are connected with the center line of the welding torch at an angle of 50 degrees and 60 degrees respectively, the first CMOS camera 1 needs to observe more welding areas, so the angle is smaller, and the second CMOS camera 2 mainly observes the molten pool and the keyhole for penetration recognition, so the angle is larger. The double camera synchronous collection scheme can more completely extract the information of the keyhole and the molten pool, can simultaneously realize penetration detection and trajectory correction, and can guarantee the precision and the welding quality in the welding tracking process.
[0132] The exposure time sequence is calibrated, and the exposure time ranges of the arc light, keyhole, molten pool, and to-be-welded / have-been-welded areas are artificially roughly selected. An equal-step exposure time sequence is generated in the respective range to perform a trial welding, and the images of the arc light, keyhole, molten pool, and to-be-welded / have-been-welded areas are intercepted. An image evaluation standard is established to automatically select the one with the highest exposure time score. In the image evaluation standard, a percentage method is used to determine the mean value of the pixel values ranked in the top 5% in brightness and the mean value of the pixel values ranked in the last 5% in brightness in the image, and then the difference between the two is calculated as the evaluation standard. The greater the difference, the higher the dynamic range of the image and the richer and clearer the content. The first CMOS camera 1 selects the exposure time for clear arc light, keyhole, molten pool, and to-be-welded area, and the second CMOS camera 2 selects the exposure time for clear arc light, keyhole, molten pool, and have-been-welded area. The square root of the product of the exposure times of the two cameras for arc light-arc light, keyhole-keyhole, molten pool-molten pool, and to-be-welded area-have-been-welded area is taken as the final four exposure times, and the four exposure times are synchronously collected through the external triggering control of the two cameras;
[0133] The stitching parameters are calibrated. A 13x9 checkerboard is designed, each grid is 3mmx3mm, and is placed in the overlapping area of the two cameras. The corner points of the checkerboard are detected, and the corner points of the two cameras are mapped to the same real world plane. The real world plane is selected as the surface of the welding plate, and the area is selected as part of the have-been-welded area, the molten pool area, and part of the to-be-welded area. There is a corresponding proportional relationship between the pixel values of the mapped images and the real world. The transformation matrix mapped to the plane is calculated respectively. wherein Unknown parameters are calculated and solved by the following equation:
[0134]
[0135] wherein and are the pixel coordinates of the i=1, 2, …, t corner points before transformation, and are the pixel coordinates of the i=1, 2, …, t corner points after transformation;
[0136] The proportional relationship between the pixel size and the actual size of the image after transformation is as follows:
[0137]
[0138] wherein , and , are the horizontal and vertical coordinates of the two points on the image after transformation, , and , are the horizontal and vertical coordinates of the corresponding two points in the real world. and are the horizontal and vertical scaling factors, respectively.
[0139] Stitching fusion, using a transformation matrix to map the images of two cameras to the same plane. By converting the pixel coordinates in each camera image to coordinates on the plane, the lockhole region is used as the overlap region to fuse by linear transition, ensuring that there is no obvious joint at the stitching. The clear images of the arc-arc, lockhole-lockhole, molten pool-molten pool, and to-be-welded area-welded area of the first CMOS camera 1 and the second CMOS camera 2 are stitched and fused respectively; during the stitching process, a linearly changing weight is used to achieve a smooth transition, the weight is calculated according to the position of the pixel in the overlap region, and an interpolation is performed between the pixel value 1 and the pixel value 2, the weight coefficient a is obtained by normalizing the distance of the pixel to the range [0, 1], and the calculation formula of the weight coefficient a is as follows: wherein d represents the distance of the pixel position to the boundary of the overlap region, is the diagonal distance of the overlap region, and the fused pixel value is obtained by linear interpolation using the weight coefficient a.
[0140] High dynamic synthesis, based on the generated 4 stitched and fused images of different exposure times, the camera response curve function is estimated, and the following formula is minimized:
[0141]
[0142] wherein, is the target loss function, is the spatial index of the pixel, is the exposure time corresponding to the stitched image index position, represents the corresponding pixel value, g( ) is the irradiance recovery function, represents the irradiance at index position m, and a hyperstatic equation is established to solve the 256 and pixel spaces in the formula;
[0143] Synthesizing a high dynamic image from the pixel value mapping to the irradiance of the 4 stitched images , the synthesis formula for the mth pixel is as follows:
[0144] In addition, each channel of the high dynamic image is linearly mapped to the range 0-255 to obtain a low dynamic image for display, and the mapping formula for the mth pixel is as follows: *255
[0145] wherein and are the maximum and minimum values in the high dynamic image respectively.
[0146] The synthesized high dynamic image is taken as input, and a CNN segmentation network is used to output a predicted feature map containing keyholes, molten pools and welds to be welded; the input of the CNN segmentation network is a 512x512x3 high dynamic image, and the output is 512x512x4, so the class of each spatial pixel point of 512x512 can be determined by the following method: the four layers of the output are compared point by point, the maximum value index is obtained, when the index is 0, it corresponds to a background point, when the index is 1, it corresponds to a keyhole point, when the index is 2, it corresponds to a molten pool point, and when the index is 3, it corresponds to a weld to be welded point. As shown in Figure 4 , the network structure is sequentially connected as follows:
[0147] Input layer: the input is 512x512x3;
[0148] First convolution layer: 64 3x3 convolution kernels are used, and the output is a 512x512x64 feature map;
[0149] First downsampling layer: a 2x2 pooling window is used for downsampling, and the feature map size is halved to 256x256x64;
[0150] First convolution block with residual structure, including: first main convolution layer, first dilated convolution layer and first residual connection, wherein the first main convolution layer: using 128 3x3 convolution kernels, generating a 256x256x128 feature map; the first dilated convolution layer: using 128 3x3 dilated convolution kernels, the expansion factor is 2, generating a 256x256x128 feature map; the first residual connection: adding the output of the first main convolution layer and the output of the first dilated convolution layer to obtain a 256x256x128 feature map;
[0151] Second downsampling layer: a 2x2 pooling window is used for downsampling, and the feature map size is halved to 128x128x128;
[0152] Second convolution block with residual structure, including: second main convolution layer, second dilated convolution layer and second residual connection, wherein the second main convolution layer: using 128 3x3 convolution kernels, generating a 128x128x128 feature map; the second dilated convolution layer: using 256 3x3 dilated convolution kernels, the expansion factor is 2, generating a 128x128x128 feature map; the second residual connection: adding the output of the second main convolution layer and the output of the second dilated convolution layer to obtain a 128x128x128 feature map;
[0153] The first up-sampling layer: using 2x2 up-pooling, the feature map size is restored to 256x256x128;
[0154] The first splicing layer: the feature map of the first up-sampling layer is spliced with the feature map of the convolution block with a residual structure, to obtain a feature map with a size of 256x256x256; the third convolution block with a residual structure comprises: a third main convolution layer, a third dilated convolution layer and a third residual connection, wherein the third main convolution layer: using 64 3x3 convolution kernels, generates a feature map of 256x256x64; the third dilated convolution layer: using 64 3x3 dilated convolution kernels, with an expansion factor of 2, generates a feature map of 256x256x64; the third residual connection: adds the output of the third main convolution layer and the output of the third dilated convolution layer to obtain a feature map of 256x256x64;
[0155] The second up-sampling layer: using 2x2 up-pooling, the feature map size is restored to 512x512x64;
[0156] The third convolution layer: using 4 3x3 convolution kernels, the output is a feature map of 512x512x4;
[0157] The output layer: the output is a feature map with a size of 512x512x4.
[0158] According to the obtained prediction feature map, the keyhole map is separated out, the center of gravity method is used to obtain the keyhole area center of gravity O, the uppermost point T, the lowermost point B, the leftmost point L and the rightmost point R of the keyhole area are calculated, and the center O is connected with the T point, the B point, the L point and the R point. Then, the four areas are divided into four areas according to the central angle of 4, and the 16 key points intersected by the dividing line and the boundary are taken as key feature points. The four areas contain the T point, the B point, the L point and the R point. The same operation is performed on the molten pool, and the 16 key feature points of the molten pool can be obtained in the same way.
[0159] The 32 key points of the keyhole and the molten pool are taken as features, and the 128 key points of the current frame and the previous 3 frames are taken as inputs. The state of the penetration is identified according to the constructed neural network. The network has 4 inputs and 1 output. The 4 inputs are respectively the 16 key points of the keyhole and the 16 key points of the molten pool of the current frame and the previous 3 frames. Each key point contains horizontal and vertical coordinates, and each input has 64 nodes. The output has 4 nodes. The maximum index of the 4 nodes is obtained. When the index is 0, it represents under-penetration; when the index is 1, it represents partial penetration; when the index is 2, it represents good penetration; and when the index is 3, it represents over-penetration. Figure 5 As shown in the figure, the network structure is sequentially connected as follows:
[0160] Input 1: 16 key points of the keyhole and 16 key points of the molten pool of the current frame, each key point contains horizontal and vertical coordinates, and input 1 has 64 nodes.
[0161] Input 2: 16 key points of the anterior first frame and 16 key points of the anterior second frame, each key point contains horizontal and vertical coordinates, and input 2 has 64 nodes;
[0162] Input 3: 16 key points of the anterior second frame and 16 key points of the anterior third frame, each key point contains horizontal and vertical coordinates, and input 3 has 64 nodes;
[0163] Input 4: 16 key points of the anterior third frame and 16 key points of the anterior fourth frame, each key point contains horizontal and vertical coordinates, and input 4 has 64 nodes;
[0164] First BP network: the output of input 1 as input, output 64 nodes, network structure contains 3 layers, the first hidden layer has 256 nodes, the second hidden layer has 1024 nodes, and the output layer has 64 nodes;
[0165] Second BP network: the output of input 2 as input, output 64 nodes, the network structure is the same as the first BP network;
[0166] Third BP network: the output of input 3 as input, output 64 nodes, the network structure is the same as the first BP network;
[0167] Fourth BP network: the output of input 4 as input, output 64 nodes, the network structure is the same as the first BP network;
[0168] Feature 1: the output of the first BP network as the feature, a total of 64 feature points;
[0169] Feature 2: the output of the second BP network as the feature, a total of 64 feature points;
[0170] Feature 3: the output of the third BP network as the feature, a total of 64 feature points;
[0171] Feature 4: the output of the fourth BP network as the feature, a total of 64 feature points;
[0172] First fully connected layer: 256 total feature points formed by connecting feature 1, feature 2, feature 3, and feature 4 in turn as input, output 2048 nodes;
[0173] Second fully connected layer: connected to the first fully connected layer, output 256 nodes;
[0174] Third fully connected layer: connected to the second fully connected layer, output 4 nodes;
[0175] Output layer: output 4 nodes.
[0176] According to the obtained prediction feature map, a molten pool image and a to-be-welded weld seam image are separated, a to-be-welded weld seam segmentation image is binarized, a weld seam region is converted into white, and a non-weld seam region is converted into black; a thinning algorithm is applied to the binarized weld seam segmentation image to extract the skeleton of the weld seam; the weld seam skeleton points are extracted through thinning, and the thinning steps are as follows: an input image is initialized, an output image with the same size as the input image is created to store the thinning result, the input image is copied to the output image, iterative processing is performed, each pixel in the output image is traversed, edge pixels are skipped, for a current pixel P, the values of 8 adjacent pixels (up, down, left, right, top-left, top-right, bottom-left and bottom-right) of P are obtained, the following conditions are checked: the value of P is foreground (white), at least two of the 8 adjacent pixels of P are background (black), and the number of transitions of adjacent pixels in the 8 adjacent pixels of P is equal to 1; if the above conditions are all met, the current pixel P is marked as a deletion candidate (background color); according to the marked deletion candidate pixel, the output image is updated, and the marked pixel is deleted; the process is repeated until no pixel can be marked as a deletion candidate, and a final thinning result image is returned; the RANSAC algorithm is used to combine a cubic polynomial to fit the skeleton point curve, and the algorithm is as follows: a small part of points in the weld seam skeleton are randomly selected as samples, a cubic polynomial curve is fitted using the sample points, the distances of other non-sample points to the fitted curve are calculated, and they are divided into inliers and outliers according to a set threshold, the above steps are repeated, and the model with the maximum number of inliers is selected as the final fitting result; based on the molten pool barycenter horizontal coordinate of the current frame as input, the fitted cubic polynomial is substituted to calculate the weld seam curve vertical coordinate of the current frame, then the current vertical deviation is the current penetration barycenter vertical coordinate minus the weld seam curve vertical coordinate, the actual vertical deviation is obtained based on the proportional conversion relationship, the original path planning between the current frame and the next frame is superimposed with the path planning of the vertical deviation, a motion superposition instruction is sent to a motion controller to control the welding deviation.
[0177] Test verification, different welding speeds, welding currents and pre-weld seam gaps are tested multiple times, and data sets are collected, the welding speed is set to 350-450 mm / min, the welding current is set to 480-540 A, the weld seam gap is set to 0.2-1 mm, and the deviation is artificially set to 0.2-0.5 mm. Under this data set, the penetration recognition method adopted by the application has a prediction accuracy of 96.3%, and compared with the current frame of the molten pool and the lock hole length and width as input features only using a single camera, the prediction accuracy increases by 7.6%, and the feature recognition accuracy of the application is higher. The trajectory deviation correction algorithm adopted by the application has an average tracking deviation of 0.35 mm and a maximum deviation of 0.53 mm, compared with the single camera fitting extraction of the molten pool barycenter and the linear fitting of the to-be-welded center, the average tracking deviation is reduced by 0.12 mm, and the maximum deviation is reduced by 0.34 mm.
[0178] The above embodiments are the preferred embodiments of the present application, but the embodiments of the present application are not limited to the above embodiments, and any changes, modifications, substitutions, combinations, simplifications, etc. made without departing from the spirit and principles of the present application should be equivalent replacement manners and should be included in the protection scope of the present application.
Claims
1. A visual inspection method based on a double camera lock hole TIG welding robot online visual inspection device, wherein, The visual detection device comprises: A first CMOS camera (1) and a second CMOS camera (2) are respectively connected with an embedded main control board (3) and are used for collecting real-time welding scene images of the forward and reverse directions of the welding advancement of the keyhole TIG welding, the two CMOS cameras synchronously collect multi-exposure images through a hard trigger signal, and image data is transmitted to the embedded main control board (3); The embedded main control board (3) is in data communication with other components of the visual detection device and integrates a software package for keyhole TIG welding penetration recognition and trajectory deviation correction; An HDMI screen (4) is connected with the embedded main control board (3) and is used for displaying and outputting local real-time low-dynamic welding images, segmentation and deviation identification images, penetration states and early warning information; A mouse (5) is connected with the embedded main control board (3) and is used for clicking and selecting a software interface; A keyboard (6) is connected with the embedded main control board (3) and is used for inputting parameter setting of a software interface; An AI computing stick (7) is connected with the embedded main control board (3), receives high-dynamic image data preprocessed by the embedded main control board (3), returns result data after inference of a segmentation model, and is in communication with the embedded main control board (3); A motion controller (8) is connected with the embedded main control board (3), receives a motion instruction of the embedded main control board (3) to control the motion of a crawling robot, feeds back real-time pose information of a welding torch, and is in communication with the embedded main control board (3) through a network port; A teach pendant (9) is connected with the motion controller (8) and is used for manually controlling the pose of the crawling robot (10) to teach a starting point, an intermediate process point and a termination point of welding, completing welding trajectory planning, and being in communication with the embedded main control board (3) through a network port; The crawling robot (10) is connected with the motion controller (8), executes a control instruction sent by the motion controller (8), and feeds back real-time axis states of the robot, and is in communication with the motion controller (8) through a network port; A WIFI module (11) is connected with the embedded main control board (3) and is used for transmitting real-time low-dynamic welding images, segmentation and deviation identification images, penetration states and early warning information to a remote server; A signal distributor (12) is connected with the embedded main control board (3), the first CMOS camera (1) and the second CMOS camera (2) IO ports, is used for generating two-way output signals from the IO port output signals of the embedded main control board (3), and simultaneously sends one signal output to the two cameras to realize synchronous triggering; The visual detection method comprises the following steps: Step S1: The first CMOS camera (1) and the second CMOS camera (2) are respectively installed at the forward and reverse directions of the welding torch advancement, the first CMOS camera (1) is at an angle of 50 degrees with the center line of the welding torch, and the second CMOS camera (2) is at an angle of 60 degrees with the center line of the welding torch; Step S2: Calibrate the exposure time sequence, artificially select the exposure time range of arc light, keyhole, molten pool, and to-be-welded / welded area, generate equal-step exposure time sequences within the respective ranges, and perform trial welding respectively, observe the trial welding image results, the first CMOS camera (1) selects the clear exposure time of the arc light, keyhole, molten pool, and to-be-welded area, and the second CMOS camera (2) selects the clear exposure time of the arc light, keyhole, molten pool, and to-be-welded area; the square root of the product of the exposure time of the two cameras arc light-arc light, keyhole-keyhole, molten pool-molten pool, and to-be-welded area-welded area is taken as the final four exposure times, and the four exposure times are synchronously collected through the external triggering of the two cameras; Step S3: Calibrate the stitching parameters, design a chessboard and place it in the overlapping area of the two cameras, detect the corner points of the chessboard, map the corner points of the two cameras to the same real-world plane, and select the welding plate surface as the real-world plane, and select part of the welded area, the molten pool area, and part of the to-be-welded area as the region. There is a corresponding proportional relationship between the pixel values of the mapped images and the real world, and the transformation matrix mapped to the plane is calculated respectively; Step S4: Stitching and fusion, map the images of the two cameras to the same plane using the transformation matrix, convert the pixel coordinates in each camera image to coordinates on the plane, and fuse the linear transition in the keyhole area as the overlapping area. The clear images of the arc light-arc light, keyhole-keyhole, molten pool-molten pool, and to-be-welded area-welded area of the first CMOS camera (1) and the second CMOS camera (2) are stitched and fused respectively; Step S5: High dynamic synthesis, estimate the camera response curve function according to the above four stitched and fused images with different exposure times, synthesize the four stitched and fused images with different exposure times into a high dynamic image according to the mapping relationship between the pixel value and the scene illumination, and linearly map each channel of the high dynamic image to the 0-255 range low dynamic image for display; Step S6: Take the high dynamic image synthesized in step S5 as input, and use the CNN segmentation network to output a predicted feature map containing the keyhole, molten pool, and to-be-welded weld; Step S7: Separate the keyhole image and the molten pool image according to the predicted feature map obtained in step S6, and extract 16 key points from each; Step S8: Take the 32 key points of the keyhole and the molten pool as features, take the features of the current frame and the previous 3 frames as input, and identify the state of the penetration according to the neural network; Step S9: Separate the molten pool image and the to-be-welded weld image according to the predicted feature map obtained in step S6, calculate the center of gravity of the molten pool, and refine the cubic curve of the to-be-welded weld center. Obtain the longitudinal deviation through the center of gravity of the molten pool and the center of the weld, and superimpose the longitudinal deviation on the original path planning to feed back to the robot controller to control the welding deviation.
2. The visual inspection method of claim 1, wherein, The AI computing stick (7) transmits data through a USB3.0 port, is used for inference of an AI model, is virtually connected to a network card after connection, and user inputs and outputs data through socket programming; The motion controller (8) integrates an EtherCAT master library, supports 1ms cycle communication, includes motion functions such as point motion, continuous trajectory, straight line and circular arc interpolation, continuous interpolation, and freely sets running speed, stopping speed, acceleration and deceleration time, and S-shaped curve smoothing parameters; The teach pendant (9) teaches joints and rectangular coordinate systems through manual keys, completes teaching of starting points, intermediate process points and ending points, and generates a program script through drag-and-drop graphical programming; during running, the program script is sent to the motion controller (8) through a network port, and the motion controller (8) analyzes and executes the script to control the robot motion; The WIFI module (11) transmits the real-time low dynamic image, the segmentation and deviation identification image, the penetration state and the early warning information to the remote server.
3. The visual inspection method of claim 1, wherein, In step S2, the exposure time ranges of the arc light, the keyhole, the molten pool and the to-be-welded / welded area are artificially coarsely selected, the equal-step exposure time sequences are generated in the respective ranges, the images of the arc light, the keyhole, the molten pool and the to-be-welded / welded area are respectively intercepted, and the image evaluation standard is set to automatically select the one with the highest exposure time score, wherein in the image evaluation standard, the percentage method is used to determine the mean value of the pixel values ranked in the top 5% in brightness and the mean value of the pixel values ranked in the last 5% in brightness, and then the difference between the two is calculated as the evaluation standard, the greater the difference, the higher the dynamic range of the image and the richer and clearer the content.
4. The visual inspection method of claim 1, wherein, In the step S3, the transformation matrix is: wherein are 8 unknown parameters to be solved, and the following equation is used to calculate the unknown parameters: wherein and are the pixel coordinates of the i = 1, 2,..., t corner points before transformation, and are the pixel coordinates of the i = 1, 2,..., t corner points after transformation. The ratio relationship between the pixel size of the transformed image and the actual size is as follows: where , and , are the horizontal and vertical coordinates of two points on the transformed image, , and , are the horizontal and vertical coordinates of the corresponding real-world actual two points, and are the horizontal and vertical scaling factors, respectively.
5. The visual inspection method of claim 1, wherein In the step S4, a linearly-varying weight is used in the splicing process to achieve smooth transition, the weight is calculated according to the position of the pixel in the overlapping area, and interpolation is performed between the pixel value 1 and the pixel value 2, the weight coefficient α is obtained by normalizing the distance of the pixel to the range of [0, 1], and the calculation formula of the weight coefficient α is as follows: wherein d represents the distance from the pixel position to the boundary of the overlapping region, is the diagonal distance of the overlapping region, and the fused pixel value is obtained by linear interpolation using the weight coefficient a.
6. The visual inspection method of claim 1, wherein, In step S5, the camera response curve function is estimated based on the four spliced and fused images with different exposure times, and the following formula is minimized to realize high dynamic synthesis: wherein, is a target loss function, is a spatial index of a pixel, is an exposure time is an index position of a corresponding stitched image, represents a corresponding pixel value, g( ) is an irradiance recovery function, represents an irradiance at index position m, hyperstatic equations are established to solve the 256 and pixel spatial ; Synthesizing high dynamic images from pixel value mapping to luminance for 4 tiled images For m points, the pixel synthesis formula is as follows: Additionally, the high dynamic image is linearly mapped to the 0-255 range to obtain a low dynamic image For display, the mapping formula for m points is as follows: * 255 wherein and are the maximum and minimum values, respectively, in a high dynamic image .
7. The visual inspection method of claim 1, wherein, The input of the CNN segmentation network is a 512x512x3 high dynamic image, and the output is 512x512x4, so the category of each spatial pixel point of 512x512 can be determined by the following method: the four layers of the output are compared point by point, the maximum value index is obtained, the index 0 corresponds to a background point, the index 1 corresponds to a keyhole point, the index 2 corresponds to a molten pool point, and the index 3 corresponds to a to-be-welded weld point; the network structure is sequentially connected as follows: The input layer: the input is 512x512x3; The first convolutional layer: 64 3x3 convolutional kernels are used, and the output is a 512x512x64 feature map; The first downsampling layer: a 2x2 pooling window is used for downsampling, and the feature map size is halved to 256x256x64; The first convolutional block with a residual structure includes a first main convolutional layer, a first dilated convolutional layer and a first residual connection, wherein the first main convolutional layer uses 128 3x3 convolutional kernels to generate a 256x256x128 feature map; the first dilated convolutional layer uses 128 3x3 dilated convolutional kernels with an expansion factor of 2 to generate a 256x256x128 feature map; and the first residual connection adds the output of the first main convolutional layer and the output of the first dilated convolutional layer to obtain a 256x256x128 feature map. The second downsampling layer: using a 2x2 pooling window for downsampling, the feature map size is reduced to 128x128x128; The second convolutional block with a residual structure includes: a second main convolutional layer, a second dilated convolutional layer and a second residual connection, wherein the second main convolutional layer: using 128 3x3 convolutional kernels, generates a feature map of 128x128x128; the second dilated convolutional layer: using 256 3x3 dilated convolutional kernels, the expansion factor is 2, generates a feature map of 128x128x128; the second residual connection: adds the output of the second main convolutional layer and the output of the second dilated convolutional layer to obtain a feature map of 128x128x128; The first upsampling layer: using 2x2 up-pooling, the feature map size is restored to 256x256x128; The first splicing layer: splicing the feature map of the first upsampling layer and the feature map of the convolutional block with a residual structure, to obtain a feature map with a size of 256x256x256; the third convolutional block with a residual structure includes: a third main convolutional layer, a third dilated convolutional layer and a third residual connection, wherein the third main convolutional layer: using 64 3x3 convolutional kernels, generates a feature map of 256x256x64; the third dilated convolutional layer: using 64 3x3 dilated convolutional kernels, the expansion factor is 2, generates a feature map of 256x256x64; the third residual connection: adds the output of the third main convolutional layer and the output of the third dilated convolutional layer to obtain a feature map of 256x256x64; The second upsampling layer: using 2x2 up-pooling, the feature map size is restored to 512x512x64; The third convolutional layer: using 4 3x3 convolutional kernels, the output is a feature map of 512x512x4; The output layer: the output is a feature map with a size of 512x512x4.
8. The vision inspection method of claim 1, wherein, In the step S7, the keyhole graph is separated out, the center of gravity method is used to obtain the keyhole area center of gravity O, the uppermost point T, the lowermost point B, the leftmost point L and the rightmost point R of the keyhole area are calculated, the line connecting the points T, B, L and R with O as the center is divided into four areas, each area is equally divided by the central angle, there are 16 key points as key feature points at the intersection of the division line and the boundary, and the four areas contain the points T, B, L and R. The same operation is performed on the molten pool, and 16 key feature points of the molten pool can be obtained in the same way.
9. The vision inspection method of claim 1, wherein, The penetration recognition network in the step S8 has 4 inputs and 1 output, and the 4 inputs are respectively 16 key points of the keyhole and 16 key points of the molten pool of the current frame and the previous 3 frames, each key point contains horizontal and vertical coordinates, and each input has 64 nodes; the output has 4 nodes, and the maximum index of the 4 nodes is obtained, the index is 0, which represents under-penetration, the index is 1, which represents partial penetration, the index is 2, which represents good penetration, and the index is 3, which represents over-penetration; the network structure is sequentially connected as follows: Input 1: 16 key points of the keyhole and 16 key points of the molten pool of the current frame, each key point contains horizontal and vertical coordinates, and input 1 has 64 nodes; Input 2: 16 key points of keyhole and 16 key points of molten pool of the first previous frame, each key point contains horizontal and vertical coordinates, input 2 is 64 nodes; Input 3: 16 key points of keyhole and 16 key points of molten pool of the second previous frame, each key point contains horizontal and vertical coordinates, input 3 is 64 nodes; Input 4: 16 key points of keyhole and 16 key points of molten pool of the third previous frame, each key point contains horizontal and vertical coordinates, input 4 is 64 nodes; First BP network: the output of input 1 as input, output 64 nodes, network structure contains 3 layers, the first hidden layer node number is 256 nodes, the second hidden layer node number is 1024 nodes, and the output layer is 64 nodes; Second BP network: the output of input 2 as input, output 64 nodes, the network structure is the same as the first BP network; Third BP network: the output of input 3 as input, output 64 nodes, the network structure is the same as the first BP network; Fourth BP network: the output of input 4 as input, output 64 nodes, the network structure is the same as the first BP network; Feature 1: the output of the first BP network as the feature, a total of 64 feature points; Feature 2: the output of the second BP network as the feature, a total of 64 feature points; Feature 3: the output of the third BP network as the feature, a total of 64 feature points; Feature 4: the output of the fourth BP network as the feature, a total of 64 feature points; First full connection layer: 256 total feature points formed by connecting feature 1, feature 2, feature 3 and feature 4 in turn as input, output 2048 nodes; Second full connection layer: connecting the first full connection layer, output 256 nodes; Third full connection layer: connecting the second full connection layer, output 4 nodes; Output layer: output is 4 nodes.
10. The visual inspection method of claim 9, wherein, The welding deviation detection is used in step S9, and the process is as follows: S91, extracting the output layer of 512x512, separating the to-be-welded weld segmentation map according to the category; S92, performing binaryzation processing on the to-be-welded weld segmentation map, converting the weld area into white and the non-weld area into black; S93, performing thinning on the binaryzation weld segmentation map to extract the weld skeleton; S94, extracting the weld skeleton point after thinning; S95, combining the cubic polynomial to fit the skeleton point curve; S96, based on the horizontal coordinate of the center of gravity of the molten pool of the current frame as input, substituting into the fitted cubic polynomial to calculate the vertical coordinate of the weld curve of the current frame, then the current vertical deviation is the vertical coordinate of the current penetration center of gravity minus the vertical coordinate of the weld curve; S97, based on the proportional conversion relationship to obtain the actual vertical deviation, superimposing the path planning of the original path planning between the current frame and the next frame with the vertical deviation, and sending the motion superposition instruction to the motion controller.
Citation Information
Patent Citations
A Multi-Information Acquisition and Monitoring System and Method for Robot Welding Process
CN109719368B
A full-position robotic deep penetration K-TIG welding system and control method
CN115106621B
Intelligent robot welding device using large-scale workpiece
CN101456182A
Embedded device and method for detecting welding quality of lockhole TIG welding in real time
CN113751920A