A laser welding keyhole detection method and system

CN122821144APending Publication Date: 2026-09-25COGNITIVE PHOTONICS (BEIJING) LASER TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611307525.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-27
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]然而,语义分割网络的训练依赖大规模像素级精确标注数据,而熔池图像存在边界模糊、受反光影响严重、形态动态多变等固有特性,导致人工像素级标注极为困难,单张图像的标注耗时数分钟至数十分钟,构建包含数千张标注样本的训练集成本极高,严重制约了语义分割模型在焊接熔池在线监测中的工程化应用

Benefits of technology

[0033]第五方面,本申请实施例还提供一种芯片模组,包括收发组件和芯片,所述芯片,用于执行第一方面中任意一种激光焊接熔池匙孔检测方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821144A_ABST
    Figure CN122821144A_ABST
Patent Text Reader

Abstract

The application discloses a laser welding molten pool crater detection method and system, and relates to the technical field of laser welding quality detection.The method comprises the following steps: acquiring time sequence image frames of a molten pool crater to be detected; inputting the time sequence image frames into a semantic segmentation model to obtain segmentation results of molten pool regions and crater regions in each image frame; constructing time sequence characteristic sequences related to the geometric sizes of the molten pool and the crater based on the segmentation results of each image frame; acquiring statistical characteristic values of the geometric size data in a first sliding window based on the position of a current first image frame and a preset sliding window length, and determining a normal fluctuation range according to the statistical characteristic values; if the geometric size of the first image frame falls within the normal fluctuation range, marking the frame segmentation result as a qualified pseudo label; and iteratively updating the semantic segmentation model based on the qualified pseudo label.Under the condition of only a small amount of manually labeled samples, the application realizes continuous optimization of the semantic segmentation model, and cooperatively improves the accuracy of welding quality detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of laser welding quality inspection technology, and in particular to a method and system for detecting the keyhole of a laser welding molten pool. Background Technology

[0002] Laser welding is widely used in high-end manufacturing fields such as automobile manufacturing, aerospace, and new energy batteries. The morphological changes of the molten pool and keyhole during the welding process directly reflect the quality of the weld. A deep learning-based semantic segmentation-based image processing method for the molten pool and keyhole can achieve pixel-level precise extraction of the molten pool region, providing key visual features for online monitoring of welding quality.

[0003] However, the training of semantic segmentation networks relies on large-scale pixel-level precise annotation data. However, molten pool images have inherent characteristics such as blurred boundaries, severe glare, and dynamic and variable shapes, making manual pixel-level annotation extremely difficult. Annotating a single image takes several minutes to tens of minutes, and building a training ensemble containing thousands of labeled samples is extremely costly, which seriously restricts the engineering application of semantic segmentation models in online monitoring of welding molten pools.

[0004] Therefore, how to provide an online detection method for laser welding weld pools that can reduce reliance on large-scale manual annotation has become a key research topic for those skilled in the art. Summary of the Invention

[0005] In a first aspect, this application provides a method for detecting keyholes in laser welding molten pools. The method includes: acquiring temporal image frames of the keyhole in the molten pool to be detected; inputting the temporal image frames into a semantic segmentation model to obtain segmentation results of the molten pool region and the keyhole region in each frame image, wherein the segmentation results include the position and contour of the molten pool region and the keyhole region in the image, wherein the initial semantic segmentation model is trained using a small amount of manually labeled sample data; based on the segmentation results of each frame image, constructing a temporal feature sequence related to the geometric dimensions of the molten pool and the keyhole, wherein the temporal feature sequence includes the geometric dimensions of the molten pool and the keyhole in the segmentation results of each frame image; and based on the position of the current first image frame and a preset sliding angle... The first sliding window is used to obtain the first geometric size data corresponding to the first sliding window in the temporal feature sequence, and the statistical feature value corresponding to the first geometric size data is determined. Based on the statistical feature value corresponding to the first geometric size data and the preset confidence interval parameter, a first normal fluctuation range is determined. The statistical feature value includes the mean and variance. The first sliding window is any segmentation window in the temporal feature sequence. If it is determined that the geometric size of the first image frame falls within the first normal fluctuation range, the segmentation result of the first image frame is marked as a qualified pseudo-label. Based on the image frame segmentation result corresponding to the qualified pseudo-label, the semantic segmentation model is iteratively updated using training data.

[0006] The laser welding molten pool keyhole detection method provided in this application transforms the semantic segmentation results of the molten pool and keyhole into a geometric dimension temporal feature sequence. A normal fluctuation range is constructed based on the mean and variance of the temporal sequence as a physical constraint criterion. Adaptive quality screening is performed on the prediction results of the semantic segmentation model, marking segmentation results whose geometric dimensions fall within the normal fluctuation range as qualified pseudo-labels. These qualified pseudo-labels are then used as training data to iteratively update the semantic segmentation model. By utilizing the inherent physical characteristic that the dimensions of the molten pool and keyhole follow statistical laws during stable welding, manual annotation is replaced as the evaluation standard for pseudo-label quality. This allows for continuous optimization of the semantic segmentation model with only a small number of manually annotated samples, eliminating the need for additional sensors such as infrared thermal imaging. This fundamentally reduces the model's dependence on large-scale pixel-level precise annotation data and effectively avoids model performance degradation caused by erroneous pseudo-labels participating in training.

[0007] On the other hand, this application organically links the quality screening of the segmentation results of the semantic segmentation model with the model iterative update process, and constructs a closed-loop self-training mechanism of "prediction -> screening -> iterative training -> re-prediction", so that the segmentation model can continuously evolve itself by using the newly added welding data after deployment.

[0008] In some possible implementations, the method further includes: marking the segmentation result of the current frame as a frame to be compensated if it is determined that the geometric size of the first image frame does not fall within the first normal fluctuation range; searching the temporal feature sequence for a forward nearest neighbor reliable frame that is located before the frame to be compensated and belongs to a qualified pseudo-label, and for a backward nearest neighbor reliable frame that is located after the frame to be compensated and belongs to a qualified pseudo-label; if both the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame exist simultaneously, calculating a first rate of change of the molten pool area and a second rate of change of the keyhole area between the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame; if both the first rate of change and the second rate of change are less than the corresponding set thresholds, then it is determined to be a quasi-compensation frame. In a steady-state condition, a pixel-level logical OR operation is performed on the segmentation results of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame to generate a compensating pseudo-label for the frame to be compensated. If the first rate of change and the second rate of change are greater than or equal to the corresponding set thresholds, it is determined to be a non-steady-state condition. Pixel-level linear interpolation is then performed on the segmentation results of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame to generate a compensating pseudo-label for the frame to be compensated. The step of iteratively updating the semantic segmentation model based on the image frame segmentation results corresponding to the qualified pseudo-labels as training data includes: iteratively updating the semantic segmentation model based on the image frame segmentation results corresponding to the candidate pseudo-label set as training data, wherein the candidate pseudo-label set includes qualified pseudo-labels and compensating pseudo-labels.

[0009] In this approach, when the geometric dimensions of the first image frame do not fall within the normal fluctuation range, the segmentation result of that frame is not directly discarded. Instead, it is marked as a frame to be compensated. By finding the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame, the area change rate is calculated, and an interpolation strategy is adaptively selected based on the area change rate criterion: under quasi-steady-state conditions, pixel-level logical "OR" union operation is used to generate compensating pseudo-labels, and under non-steady-state conditions, pixel-level linear interpolation is used to generate compensating pseudo-labels. Finally, the compensating pseudo-labels are merged with qualified pseudo-labels as training data to iteratively update the model.

[0010] This method transforms anomalous frames, which are discarded directly due to unreliable predictions in traditional methods, into effective compensatory pseudo-labels. This not only expands the diversity of training samples but also significantly enhances the model's ability to perceive precursors or boundary conditions of welding instability. Frames marked as unreliable often correspond to moments of drastic changes in the molten pool morphology, which are precisely the samples the model most needs to learn to improve its generalization ability. At the same time, the compensatory pseudo-labels and qualified pseudo-labels together form a candidate pseudo-label set to participate in iterative training. This allows the segmentation model to continuously absorb abnormal morphological information in the welding process with minimal manual annotation, avoiding the loss of the model's ability to perceive welding anomalies due to the discarding of anomalous samples.

[0011] On the other hand, existing interpolation methods based on pixel-value weighted averaging produce soft labels between 0 and 1 when generating pseudo-labels, leading to blurred melt pool boundaries and ghosting issues, and requiring additional binarization threshold selection. This application employs pixel-level logical OR union operations under quasi-steady-state conditions to directly generate binary hard pseudo-labels with values ​​of 0 or 1, avoiding the boundary value blurring and ghosting introduced by weighted averaging. Furthermore, the logical OR operation only involves binary logic operations, resulting in computational efficiency far exceeding floating-point multiplication and addition operations, making it more suitable for high-speed online industrial processing. Simultaneously, the union strategy comprehensively utilizes the boundary information of the preceding and following frames. When the segmentation of the preceding frame is too small and the segmentation of the following frame is too large, the union precisely compensates for the shortcomings, effectively compensating for potential edge shrinkage or expansion errors in single-frame segmentation, thus improving the computational efficiency of interpolation compensation while ensuring the geometric accuracy of the pseudo-labels.

[0012] Furthermore, this application adaptively distinguishes between quasi-steady-state and non-steady-state conditions using an area change rate criterion. This approach considers both the efficiency and clear boundaries of union operations in quasi-steady-state conditions, and the smooth transition capability of linear interpolation for rapid molten pool deformation in non-steady-state conditions. This avoids the problems of a single interpolation strategy leading to an inflated area in non-steady-state conditions or introducing unnecessary floating-point operations in quasi-steady-state conditions through linear interpolation. Thus, by organically combining interpolation compensation for defective frames with closed-loop self-training, the segmentation model continuously absorbs abnormal condition information during iteration, thereby continuously improving segmentation accuracy and generalization ability.

[0013] In some possible implementations, the first rate of change and the second rate of change are calculated based on the following formula:

[0014] in, , These are the times of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame, respectively; in In the case of the first rate of change, , The areas of the melt pools for the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame are respectively; in In the case of the second rate of change, , The areas of the melt pools for the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame are respectively. The pixel-level linear interpolation process uses the following formula:

[0015] Where t is the time of the frame to be compensated. , The values ​​are the segmentation results at pixel position (x, y) for the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame, respectively. The value of the compensation pseudo-label is assigned to the frame to be compensated at pixel position (x,y).

[0016] This method quantitatively calculates the relative change in the area of ​​the molten pool and keyhole between consecutive reliable frames using an area change rate formula. Normalization is performed using A as a benchmark, making the area change rate a dimensionless relative indicator. This eliminates the influence of differences in the absolute size of the molten pool under different welding conditions on the threshold setting. Simultaneously, the pixel-level linear interpolation formula weights the frame to be compensated with respect to the time distance between it and the consecutive reliable frames, with closer frames having higher weights. This aligns with the physical law that the molten pool morphology changes continuously over a short period. After interpolation, a hard label is generated using a threshold of 0.5, avoiding the boundary ambiguity problem introduced by soft labels.

[0017] In some possible implementations, the method further includes: if only the forward nearest neighbor reliable frame or only the backward nearest neighbor reliable frame exists, then the segmentation result of the existing frame is used as the compensatory pseudo-label of the frame to be compensated.

[0018] This approach, when the frame to be compensated is located at the beginning or end of the temporal sequence and only the forward or backward nearest neighbor reliable frame exists, uses the segmentation result of that existing frame as the compensatory pseudo-label. This avoids the problem of not being able to generate compensatory pseudo-labels due to the lack of bilateral reference frames, ensuring that the interpolation compensation mechanism can take effect at any position in the temporal sequence, guaranteeing the integrity of abnormal frame compensation and full temporal coverage. This boundary processing strategy, together with the bilateral interpolation mechanism, constitutes a complete pseudo-label compensation link from the inside of the sequence to the edge, further expanding the number of pseudo-label samples that can be used for iterative training.

[0019] In some possible implementations, the step of iteratively updating the semantic segmentation model based on the image frame segmentation results corresponding to the candidate pseudo-label set as training data includes: reserving one frame as the target pseudo-label every interval Δ frames in the candidate pseudo-label set, where Δ satisfies the formula: f is the current camera frame rate. The target effective frame rate is set as follows: and / or, in the candidate pseudo-label set, the intersection-union ratio (IUR) of the segmentation results of adjacent frames is calculated. If the IUR is greater than a set similarity threshold, the segmentation results of adjacent frames are determined to be redundant, and only one frame is retained as the target pseudo-label. The segmentation results of the image frames corresponding to the retained target pseudo-labels are used as training data to iteratively update the semantic segmentation model.

[0020] This approach removes redundancy from the candidate pseudo-label set through adaptive frame extraction and / or inter-frame intersection-union (ICUUID) deduplication: adaptive frame extraction dynamically adjusts the retention interval based on the camera frame rate, with higher frame rates resulting in sparser retention, thus uniformly compressing the sample size over time; ICUUI deduplication directly eliminates adjacent frames with highly similar morphologies. Using either approach alone or in combination can effectively compress the training set size while maintaining the diversity of melt pool morphology distribution, suppressing the model's tendency to overfit specific morphologies, reducing the computational cost of iterative training, and improving model training efficiency and generalization ability.

[0021] In some possible implementations, the preset sliding window length is determined based on the current image acquisition frame rate of the time-series image frame and a preset physical time span.

[0022] Using this method, the preset sliding window length is determined based on the current image acquisition frame rate and the preset physical time span. The window length automatically increases when the frame rate increases and automatically decreases when the frame rate decreases. This ensures that the mean and variance statistically obtained within the window are physically comparable under different frame rate conditions, avoids the problem of statistical distortion when the frame rate changes in a fixed-length window, and significantly improves the robustness of the detection method under variable frame rate conditions in industrial settings.

[0023] In some possible implementations, the method further includes: obtaining a segmentation mask sequence output by the iteratively optimized semantic segmentation model; constructing a reference temporal feature sequence related to the reference geometric dimensions of the molten pool and keyhole based on the segmentation results of each frame image in the segmentation mask sequence, wherein the reference temporal feature sequence includes the first-order difference and second-order difference of the reference geometric dimensions of the molten pool and keyhole in the segmentation results of each frame image, wherein the reference geometric dimensions include at least one of the following: molten pool area, molten pool length, molten pool width, and molten pool aspect ratio; extracting temporal feature data within a first temporal window from the reference temporal feature sequence based on the current frame position and a preset temporal window length; inputting the temporal feature data within the first temporal window into a preset temporal prediction model, and outputting a welding quality status discrimination result corresponding to the first temporal window; wherein the preset temporal prediction model is trained using welding process temporal data with quality labels.

[0024] This approach involves acquiring the segmentation mask sequence output by an iteratively optimized semantic segmentation model. A reference temporal feature sequence is then constructed, containing the geometric dimensions of the molten pool and keyhole (at least one of area, length, width, and aspect ratio) and their first-order and second-order differences. This sequence is input into a pre-defined temporal prediction model in units of temporal windows, outputting welding quality status discrimination results. The first-order and second-order differences reflect the rate and acceleration of change in the molten pool and keyhole dimensions, respectively, capturing the dynamic evolution trend of the molten pool morphology during welding and providing richer dynamic feature information for quality discrimination. Since the input features of the temporal prediction model come from the iteratively optimized segmentation model, the continuous improvement in segmentation accuracy directly improves the quality of the temporal features, thereby increasing the accuracy of quality detection.

[0025] On the other hand, this application organically links the iterative optimization of the semantic segmentation model with temporal quality detection. Each time the segmentation model completes an iterative update, it can provide a higher quality segmentation mask sequence as input to the temporal prediction model. The quality detection results output by the temporal prediction model can then be used to verify the optimization effect of the segmentation model, forming a collaborative positive feedback mechanism of 'segmentation optimization → temporal detection → feedback verification'. This achieves the collaborative optimization of improving the performance of the segmentation model and increasing the accuracy of quality detection.

[0026] Secondly, this application also provides a laser welding molten pool keyhole detection system, the system comprising: an image acquisition module for acquiring temporal image frames of the molten pool keyhole to be detected; a segmentation processing module for inputting the temporal image frames into a semantic segmentation model to obtain segmentation results of the molten pool region and keyhole region in each frame image, the segmentation results including the position and contour of the molten pool region and keyhole region in the image, wherein the initial semantic segmentation model is trained using a small amount of manually labeled sample data; a size calculation module for constructing a temporal feature sequence related to the geometric dimensions of the molten pool and keyhole based on the segmentation results of each frame image, the temporal feature sequence including the geometric dimensions of the molten pool and keyhole in the segmentation results of each frame image; and a statistical evaluation module for obtaining the size of the molten pool and keyhole based on the position of the current first image frame and a preset sliding window length. The first geometric dimension data corresponding to the first sliding window in the time-series feature sequence is described, and the statistical feature value corresponding to the first geometric dimension data is determined. Based on the statistical feature value corresponding to the first geometric dimension data and a preset confidence interval parameter, a first normal fluctuation range is determined. The statistical feature value includes the mean and variance. The first sliding window is any segmentation window in the time-series feature sequence. The statistical evaluation module is further used to mark the segmentation result of the first image frame as a qualified pseudo-label when it is determined that the geometric dimension of the first image frame falls within the first normal fluctuation range. The model update module is used to iteratively update the semantic segmentation model based on the segmentation result of the image frame corresponding to the qualified pseudo-label as training data, including a unit for executing any one of the laser welding molten pool keyhole detection methods in the first aspect.

[0027] In some possible implementations, the system further includes: an interpolation compensation module, which is used to: mark the segmentation result of the current frame as a frame to be compensated when it is determined that the geometric size of the first image frame does not fall within the first normal fluctuation range; search the temporal feature sequence for whether there is a forward nearest neighbor reliable frame that is located before the frame to be compensated and belongs to a qualified pseudo-label, and whether there is a backward nearest neighbor reliable frame that is located after the frame to be compensated and belongs to a qualified pseudo-label; if both the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame exist, calculate a first rate of change of the molten pool area and a second rate of change of the keyhole area between the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame; if the first rate of change and If the second rate of change is less than the corresponding set threshold, it is determined to be a quasi-steady-state condition. A pixel-level logical OR operation is performed on the segmentation results of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame to generate a compensating pseudo-label for the frame to be compensated. If the first rate of change and the second rate of change are greater than or equal to the corresponding set threshold, it is determined to be a non-steady-state condition. Pixel-level linear interpolation is performed on the segmentation results of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame to generate a compensating pseudo-label for the frame to be compensated. The model update module is specifically used to iteratively update the semantic segmentation model based on the image frame segmentation results corresponding to the candidate pseudo-label set as training data. The candidate pseudo-label set includes qualified pseudo-labels and compensating pseudo-labels.

[0028] In some possible implementations, the interpolation compensation module is further configured to use the segmentation result of the existing frame as the compensatory pseudo-label of the frame to be compensated if it is determined that only the forward nearest neighbor reliable frame or only the backward nearest neighbor reliable frame exists.

[0029] In some possible implementations, the system further includes: a timing quality detection module, configured to: acquire a segmentation mask sequence output by the iteratively optimized semantic segmentation model; construct a reference timing feature sequence related to the reference geometric dimensions of the molten pool and keyhole based on the segmentation results of each frame image in the segmentation mask sequence, wherein the reference timing feature sequence includes the first-order difference and second-order difference of the reference geometric dimensions of the molten pool and keyhole in the segmentation results of each frame image, wherein the reference geometric dimensions include at least one of the following: molten pool area, molten pool length, molten pool width, and molten pool aspect ratio; extract timing feature data within a first timing window from the reference timing feature sequence based on the current frame position and a preset timing window length; input the timing feature data within the first timing window into a preset timing prediction model, and output a welding quality status discrimination result corresponding to the first timing window; wherein the preset timing prediction model is trained using welding process timing data with quality labels.

[0030] Any of the implementations shown in the first aspect can be applied to the laser welding molten pool keyhole detection system in this application, and will not be described in detail here.

[0031] Thirdly, this application also provides a computer storage medium that can store multiple instructions, which are adapted to be loaded and executed by a processor using any one of the laser welding molten pool keyhole detection methods in the first aspect.

[0032] Fourthly, embodiments of this application also provide a computer program product containing instructions, which, when run on an electronic device, causes the electronic device to execute any one of the laser welding molten pool keyhole detection methods in the first aspect.

[0033] Fifthly, embodiments of this application also provide a chip module, including a transceiver component and a chip, wherein the chip is used to perform any of the laser welding molten pool keyhole detection methods in the first aspect.

[0034] In a sixth aspect, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method described in any one of the first aspects is executed when the processor executes the program.

[0035] It is understood that the laser welding molten pool keyhole detection system, computer storage medium, computer program, computer program product, chip system, and electronic device provided above are all used to execute the method shown in any implementation of the first aspect of the embodiments of this application. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description

[0036] Figure 1 This is a schematic flowchart of a laser welding molten pool keyhole detection method provided in an embodiment of this application; Figure 2 This is a schematic flowchart of an interpolation compensation method for defective image frames provided in an embodiment of this application; Figure 3 This is a schematic diagram of a laser welding molten pool keyhole detection system provided in an embodiment of this application; Figure 4 This is a schematic diagram of a timing quality detection model structure provided in an embodiment of this application; Figure 5 This is a system overall architecture block diagram provided in an embodiment of this application. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described below in conjunction with the accompanying drawings.

[0038] It should be noted that the embodiments of this application use an electronic device as an example to illustrate the execution subject of the laser welding molten pool keyhole detection method provided in this application. This electronic device can also be understood as the laser welding molten pool keyhole detection system shown in the embodiments of this application. In the embodiments of this application, the electronic device can be an industrial computer, embedded processing system, or GPU computing platform, etc., for executing program code. Any electronic device that can be used to execute the method provided in the embodiments of this application falls within the protection scope of the embodiments of this application, and this application does not impose any limitations. For example, the electronic device can be an industrial control computer, an embedded ARM processor, an FPGA-based processing system, or a digital signal processor, etc., and the embodiments of this application do not limit this.

[0039] Please see Figure 1 , Figure 1 This is a flowchart illustrating a laser welding molten pool keyhole detection method provided in an embodiment of this application. Figure 1 As shown, the laser welding molten pool keyhole detection method includes the following steps: S101, the electronic device acquires a timing image frame of the keyhole of the molten pool to be detected.

[0040] In this embodiment, the electronic device uses an industrial camera with a narrowband filter to acquire time-series image frames of the molten pool area during laser welding. The camera can use a CMOS or CCD sensor, with a frame rate f ranging from 50 to 1000 fps and a resolution of no less than 896×896 pixels. The acquired images are transmitted to an industrial computer or embedded processing platform. The time-series image frames are arranged in chronological order of acquisition, and the time interval between adjacent frames is determined by the camera frame rate.

[0041] S102, the electronic device inputs the time-series image frames into the semantic segmentation model to obtain the segmentation results of the molten pool region and keyhole region in each frame image.

[0042] In this embodiment, the electronic device inputs the temporal image frames obtained in step S101 into a semantic segmentation model. This model classifies each pixel in each frame and outputs a segmentation result, which includes the position and contour of the molten pool region and keyhole region in the image. For example, a pixel-level segmentation mask M_t is output. For instance, the value of each pixel in the segmentation mask indicates whether the pixel belongs to the molten pool region (value 1), the keyhole region (value 2), or the background region (value 0), or vice versa. The initial semantic segmentation model is obtained through supervised pre-training on a small number of manually pixel-level annotated melt pool keyhole images (e.g., 200-300 images). The initial semantic segmentation model can employ an encoder-decoder structure based on a deep convolutional neural network, such as DDRNet, U-Net, or variants thereof. After training, the model is deployed in an online monitoring system to perform forward inference on each real-time input frame of the melt pool image.

[0043] S103, the electronic device constructs a temporal feature sequence related to the geometric dimensions of the molten pool and keyhole based on the segmentation results of each frame image.

[0044] In this embodiment, the electronic device calculates the geometric dimensions of the molten pool and keyhole in each frame of the image based on the segmentation mask M_t output in step S102, and constructs a temporal feature sequence of the geometric dimensions of the molten pool and keyhole. The geometric dimensions include, but are not limited to, one or more of the following: area, length, width, and aspect ratio.

[0045] For example, the meaning of area: Taking the weld pool area A_weldpool_t as an example: Let the weld pool mask value be 1. Count all pixels with a value of 1 in the segmentation mask, and multiply by the physical area corresponding to a single pixel (obtained through camera calibration), in mm. 2 Taking the keyhole area A_keyhole_t as an example: Let the keyhole mask value be 2. Count all pixels with a value of 2 in the segmentation mask and multiply by the physical area corresponding to a single pixel.

[0046] The meaning of length: the maximum span of the corresponding area in the segmentation mask in the welding direction.

[0047] Width refers to the maximum span of the corresponding area in the segmentation mask in the direction perpendicular to the welding direction.

[0048] The meaning of aspect ratio: the ratio of length to width.

[0049] The constructed temporal feature sequence can be represented as: Seq_t=[A_{t-N+1}, A_{t-N+2}, ..., A_t], where N is the length of the sequence, which is sufficient to cover the preset sliding window length L in subsequent steps.

[0050] In the embodiments of this application, the length, width, aspect ratio, etc. of the molten pool and keyhole can be extracted simultaneously to construct multi-dimensional time-series features, thereby improving the robustness of subsequent statistical criteria.

[0051] S104, the electronic device acquires the first geometric dimension data corresponding to the first sliding window in the time-series feature sequence based on the position of the current first image frame and the preset sliding window length, and determines the statistical feature value corresponding to the first geometric dimension data, and determines the first normal fluctuation range based on the statistical feature value corresponding to the first geometric dimension data and the preset confidence interval parameter.

[0052] For example, the preset sliding window length is determined based on the current image acquisition frame rate of the time-series image frames and a preset physical time span. For instance, the electronic device acquires the current real-time acquisition frame rate f (frames / second) of the camera, sets the fixed physical time span covered by the window to T (seconds), preferably in the range of 0.1 to 1.0 seconds, and dynamically adjusts it based on the welding speed. The current sliding window length L is calculated using the following formula: ,in To round down, L is guaranteed to be a positive integer.

[0053] When the frame rate increases, the window length automatically increases to include more frames; when the frame rate decreases, the window length automatically decreases. Regardless of the frame rate change, the window always covers the data within a fixed physical time span T, ensuring the physical comparability of the statistics. The physical time span T can be dynamically adjusted based on the welding speed: when the welding speed increases, the molten pool size changes faster, and the T value can be appropriately decreased (e.g., 0.1~0.3s) to improve the timeliness of the statistics; when the welding speed decreases, the T value can be appropriately increased (e.g., 0.5~1.0s) to smooth out random noise.

[0054] For example, for the image frame corresponding to time t (the first image frame), the geometric dimension data {A_{t-L+1}, A_{t-L+2}, ..., A_t} within the interval [t-L+1,t] of the temporal feature sequence are extracted, and their sample mean is calculated. and sample variance :

[0055]

[0056] Based on the mean and variance mentioned above, a normal fluctuation range is constructed. A preset confidence interval parameter k is set (e.g., a value of 2~3), and the normal fluctuation range is [...]. ], The standard deviation is denoted as .

[0057] It should be noted that the above process is performed independently for each different geometric dimension, calculating the mean, variance, and normal fluctuation range for each corresponding geometric dimension. For example, if the geometric dimension includes area (molten pool area and keyhole area), then the mean, variance, and normal fluctuation range for the molten pool area and keyhole area are calculated separately. It should also be noted that this geometric dimension including the molten pool and keyhole is only an example; constraints such as the aspect ratio of the molten pool can also be introduced to form multi-dimensional criteria, further reducing the probability of misjudgment.

[0058] It should be noted that, in the embodiments of this application, the current first image frame refers to any one of the timing image frames of the keyhole of the molten pool to be detected. Based on S104-S106, it is determined whether any one of the timing image frames of the keyhole of the molten pool to be detected is a qualified frame. If it is an unqualified frame, interpolation compensation can be performed based on the steps of S201-S206.

[0059] It should be noted that to calculate the mean and variance, at least two frames of data (with a denominator of L-1) are required within the first sliding window. In practice, to ensure statistical stability, the window length L is usually much greater than 2 (e.g., L=100, corresponding to T=0.5s, f=200fps). In some possible implementations, after acquiring time-series image frames, the electronic device needs to accumulate segmentation result data for at least L frames before starting the effective statistical filtering in S104. That is, the initial frames are only used to fill the window to establish a statistical baseline, do not generate pseudo-labels, and only temporarily store the segmentation results. After the initial window is filled, the current frame is used as the last frame of the window before the effective statistical filtering in S104 begins.

[0060] S105, the electronic device determines whether the geometry of the first image frame falls within the first normal fluctuation range.

[0061] S106, if it is determined that the geometric size of the first image frame falls within the first normal fluctuation range, the segmentation result of the first image frame is marked as a qualified pseudo-label.

[0062] For example, if the geometry includes area (molten pool area and keyhole area), the molten pool area and keyhole area of ​​the current frame simultaneously satisfy: If the melt pool area and keyhole area are within the normal fluctuation range, the segmentation mask M_t of the current frame is deemed reliable and directly enters the pseudo-label pool as a candidate pseudo-label. If either the melt pool area or the keyhole area does not meet the above conditions, it indicates that the segmentation mask of the current frame is likely to have oversegmentation or undersegmentation errors, and it is marked as an unqualified frame (abnormal frame) and enters the interpolation compensation process.

[0063] In some possible implementations, the electronic device performs further interpolation compensation on unqualified frames that meet certain conditions and then includes them in candidate pseudo-labels. Specifically, as Figure 2 shown, after said S105, the compensation procedure for unqualified frames includes: S201: in a case where it is determined that the geometric dimension of the first image frame does not fall within a first normal fluctuation range, marking the segmentation result of the current frame as a frame to be compensated.

[0064] For example, the geometric dimensions include an area (a molten pool area and a keyhole area), when any one of the molten pool area or the keyhole area does not fall within a corresponding normal fluctuation range, it indicates that the segmentation result of the current frame may have errors such as over-segmentation, under-segmentation or boundary offset, and this frame is marked as a frame to be compensated.

[0065] S202: searching whether there is a forward nearest neighbor reliable frame located before the frame to be compensated and belonging to a qualified pseudo-label, and whether there is a backward nearest neighbor reliable frame located after the frame to be compensated and belonging to a qualified pseudo-label in the time-series feature sequence.

[0066] In the embodiment of the present application, the electronic device backtracks and searches for a time t1 of the forward nearest neighbor reliable frame in the time series (satisfying t1 < t, and this frame is determined as a qualified pseudo-label in step S105) and a time t2 of the backward nearest neighbor reliable frame (satisfying t2 > t, and this frame is determined as reliable in step S105).

[0067] S203: if both the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame exist, calculating a first change rate for the molten pool area and a second change rate for the keyhole area between the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame.

[0068] For example, the first change rate or the second change rate is calculated based on the following formula:

[0069] wherein, , are respectively the moments of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame; when represents said first change rate, , are respectively the molten pool areas of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame; when represents said second change rate, , are respectively the molten pool areas of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame; S204. If both the first rate of change and the second rate of change are less than the corresponding set threshold, it is determined to be a quasi-steady state. A pixel-level logical "OR" operation is performed on the segmentation results of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame to generate a compensating pseudo-label for the frame to be compensated.

[0070] For example, a threshold θ_area for the rate of change of the area is set (preferably within the range of 15% to 25%). When both the rate of change of the molten pool area and the rate of change of the keyhole area are less than the corresponding set threshold, it is determined that both the molten pool and the keyhole are in a quasi-steady-state change process, and strategy A (logical "OR" union operation) is adopted.

[0071] For quasi-steady-state conditions, a pixel-level logical OR operation is directly performed on the masks of the preceding and following frames to generate compensating pseudo-labels for the current frame:

[0072] That is, for any pixel position (x, y), if it is in M t1 Or M t2 If any pixel in the mask belongs to the melt pool region (value 1), then in the generated compensating pseudo-label, that pixel also belongs to the melt pool region (value 1). The compensation pseudo-label value is assigned to the frame to be compensated at pixel position (x, y). The keyhole region is generated in the same way.

[0073] S205, if the first rate of change and the second rate of change are greater than or equal to the corresponding set threshold, it is determined to be a non-steady-state condition. Pixel-level linear interpolation is performed on the segmentation results of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame to generate a compensating pseudo-label for the frame to be compensated.

[0074] In this embodiment, when the rate of change of the molten pool area or the rate of change of the keyhole area is greater than or equal to the corresponding set threshold, it is determined that the molten pool or keyhole is in a non-steady-state process of rapid contraction or expansion, and strategy B (pixel-level linear interpolation) is adopted.

[0075] Pixel-level linear interpolation generates compensatory pseudo-labels for the frame to be compensated using the following formula:

[0076] Where t is the time of the frame to be compensated. , These are the times of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame, respectively. , These are the segmentation results at pixel position (x, y) for the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame, respectively. The compensation pseudo-label value is assigned to the frame to be compensated at pixel position (x, y). The above formula is applied to the segmentation results of the molten pool region and the keyhole region, respectively, with independent interpolation for each region.

[0077] S206. If only a forward nearest neighbor reliable frame or only a backward nearest neighbor reliable frame exists, then the segmentation result of the existing frame is used as the compensatory pseudo-label of the frame to be compensated.

[0078] In this embodiment of the application, if only a forward reliable frame t1 exists but a backward reliable frame t2 does not exist (i.e., the abnormal frame occurs at the end of the sequence), then it is directly used as... As a compensation pseudo-label for the current frame; if only a backward reliable frame t2 exists, then... As a compensation pseudo-label. If no reliable frames exist before or after, the pseudo-label generation for that frame is abandoned.

[0079] Through the above interpolation compensation mechanism, the frames marked as unreliable are often samples with drastic changes in the molten pool morphology (corresponding to the precursors of welding instability). These are converted into compensatory pseudo-labels and incorporated into the training, effectively enhancing the model's ability to perceive abnormal working conditions.

[0080] In this embodiment, pseudo-labels refer to training targets generated from model predictions after filtering or interpolation compensation, used to replace manual labels; they are understood as being distinct from manually labeled real labels. Qualified pseudo-labels refer to image frames used to record reliable segmentation results that meet confidence criteria. Compensatory pseudo-labels are pseudo-labels generated after interpolation compensation processing of unqualified image frames, distinct from directly retained reliable pseudo-labels.

[0081] In this embodiment of the application, if an unqualified frame has neither a forward nearest neighbor reliable frame nor a backward nearest neighbor reliable frame, it is directly discarded.

[0082] S107, the electronic device uses the image frame segmentation results corresponding to the candidate pseudo-label set as training data to iteratively update the semantic segmentation model.

[0083] In this embodiment of the application, the candidate pseudo-label set includes qualified pseudo-labels generated in step S106 and compensatory pseudo-labels generated in steps S201 to S206.

[0084] In some possible implementations, redundancy removal can be performed on the candidate pseudo-label set before it is used for model training. Specifically: Method 1: Adaptive frame dropping based on fixed intervals: One frame is reserved as a pseudo-label every interval Δ frames, and Δ is adaptively calculated based on the camera frame rate. ), where f is the current frame rate and f_target is the target effective frame rate (preferred range is 30~60fps). For example, when f=200fps and f_target=50fps, Δ=4, that is, 1 frame is reserved every 4 frames.

[0085] Method 2: Adaptive deduplication based on inter-frame intersection-over-union (IoU): Adjacent pseudo-labels in the candidate pseudo-label set and The cross-union ratio (CUI) between segmentation masks is calculated for adjacent pseudo-labels to measure the spatial overlap of the foreground regions of the molten pool keyhole in two frames:

[0086] Set a similarity threshold θ_IoU (preferably within the range of 0.85~0.95). When S≥θ_IoU, the two frame segmentation masks are determined to be highly redundant, and only one frame is retained while the other frame is discarded.

[0087] In some possible implementations, a two-stage screening process of "frame extraction + deduplication" can be adopted: the first stage extracts frames at a fixed interval Δ to compress the number of candidate pseudo-labels to the target level; the second stage calculates the IoU between adjacent frames for the remaining pseudo-labels after frame extraction and removes redundant frames with similarity greater than the threshold similarity threshold.

[0088] The simplified pseudo-label set after redundancy removal is merged with the initial manually labeled training set, and the current semantic segmentation model is iteratively trained. After iterative training, the updated model is redeployed for online segmentation in step S102, forming a closed-loop self-training mechanism of "prediction → filtering / interpolation → deduplication → training → update". Multiple iterations can be performed until the model's segmentation accuracy on the validation set converges.

[0089] Preferably, a "student-teacher" framework can be introduced to further stabilize the iterative training process: the online running model is used as the teacher model, and the new model after each iteration of training is used as the student model. The student model updates the parameters of the teacher model through an exponential moving average, avoiding noise disturbances in a single training session that could lead to model performance degradation.

[0090] In some possible implementations, the electronic device also acquires the segmentation mask sequence output by the iteratively optimized semantic segmentation model, constructs a reference temporal feature sequence based on the segmentation mask sequence, inputs it into the temporal prediction model, and outputs the welding quality status discrimination result.

[0091] In this embodiment, the electronic device acquires the semantic segmentation model iteratively optimized in step S107, and outputs a segmentation mask sequence in subsequent applications. Based on the segmentation results of each frame in the segmentation mask sequence, temporal features of each frame are extracted to construct a reference temporal feature sequence. The temporal features include at least one of the following: melt pool area, melt pool length, melt pool width, melt pool aspect ratio, and the first-order and second-order differences of the above geometric dimensions; and / or, keyhole area, keyhole length, keyhole width, keyhole aspect ratio, and the first-order and second-order differences of the above geometric dimensions.

[0092] Based on the current frame position and a preset timing window length, timing feature data within the current timing window is extracted from the reference timing feature sequence. This data is then input into a preset timing prediction model (such as InceptionTime, LSTM, or a Transformer encoder), and the output is the welding quality status judgment result at the current moment. The judgment result includes: penetration status (normal penetration / incomplete penetration / overpenetration) and / or defect type (spatter, porosity, dents, etc.). The preset timing prediction model is trained using quality-labeled welding process timing data.

[0093] The input features of the time series prediction model come from the segmentation model that has been iteratively optimized. The improvement of segmentation accuracy directly improves the quality of the input features of the time series model, thereby improving the accuracy of quality detection.

[0094] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of a timing quality detection model provided in an embodiment of this application. Figure 3 As shown, the temporal quality detection model adopts the InceptionTime architecture. The input is the temporal feature vector sequence extracted from the temporal segmentation mask sequence output by the semantic segmentation model after iterative optimization. The input time series is sequentially passed through convolutional layers and multiple cascaded Inception modules to extract multi-scale temporal features. Shallow features are passed between Inception modules through residual connections. The extracted features are reduced in dimensionality by global average pooling layers and then mapped to the category space by fully connected layers. Finally, the model outputs the discrimination results of K welding quality status categories.

[0095] The laser welding molten pool keyhole detection method provided in this application trains an initial semantic segmentation model using a small number of manually labeled samples. Leveraging the physical constraint that the molten pool and keyhole dimensions follow statistical laws during stable welding, a normal fluctuation range is constructed based on the mean and variance of the time series. Adaptive quality screening is then applied to the prediction results of the semantic segmentation model. Qualified frames are directly used as pseudo-labels, while unqualified frames are converted into valid pseudo-labels through multi-strategy interpolation compensation (using logical "OR" union operations to generate binary hard labels in quasi-steady-state conditions, and pixel-level linear interpolation in non-steady-state conditions). After redundancy removal, these pseudo-labels are merged with the initial manually labeled samples for iterative model training. Thus, this application achieves continuous optimization of the semantic segmentation model with only a small number of manually labeled samples and without relying on additional sensors.

[0096] Laser welding, with its advantages of high energy density, high precision, and low heat input, is widely used in high-end manufacturing fields such as automobile manufacturing, aerospace, and new energy batteries. The morphological changes of the molten pool during welding directly reflect the quality of the weld—the stability of the molten pool size, the opening and closing state of the keyhole, and the generation of spatter are all closely related to the weld penetration state and defect formation. Therefore, real-time visual monitoring of the molten pool and keyhole during laser welding is a key technical means to achieve online control of welding quality.

[0097] With the rapid development of machine vision and deep learning technologies, semantic segmentation-based keyhole image processing methods have become a research hotspot. Semantic segmentation can achieve pixel-level accurate extraction of the keyhole region, providing high-quality feature input for subsequent calculations of keyhole geometry, morphological analysis, and quality status determination. Researchers have attempted to use semantic segmentation networks such as U-Net, PSPNet, and DDRNet to achieve automatic segmentation of the molten pool and keyhole regions. Although deep learning semantic segmentation methods have shown superior performance in molten pool image processing, their training process faces a core bottleneck: semantic segmentation networks require pixel-level accurately labeled image data for supervised training. Molten pool keyhole images have the following characteristics, making annotation difficult: 1. Blurred boundary of molten pool: The transition between the molten pool area and the surrounding base material in terms of grayscale is smooth and lacks a clear edge, making it difficult for the marking personnel to accurately delineate the boundary of the molten pool; 2. The keyhole is easily affected by the reflected light signal of the molten pool: In laser welding, due to the influence of strong light during welding, an additional lighting source is used to improve the clarity of the molten pool. The lighting source will form a strong reflected light signal on the molten pool, affecting the boundary definition of the keyhole. 3. Dynamic and variable morphology: The size and shape of the molten pool and keyhole vary significantly under different welding parameters, making it difficult to standardize the labeling. 4. The annotation workload is huge: Manual pixel-level annotation of a single melt pool keyhole image takes several minutes to tens of minutes, and the cost of building a training ensemble containing thousands of annotated samples is extremely high.

[0098] The high cost of annotation mentioned above severely restricts the engineering application of semantic segmentation models in online monitoring of weld pool keyholes.

[0099] Furthermore, in practical industrial applications, to achieve high temporal resolution capture of dynamic changes in the molten pool, camera frame rates are typically set high (e.g., 200-1000 fps). However, a direct consequence of high frame rates is that during relatively stable welding periods, the differences between consecutive frames of molten pool images are minimal, resulting in highly similar segmentation masks. In model training, these highly redundant samples do not increase the amount of effective information; instead, they amplify the model's tendency to overfit specific samples, reducing its generalization ability to changes in welding conditions. This further exacerbates the difficulty of obtaining efficient, high-quality training samples.

[0100] The following describes the existing technologies related to this solution: In recent years, researchers have introduced deep learning semantic segmentation networks into the field of molten pool image processing. Typical approaches include: using U-Net networks to achieve high-precision segmentation and localization of the molten pool region; combining GANs and improved PSPNet to achieve dynamic tracking and accurate detection of the molten pool; and real-time semantic segmentation networks based on DDRNet. The core workflow of these methods is: acquiring a large number of molten pool keyhole images → manual pixel-level annotation → training the semantic segmentation network → deployment for online segmentation.

[0101] The closest existing implementation to this invention is the "Visual Monitoring Method for Welding Molten Pool Based on Multimodal Semi-Supervised Semantic Segmentation" (Publication No. CN119027672A) applied for by Dongfang Turbine Co., Ltd. of Dongfang Electric Corporation. Its technical solution involves constructing a semantic segmentation network (DDRNet), training the network using a small amount of labeled data and pseudo-label data from different modalities (such as pseudo-labels generated by infrared thermal imaging) based on a semi-supervised learning strategy, and finally using the trained network for molten pool image segmentation. The core innovation of this solution lies in adopting a "student-teacher" semi-supervised framework. It utilizes infrared thermal imaging modality to generate the first pseudo-label, combines it with the second pseudo-label generated by the teacher model, and supervises the unlabeled data through multimodal pseudo-labels, thereby reducing the reliance on manual annotation.

[0102] In the broader field of computer vision, various pseudo-label generation techniques exist. For example, by using two teacher models with different structures to perform object detection on the same unlabeled image, the independent detection box (i.e., the object box detected by only one of the two models) is identified as a pseudo-label to recall missed targets. Such methods are designed for general object detection scenarios and do not consider the specific physical characteristics of weld pools.

[0103] The following describes the main disadvantages of existing technologies. (1) The semi-supervised approach failed to address the issue of counterfeit label quality. Semi-supervised melt pool segmentation schemes, exemplified by CN119027672A, reduce annotation costs by generating pseudo-labels from multimodal data. However, this pseudo-label generation relies on auxiliary modal data acquired by additional sensors such as infrared thermal imaging, increasing the system's hardware cost and deployment complexity. More importantly, this scheme lacks an effective screening and quality assurance mechanism for pseudo-labels—pseudo-labels predicted incorrectly by the teacher model directly participate in training, causing error accumulation and leading to model performance degradation.

[0104] (2) False labels only "qualified" and "discard abnormal" Existing pseudo-label-based self-training methods typically only set a confidence threshold for the model's predictions, using only predictions above the threshold as pseudo-labels and discarding those below. This "discarding anomalies" strategy leads to two problems: first, a large number of samples carrying useful information are wasted; second, the discarded samples are often boundary conditions or abnormal states, precisely the samples the model most needs to learn to improve its generalization ability. In the scenario of welding pool keyhole monitoring, abnormal keyhole morphology (such as excessively large, excessively small, or drastic fluctuations) corresponds to precursors of welding process instability. The lack of such samples will severely weaken the model's ability to perceive welding anomalies.

[0105] (3) Fixed window statistical method is not suitable for strained frame rate conditions Existing time-series statistical methods typically employ a fixed-length sliding window. However, in actual industrial welding processes, the camera's frame rate may be adjusted due to hardware performance limitations, variations in transmission bandwidth, or process requirements. When the frame rate changes, the physical time span covered by the fixed-length window also changes, causing the statistically derived mean and variance to lose physical comparability and failing to accurately characterize the dynamic fluctuations in the keyhole size of the molten pool.

[0106] (4) Existing interpolation methods are prone to producing boundary blurring. Traditional pixel-weighted averaging interpolation methods require floating-point multiplication and addition operations on the masks of consecutive frames when generating pseudo-labels. The interpolation result will form soft label values ​​between 0 and 1 at the melt pool boundary. Although soft labels retain probabilistic information, additional binarization threshold selection is required in actual training. Furthermore, when the melt pool is undergoing rapid deformation or displacement, the weighted averaging method is prone to introducing "ghosting" or area shrinkage, affecting the geometric accuracy of the pseudo-labels.

[0107] (5) High frame rate acquisition leads to redundant pseudo-labels, causing model overfitting. Existing pseudo-label generation schemes do not consider the high similarity between consecutive frames under high frame rate acquisition conditions. During the stable welding stage, the differences between dozens or even hundreds of consecutive frames captured by a high frame rate camera are minimal. If all reliable frames are included in the training set as pseudo-labels, it will result in a large amount of redundant information in the training samples. This will cause the model to overfit to specific morphologies of the molten pool, reduce its generalization ability to changes in welding conditions, and increase unnecessary training computational overhead.

[0108] (6) Separation of semantic segmentation and quality detection In existing technologies, the training and optimization of semantic segmentation models are independent of downstream quality detection tasks. The optimization goal of the segmentation model is only to improve segmentation accuracy, without considering the direct impact of segmentation results on the final quality detection accuracy. There is a lack of an integrated closed-loop design of "segmentation → screening → detection".

[0109] In view of the shortcomings of the prior art, the objectives of the present invention include: 1. A self-training method is provided that does not rely on additional sensor modalities and only uses the temporal statistical features of the molten pool keyhole image itself for pseudo-label quality screening and compensation, and continuously optimizes the semantic segmentation model under the condition of a small number of manually labeled samples; 2. By using a multi-strategy interpolation compensation mechanism, the abnormal segmentation frames that are discarded in the traditional method are transformed into effective pseudo-labels. In particular, sharp binary hard pseudo-labels are generated by logical "OR" union operation, avoiding the boundary blurring and ghosting problems introduced by the traditional weighted average method, and enhancing the model's ability to perceive welding instability precursors and abnormal working conditions. 3. An adaptive window length calculation method linked to the camera frame rate is proposed to ensure that the temporal statistical features maintain physical consistency under different frame rate conditions; 4. An adaptive removal mechanism for redundant pseudo-labels is provided. By deduplicating between frames and / or adaptively extracting frames, the size of the training set can be effectively controlled while ensuring sample diversity, suppressing model overfitting and improving model training efficiency. 5. Construct an integrated online monitoring system for "semantic segmentation → pseudo-label filtering / interpolation → redundancy removal → iterative training → temporal quality detection" to achieve synergistic optimization of segmentation model performance improvement and quality detection accuracy improvement.

[0110] This application also provides a laser welding molten pool keyhole detection system, including a unit for performing any of the laser welding molten pool keyhole detection methods described in the above method embodiments.

[0111] Please refer to Figure 4 This is a schematic diagram of a laser welding molten pool keyhole detection system provided in an embodiment of this application. Figure 4 As shown, the laser welding molten pool keyhole detection system of this application embodiment may include: Image acquisition module 401 is used to acquire timing image frames of the keyhole of the molten pool to be detected; The segmentation processing module 402 is used to input the temporal image frames into the semantic segmentation model to obtain the segmentation results of the molten pool region and keyhole region in each frame image. The segmentation results include the position and contour of the molten pool region and keyhole region in the image. The initial semantic segmentation model is trained by a small amount of manually labeled sample data. The size calculation module 403 is used to construct a temporal feature sequence related to the geometric dimensions of the molten pool and keyhole based on the segmentation results of each frame image. The temporal feature sequence includes the geometric dimensions of the molten pool and keyhole in the segmentation results of each frame image. The statistical evaluation module 404 is used to obtain the first geometric size data corresponding to the first sliding window in the time-series feature sequence based on the position of the current first image frame and the preset sliding window length, and to determine the statistical feature value corresponding to the first geometric size data. Based on the statistical feature value corresponding to the first geometric size data and the preset confidence interval parameter, a first normal fluctuation range is determined. The statistical feature value includes the mean and variance. The first sliding window is any segmentation window in the time-series feature sequence. The statistical evaluation module 404 is further configured to mark the segmentation result of the first image frame as a qualified pseudo-label when it is determined that the geometric size of the first image frame falls within the first normal fluctuation range. The model update module 405 is used to iteratively update the semantic segmentation model based on the image frame segmentation results corresponding to the qualified pseudo-labels as training data.

[0112] In some possible implementations, the system further includes: an interpolation compensation module, which is used to: mark the segmentation result of the current frame as a frame to be compensated when it is determined that the geometric size of the first image frame does not fall within the first normal fluctuation range; search the temporal feature sequence for whether there is a forward nearest neighbor reliable frame that is before the frame to be compensated and belongs to a qualified pseudo-label, and whether there is a backward nearest neighbor reliable frame that is after the frame to be compensated and belongs to a qualified pseudo-label; if both the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame exist, calculate a first rate of change of the molten pool area and a second rate of change of the keyhole area between the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame; if both the first rate of change and the second rate of change are less than the corresponding set values... If the threshold is reached, it is determined to be a quasi-steady-state condition. A pixel-level logical OR operation is performed on the segmentation results of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame to generate a compensating pseudo-label for the frame to be compensated. If the first rate of change and the second rate of change are greater than or equal to the corresponding set threshold, it is determined to be a non-steady-state condition. Pixel-level linear interpolation is performed on the segmentation results of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame to generate a compensating pseudo-label for the frame to be compensated. The step of iteratively updating the semantic segmentation model based on the image frame segmentation results corresponding to the qualified pseudo-labels as training data includes: iteratively updating the semantic segmentation model based on the image frame segmentation results corresponding to the candidate pseudo-label set as training data, wherein the candidate pseudo-label set includes qualified pseudo-labels and compensating pseudo-labels.

[0113] Please see Figure 5 , Figure 5 This is a schematic diagram of the architecture of a laser welding molten pool keyhole detection system provided in an embodiment of this application. Figure 5 As shown, the system consists of two parts: an industrial control computer and a model training server. The industrial control computer is responsible for real-time image acquisition, semantic segmentation inference, extraction of physical features of the melt pool keyhole, temporal quality detection, and display of monitoring results. The model training server is responsible for melt pool keyhole feature statistics, screening of qualified pseudo-labels and interpolation compensation of abnormal frames, pseudo-label deduplication, merging of pseudo-labels and manual annotations, iterative training and update deployment of the model. After training, the model is deployed to the industrial control computer through the model deployment module for real-time inference.

[0114] This application also provides a computer storage medium that can store multiple instructions. These instructions are adapted to be loaded and executed by a processor using the laser welding molten pool keyhole detection method provided in this application. For details of the execution process, please refer to the specific description of the method embodiments shown above, which will not be elaborated here.

[0115] This application also provides a computer program product containing instructions that, when run on an electronic device, cause the electronic device to execute the method steps of the method embodiments shown above.

[0116] This application also provides a chip module, including a transceiver component and a chip, wherein the chip is used to execute the method steps of the above-described method embodiments.

[0117] It is understood that the laser welding molten pool keyhole detection system, computer storage medium, computer program, computer program product, and chip provided above are all used to execute the method shown in any implementation of the method embodiments of this application. Therefore, the specific implementation and the beneficial effects that can be achieved can be referred to the beneficial effects in the corresponding method, and will not be detailed here.

[0118] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it includes the processes of the embodiments of the above methods.

[0119] The term "at least one" in this application refers to one or more items. "More than one item" means two or more items. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, it should be understood that although the terms "first," "second," etc., may be used to describe objects in this application, these objects should not be limited to these terms. These terms are only used to distinguish the objects from each other.

[0120] The terms “including” and “having” mentioned above, and any variations thereof, are intended to cover non-exclusive inclusion.

[0121] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for detecting the keyhole in a laser welding weld pool, characterized in that, The method includes: Acquire timing image frames of the keyhole in the molten pool to be detected; The temporal image frames are input into the semantic segmentation model to obtain the segmentation results of the molten pool region and keyhole region in each frame image. The segmentation results include the position and contour of the molten pool region and keyhole region in the image. The initial semantic segmentation model is trained using a small amount of manually labeled sample data. Based on the segmentation results of each frame image, a temporal feature sequence related to the geometric dimensions of the molten pool and keyhole is constructed. The temporal feature sequence includes the geometric dimensions of the molten pool and keyhole in the segmentation results of each frame image. Based on the position of the current first image frame and the preset sliding window length, the first geometric size data corresponding to the first sliding window in the time-series feature sequence is obtained, and the statistical feature value corresponding to the first geometric size data is determined. Based on the statistical feature value corresponding to the first geometric size data and the preset confidence interval parameter, the first normal fluctuation range is determined. The statistical feature value includes the mean and variance. The first sliding window is any segmentation window in the time-series feature sequence. If it is determined that the geometric dimensions of the first image frame fall within the first normal fluctuation range, the segmentation result of the first image frame is marked as a qualified pseudo-label. The semantic segmentation model is iteratively updated using the image frame segmentation results corresponding to the qualified pseudo-labels as training data.

2. The method as described in claim 1, characterized in that, The method further includes: If it is determined that the geometric dimensions of the first image frame do not fall within the first normal fluctuation range, the segmentation result of the current frame is marked as a frame to be compensated. The system searches the temporal feature sequence for a forward nearest neighbor reliable frame that is located before the frame to be compensated and belongs to a qualified pseudo-label, and for a backward nearest neighbor reliable frame that is located after the frame to be compensated and belongs to a qualified pseudo-label. If both the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame exist simultaneously, calculate the first rate of change of the molten pool area and the second rate of change of the keyhole area between the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame; If both the first rate of change and the second rate of change are less than the corresponding set threshold, it is determined to be a quasi-steady state condition. Then, a pixel-level logical OR operation is performed on the segmentation results of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame to generate a compensatory pseudo-label for the frame to be compensated. If the first rate of change and the second rate of change are greater than or equal to the corresponding set threshold, it is determined to be a non-steady-state condition. Pixel-level linear interpolation is performed on the segmentation results of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame to generate a compensating pseudo-label for the frame to be compensated. The semantic segmentation model is iteratively updated using the image frame segmentation results corresponding to the qualified pseudo-labels as training data, including: The semantic segmentation model is iteratively updated based on the image frame segmentation results corresponding to the candidate pseudo-label set as training data. The candidate pseudo-label set includes qualified pseudo-labels and compensatory pseudo-labels.

3. The method as described in claim 2, characterized in that, The first rate of change and the second rate of change are calculated based on the following formula: in, , These are the times of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame, respectively; in In the case of the first rate of change, , The areas of the melt pools for the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame are respectively; in In the case of the second rate of change, , The areas of the melt pools for the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame are respectively. The pixel-level linear interpolation process uses the following formula: Where t is the time of the frame to be compensated. , The values ​​are the segmentation results at pixel position (x, y) for the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame, respectively. The value of the compensation pseudo-label is assigned to the frame to be compensated at pixel position (x,y).

4. The method as described in claim 2 or 3, characterized in that, The method further includes: If only the forward nearest neighbor reliable frame or only the backward nearest neighbor reliable frame exists, then the segmentation result of the existing frame is used as the compensatory pseudo-label of the frame to be compensated.

5. The method as described in claim 2 or 3, characterized in that, The process of using image frame segmentation results corresponding to the candidate pseudo-label set as training data to iteratively update the semantic segmentation model includes: In the candidate pseudo-label set, one frame is reserved as the target pseudo-label every interval Δ frames, where Δ satisfies the formula: f is the current camera frame rate. The target effective frame rate is set as follows: and / or, in the candidate pseudo-label set, the cross-union ratio (CUI) of the segmentation results of adjacent frames is calculated. If the CUI is greater than a set similarity threshold, the segmentation results of adjacent frames are determined to be redundant, and only one frame is retained as the target pseudo-label. The image frame segmentation results corresponding to the retained target pseudo-labels are used as training data to iteratively update the semantic segmentation model.

6. The method according to any one of claims 1-3, characterized in that, The preset sliding window length is determined based on the current image acquisition frame rate of the time-series image frame and the preset physical time span.

7. The method according to any one of claims 1-3, characterized in that, The method further includes: Obtain the segmentation mask sequence output by the semantic segmentation model after iterative optimization; Based on the segmentation results of each frame image in the segmentation mask sequence, a reference temporal feature sequence related to the reference geometry of the molten pool and keyhole is constructed. The reference temporal feature sequence includes the first-order difference and the second-order difference of the reference geometry of the molten pool and keyhole in the segmentation results of each frame image. The reference geometry includes at least one of the following: molten pool area, molten pool length, molten pool width, and molten pool aspect ratio. Based on the current frame position and the preset timing window length, timing feature data within the first timing window is extracted from the reference timing feature sequence; The timing feature data within the first timing window is input into a preset timing prediction model, and the welding quality status discrimination result corresponding to the first timing window is output; wherein, the preset timing prediction model is trained on welding process timing data with quality labels.

8. A laser welding molten pool keyhole detection system, characterized in that, The system includes: The image acquisition module is used to acquire timing image frames of the keyhole of the molten pool to be detected; The segmentation processing module is used to input the temporal image frames into the semantic segmentation model to obtain the segmentation results of the molten pool region and keyhole region in each frame image. The segmentation results include the position and contour of the molten pool region and keyhole region in the image. The initial semantic segmentation model is trained by a small amount of manually labeled sample data. The size calculation module is used to construct a temporal feature sequence related to the geometric dimensions of the molten pool and keyhole based on the segmentation results of each frame image. The temporal feature sequence includes the geometric dimensions of the molten pool and keyhole in the segmentation results of each frame image. The statistical evaluation module, based on the position of the current first image frame and the preset sliding window length, obtains the first geometric size data corresponding to the first sliding window in the time-series feature sequence, determines the statistical feature value corresponding to the first geometric size data, and determines the first normal fluctuation range based on the statistical feature value corresponding to the first geometric size data and the preset confidence interval parameter. The statistical feature value includes the mean and variance, and the first sliding window is any segmentation window in the time-series feature sequence. The statistical evaluation module is also used to mark the segmentation result of the first image frame as a qualified pseudo-label when it is determined that the geometric size of the first image frame falls within the first normal fluctuation range. The model update module is used to iteratively update the semantic segmentation model based on the image frame segmentation results corresponding to the qualified pseudo-labels as training data.

9. The system as described in claim 8, characterized in that, The system also includes: an interpolation compensation module. The interpolation compensation module is used for: If it is determined that the geometric dimensions of the first image frame do not fall within the first normal fluctuation range, the segmentation result of the current frame is marked as a frame to be compensated. The system searches the temporal feature sequence for a forward nearest neighbor reliable frame that is located before the frame to be compensated and belongs to a qualified pseudo-label, and for a backward nearest neighbor reliable frame that is located after the frame to be compensated and belongs to a qualified pseudo-label. If both the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame exist simultaneously, calculate the first rate of change of the molten pool area and the second rate of change of the keyhole area between the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame; If both the first rate of change and the second rate of change are less than the corresponding set threshold, it is determined to be a quasi-steady state condition. Then, a pixel-level logical OR operation is performed on the segmentation results of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame to generate a compensatory pseudo-label for the frame to be compensated. If the first rate of change and the second rate of change are greater than or equal to the corresponding set threshold, it is determined to be a non-steady-state condition. Pixel-level linear interpolation is performed on the segmentation results of the forward nearest neighbor reliable frame and the backward nearest neighbor reliable frame to generate a compensating pseudo-label for the frame to be compensated. The model update module is specifically used to iteratively update the semantic segmentation model based on the image frame segmentation results corresponding to the candidate pseudo-label set as training data, wherein the candidate pseudo-label set includes qualified pseudo-labels and compensating pseudo-labels.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, which, when executed, performs the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Welding pool visual monitoring method based on multi-mode semi-supervised semantic segmentation

    CN119027672A