Ultrasonic fin brazing identification method based on improved Mask2Former network

By improving the Mask2Former network, combining feature extraction and differential evolution algorithms, the problem of hollow area identification in fin brazing is solved, and the accuracy of the hole position is accurately positioned in ultrasonic images is achieved, and the calculation accuracy and training stability of the brazing rate are improved.

CN120411087AActive Publication Date: 2025-08-01广州多浦乐电子科技股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510905514.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-01
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

In the evaluation of fin brazing quality, due to the inherent non-weldable hollow areas in the fin structure, the solder overflow and fills and plugs in the hollow areas, the contour in the ultrasonic image is not obvious, and it is difficult to accurately calculate the brazing rate. It is difficult to effectively identify and eliminate interference in the hollow areas in the prior art.

Method used

Using the improved Mask2Former network, through feature extraction, query vector selection module and noise addition module, combined with differential evolution algorithm, the void position is accurately positioned and the brazing rate is calculated, including feature extraction, multi-scale feature processing, Transformer decoder training, differential evolution algorithm optimization and template matching.

Benefits of technology

Even when the hollow contour is not visible, the hollow position can be accurately positioned, and the brazing rate calculation accuracy can be improved, which solves the problem of matching errors in the fins when the fins are tilted or offset by traditional template matching methods, and improves training stability and matching accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411087A_ABST
    Figure CN120411087A_ABST
Patent Text Reader

Abstract

The invention discloses an ultrasonic fin brazing identification method based on an improved Mask2Former network. The ultrasonic fin brazing identification method comprises the following steps of 1, segmenting a fin welding area in an ultrasonic image by adopting the improved Mask2Former network; secondly, the welding area template and the fin welding area are matched in an optimized mode through a differential evolution algorithm, the intersection-to-union ratio of the welding area template and the fin welding area is made to be minimum, and optimal matching parameters including the rotation angle of the welding area template and the horizontal and vertical coordinate translation distance are obtained; 3, the theoretical welding area template and the inherent cavity structure template are synchronously applied to the obtained optimal matching parameters for spatial transformation; and 4, based on the transformed template position, calculating the quotient of the number of pixels exceeding a defect threshold value in the welding area template and outside the cavity template and the difference value of the number of the pixels of the two templates, and obtaining the brazing rate. Under the condition that the outline of the cavity is invisible, the position of the cavity can be accurately positioned, interference is eliminated when the brazing rate is calculated, and the calculation precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of ultrasonic non-destructive testing, and specifically relates to an ultrasonic fin brazing identification method based on an improved Mask2Former network. Background Art

[0002] A fin is an extended surface structure, usually made of highly thermally conductive metals such as aluminum and copper. By setting fins on the substrate, the heat transfer area can be significantly increased, thereby improving the heat dissipation efficiency. Therefore, it is widely used in heat dissipation scenarios such as batteries, chips, and data centers. In the evaluation of fin welding quality, the brazing rate is a key indicator, which is defined as the ratio of the actual brazed area to the theoretical contact area. The industry usually requires the brazing rate to be not less than 80% to ensure good thermal conductivity and structural reliability.

[0003] However, in actual production, due to the existence of inherently non-weldable void areas in the fin structure, these areas are filled with solder overflow in the inherent void areas, resulting in unclear contours in the ultrasonic image and being easily confused with the welding area, which brings difficulties to the calculation of the brazing rate. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide an ultrasonic fin brazing identification method based on an improved Mask2Former network, which can accurately locate the void position even when the void contour is invisible in the ultrasonic image, and exclude its interference when calculating the brazing rate, thereby improving the calculation accuracy of the brazing rate.

[0005] To achieve the above purpose, the present invention provides the following technical solutions: An ultrasonic fin brazing identification method based on an improved Mask2Former network, comprising the following steps: Step 1: Use an improved Mask2Former network to segment the fin welding area in the ultrasonic image, including: 11) Extract multi-scale features from the input fin ultrasonic image through a feature extraction network; 12) The multi-scale features obtain a selection mask feature and selection category information through a query vector selection module, and are used as the query vectors of the Transformer decoder; 13) Introduce a noise addition module during the training process of the Transformer decoder. Use the noise addition module to convert the category of the labeled target into another category among all detection categories with a set probability, and encode it with a learnable word embedding vector; use the noise addition module to perform noise addition processing on the mask of the labeled target with a set probability through random transformation; 14) Use the samples with the intersection over union (IoU) of the noisy mask and the true mask greater than or equal to 0.5 as positive samples, and the samples with the IoU less than 0.5 as negative samples; (15) Input the positive samples into the Transformer decoder for denoising training to reconstruct the real mask, and obtain the labeled targets. Fit the negative samples to the empty set. Step 2: Optimize the matching fin welding area template and the fin welding area obtained in Step 1 through the differential evolution algorithm to minimize the intersection over union (IoU) between the fin welding area template and the fin welding area predicted by the network, and obtain the optimal matching parameters including the rotation angle of the fin welding area template, the horizontal translation distance, and the vertical translation distance. Step 3: Synchronously apply the optimal matching parameters obtained in Step 2 to perform spatial transformation on the theoretical welding area template and the template of the inherent void structure within the fin welding area. Step 4: Based on the transformed template positions, calculate the quotient of the number of pixels below the defect threshold within the fin welding area template and outside the void template and the difference in the number of pixels between the two templates to obtain the brazing rate.

[0006] Furthermore, in the above Step (12), the method for obtaining the selection mask feature and the selection category feature through the query vector selection module is as follows: Extract multi-scale features from the input fin ultrasonic image through the feature extraction network. Reduce the multi-scale features to a unified channel dimension through 1×1 convolution to obtain multi-scale features ; Convolve the multi-scale features to obtain the mask feature , and generate a mask feature at each position on the multi-scale features ; After flattening and concatenating the mask features , obtain the mask feature ; ; Obtain the category feature by flattening and concatenating the multi-scale features ; Output the category information from the category feature through a multi-layer perceptron and a Softmax layer ; Apply the Varifocal loss function to the category information and the mask feature ; Calculate the maximum value for each row of the category information along the category dimension respectively to obtain the maximum value vector , sort the maximum value vector in descending order, and select the original row indices corresponding to the top K maximum values to form the output index set ; Select the index set from the category information The corresponding rows constitute the selection category information , select an index set from the mask features ; the corresponding rows of which constitute the selection mask features ; use the selection category information and the mask features as part of the query vector of the Transformer decoder .

[0007] Further, in step 13), at least one of translation, scaling, rotation, elastic deformation, or Gaussian noise is applied to the mask of the labeled target to generate a noisy mask

[0008] Further, the noisy mask features output by the noise addition module and the noisy category features are concatenated to form a noisy query vector, and the selection mask features output by the query vector selection module and the selection category information are concatenated to form a selection query vector In the Transformer decoder, an attention matrix is added between the self-attention Transformer transformations. The values of the selection query vector of the attention matrix and the selection query vector region are 0, the values of the noisy query vector of the attention matrix and the noisy query vector region are 0, and the values of the noisy query vector of the attention matrix and the selection query vector region are negative infinity; so that when performing the Transformer transformation, the selection query vector only performs attention calculation with the selection query vector, the noisy query vector only performs attention calculation with the noisy query vector, and the noisy query vector does not perform attention calculation with the selection query vector

[0009] The beneficial effects of the present invention are as follows The ultrasonic fin brazing recognition method based on the improved Mask2Former network of the present invention has the following advantages (1) In the pure template matching technology without AI, when the fins are placed obliquely or deviated greatly, the traditional template matching technology will have a situation of matching errors. However, the method of the present invention using AI can ignore the influence of rotation and translation, and can achieve pixel-level detection even with large offsets and translations (2) When there are inherent non-weldable cavity regions in the fin structure, and due to the filling of solder, the contours of the cavity regions are not obvious in the ultrasonic image and are confused with the welding regions, which brings difficulties to the calculation of the brazing rate, the method of the present invention can still accurately find out the invisible non-weldable cavity regions in the ultrasonic image, so as to accurately calculate the fin brazing rate (3) The randomly initialized queries in the original Mask2Former lack spatial priors, resulting in an unstable training process and slow convergence. Therefore, the method of the present invention uses a query vector selection module to obtain the selected mask features and selected class information, and uses them as the query vectors for the position query and content query of the Transformer decoder, effectively improving the positioning accuracy and training stability. (4) To solve the problem of the instability of the bipartite graph matching based on DETR in Mask2Former at the initial stage of training, the method of the present invention further introduces a mask noise training mechanism through a noise addition module. By using the noisy mask to guide the decoder to reconstruct the real mask, the consistency between the query and the target is significantly enhanced, the convergence speed is accelerated, and the matching accuracy is improved.

[0010] In summary, the ultrasonic fin brazing identification method based on the improved Mask2Former network of the present invention can accurately locate the position of the void even when the void contour is invisible in the ultrasonic image, and exclude its interference when calculating the brazing rate, improving the calculation accuracy of the brazing rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To make the objectives, technical solutions and beneficial effects of the present invention clearer, the present invention provides the following drawings for illustration: Figure 1 It is the structural schematic diagram of the improved Mask2Former network; Figure 2 It is the schematic diagram of the query vector selection module; Figure 3 It is the schematic diagram of the theoretical welding area template; Figure 4 It is the schematic diagram of the detection result; Figure 5 It is the original image of the fin.

[0012] In the schematic diagram of the detection result, the red area is the detected defect, the light green small block is the void and is the non-detection area, and the light blue area is the welding area detected by the improved Mask2Former model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] The following further describes the present invention in conjunction with the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and implement it, but the examples given are not intended to limit the present invention.

[0014] The ultrasonic fin brazing identification method based on the improved Mask2Former network in this embodiment includes the following steps.

[0015] Step 1: Use the improved Mask2Former network to segment the fin welding area in the ultrasonic image.

[0016] Specifically, as Figure 1 shown, this embodiment mainly improves the query vector based on Mask2Former. The input fin ultrasonic image passes through the feature extraction network to obtain multi-scale features : , , Among them: , and respectively represent features of different resolutions, and respectively represent the height and width of the network input resolution.

[0017] The multi-scale features pass through the query vector selection module to obtain the selected mask feature and the selected category feature for the input of the Transformer decoder. The noise addition module is used to add noise to the labeled target, enabling the model Transformer decoder to output the true value based on the input noisy labeled target, thereby helping the model learn and reducing the learning difficulty of bipartite graph matching. The pixel decoder is used to fuse the multi-scale features output by the feature extraction network and upsample them to a higher resolution to obtain the final feature for segmentation. The Transformer decoder then performs feature query and selection based on the input query vector and the multi-scale image features, and outputs the prediction of N pairs of category-mask pairs.

[0018] The existing query vectors of Mask2Former are all learnable, initialized with 0 parameters or randomly, and there are the following problems: First, due to the random or zero initialization of the initial query vector, it is difficult for the model to align image regions in the early stage, and problems such as oscillation or gradient explosion and disappearance often occur, resulting in unstable training processes and large differences in results; Second, the initial query vector has no explicit spatial prior. The static query vector treats all pixels equally and it is difficult to quickly locate the target center. The query vector initialized with zero has no position information, and it is necessary to learn to assign an image region to each query vector. Lacking spatial information as a guide, the attention mechanism almost equally covers the entire image in the initial stage, with low efficiency.

[0019] This embodiment addresses the above-mentioned shortcomings of Mask2Former, improves the query vector generation method of Mask2Former, and directly predicts and generates the target mask and category prediction using the network based on the multi-scale features generated by the Mask2Former encoder. Then, the Varifocal loss function is used to select the top The mask and the corresponding category information are used as a rough candidate region and as the Transformer decoder The data-driven candidate regions replace the “blindly initialized” query vectors, providing accurate spatial and content information to the Transformer decoder.

[0020] Specifically, the method for segmenting the fin welding area in the ultrasonic image using the improved Mask2Former network includes the following steps.

[0021] 11) Extract multi-scale features from the input fin ultrasonic image through the feature extraction network Specifically, multi-scale features Expressed as: , , in: 、 and Represent multi-scale features Different resolution features in and Represent the height and width of the network input resolution respectively.

[0022] 12) Multi-scale features The selection mask features and selection category information are obtained through the query vector selection module and used as the query vector of the Transformer decoder.

[0023] like Figure 2 As shown, the method for obtaining the selection mask feature and the selection category feature through the query vector selection module is as follows.

[0024] Extract multi-scale features from the input fin ultrasonic image through the feature extraction network , multi-scale features After 1×1 convolution, the dimension is reduced to a unified channel dimension of 256 to obtain multi-scale features. : , , in: 、 and Representing multi-scale features Different resolution features in .

[0025] Multi-scale features Obtain the mask feature through convolution , to generate a mask feature at each position on the multi-scale feature . The mask feature is expressed as: , , where: is expressed as ; , and represent different resolution features in the mask feature .The purpose of this operation is to generate a mask feature at each position on the multi-scale feature map. The mask feature does not require high precision. At the same time, in order to balance the inference speed, the dimension of the mask feature is designed to be the same as the minimum resolution of the multi-scale feature.

[0026] After flattening and concatenating the mask feature

[0027] , the mask feature is obtained: : where: represents the number of queries.

[0028] A similar method can be used to obtain the class feature. Specifically, the multi-scale feature is flattened and concatenated to obtain the class feature ; the class feature is output through a multi-layer perceptron and a Softmax layer to obtain the class information ; where: represents the number of classes to be detected.

[0029] The number of all candidate query vectors is too large. Select the top most representative ones for the inference of the Transformer decoder. In this embodiment, 300 are selected. When selecting the most representative query vectors, the class information predicted by the network with high scores and the mask feature are associated with the high IoU between the true mask. Specifically, for the class information Filter according to the scores of category information. For each piece of category information, first find the category with the highest category score as its predicted category, and then select the 300 with the highest scores from the predicted categories of all category information as the selected and most excellent category features, and obtain the Top 300 indices. At the same time, use the Varifocal loss function so that the IoU between the mask feature corresponding to the Top 300 index of the category feature with a high category information score and the true mask is also high. A query vector with a high category score and an accurate mask can be considered a better query vector.

[0030] In this embodiment, for the category information and the mask feature apply the Varifocal loss function; for each row in the category information calculate the maximum value along the category dimension respectively to obtain the maximum value vector , sort the maximum value vector in descending order, and select the original row indices corresponding to the top K maximum values to form the output index set ; select the rows corresponding to the index set from the category information to form the selected category information , and select the rows corresponding to the index set from the mask feature to form the selected mask feature ; use the selected category information and the mask feature as part of the query vector of the Transformer decoder.

[0031] 13) Introduce a noise addition module during the training of the Transformer decoder. Use the noise addition module to convert the category of the labeled target into another category among all detection categories with a set probability, and encode it with a learnable word embedding vector; use the noise addition module to perform noise addition processing on the mask of the labeled target with a set probability through random transformation. In this embodiment, at least one of translation, scaling, rotation, elastic deformation, or Gaussian noise is applied to the mask of the labeled target to generate a noisy mask.

[0032] Mask2Former follows the DETR set prediction paradigm. During training, the predicted mask targets of a fixed number of query vectors output by the model are paired with the ground-truth mask targets through bipartite graph matching, and then the classification and mask losses are calculated for the matched results respectively. Due to the discreteness of Hungarian matching and the randomness of model training, the matching of predicted targets to ground-truth targets becomes a dynamic and unstable process. During different training epochs, the same predicted query vector usually matches different ground-truth targets, that is, its matching target switches frequently, especially in the early stage of training. This makes the optimization of the model ambiguous, resulting in difficult and unstable optimization, and ultimately the slow convergence. Therefore, in this embodiment, a target noise addition and denoising process is introduced during the training of Mask2Former. The input of some query vectors of the Transformer decoder is changed to the target after adding noise to the annotation target, and the output reconstructs the original annotation target. Let the model predict the ground-truth target based on the noisy target, so as to help the model learn and reduce the learning difficulty. The learning goal is very clear. The input obtained by adding noise to which annotation target will be responsible for predicting the corresponding annotation target, which also avoids the ambiguity phenomenon in Hungarian matching.

[0033] In this embodiment, a target noise addition and denoising process is introduced during the training of Mask2Former. The input of some query vectors of the Transformer decoder is changed to the target after adding noise to the annotation target, and the output reconstructs the original annotation target. Let the model predict the ground-truth target based on the noisy target, so as to help the model learn and reduce the learning difficulty. The learning goal is very clear. The input obtained by adding noise to which annotation target will be responsible for predicting the corresponding annotation target, which also avoids the ambiguity phenomenon in Hungarian matching.

[0034] Specifically, an annotation target has two parts of information: category and mask, and noise is added to these two parts respectively. When adding noise to the category, with a certain probability, it is converted to another category among all detection categories, and then it is encoded into 256 dimensions using a learnable word embedding vector. When adding noise to the ground-truth mask, it is randomly transformed with a certain probability, and the transformations are as follows.

[0035] (1) Translation: Translate the foreground region in a random direction of the horizontal and vertical coordinates.

[0036] (2) Scaling: Randomly scale the mask within a certain range, changing the length and width of the minimum bounding rectangle of the mask region.

[0037] (3) Rotation: Rotate the mask by a random angle.

[0038] (4)Elastic deformation: Apply the global elastic distortion transformation ElasticTransform and the grid local control elastic distortion GridElasticDeform to the masked area.

[0039] (5)Gaussian noise: Add noise to the mask to generate random jagged edges or small holes.

[0040] The above transformations are applied to the same real mask in sequence with a certain probability for noise addition.

[0041] 14) Samples with an intersection over union (IoU) greater than or equal to 0.5 between the noisy mask and the real mask are used as positive samples, and samples with an IoU less than 0.5 are used as negative samples.

[0042] Specifically, in order to both guide the model to recover accurate predictions under slight perturbations and let the model learn to reject overly deviated prediction masks, the target after noise addition is divided into positive and negative samples, so as to accelerate convergence, stabilize bipartite graph matching, and suppress repeated detections. In this embodiment, masks with an IoU greater than or equal to 0.5 between the noisy mask and the real mask are regarded as positive samples, and those less than or equal to 0.5 are regarded as negative samples. Adding noise to the category does not affect the positivity or negativity of the samples, and the positive and negative samples are only determined by the degree of noise addition to the mask. When adding noise, 100 positive and negative samples are generated in the same training batch, so the total number of noisy query vectors is 200. The positive samples are used to input into the Transformer decoder of the network to denoise and obtain the labeled target, while the negative samples let the network fit to the empty set ∅ "no target".

[0043] 15) Input the positive samples into the Transformer decoder for denoising training to reconstruct the real mask and obtain the labeled target, and fit the negative samples to the empty set.

[0044] In this embodiment, the noisy mask feature output by the noise addition module and the noisy category feature are concatenated to form a noisy query vector, and the selected mask feature output by the query vector selection module and the selected category information are concatenated to form a selected query vector.

[0045] During training, the query vector composition of the input to the Transformer decoder is set to consist of 300 selected query vectors and 200 noisy query vectors. To prevent the selected query vectors from seeing the labeled target information in the noisy query vectors during training, an additional attention matrix is added between the self-attention Transformer transformations of the decoder. By setting the values of the selected query vectors and the selected query vector region of the attention matrix to 0, setting the values of the noisy query vectors and the noisy query vector region of the attention matrix to 0, and setting the values of the noisy query vectors and the selected query vector region of the attention matrix to negative infinity, when performing the Transformer transformation, the selected query vectors only perform attention calculations with the selected query vectors, the noisy query vectors only perform attention calculations with the noisy query vectors, and the noisy query vectors do not perform attention calculations with the selected query vectors. In this way, information leakage can be prevented, and it can be prevented that the selected query vectors in the bipartite graph matching part see the labeled target information in the denoising query vectors in the denoising part, resulting in learning a shortcut solution and learning failure. In the test phase, since there is no labeled target information for inference, the input query vectors of the network only use the selected query vectors, and the noisy query vectors are only used to stabilize the training process, as Figure 1 shown.

[0046] Step 2: Optimize the matching fin welding area template with the fin welding area obtained in Step 1 by using the differential evolution algorithm to minimize the intersection over union (IoU) between the fin welding area template and the fin welding area, and obtain the optimal matching parameters including the rotation angle of the fin welding area template, the translation distance in the horizontal coordinate direction, and the translation distance in the vertical coordinate direction.

[0047] When calculating the brazing rate of the fins, first detect the welding area of the fins by the method described in Step 1, and then perform template matching between the template of the welding area (as Figure 3 shown) and the welding area detected by the method described in Step 1 to make the IoU between the two the smallest, and determine the rotation angle of the template and the translation distances in the horizontal and vertical coordinate directions at this time. When performing template matching, this embodiment applies the differential evolution algorithm for optimization. During the optimization process, the population individuals are set to 10, and each individual has 3 dimensions, which are respectively mapped to the offset in the horizontal coordinate direction, the offset in the vertical coordinate direction, and the rotation angle. The fitness function is represented by the IoU between the template of the welding area and the detected welding area. The optimal population obtained after performing the differential evolution algorithm is the optimal translation distance and rotation angle.

[0048] Step 3: Synchronously apply the optimal matching parameters obtained in Step 2 to the theoretical welding area template and the inherent cavity structure template for spatial transformation. In this way, although the inherent cavity structure and contour cannot be seen in the ultrasonic image, the position of the inherent cavity template in the ultrasonic image can be determined according to the position after rotation and translation of the inherent cavity template.

[0049] Step 4: Based on the transformed template position, calculate the quotient of the number of pixels within the fin welding area template and outside the void template that are lower than the defect threshold and the difference in the number of pixels between the two templates, obtaining a brazing rate of 89.2%, as Figure 4 and Figure 5 shown.

[0050] The above-described embodiments are merely preferred embodiments cited to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the protection scope of the present invention. The protection scope of the present invention is subject to the claims.

Claims

1. An ultrasonic fin brazing recognition method based on an improved Mask2Former network, characterized in that: It includes the following steps: Step 1: Use the improved Mask2Former network to segment the fin welding area in the ultrasonic image, including: 11) Extract multi-scale features from the input fin ultrasonic image through the feature extraction network; 12) The multi-scale features obtain the selection mask feature and the selection category information through the query vector selection module and serve as the query vectors of the Transformer decoder; 13) During the training process of the Transformer decoder, introduce a noise addition module. Use the noise addition module to convert the category of the labeled target into another category among all detection categories with a set probability and encode it with a learnable word embedding vector; use the noise addition module to add noise to the mask of the labeled target through random transformation with a set probability; 14) Use the samples with the intersection over union (IoU) between the noise-added mask and the true mask greater than or equal to 0.5 as positive samples, and the samples with the IoU less than 0.5 as negative samples; 15) Input the positive samples into the Transformer decoder for denoising training to reconstruct the true mask and obtain the labeled target, and fit the negative samples to the empty set; Step 2: Optimize the matching of the fin welding area template and the fin welding area obtained in Step 1 through the differential evolution algorithm to minimize the intersection over union between the fin welding area template and the fin welding area predicted by the network, and obtain the optimal matching parameters including the rotation angle of the fin welding area template, the horizontal translation distance, and the vertical translation distance; Step 3: Synchronously apply the optimal matching parameters obtained in Step 2 to perform spatial transformation on the theoretical welding area template and the template of the inherent cavity structure in the fin welding area; Step 4: Based on the position of the transformed template, calculate the quotient of the number of pixels lower than the defect threshold inside the fin welding area template and outside the cavity template and the difference in the number of pixels between the two templates to obtain the brazing rate.

2. The ultrasonic fin brazing recognition method based on the improved Mask2Former network according to claim 1, wherein: In the above Step 12), the method for obtaining the selection mask feature and the selection category feature through the query vector selection module is: Extract multi-scale features from the input fin ultrasonic image through the feature extraction network , and reduce the multi-scale features to a unified channel dimension through 1×1 convolution to obtain multi-scale features ; Multiscale features are convolved to obtain mask features , and a mask feature is generated at each position on the multiscale features ; after flattening and concatenating the mask features , the mask features are obtained ; The multi-scale features are used to obtain the class features through flattening and concatenation operations ; The category features Output category information through a multi-layer perceptron and a Softmax layer ; For category information and mask features apply the Varifocal loss function; for the category information calculate the maximum value for each row along the category dimension respectively to obtain the maximum value vector , sort the maximum value vector in descending order, and select the original row indices corresponding to the top K maximum values to form the output index set ; select the rows corresponding to the index set from the category information to form the selected category information , and select the rows corresponding to the index set from the mask features to form the selected mask features ; use the selected category information and the mask features as part of the query vector of the Transformer decoder.

3. The ultrasonic fin brazing recognition method based on the improved Mask2Former network according to claim 1, characterized in that: In the above Step 13), apply at least one of the perturbations of translation, scaling, rotation, elastic deformation, or Gaussian noise to the mask of the labeled target to generate a noise-added mask.

4. The ultrasonic fin brazing recognition method based on the improved Mask2Former network according to claim 1, characterized in that: Concatenate the noise-added mask feature output by the noise addition module and the noise-added category feature to form a noise-added query vector, and concatenate the selection mask feature output by the query vector selection module and the selection category information to form a selection query vector; In the above Transformer decoder, add an attention matrix between the self-attention Transformer transformations. Set the values of the selection query vector of the attention matrix and the selection query vector area to 0, set the values of the noise-added query vector of the attention matrix and the noise-added query vector area to 0, and set the values of the noise-added query vector of the attention matrix and the selection query vector area to negative infinity; so that when performing the Transformer transformation, the selection query vector only performs attention calculation with the selection query vector, the noise-added query vector only performs attention calculation with the noise-added query vector, and the noise-added query vector does not perform attention calculation with the selection query vector.

Citation Information

Patent Citations

  • Welding spot defect detection method, system and equipment

    CN110880175A

  • Pad voidage calculation method based on automatic X-ray imaging

    CN114119445A