Casting Appearance Defect Detection Method, Device, Electronic Device and Storage Medium

Through the integration and blocking processing of the casting appearance image and design reference image, combined with the defect semantic segmentation model, the problem of low detection speed and accuracy of casting appearance defects is solved, and the rapid, accurate detection and automated classification of casting defects are achieved.

CN120182249BActive Publication Date: 2025-07-29SHENZHEN XINRUN FULIAN DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510631397.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-29
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

In the prior art, the detection speed and accuracy of casting appearance defects are relatively low, making it difficult to meet the demand for efficient production.

Method used

The fusion and blocking processing of three-channel appearance images and single-channel design reference images are adopted, combined with the defect semantic segmentation model, and the potential semantic correlation between the appearance images and the design reference images is obtained through the embedding module, the multi-scale convolution module and the feature extraction module, and the defect type and location of the casting are determined.

Benefits of technology

It realizes rapid and accurate detection of casting appearance defects, improves detection speed and accuracy, supports automatic casting classification storage, and improves production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182249B_ABST
    Figure CN120182249B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, electronic device, and storage medium for detecting appearance defects of castings. The method includes: obtaining an appearance image and a design reference image corresponding to a casting to be detected, where the appearance image is a three-channel image and the design reference image is a single-channel image; aligning and performing channel fusion on the appearance image and the design reference image to obtain a fused image with four channels; partitioning the fused image to obtain an image block sequence, where the image block sequence includes a plurality of image blocks, and each of the plurality of image blocks has the same size and number of channels; inputting the image block sequence into a pre-trained defect semantic segmentation model, and outputting a defect semantic segmentation result of the casting to be detected, where the defect semantic segmentation result is used to characterize the defect type and defect location of the casting to be detected. In this way, the defect type and defect location of the casting to be detected can be determined quickly and accurately, thereby improving the detection speed and detection accuracy of the appearance defects of the casting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of casting production, and in particular, to a method, device, electronic device, and storage medium for detecting appearance defects of castings. Background Art

[0002] Currently, in the production process of castings (such as automobile wheels, steering knuckles, etc.), it is easily affected by many factors such as production environment, process flow, raw materials, and production equipment, resulting in appearance defects in castings, such as depressions, holes, burrs, deformations, dirt, etc. Therefore, after the production of castings, it is usually necessary to detect the appearance defects of castings to ensure the quality of castings.

[0003] In the related art, the manual inspection method is usually used to detect the appearance defects of castings, so there are problems of low detection speed and low detection accuracy. Therefore, how to improve the detection speed and detection accuracy of the appearance defects of castings has become an urgent technical problem to be solved. Summary of the Invention

[0004] The present application provides a method, device, electronic device, and storage medium for detecting appearance defects of castings to solve the problems of low detection speed and low detection accuracy of appearance defects of castings in the related art.

[0005] In a first aspect, an embodiment of the present application provides a method for detecting appearance defects of castings, the method including:

[0006] Obtain an appearance image and a design reference image corresponding to the casting to be detected, where the appearance image is a three-channel image and the design reference image is a single-channel image;

[0007] Align and perform channel fusion on the appearance image and the design reference image to obtain a fused image with four channels;

[0008] Divide the fused image into blocks to obtain an image block sequence, where the image block sequence includes a plurality of image blocks, and each image block in the plurality of image blocks has the same size and number of channels;

[0009] Input the image block sequence into a pre-trained defect semantic segmentation model, and output a defect semantic segmentation result of the casting to be detected, where the defect semantic segmentation result is used to characterize the defect type and defect position of the casting to be detected, and the defect semantic segmentation model is used to obtain the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and determine the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image.

[0010] Optionally, the defect semantic segmentation model includes an embedding module, a multi-scale convolution module, and a feature extraction module;

[0011] The step of inputting the image patch sequence into a pre-trained defect semantic segmentation model to output the defect semantic segmentation result of the casting to be detected includes:

[0012] Input the image patch sequence into the embedding module to output a first vector sequence corresponding to the image patch sequence, where the first vector sequence is obtained by concatenating the embedding vectors and offset vectors of each image patch in the image patch sequence;

[0013] Input the image patch sequence into the multi-scale convolution module to output a second vector sequence corresponding to the image patch sequence, where the second vector sequence is obtained by unfolding the feature maps of each image patch in the image patch sequence;

[0014] Concatenate the first vector sequence and the second vector sequence respectively to obtain a third vector sequence;

[0015] Input the third vector sequence into the feature extraction module to output the defect semantic segmentation result.

[0016] Optionally, the feature extraction module includes a plurality of feature extraction sub-modules connected in series in sequence, and each feature extraction sub-module includes at least a multi-head attention layer, a cross-attention layer, and a multi-layer perceptron layer;

[0017] The step of inputting the third vector sequence into the feature extraction module to output the defect semantic segmentation result includes:

[0018] Use multiple attention mechanisms running in parallel in the multi-head attention layer to obtain a first attention distribution from the third vector sequence, and obtain a first context vector according to the first attention analysis, where the first context vector is used to represent the potential semantic association between each image patch in the image patch sequence;

[0019] Use the cross-attention mechanism running in the cross-attention layer to obtain a second attention distribution from the third vector sequence and the first context vector, and obtain a second context vector according to the second attention distribution, where the second context vector is used to represent the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image;

[0020] Use the multi-layer perceptron layer to perceive the first context vector and the second context vector to obtain the defect semantic segmentation result.

[0021] Optionally, the cross-attention mechanism operating using the cross-attention layer obtains a second attention distribution from the third vector sequence and the first context vector, including:

[0022] Extract the first vector sequence and the second vector sequence from the third vector sequence;

[0023] Use the first vector sequence as the query, the second vector sequence as the key, and the first context vector as the value to operate the cross-attention mechanism, and obtain the second attention distribution from the third vector sequence and the first context vector.

[0024] Optionally, the embedding module includes an embedding layer and a position encoding layer;

[0025] Inputting the image patch sequence into the embedding module and outputting a first vector sequence corresponding to the image patch sequence includes:

[0026] Input the image patch sequence into the embedding layer and output an embedding vector sequence corresponding to the image patch sequence, where the embedding vector sequence includes a plurality of embedding vectors, and each embedding vector corresponds to an image patch in the image patch sequence;

[0027] Input the image patch sequence into the position encoding layer and output an offset vector sequence corresponding to the image patch sequence, where the offset vector sequence includes a plurality of offset vectors, and each offset vector corresponds to an image patch in the image patch sequence;

[0028] According to the arrangement order of each image patch in the image patch sequence, splice the embedding vector and the offset vector of each image patch in sequence to obtain the first vector sequence.

[0029] Optionally, the multi-scale convolution module includes a plurality of convolution layers and an unfolding layer connected in series in sequence, where the convolution kernel sizes of the convolution layers in the plurality of convolution layers are different;

[0030] Inputting the image patch sequence into the multi-scale convolution module and outputting a second vector sequence corresponding to the image patch sequence includes:

[0031] Input the image patch sequence into the plurality of convolution layers and output a multi-dimensional feature map corresponding to the image patch sequence, where each dimension of the multi-dimensional feature map represents the convolution result of performing multiple convolution operations on each image patch in the image patch sequence;

[0032] Input the multi-dimensional feature map into the unfolding layer to output the second vector sequence, where the second vector sequence is a one-dimensional vector sequence obtained by transforming the multi-dimensional feature map.

[0033] Optionally, after inputting the image patch sequence into a pre-trained defect semantic segmentation model to output the defect semantic segmentation result of the casting to be detected, the method further includes:

[0034] Based on the defect semantic segmentation result, calculate the actual defect index of the casting to be detected;

[0035] Compare the actual defect index of the casting to be detected with the preset defect index of the casting to be detected;

[0036] According to the comparison result, determine whether the casting to be detected is a scrap casting;

[0037] When it is determined that the casting to be detected is a scrap casting, control the production line track to transfer the casting to be detected to the scrap casting storage area.

[0038] In a second aspect, an embodiment of the present application further provides a casting appearance defect detection device, and the device includes:

[0039] An acquisition module, configured to acquire an appearance image and a design reference image corresponding to a casting to be detected, where the appearance image is a three-channel image and the design reference image is a single-channel image;

[0040] A fusion module, configured to align and perform channel fusion on the appearance image and the design reference image to obtain a fused image with four channels;

[0041] A blocking module, configured to block the fused image to obtain an image patch sequence, where the image patch sequence includes a plurality of image patches, and each image patch in the plurality of image patches has the same size and number of channels;

[0042] An output module, configured to input the image patch sequence into a pre-trained defect semantic segmentation model to output the defect semantic segmentation result of the casting to be detected, where the defect semantic segmentation result is used to characterize the defect type and defect position of the casting to be detected, and the defect semantic segmentation model is used to obtain the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and determine the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image.

[0043] In a third aspect, an embodiment of the present application further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0044] The memory is used to store a computer program;

[0045] The processor is used to implement the casting appearance defect detection method described in the first aspect when executing the program stored on the memory.

[0046] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the casting appearance defect detection method described in the first aspect.

[0047] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art: In the method provided by the embodiments of the present application, by obtaining an appearance image and a design reference image corresponding to a casting to be detected, where the appearance image is a three-channel image and the design reference image is a single-channel image; aligning the appearance image and the design reference image and performing channel fusion to obtain a fused image with four channels; dividing the fused image into blocks to obtain an image block sequence, where the image block sequence includes a plurality of image blocks, and each image block in the plurality of image blocks has the same size and number of channels; inputting the image block sequence into a pre-trained defect semantic segmentation model, and outputting a defect semantic segmentation result of the casting to be detected, where the defect semantic segmentation result is used to characterize the defect type and defect location of the casting to be detected, and the defect semantic segmentation model is used to obtain the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and determine the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image. Through the above method, the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image can be obtained by using the defect semantic segmentation model, and then the defect type and defect location of the casting to be detected can be quickly and accurately determined according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, thereby improving the detection speed and detection accuracy of the casting appearance defect. Description of the Drawings

[0048] The drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with the present application, and are used together with the description to explain the principles of the present application.

[0049] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0050] One or more embodiments are exemplarily illustrated by the pictures in the corresponding accompanying drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the drawings in the figures do not constitute a proportional limitation.

[0051] Figure 1 It is a schematic flowchart of a method for detecting the appearance defects of castings provided by an embodiment of the present application;

[0052] Figure 2 It is a schematic diagram of a defect semantic segmentation model provided by an embodiment of the present application;

[0053] Figure 3 It is a schematic diagram of a feature extraction sub-module provided by an embodiment of the present application;

[0054] Figure 4 It is a schematic flowchart of another method for detecting the appearance defects of castings provided by an embodiment of the present application;

[0055] Figure 5 It is a schematic structural diagram of a device for detecting the appearance defects of castings provided by an embodiment of the present application;

[0056] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0058] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0059] To solve the problem of low detection speed and detection accuracy of casting appearance defects in related technologies, the present application provides a casting appearance defect detection method, device, electronic device and storage medium, which can improve the detection speed and detection accuracy of casting appearance defects.

[0060] See Figure 1 , Figure 1 which is a schematic flowchart of a casting appearance defect detection method provided by an embodiment of the present application. As Figure 1 shown, the casting appearance defect detection method may include the following steps:

[0061] Step S102, obtain an appearance image and a design reference image corresponding to the casting to be detected, where the appearance image is a three-channel image and the design reference image is a single-channel image.

[0062] Specifically, the casting to be detected may be any casting that needs to be detected for appearance defects, such as automobile wheels, steering knuckles, etc. The casting to be detected may be made of any metal material, such as aluminum alloy, iron, copper, etc. The appearance image corresponding to the casting to be detected may be a color image obtained by camera shooting, which has red, green, and blue three-channel data. The design reference image corresponding to the casting to be detected may be a grayscale image generated by design software, which has single-channel data. It can be understood that the casting to be detected is produced based on the design reference image corresponding to the casting to be detected, and the design reference image corresponding to the casting to be detected has no defects.

[0063] Step S104, align the appearance image and the design reference image and perform channel fusion to obtain a fused image with four channels.

[0064] In this step, after obtaining the appearance image and the design reference image corresponding to the casting to be detected, the appearance image and the design reference image corresponding to the casting to be detected may be aligned so that the pixel points at the same position on the appearance image and the design reference image all correspond to the same actual position of the casting to be detected. Then, the aligned appearance image and design reference image are subjected to channel fusion to obtain a fused image. Each pixel point position in the fused image here contains four-channel data.

[0065] Step S106, divide the fused image into blocks to obtain an image block sequence, where the image block sequence includes a plurality of image blocks, and each image block in the plurality of image blocks has the same size and number of channels.

[0066] Specifically, the above image block sequence refers to a sequence formed by arranging multiple image blocks, and each image block in the multiple image blocks has the same size and number of channels. For example, assume that the appearance image is a 1024*1024 RGB color image and the design reference image is a 1024*1024 grayscale image. After aligning and fusing the two, a 1024*1024 mixed image is obtained. At this time, the mixed image can be sliced into image blocks of size P*P in the order from top to bottom and from left to right, obtaining B = (1024 * 1024) / (P * P) image blocks, and then combining the B image blocks into an image block sequence. The value of P here can be 8, 16, 32 and other values. If there are fewer remaining pixels than P when the mixed image is segmented, four-channel data of (255, 255, 255, 255) can be used for filling.

[0067] Step S108: Input the image block sequence into a pre-trained defect semantic segmentation model, and output the defect semantic segmentation result of the casting to be detected. Among them, the defect semantic segmentation result is used to represent the defect type and defect location of the casting to be detected. The defect semantic segmentation model is used to obtain the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and determine the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image.

[0068] Specifically, the above defect semantic segmentation model can be implemented based on a modified Vision Transformer (ViT for short), or can be implemented based on other models such as Convolutional Neural Networks (CNN for short), and the embodiments of the present application do not make specific limitations. The defect semantic segmentation model can extract the feature information in the image block sequence, thereby obtaining the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and determining the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, that is, the defect type (such as depression, hole, burr, deformation, dirt, etc.) and defect location of the casting to be detected.

[0069] In this embodiment, the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image can be obtained by using the defect semantic segmentation model, and then the defect type and defect location of the casting to be detected can be quickly and accurately determined according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, thereby improving the detection speed and detection accuracy of the casting appearance defects.

[0070] As an alternative embodiment, refer to Figure 2 , the defect semantic segmentation model includes an embedding module, a multi-scale convolution module, and a feature extraction module;

[0071] In step S108 above, input the image patch sequence into the pre-trained defect semantic segmentation model, and output the defect semantic segmentation result of the casting to be detected, including:

[0072] Input the image patch sequence into the embedding module, and output the first vector sequence corresponding to the image patch sequence, where the first vector sequence is obtained by concatenating the embedding vectors and offset vectors of each image patch in the image patch sequence;

[0073] Input the image patch sequence into the multi-scale convolution module, and output the second vector sequence corresponding to the image patch sequence, where the second vector sequence is obtained by unfolding the feature maps of each image patch in the image patch sequence;

[0074] Concatenate the first vector sequence and the second vector sequence respectively to obtain a third vector sequence;

[0075] Input the third vector sequence into the feature extraction module, and output the defect semantic segmentation result.

[0076] Specifically, when using the defect semantic segmentation model to detect the image patch sequence, the image patch sequence can be input into the embedding module and the multi-scale convolution module respectively, so as to use the embedding module to perform linear projection and offset calculation on the image patch sequence to obtain the first vector sequence corresponding to the image patch sequence, and use the multi-scale convolution module to perform convolution calculation and flattening on the image patch sequence to obtain the second vector sequence corresponding to the image patch sequence, then concatenate the first vector sequence and the second vector sequence respectively to obtain a third vector sequence, and then input the third vector sequence into the feature extraction module to output the defect semantic segmentation result. Among them, the feature extraction module can use the attention mechanism to extract features from the third vector sequence to obtain the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and can also determine the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image. It should be noted that the embodiments of the present application do not limit the structures in the embedding module, the multi-scale convolution module, and the feature extraction module, as long as their corresponding functions can be realized.

[0077] In this embodiment, the defect type and defect location of the casting to be detected can be quickly and accurately determined by using the defect semantic segmentation model, thereby improving the detection speed and detection accuracy of the casting appearance defects.

[0078] In an alternative embodiment, the feature extraction module includes a plurality of feature extraction sub-modules connected in series in sequence, and each feature extraction sub-module includes at least a multi-head attention layer, a cross-attention layer, and a multi-layer perceptron layer;

[0079] The above steps of inputting the third vector sequence into the feature extraction module and outputting the defect semantic segmentation result include:

[0080] Using a plurality of attention mechanisms running in parallel in the multi-head attention layer to obtain a first attention distribution from the third vector sequence, and obtaining a first context vector according to the first attention analysis, where the first context vector is used to represent the potential semantic association between each image block in the image block sequence;

[0081] Using the cross-attention mechanism running in the cross-attention layer to obtain a second attention distribution from the third vector sequence and the first context vector, and obtaining a second context vector according to the second attention distribution, where the second context vector is used to represent the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image;

[0082] Using the multi-layer perceptron layer to perceive the first context vector and the second context vector to obtain the defect semantic segmentation result.

[0083] Specifically, the feature extraction module may include a plurality of feature extraction sub-modules connected in series in sequence, where the number of feature extraction sub-modules can be set according to actual needs and will not be specifically limited here.

[0084] The feature extraction sub-module may at least include a multi-head attention layer, a cross-attention layer, and a multi-layer perceptron layer. Among them, the multi-head attention layer can calculate multiple attention heads in parallel, and each attention head learns the dependence pattern in different sub-spaces of the third vector sequence to enhance the expression ability of the model. It can also split the third vector sequence into multiple attention heads and reduce the dimension of a single attention head to improve the parallel processing efficiency. Each attention head can be independently optimized, avoiding local optima through a competition mechanism and capturing diverse association patterns in the third vector sequence. It can splice the outputs of all attention heads into a high-dimensional matrix and then integrate them into the final output through a linear transformation. The cross-attention layer can capture long-distance dependence relationships by calculating the weights between different vector sequences, improving the data modeling ability. The multi-layer perceptron layer can recombine information layer by layer, and the information recombined in each layer enters the data recombination of the next layer after being amplified or suppressed by the activation function, so as to realize defect recognition.

[0085] When using the feature extraction module to perform feature extraction and defect recognition on the third vector sequence, multiple attention mechanisms running in parallel in the multi-head attention layer can be used to obtain the first attention distribution from the third vector sequence, and based on the first attention analysis, the first context vector can be obtained. The first context vector here can be used to represent the potential semantic associations between the image patches in the image patch sequence. The cross-attention mechanism running in the cross-attention layer can also be used to obtain the second attention distribution from the third vector sequence and the first context vector, and based on the second attention distribution, the second context vector can be obtained. The second context vector here can be used to represent the potential semantic associations between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image. Then, the multi-layer perceptron layer can be used to perceive the first context vector and the second context vector to obtain the defect semantic segmentation result.

[0086] Of course, in addition to the multi-head attention layer, the cross-attention layer, and the multi-layer perceptron layer, the feature extraction sub-module can also include a feed-forward network layer and multiple normalization layers, etc., and the structure is as Figure 3 shown. The feed-forward network layer can be set between the multi-head attention layer and the cross-attention layer, and is used to perform non-linear mapping on the features output by the multi-head attention layer to enhance the expression ability of the model. The feature sequence after non-linear transformation has the same dimension as the input. Considering the efficiency of the entire network, the weight coefficient of the feed-forward network layer can be set to 1 / 32 to 1 / 8, and the feed-forward network layer strengthens the continuity of local feature flow through convolution in the longitudinal direction. Multiple normalization layers can be set on the input side and the output side of the multi-head attention layer, the cross-attention layer, and the feed-forward network layer, and it can normalize the input data to alleviate the problem of gradient disappearance or explosion.

[0087] It should be noted that the feature extraction module can be used to perform feature extraction and defect recognition on the third vector sequence, and quickly and accurately obtain the defect semantic segmentation result, thereby improving the detection speed and detection accuracy of casting appearance defects.

[0088] In an optional embodiment, the above step of using the cross-attention mechanism running in the cross-attention layer to obtain the second attention distribution from the third vector sequence and the first context vector includes:

[0089] Extract the first vector sequence and the second vector sequence from the third vector sequence;

[0090] Use the first vector sequence as the query, the second vector sequence as the key, and the first context vector as the value to run the cross-attention mechanism to obtain the second attention distribution from the third vector sequence and the first context vector.

[0091] Specifically, when obtaining the second attention distribution, the first vector sequence and the second vector sequence can be extracted from the third vector sequence first, and then the first vector sequence is used as the query Query (denoted as Q), the second vector sequence is used as the key Key (denoted as K), and the first context vector is used as the value Value (denoted as V) to run the cross-attention mechanism, calculate the dot product similarity between Q and K, and stabilize the value through a scaling factor (wherein, is the dimension of Key) to obtain the original attention scores. Then, the original attention scores are normalized by Softmax to generate the second attention distribution. This second attention distribution can represent the contribution weights of each position of the input feature vector to the current decoding step. Using this second attention distribution, a weighted sum of Value can be calculated to generate the second context vector, which integrates the information most relevant to the current generation target. That is, this second context vector can be used to characterize the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image.

[0092] In this way, the cross-attention layer can capture long-range dependencies to obtain the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image.

[0093] In an optional embodiment, the embedding module includes an embedding layer and a position encoding layer;

[0094] The above step of inputting the image patch sequence into the embedding module and outputting the first vector sequence corresponding to the image patch sequence includes:

[0095] Input the image patch sequence into the embedding layer to output the embedding vector sequence corresponding to the image patch sequence, where the embedding vector sequence includes multiple embedding vectors, and each embedding vector corresponds to an image patch in the image patch sequence;

[0096] Input the image patch sequence into the position encoding layer to output the offset vector sequence corresponding to the image patch sequence, where the offset vector sequence includes multiple offset vectors, and each offset vector corresponds to an image patch in the image patch sequence;

[0097] According to the arrangement order of each image patch in the image patch sequence, the embedding vector and the offset vector of each image patch are sequentially concatenated to obtain the first vector sequence.

[0098] Specifically, when using the embedding module to obtain the vector representation of the image patch sequence, the image patch sequence can be first input into the embedding layer for linear transformation to output the embedding vector sequence corresponding to the image patch sequence. And since image segmentation is very likely to divide a defect or a key area into multiple different image patches, the image patch sequence can be input into the position encoding layer to output the offset vector sequence corresponding to the image patch sequence. The offset vector can be a trainable offset vector, which can be subtracted from the true position (Gx, Gy) of each image patch to form a gradient, and then the gradient is minimized by gradient descent training. After training, during inference, each image patch is translated according to the coordinates of each offset vector to reorganize the separated key information.

[0099] Next, according to the arrangement order of each image patch in the image patch sequence, the embedding vector and the offset vector of each image patch are sequentially concatenated to obtain the first vector sequence. For example, assuming that the image patch sequence includes B image patches, each image patch can be linearly transformed into a D-dimensional embedding vector through the embedding layer, so that B D-dimensional embedding vectors can be obtained. The offset vector corresponding to each image patch line can also be calculated through the position encoding layer, so that B offset vectors can be obtained. Then, an offset vector is concatenated in front of each D-dimensional embedding vector to obtain the first vector sequence.

[0100] Through the above method, not only can the embedding vectors of each image patch be obtained, but also the offset vectors of each image patch can be obtained, so that the key information on different image patches can be reorganized to avoid loss of defect features.

[0101] In an alternative embodiment, the multi-scale convolution module includes a plurality of convolution layers and an unfolding layer connected in series in sequence, wherein the convolution kernel sizes of the convolution layers in the plurality of convolution layers are different;

[0102] The above step of inputting the image patch sequence into the multi-scale convolution module and outputting the second vector sequence corresponding to the image patch sequence includes:

[0103] Input the image patch sequence into the plurality of convolution layers to output a multi-dimensional feature map corresponding to the image patch sequence, wherein each dimension of the multi-dimensional feature map represents the convolution result after performing multiple convolution operations on each image patch in the image patch sequence;

[0104] Input the multi-dimensional feature map into the unfolding layer to output the second vector sequence, wherein the second vector sequence is a one-dimensional vector sequence obtained after transforming the multi-dimensional feature map.

[0105] It should be noted that the convolution kernel sizes of the convolution layers in the above-mentioned plurality of convolution layers and the number of convolution layers can be set according to actual situations and are not limited herein.

[0106] Specifically, when obtaining the vector representation of the image patch sequence by using the multi-scale convolution module, the image patch sequence can be input into multiple convolutional layers for convolutional operations to extract the image features in each image patch and generate a multi-dimensional feature map. Each dimension of the multi-dimensional feature map can correspond to an image patch in the fused image. Then, the multi-dimensional feature map is input into the unfolding layer for unfolding, and a second vector sequence is output. The second vector sequence is a one-dimensional vector sequence obtained after transforming the multi-dimensional feature map.

[0107] Through the above method, by combining multi-scale feature extraction and the attention mechanism, it is possible to focus on the important regions of the image at different scales, thereby improving the performance and accuracy of the model.

[0108] In an alternative embodiment, referring to Figure 4 , after the above step S108 of inputting the image patch sequence into the pre-trained defect semantic segmentation model and outputting the defect semantic segmentation result of the casting to be detected, the method further includes:

[0109] Step S110: Calculate the actual defect index of the casting to be detected based on the defect semantic segmentation result;

[0110] Step S112: Compare the actual defect index of the casting to be detected with the preset defect index of the casting to be detected;

[0111] Step S114: Determine whether the casting to be detected is a scrapped casting according to the comparison result;

[0112] Step S116: When it is determined that the casting to be detected is a scrapped casting, control the production line track to transfer the casting to be detected to the scrapped casting storage area.

[0113] Specifically, after outputting the defect semantic segmentation result of the casting to be detected, the actual defect index of the casting to be detected can be calculated based on the defect semantic segmentation result, and then the actual defect index of the casting to be detected is compared with the preset defect index of the casting to be detected. Here, the preset defect index can be set in advance according to human experience. Then, according to the comparison result, it is determined whether the casting to be detected is a scrapped casting. If the casting to be detected is a scrapped casting, the production line track is controlled to transfer the casting to be detected to the scrapped casting storage area. In this way, it is beneficial to automatically control the production line track to classify and store the castings according to the casting appearance defect detection results, thereby improving the production efficiency.

[0114] Referring to Figure 5 , Figure 5 is a schematic structural diagram of a casting appearance defect detection device provided by an embodiment of the present application. As Figure 5 shown, the casting appearance defect detection device 500 includes:

[0115] An acquisition module 502, configured to acquire an appearance image and a design reference image corresponding to a casting to be detected, wherein the appearance image is a three-channel image and the design reference image is a single-channel image;

[0116] A fusion module 504, configured to align the appearance image and the design reference image and perform channel fusion to obtain a fused image with four channels;

[0117] A block division module 506, configured to divide the fused image into blocks to obtain a sequence of image blocks, wherein the sequence of image blocks includes a plurality of image blocks, and each image block in the plurality of image blocks has the same size and number of channels;

[0118] An output module 508, configured to input the sequence of image blocks into a pre-trained defect semantic segmentation model, and output a defect semantic segmentation result of the casting to be detected, wherein the defect semantic segmentation result is used to characterize the defect type and defect location of the casting to be detected, and the defect semantic segmentation model is used to obtain the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and determine the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image.

[0119] Further, the defect semantic segmentation model includes an embedding module, a multi-scale convolution module, and a feature extraction module; the output module 508 includes:

[0120] A first output sub-module, configured to input the sequence of image blocks into the embedding module, and output a first vector sequence corresponding to the sequence of image blocks, wherein the first vector sequence is obtained by splicing the embedding vectors and offset vectors of each image block in the sequence of image blocks;

[0121] A second output sub-module, configured to input the sequence of image blocks into the multi-scale convolution module, and output a second vector sequence corresponding to the sequence of image blocks, wherein the second vector sequence is obtained by unfolding the feature maps of each image block in the sequence of image blocks;

[0122] A splicing sub-module, configured to splice the first vector sequence and the second vector sequence respectively to obtain a third vector sequence;

[0123] A third output sub-module, configured to input the third vector sequence into the feature extraction module, and output a defect semantic segmentation result.

[0124] Further, the feature extraction module includes a plurality of feature extraction sub-modules connected in series in sequence, and each feature extraction sub-module includes at least a multi-head attention layer, a cross-attention layer, and a multi-layer perceptron layer; the third output sub-module includes:

[0125] The first acquisition unit is used to obtain the first attention distribution from the third vector sequence by using multiple attention mechanisms running in parallel in the multi-head attention layer, and obtain the first context vector according to the first attention analysis, where the first context vector is used to represent the potential semantic association between the image patches in the image patch sequence;

[0126] The second acquisition unit is used to obtain the second attention distribution from the third vector sequence and the first context vector by using the cross-attention mechanism running in the cross-attention layer, and obtain the second context vector according to the second attention distribution, where the second context vector is used to represent the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image;

[0127] The perception unit is used to perceive the first context vector and the second context vector by using the multi-layer perceptron layer to obtain the defect semantic segmentation result.

[0128] Further, the second acquisition unit is specifically used for:

[0129] Extract the first vector sequence and the second vector sequence from the third vector sequence;

[0130] Use the first vector sequence as the query, the second vector sequence as the key, and the first context vector as the value to run the cross-attention mechanism to obtain the second attention distribution from the third vector sequence and the first context vector.

[0131] Further, the embedding module includes an embedding layer and a position encoding layer; the first output sub-module includes:

[0132] The first output unit is used to input the image patch sequence into the embedding layer and output the embedding vector sequence corresponding to the image patch sequence, where the embedding vector sequence includes multiple embedding vectors, and each embedding vector corresponds to an image patch in the image patch sequence;

[0133] The second output unit is used to input the image patch sequence into the position encoding layer and output the offset vector sequence corresponding to the image patch sequence, where the offset vector sequence includes multiple offset vectors, and each offset vector corresponds to an image patch in the image patch sequence;

[0134] The splicing unit is used to splice the embedding vector and the offset vector of each image patch in sequence according to the arrangement order of the image patches in the image patch sequence to obtain the first vector sequence.

[0135] Further, the multi-scale convolution module includes a plurality of convolution layers and an unfolding layer connected in series in sequence, where the convolution kernel sizes of the convolution layers in the plurality of convolution layers are different; the second output sub-module includes:

[0136] A third output unit, configured to input the image block sequence into multiple convolutional layers and output a multi-dimensional feature map corresponding to the image block sequence, where each dimension of the multi-dimensional feature map represents the convolution result after performing multiple convolution operations on each image block in the image block sequence;

[0137] A fourth output unit, configured to input the multi-dimensional feature map into an unfolding layer and output a second vector sequence, where the second vector sequence is a one-dimensional vector sequence obtained by transforming the multi-dimensional feature map.

[0138] Further, the casting appearance defect detection device 500 further includes:

[0139] A calculation module, configured to calculate the actual defect index of the casting to be detected based on the defect semantic segmentation result;

[0140] A comparison module, configured to compare the actual defect index of the casting to be detected with the preset defect index of the casting to be detected;

[0141] A determination module, configured to determine whether the casting to be detected is a scrapped casting according to the comparison result;

[0142] A control module, configured to control the production line track to transfer the casting to be detected to the scrapped casting storage area when it is determined that the casting to be detected is a scrapped casting.

[0143] It should be noted that the casting appearance defect detection device 500 provided in the embodiments of the present application can implement the casting appearance defect detection method provided in any of the foregoing method embodiments and achieve the same technical effects, which will not be elaborated herein.

[0144] As Figure 6 shown, the embodiments of the present application further provide an electronic device, including a processor 611, a communication interface 612, a memory 613, and a communication bus 614. Among them, the processor 611, the communication interface 612, and the memory 613 communicate with each other through the communication bus 614.

[0145] The memory 613 is used to store a computer program;

[0146] In an embodiment of the present application, when the processor 611 executes the program stored on the memory 613, it implements the casting appearance defect detection method provided in any of the foregoing method embodiments.

[0147] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the casting appearance defect detection method provided in any of the foregoing method embodiments.

[0148] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0149] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0150] It should be understood that the terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "include", "comprise", "contain", and "have" are inclusive and thus specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the particular order described or illustrated, unless the order of execution is explicitly stated. It should also be understood that additional or alternative steps may be used.

[0151] The above description is only the specific implementation manners of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for detecting appearance defects of castings, characterized in that, The method includes: Obtaining an appearance image and a design reference image corresponding to the casting to be detected, wherein the appearance image is a three-channel image and the design reference image is a single-channel image; Aligning and channel fusing the appearance image and the design reference image to obtain a fused image with four channels; Partitioning the fused image to obtain an image block sequence, wherein the image block sequence includes a plurality of image blocks, and each image block in the plurality of image blocks has the same size and number of channels; Inputting the image block sequence into a pre-trained defect semantic segmentation model, and outputting a defect semantic segmentation result of the casting to be detected, wherein the defect semantic segmentation result is used to characterize the defect type and defect location of the casting to be detected, and the defect semantic segmentation model is used to obtain a potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and determine the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image; Wherein, the defect semantic segmentation model includes an embedding module, a multi-scale convolution module and a feature extraction module; The step of inputting the image block sequence into a pre-trained defect semantic segmentation model and outputting a defect semantic segmentation result of the casting to be detected includes: Inputting the image block sequence into the embedding module, and outputting a first vector sequence corresponding to the image block sequence, wherein the first vector sequence is obtained by splicing the embedding vectors and offset vectors of each image block in the image block sequence; Inputting the image block sequence into the multi-scale convolution module, and outputting a second vector sequence corresponding to the image block sequence, wherein the second vector sequence is obtained by unfolding the feature maps of each image block in the image block sequence; Splicing the first vector sequence and the second vector sequence respectively to obtain a third vector sequence; Inputting the third vector sequence into the feature extraction module, and outputting the defect semantic segmentation result; Wherein, the feature extraction module includes a plurality of feature extraction sub-modules connected in series in sequence, and each feature extraction sub-module includes at least a multi-head attention layer, a cross-attention layer and a multi-layer perceptron layer; The step of inputting the third vector sequence into the feature extraction module and outputting the defect semantic segmentation result includes: Using a plurality of attention mechanisms operating in parallel in the multi-head attention layer to obtain a first attention distribution from the third vector sequence, and obtaining a first context vector according to the first attention distribution, wherein the first context vector is used to characterize the potential semantic association between the image blocks in the image block sequence; The cross-attention mechanism operating using the cross-attention layer obtains a second attention distribution from the third vector sequence and the first context vector, and based on the second attention distribution, obtains a second context vector, where the second context vector is used to characterize the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image; The multi-layer perceptron layer is used to perceive the first context vector and the second context vector to obtain the defect semantic segmentation result.

2. The method for detecting appearance defects of a casting according to claim 1, characterized in that, The cross-attention mechanism operating using the cross-attention layer obtains a second attention distribution from the third vector sequence and the first context vector, including: Extract the first vector sequence and the second vector sequence from the third vector sequence; Use the first vector sequence as the query, the second vector sequence as the key, and the first context vector as the value to run the cross-attention mechanism to obtain the second attention distribution from the third vector sequence and the first context vector.

3. The method for detecting the appearance defects of a casting according to claim 1, wherein, The embedding module includes an embedding layer and a position encoding layer; The step of inputting the image patch sequence into the embedding module and outputting the first vector sequence corresponding to the image patch sequence includes: Input the image patch sequence into the embedding layer to output an embedding vector sequence corresponding to the image patch sequence, where the embedding vector sequence includes multiple embedding vectors, and each embedding vector corresponds to an image patch in the image patch sequence; Input the image patch sequence into the position encoding layer to output an offset vector sequence corresponding to the image patch sequence, where the offset vector sequence includes multiple offset vectors, and each offset vector corresponds to an image patch in the image patch sequence; According to the arrangement order of each image patch in the image patch sequence, sequentially splice the embedding vector and the offset vector of each image patch to obtain the first vector sequence.

4. The method for detecting the appearance defects of the casting according to claim 1, wherein The multi-scale convolution module includes a plurality of convolution layers and an unfolding layer connected in series in sequence, where the convolution kernel sizes of the convolution layers in the plurality of convolution layers are different; The step of inputting the image patch sequence into the multi-scale convolution module and outputting the second vector sequence corresponding to the image patch sequence includes: Input the image patch sequence into the plurality of convolution layers to output a multi-dimensional feature map corresponding to the image patch sequence, where each dimension of the multi-dimensional feature map represents the convolution result after performing multiple convolution operations on each image patch in the image patch sequence; Input the multi-dimensional feature map into the unfolding layer to output the second vector sequence, where the second vector sequence is a one-dimensional vector sequence obtained after transforming the multi-dimensional feature map.

5. The method for detecting appearance defects of castings according to claim 1, characterized in that, After inputting the image patch sequence into the pre-trained defect semantic segmentation model and outputting the defect semantic segmentation result of the casting to be detected, the method further includes: Based on the defect semantic segmentation result, calculate the actual defect index of the casting to be detected; Compare the actual defect index of the casting to be detected with the preset defect index of the casting to be detected; Determine whether the casting to be detected is a scrap casting according to the comparison result; When it is determined that the casting to be detected is a scrap casting, control the production line track to transfer the casting to be detected to the scrap casting storage area.

6. A casting appearance defect detection device, characterized in that, The device includes: An acquisition module, configured to acquire an appearance image and a design reference image corresponding to a casting to be detected, wherein the appearance image is a three-channel image and the design reference image is a single-channel image; A fusion module, configured to align and perform channel fusion on the appearance image and the design reference image to obtain a fused image with four channels; A block division module, configured to divide the fused image into blocks to obtain an image block sequence, wherein the image block sequence includes a plurality of image blocks, and each image block in the plurality of image blocks has the same size and number of channels; An output module, configured to input the image block sequence into a pre-trained defect semantic segmentation model and output a defect semantic segmentation result of the casting to be detected, wherein the defect semantic segmentation result is used to characterize the defect type and defect location of the casting to be detected, and the defect semantic segmentation model is used to obtain a potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and determine the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image; Wherein, the defect semantic segmentation model includes an embedding module, a multi-scale convolution module and a feature extraction module; the output module includes: A first output sub-module, configured to input the image block sequence into the embedding module and output a first vector sequence corresponding to the image block sequence, wherein the first vector sequence is obtained by splicing the embedding vectors and offset vectors of each image block in the image block sequence; A second output sub-module, configured to input the image block sequence into the multi-scale convolution module and output a second vector sequence corresponding to the image block sequence, wherein the second vector sequence is obtained by unfolding the feature maps of each image block in the image block sequence; A splicing sub-module, configured to splice the first vector sequence and the second vector sequence respectively to obtain a third vector sequence; A third output sub-module, configured to input the third vector sequence into the feature extraction module and output the defect semantic segmentation result; Wherein, the feature extraction module includes a plurality of feature extraction sub-modules connected in series in sequence, and each feature extraction sub-module includes at least a multi-head attention layer, a cross-attention layer and a multi-layer perceptron layer; the third output sub-module includes: A first acquisition unit, configured to use a plurality of attention mechanisms running in parallel by the multi-head attention layer to obtain a first attention distribution from the third vector sequence, and obtain a first context vector according to the first attention distribution, wherein the first context vector is used to characterize the potential semantic association between the image blocks in the image block sequence; A second acquisition unit, configured to obtain a second attention distribution from the third vector sequence and the first context vector by using the cross-attention mechanism of the cross-attention layer, and obtain a second context vector according to the second attention distribution, where the second context vector is used to characterize the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image; A perception unit, configured to perceive the first context vector and the second context vector by using the multi-layer perceptron layer to obtain the defect semantic segmentation result.

7. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store a computer program; The processor is configured to implement the casting appearance defect detection method according to any one of claims 1-5 when executing the program stored on the memory.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the casting appearance defect detection method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Express outer package defect detection method and device based on deep learning

    CN111862092A

  • Transform-based defect detection method and electronic equipment

    CN114359283A