Casting appearance defect detection method and device, electronic equipment and storage medium

By fusing and chunking the casting appearance image and design reference image and inputting it into the defect semantic segmentation model, the problem of low detection speed and accuracy in the prior art is solved, and fast and accurate defect detection is achieved.

CN120182249AActive Publication Date: 2025-06-20SHENZHEN XINRUN FULIAN DIGITAL TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510631397.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-20
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

In the prior art, the detection speed and detection accuracy of casting appearance defects are relatively low, making it difficult to meet production needs.

Method used

By acquiring the appearance image and design reference image of the casting to be detected, aligning and channel fusion are performed to obtain the fused image. The fusion image is then chunked and inputted into the pre-trained defect semantic segmentation model, output the defect semantic segmentation result, and characterize the defect type and location.

Benefits of technology

The speed and accuracy of the casting appearance defect detection are improved, and the defect type and position of the casting to be detected can be quickly and accurately determined.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182249A_ABST
    Figure CN120182249A_ABST
Patent Text Reader

Abstract

The invention relates to a casting appearance defect detection method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining an appearance image and a design reference image corresponding to a to-be-detected casting, the appearance image being a three-channel image, and the design reference image being a single-channel image; aligning the appearance image and the design reference image and carrying out channel fusion to obtain a fusion image fused with four channels; the fused image is partitioned to obtain an image block sequence, the image block sequence comprises a plurality of image blocks, and the image blocks in the plurality of image blocks have the same size and channel number; and inputting the image block sequence into a pre-trained defect semantic segmentation model, and outputting a defect semantic segmentation result of the to-be-detected casting, the defect semantic segmentation result being used for representing a defect type and a defect position of the to-be-detected casting. Therefore, the defect type and the defect position of the casting to be detected can be quickly and accurately determined, so that the detection speed and the detection precision of the appearance defect of the casting are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of casting production, and particularly to a method, device, electronic device and storage medium for detecting appearance defects of castings. Background Art

[0002] Currently, during the production process of castings (such as automobile wheels, steering knuckles, etc.), it is easily affected by many factors such as production environment, technological process, raw materials, production equipment, etc., resulting in appearance defects in the castings, such as depressions, holes, burrs, deformations, dirt, etc. Therefore, after the castings are produced, it is usually necessary to detect the appearance defects of the castings to ensure the quality of the castings.

[0003] In the related art, usually the manual detection method is adopted to detect the appearance defects of castings, so there are problems of low detection speed and detection accuracy. Therefore, how to improve the detection speed and detection accuracy of the appearance defects of castings has become an urgent technical problem to be solved. Summary of the Invention

[0004] The present application provides a method, device, electronic device and storage medium for detecting appearance defects of castings to solve the problems of low detection speed and detection accuracy of the appearance defects of castings in the related art.

[0005] In a first aspect, an embodiment of the present application provides a method for detecting appearance defects of castings, and the method includes: Obtain an appearance image and a design reference image corresponding to a casting to be detected, wherein the appearance image is a three-channel image, and the design reference image is a single-channel image; Align the appearance image and the design reference image and perform channel fusion to obtain a fused image with four channels; Divide the fused image into blocks to obtain an image block sequence, wherein the image block sequence includes a plurality of image blocks, and each image block in the plurality of image blocks has the same size and number of channels; Input the image block sequence into a pre-trained defect semantic segmentation model, and output a defect semantic segmentation result of the casting to be detected, wherein the defect semantic segmentation result is used to characterize the defect type and defect position of the casting to be detected, and the defect semantic segmentation model is used to obtain a potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and determine the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image.

[0006] Optionally, the defect semantic segmentation model includes an embedding module, a multi-scale convolution module and a feature extraction module; Inputting the sequence of image patches into a pre-trained defect semantic segmentation model to output the defect semantic segmentation result of the casting to be detected, including: Inputting the sequence of image patches into the embedding module to output a first vector sequence corresponding to the sequence of image patches, where the first vector sequence is obtained by splicing the embedding vectors and offset vectors of each image patch in the sequence of image patches; Inputting the sequence of image patches into the multi-scale convolution module to output a second vector sequence corresponding to the sequence of image patches, where the second vector sequence is obtained by unfolding the feature maps of each image patch in the sequence of image patches; Splicing the first vector sequence and the second vector sequence respectively to obtain a third vector sequence; Inputting the third vector sequence into the feature extraction module to output the defect semantic segmentation result.

[0007] Optionally, the feature extraction module includes a plurality of feature extraction sub-modules connected in series in sequence, and each feature extraction sub-module includes at least a multi-head attention layer, a cross-attention layer and a multi-layer perceptron layer; Inputting the third vector sequence into the feature extraction module to output the defect semantic segmentation result, including: Using a plurality of attention mechanisms running in parallel by the multi-head attention layer to obtain a first attention distribution from the third vector sequence, and obtaining a first context vector according to the first attention analysis, where the first context vector is used to represent the potential semantic association between each image patch in the sequence of image patches; Using the cross-attention mechanism running by the cross-attention layer to obtain a second attention distribution from the third vector sequence and the first context vector, and obtaining a second context vector according to the second attention distribution, where the second context vector is used to represent the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image; Using the multi-layer perceptron layer to perceive the first context vector and the second context vector to obtain the defect semantic segmentation result.

[0008] Optionally, using the cross-attention mechanism running by the cross-attention layer to obtain a second attention distribution from the third vector sequence and the first context vector, including: Extracting the first vector sequence and the second vector sequence from the third vector sequence; Run a cross-attention mechanism with the first vector sequence as the query, the second vector sequence as the key, and the first context vector as the value, and obtain the second attention distribution from the third vector sequence and the first context vector.

[0009] Optionally, the embedding module includes an embedding layer and a position encoding layer; The inputting the image patch sequence into the embedding module and outputting the first vector sequence corresponding to the image patch sequence includes: Input the image patch sequence into the embedding layer and output an embedding vector sequence corresponding to the image patch sequence, where the embedding vector sequence includes multiple embedding vectors, and each embedding vector corresponds to an image patch in the image patch sequence; Input the image patch sequence into the position encoding layer and output an offset vector sequence corresponding to the image patch sequence, where the offset vector sequence includes multiple offset vectors, and each offset vector corresponds to an image patch in the image patch sequence; According to the arrangement order of each image patch in the image patch sequence, sequentially splice the embedding vector and the offset vector of each image patch to obtain the first vector sequence.

[0010] Optionally, the multi-scale convolution module includes a plurality of convolution layers and an unfolding layer connected in series in sequence, where the convolution kernel sizes of the convolution layers in the plurality of convolution layers are different; The inputting the image patch sequence into the multi-scale convolution module and outputting the second vector sequence corresponding to the image patch sequence includes: Input the image patch sequence into the plurality of convolution layers and output a multi-dimensional feature map corresponding to the image patch sequence, where each dimension of the multi-dimensional feature map represents the convolution result after performing multiple convolution operations on each image patch in the image patch sequence; Input the multi-dimensional feature map into the unfolding layer and output the second vector sequence, where the second vector sequence is a one-dimensional vector sequence obtained by transforming the multi-dimensional feature map.

[0011] Optionally, after inputting the image patch sequence into a pre-trained defect semantic segmentation model and outputting the defect semantic segmentation result of the casting to be detected, the method further includes: Based on the defect semantic segmentation result, calculate the actual defect index of the casting to be detected; Compare the actual defect index of the casting to be detected with the preset defect index of the casting to be detected; According to the comparison result, determine whether the casting to be detected is a scrapped casting; In the case where it is determined that the casting to be detected is a scrap casting, control the production line track to transfer the casting to be detected to the scrap casting storage area.

[0012] In a second aspect, an embodiment of the present application further provides a casting appearance defect detection device, and the device includes: An acquisition module, configured to acquire an appearance image and a design reference image corresponding to a casting to be detected, where the appearance image is a three-channel image, and the design reference image is a single-channel image; A fusion module, configured to align and perform channel fusion on the appearance image and the design reference image to obtain a fused image with four channels; A block division module, configured to divide the fused image into blocks to obtain an image block sequence, where the image block sequence includes a plurality of image blocks, and each image block in the plurality of image blocks has the same size and number of channels; An output module, configured to input the image block sequence into a pre-trained defect semantic segmentation model, and output a defect semantic segmentation result of the casting to be detected, where the defect semantic segmentation result is used to characterize the defect type and defect position of the casting to be detected, and the defect semantic segmentation model is used to obtain a potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and determine the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image.

[0013] In a third aspect, an embodiment of the present application further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store a computer program; The processor is configured to implement the casting appearance defect detection method described in the first aspect when executing the program stored in the memory.

[0014] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and the computer program implements the casting appearance defect detection method described in the first aspect when executed by a processor.

[0015] The above technical solution provided by the embodiments of the present application has the following advantages compared with the prior art: In the method provided by the embodiments of the present application, by obtaining an appearance image and a design reference image corresponding to a casting to be detected, wherein the appearance image is a three-channel image and the design reference image is a single-channel image; aligning and channel-fusing the appearance image and the design reference image to obtain a fused image with four channels; partitioning the fused image to obtain an image block sequence, wherein the image block sequence includes a plurality of image blocks, and each image block in the plurality of image blocks has the same size and number of channels; inputting the image block sequence into a pre-trained defect semantic segmentation model, and outputting a defect semantic segmentation result of the casting to be detected, wherein the defect semantic segmentation result is used to characterize the defect type and defect location of the casting to be detected, and the defect semantic segmentation model is used to obtain a potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and determine the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image. By the above method, a potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image can be obtained by using the defect semantic segmentation model, and further, the defect type and defect location of the casting to be detected can be quickly and accurately determined according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, thereby improving the detection speed and detection accuracy of the casting appearance defects. Description of the Drawings

[0016] The drawings here are incorporated into the specification and constitute a part of this specification, showing the embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0018] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the figures do not constitute a proportional limitation.

[0019] Figure 1 It is a schematic flowchart of a method for detecting casting appearance defects provided by the embodiments of the present application; Figure 2Schematic diagram of a defect semantic segmentation model provided by an embodiment of the present application; Figure 3 Schematic diagram of a feature extraction sub-module provided by an embodiment of the present application; Figure 4 Flow schematic diagram of another casting appearance defect detection method provided by an embodiment of the present application; Figure 5 Schematic diagram of the structure of a casting appearance defect detection device provided by an embodiment of the present application; Figure 6 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0021] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplicity and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0022] To solve the problem of low detection speed and detection accuracy of casting appearance defects in the related art, the present application provides a casting appearance defect detection method, device, electronic device, and storage medium, which can improve the detection speed and detection accuracy of casting appearance defects.

[0023] See Figure 1 , Figure 1 is a flow schematic diagram of a casting appearance defect detection method provided by an embodiment of the present application. As Figure 1 shown, the casting appearance defect detection method may include the following steps: Step S102, obtain an appearance image and a design reference image corresponding to the casting to be detected, where the appearance image is a three-channel image and the design reference image is a single-channel image.

[0024] Specifically, the above-mentioned casting to be detected can be any casting that needs to be detected for appearance defects, such as automobile wheels, steering knuckles, etc. The casting to be detected can be made of any metal material, such as aluminum alloy, iron, copper, etc. The corresponding appearance image of the casting to be detected can be a color image obtained by camera shooting, which has three-channel data of red, green, and blue. The corresponding design reference image of the casting to be detected can be a grayscale image generated by design software, which has single-channel data. It can be understood that the casting to be detected is produced based on the corresponding design reference image of the casting to be detected, and the corresponding design reference image of the casting to be detected has no defects.

[0025] Step S104: Align the appearance image and the design reference image and perform channel fusion to obtain a fused image with four channels.

[0026] In this step, after obtaining the appearance image and the design reference image corresponding to the casting to be detected, the appearance image and the design reference image corresponding to the casting to be detected can be aligned so that the pixel points at the same position on the appearance image and the design reference image all correspond to the same actual position of the casting to be detected. Then, the aligned appearance image and design reference image are subjected to channel fusion to obtain a fused image. Each pixel point position in the fused image here contains four-channel data.

[0027] Step S106: Divide the fused image into blocks to obtain an image block sequence, where the image block sequence includes a plurality of image blocks, and each of the plurality of image blocks has the same size and number of channels.

[0028] Specifically, the above-mentioned image block sequence refers to a sequence formed by arranging a plurality of image blocks, and each of the plurality of image blocks has the same size and number of channels. For example, assume that the appearance image is a 1024*1024 RGB color image, and the design reference image is a 1024*1024 grayscale image. After aligning and fusing the two, a 1024*1024 mixed image is obtained. At this time, the mixed image can be divided into image blocks of size P*P in the order from top to bottom and from left to right to obtain B = (1024 * 1024) / (P * P) image blocks, and then the B image blocks are combined into an image block sequence. The value of P here can be 8, 16, 32, etc. If there are remaining pixel points less than P when dividing the mixed image into blocks, four-channel data of (255, 255, 255, 255) can be used for filling.

[0029] Step S108: Input the image patch sequence into a pre-trained defect semantic segmentation model, and output the defect semantic segmentation result of the casting to be detected. The defect semantic segmentation result is used to characterize the defect type and defect location of the casting to be detected. The defect semantic segmentation model is used to obtain the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and determine the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image.

[0030] Specifically, the above defect semantic segmentation model can be implemented based on a modified Vision Transformer (ViT), or can be implemented based on other models such as Convolutional Neural Networks (CNN), and the embodiments of the present application do not make specific limitations. The defect semantic segmentation model can extract the feature information in the image patch sequence, thereby obtaining the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and determining the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, that is, the defect type (such as depression, hole, burr, deformation, dirt, etc.) and defect location of the casting to be detected.

[0031] In this embodiment, the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image can be obtained by using the defect semantic segmentation model, and then the defect type and defect location of the casting to be detected can be quickly and accurately determined according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, thereby improving the detection speed and detection accuracy of the casting appearance defects.

[0032] As an alternative embodiment, refer to Figure 2 , the defect semantic segmentation model includes an embedding module, a multi-scale convolution module, and a feature extraction module; The above step S108: Input the image patch sequence into a pre-trained defect semantic segmentation model, and output the defect semantic segmentation result of the casting to be detected, includes: Input the image patch sequence into the embedding module, and output the first vector sequence corresponding to the image patch sequence, where the first vector sequence is obtained by splicing the embedding vectors and offset vectors of each image patch in the image patch sequence; Input the image patch sequence into the multi-scale convolution module, and output the second vector sequence corresponding to the image patch sequence, where the second vector sequence is obtained by unfolding the feature maps of each image patch in the image patch sequence; Concatenate the first vector sequence and the second vector sequence respectively to obtain a third vector sequence; Input the third vector sequence into the feature extraction module to output the defect semantic segmentation result.

[0033] Specifically, when using the defect semantic segmentation model to detect the image patch sequence, the image patch sequence can be input into the embedding module and the multi-scale convolution module respectively, so as to use the embedding module to perform linear projection and offset calculation on the image patch sequence to obtain the first vector sequence corresponding to the image patch sequence, and use the multi-scale convolution module to perform convolution calculation and flattening on the image patch sequence to obtain the second vector sequence corresponding to the image patch sequence. Then, concatenate the first vector sequence and the second vector sequence respectively to obtain a third vector sequence, and then input the third vector sequence into the feature extraction module to output the defect semantic segmentation result. Among them, the feature extraction module can use the attention mechanism to extract features from the third vector sequence to obtain the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and can also determine the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image. It should be noted that the embodiments of the present application do not limit the structures in the embedding module, the multi-scale convolution module and the feature extraction module, as long as their corresponding functions can be realized.

[0034] In this embodiment, the defect semantic segmentation model can be used to quickly and accurately determine the defect type and defect location of the casting to be detected, thereby improving the detection speed and detection accuracy of the casting appearance defects.

[0035] In an alternative embodiment, the feature extraction module includes a plurality of feature extraction sub-modules connected in series in sequence, and each feature extraction sub-module includes at least a multi-head attention layer, a cross-attention layer and a multi-layer perceptron layer; The above step of inputting the third vector sequence into the feature extraction module to output the defect semantic segmentation result includes: Use multiple attention mechanisms running in parallel in the multi-head attention layer to obtain a first attention distribution from the third vector sequence, and obtain a first context vector according to the first attention analysis, where the first context vector is used to represent the potential semantic association between the image patches in the image patch sequence; Use the cross-attention mechanism running in the cross-attention layer to obtain a second attention distribution from the third vector sequence and the first context vector, and obtain a second context vector according to the second attention distribution, where the second context vector is used to represent the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image; The first context vector and the second context vector are perceived by a multi-layer perceptron layer to obtain a defect semantic segmentation result.

[0036] Specifically, the feature extraction module may include a plurality of feature extraction sub-modules connected in series in sequence. Among them, the number of feature extraction sub-modules can be set according to actual needs and will not be specifically limited here.

[0037] The feature extraction sub-module can at least include a multi-head attention layer, a cross-attention layer, and a multi-layer perceptron layer. Among them, the multi-head attention layer can calculate multiple attention heads in parallel. Each attention head learns the dependence patterns in different sub-spaces of the third vector sequence, enhancing the expression ability of the model. It can also split the third vector sequence into multiple attention heads and reduce the dimension of a single attention head, improving the parallel processing efficiency. Each attention head can be independently optimized, avoiding local optima through a competition mechanism and capturing diverse association patterns in the third vector sequence. It can splice the outputs of all attention heads into a high-dimensional matrix and then integrate them into the final output through a linear transformation. The cross-attention layer can capture long-distance dependence relationships by calculating the weights between different vector sequences, enhancing the data modeling ability. The multi-layer perceptron layer can recombine information layer by layer. The information recombined in each layer enters the data recombination of the next layer after being amplified or inhibited by an activation function, thereby enabling defect recognition.

[0038] When using the feature extraction module to perform feature extraction and defect recognition on the third vector sequence, multiple attention mechanisms running in parallel in the multi-head attention layer can be used to obtain a first attention distribution from the third vector sequence, and based on the first attention analysis, a first context vector can be obtained. Here, the first context vector can be used to represent the potential semantic associations between the image patches in the image patch sequence. The cross-attention mechanism running in the cross-attention layer can also be used to obtain a second attention distribution from the third vector sequence and the first context vector, and based on the second attention distribution, a second context vector can be obtained. Here, the second context vector can be used to represent the potential semantic associations between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image. Then, the first context vector and the second context vector can be perceived by the multi-layer perceptron layer to obtain a defect semantic segmentation result.

[0039] Of course, in addition to including a multi-head attention layer, a cross-attention layer, and a multi-layer perceptron layer, the feature extraction sub-module may also include a feed-forward network layer and multiple normalization layers, etc. The structure is as Figure 3As shown. This feed - forward network layer can be set between the multi - head attention layer and the cross - attention layer, and is used to perform a non - linear mapping on the features output by the multi - head attention layer to enhance the expression ability of the model. The feature sequence after non - linear transformation has the same dimension as the input. Considering the efficiency of the entire network, the weight coefficient of the feed - forward network layer can be set to 1 / 32 - 1 / 8. The feed - forward network layer strengthens the continuity of local feature flow through convolution in the longitudinal direction. Multiple normalization layers can be set on the input side and output side of the multi - head attention layer, cross - attention layer, and feed - forward network layer. It can normalize the input data to alleviate the problem of gradient disappearance or explosion.

[0040] It should be noted that the feature extraction module can be used to extract features and identify defects from the third vector sequence, quickly and accurately obtain the defect semantic segmentation result, thereby improving the detection speed and detection accuracy of casting appearance defects.

[0041] In an optional embodiment, the above - mentioned step of using the cross - attention mechanism running in the cross - attention layer to obtain the second attention distribution from the third vector sequence and the first context vector includes: Extract the first vector sequence and the second vector sequence from the third vector sequence; Use the first vector sequence as the query, the second vector sequence as the key, and the first context vector as the value to run the cross - attention mechanism, and obtain the second attention distribution from the third vector sequence and the first context vector.

[0042] Specifically, when obtaining the second attention distribution, the first vector sequence and the second vector sequence can be first extracted from the third vector sequence, and then the first vector sequence is used as the query Query (denoted as Q), the second vector sequence is used as the key Key (denoted as K), and the first context vector is used as the value Value (denoted as V) to run the cross - attention mechanism. Calculate the dot - product similarity between Q and K, and stabilize the value through a scaling factor (where is the dimension of Key) to obtain the original attention scores. Then perform Softmax normalization on the original attention scores to generate the second attention distribution. This second attention distribution can represent the contribution weights of each position of the input feature vector to the current decoding step. Using this second attention distribution, weighted summation can be performed on Value to generate the second context vector, which integrates the information most relevant to the current generation target. That is, this second context vector can be used to characterize the potential semantic association between the three - channel data corresponding to the appearance image and the single - channel data corresponding to the design reference image.

[0043] In this way, the cross - attention layer can capture long - distance dependencies and obtain the potential semantic association between the three - channel data corresponding to the appearance image and the single - channel data corresponding to the design reference image.

[0044] In an alternative embodiment, the embedding module includes an embedding layer and a positional encoding layer; The above step of inputting the image patch sequence into the embedding module and outputting the first vector sequence corresponding to the image patch sequence includes: Inputting the image patch sequence into the embedding layer to output an embedding vector sequence corresponding to the image patch sequence, where the embedding vector sequence includes a plurality of embedding vectors, and each embedding vector corresponds to an image patch in the image patch sequence; Inputting the image patch sequence into the positional encoding layer to output an offset vector sequence corresponding to the image patch sequence, where the offset vector sequence includes a plurality of offset vectors, and each offset vector corresponds to an image patch in the image patch sequence; Sequentially concatenating the embedding vectors and offset vectors of each image patch according to the arrangement order of the image patches in the image patch sequence to obtain the first vector sequence.

[0045] Specifically, when using the embedding module to obtain the vector representation of the image patch sequence, the image patch sequence can be first input into the embedding layer for linear transformation to output the embedding vector sequence corresponding to the image patch sequence. And since image partitioning is very likely to divide a defect or a key region into multiple different image patches, the image patch sequence can be input into the positional encoding layer to output the offset vector sequence corresponding to the image patch sequence. The offset vector can be a trainable offset vector, and this offset vector can be subtracted from the true position (Gx, Gy) of each image patch to form a gradient, and then the gradient is minimized by gradient descent training. After training, during inference, each image patch is translated according to the coordinates of each offset vector to recombine the separated key information.

[0046] Next, sequentially concatenating the embedding vectors and offset vectors of each image patch according to the arrangement order of the image patches in the image patch sequence to obtain the first vector sequence. For example, assuming the image patch sequence includes B image patches, each image patch can be linearly transformed into a D-dimensional embedding vector through the embedding layer, thus obtaining B D-dimensional embedding vectors. Also, the offset vector corresponding to each image patch line can be calculated through the positional encoding layer, thus obtaining B offset vectors. Then, an offset vector is concatenated in front of each D-dimensional embedding vector to obtain the first vector sequence.

[0047] Through the above method, not only can the embedding vectors of each image patch be obtained, but also the offset vectors of each image patch can be obtained, enabling the recombination of key information on different image patches and avoiding the loss of defect features.

[0048] In an alternative embodiment, the multi-scale convolution module includes a plurality of convolutional layers and an unfolding layer connected in series in sequence, wherein the convolutional kernels of the convolutional layers in the plurality of convolutional layers have different sizes; The above step of inputting the sequence of image patches into the multi-scale convolution module and outputting the corresponding second vector sequence of the sequence of image patches includes: Inputting the sequence of image patches into the plurality of convolutional layers to output a multi-dimensional feature map corresponding to the sequence of image patches, wherein each dimension of the multi-dimensional feature map represents the convolution result after performing multiple convolution operations on each image patch in the sequence of image patches; Inputting the multi-dimensional feature map into the unfolding layer to output a second vector sequence, wherein the second vector sequence is a one-dimensional vector sequence obtained by transforming the multi-dimensional feature map.

[0049] It should be noted that the sizes of the convolutional kernels of the convolutional layers in the above-mentioned plurality of convolutional layers and the number of convolutional layers can be set according to actual situations and are not limited herein.

[0050] Specifically, when using the multi-scale convolution module to obtain the vector representation of the sequence of image patches, the sequence of image patches can be input into the plurality of convolutional layers for convolution operations to extract the image features in each image patch and generate a multi-dimensional feature map. Each dimension of the multi-dimensional feature map can correspond to an image patch in the fused image. Then, the multi-dimensional feature map is input into the unfolding layer for unfolding to output a second vector sequence, which is a one-dimensional vector sequence obtained by transforming the multi-dimensional feature map.

[0051] In the above manner, by combining multi-scale feature extraction and the attention mechanism, it is possible to focus on the important regions of the image at different scales, thereby improving the performance and accuracy of the model.

[0052] In an alternative embodiment, referring to Figure 4 , after the above step S108 of inputting the sequence of image patches into the pre-trained defect semantic segmentation model and outputting the defect semantic segmentation result of the casting to be detected, the method further includes: Step S110: Calculating the actual defect index of the casting to be detected based on the defect semantic segmentation result; Step S112: Comparing the actual defect index of the casting to be detected with the preset defect index of the casting to be detected; Step S114: Determining whether the casting to be detected is a scrapped casting according to the comparison result; Step S116: When it is determined that the casting to be detected is a scrapped casting, controlling the production line track to transfer the casting to be detected to the scrapped casting storage area.

[0053] Specifically, after outputting the defect semantic segmentation result of the casting to be detected, the actual defect index of the casting to be detected can be calculated based on the defect semantic segmentation result, and then the actual defect index of the casting to be detected is compared with the preset defect index of the casting to be detected. Here, the preset defect index can be set in advance according to human experience. Then, according to the comparison result, it is determined whether the casting to be detected is a scrapped casting. If the casting to be detected is a scrapped casting, the production line track is controlled to transfer the casting to be detected to the scrapped casting storage area. In this way, it is beneficial to automatically control the production line track to classify and store the castings according to the casting appearance defect detection results, thereby improving production efficiency.

[0054] See Figure 5 , Figure 5 is a schematic structural diagram of a casting appearance defect detection device provided by an embodiment of the present application. As Figure 5 shown, the casting appearance defect detection device 500 includes: An acquisition module 502, configured to acquire an appearance image and a design reference image corresponding to the casting to be detected, where the appearance image is a three-channel image and the design reference image is a single-channel image; A fusion module 504, configured to align and perform channel fusion on the appearance image and the design reference image to obtain a fused image with four channels; A block division module 506, configured to divide the fused image into blocks to obtain an image block sequence, where the image block sequence includes a plurality of image blocks, and each image block in the plurality of image blocks has the same size and number of channels; An output module 508, configured to input the image block sequence into a pre-trained defect semantic segmentation model and output a defect semantic segmentation result of the casting to be detected, where the defect semantic segmentation result is used to characterize the defect type and defect position of the casting to be detected, and the defect semantic segmentation model is used to obtain the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and determine the defect semantic segmentation result according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image.

[0055] Further, the defect semantic segmentation model includes an embedding module, a multi-scale convolution module, and a feature extraction module; the output module 508 includes: A first output sub-module, configured to input the image block sequence into the embedding module and output a first vector sequence corresponding to the image block sequence, where the first vector sequence is obtained by splicing the embedding vectors and offset vectors of each image block in the image block sequence; A second output sub-module, configured to input the image block sequence into the multi-scale convolution module and output a second vector sequence corresponding to the image block sequence, where the second vector sequence is obtained by unfolding the feature maps of each image block in the image block sequence; The splicing sub-module is used to splice the first vector sequence and the second vector sequence respectively to obtain a third vector sequence; The third output sub-module is used to input the third vector sequence into the feature extraction module and output the defect semantic segmentation result.

[0056] Furthermore, the feature extraction module includes a plurality of feature extraction sub-modules connected in series in sequence. Each feature extraction sub-module includes at least a multi-head attention layer, a cross-attention layer, and a multi-layer perceptron layer; the third output sub-module includes: The first acquisition unit is used to use a plurality of attention mechanisms running in parallel in the multi-head attention layer to obtain a first attention distribution from the third vector sequence, and obtain a first context vector according to the first attention analysis. The first context vector is used to represent the potential semantic association between each image patch in the image patch sequence; The second acquisition unit is used to use the cross-attention mechanism running in the cross-attention layer to obtain a second attention distribution from the third vector sequence and the first context vector, and obtain a second context vector according to the second attention distribution. The second context vector is used to represent the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image; The perception unit is used to use the multi-layer perceptron layer to perceive the first context vector and the second context vector to obtain the defect semantic segmentation result.

[0057] Furthermore, the second acquisition unit is specifically used for: Extract the first vector sequence and the second vector sequence from the third vector sequence; Use the first vector sequence as the query, the second vector sequence as the key, and the first context vector as the value to run the cross-attention mechanism to obtain the second attention distribution from the third vector sequence and the first context vector.

[0058] Furthermore, the embedding module includes an embedding layer and a position encoding layer; the first output sub-module includes: The first output unit is used to input the image patch sequence into the embedding layer and output the embedding vector sequence corresponding to the image patch sequence. The embedding vector sequence includes a plurality of embedding vectors, and each embedding vector corresponds to an image patch in the image patch sequence; The second output unit is used to input the image patch sequence into the position encoding layer and output the offset vector sequence corresponding to the image patch sequence. The offset vector sequence includes a plurality of offset vectors, and each offset vector corresponds to an image patch in the image patch sequence; The splicing unit is used to splice the embedding vector and the offset vector of each image patch in sequence according to the arrangement order of each image patch in the image patch sequence to obtain the first vector sequence.

[0059] Further, the multi-scale convolution module includes a plurality of convolution layers and unfolding layers connected in series in sequence, wherein the convolution kernels of the convolution layers in the plurality of convolution layers have different sizes; the second output sub-module includes: A third output unit, configured to input the image block sequence into the plurality of convolution layers and output a multi-dimensional feature map corresponding to the image block sequence, wherein each dimension of the multi-dimensional feature map represents the convolution result after performing multiple convolution operations on each image block in the image block sequence; A fourth output unit, configured to input the multi-dimensional feature map into the unfolding layer and output a second vector sequence, wherein the second vector sequence is a one-dimensional vector sequence obtained by transforming the multi-dimensional feature map.

[0060] Further, the casting appearance defect detection device 500 further includes: A calculation module, configured to calculate the actual defect index of the casting to be detected based on the defect semantic segmentation result; A comparison module, configured to compare the actual defect index of the casting to be detected with the preset defect index of the casting to be detected; A determination module, configured to determine whether the casting to be detected is a scrapped casting according to the comparison result; A control module, configured to control the production line track to transfer the casting to be detected to the scrapped casting storage area when it is determined that the casting to be detected is a scrapped casting.

[0061] It should be noted that the casting appearance defect detection device 500 provided in the embodiments of the present application can implement the casting appearance defect detection method provided in any one of the foregoing method embodiments and achieve the same technical effects, which will not be elaborated herein.

[0062] As Figure 6 shown, the embodiments of the present application further provide an electronic device, including a processor 611, a communication interface 612, a memory 613, and a communication bus 614. Among them, the processor 611, the communication interface 612, and the memory 613 communicate with each other through the communication bus 614. The memory 613 is used to store a computer program; In an embodiment of the present application, when the processor 611 executes the program stored on the memory 613, it implements the casting appearance defect detection method provided in any one of the foregoing method embodiments.

[0063] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the casting appearance defect detection method provided in any one of the foregoing method embodiments.

[0064] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0065] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0066] It should be understood that the terms used herein are only for the purpose of describing specific example embodiments and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or their combinations. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be executed in the specific order described or illustrated, unless the execution order is explicitly stated. It should also be understood that additional or alternative steps can be used.

[0067] The above description is only the specific implementation manners of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for detecting appearance defects of castings, characterized in that: The method comprises: Acquire an appearance image and a design reference image corresponding to the casting to be inspected, wherein the appearance image is a three-channel image and the design reference image is a single-channel image; Aligning the appearance image and the design reference image and performing channel fusion to obtain a fused image with four channels; The fused image is divided into blocks to obtain an image block sequence, wherein the image block sequence includes multiple image blocks, and each image block in the multiple image blocks has the same size and number of channels; the image block sequence is input into a pre-trained defect semantic segmentation model, and a defect semantic segmentation result of the casting to be inspected is output, wherein the defect semantic segmentation result is used to characterize the defect type and defect position of the casting to be inspected, and the defect semantic segmentation model is used to obtain the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and the defect semantic segmentation result is determined according to the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image.

2. The casting appearance defect detection method according to claim 1, characterized in that: The defect semantic segmentation model includes an embedding module, a multi-scale convolution module and a feature extraction module; The step of inputting the image block sequence into a pre-trained defect semantic segmentation model and outputting the defect semantic segmentation result of the casting to be inspected comprises: Input the image block sequence into the embedding module, and output a first vector sequence corresponding to the image block sequence, wherein the first vector sequence is obtained by concatenating the embedding vectors and offset vectors of each image block in the image block sequence; Inputting the image block sequence into the multi-scale convolution module, and outputting a second vector sequence corresponding to the image block sequence, wherein the second vector sequence is obtained by expanding the feature graphs of each image block in the image block sequence; respectively concatenating the first vector sequence and the second vector sequence to obtain a third vector sequence; The third vector sequence is input into the feature extraction module, and the defect semantic segmentation result is output.

3. The casting appearance defect detection method according to claim 2, characterized in that: The feature extraction module includes a plurality of feature extraction submodules connected in series, each of which includes at least a multi-head attention layer, a cross attention layer and a multi-layer perceptron layer; The step of inputting the third vector sequence into the feature extraction module and outputting the defect semantic segmentation result includes: Utilizing the multiple attention mechanisms of the multi-head attention layer running in parallel, obtaining a first attention distribution from the third vector sequence, and obtaining a first context vector according to the first attention analysis, wherein the first context vector is used to represent the potential semantic association between the image blocks in the image block sequence; Obtaining a second attention distribution from the third vector sequence and the first context vector using the cross attention mechanism operated by the cross attention layer, and obtaining a second context vector according to the second attention distribution, wherein the second context vector is used to represent the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image; The first context vector and the second context vector are perceived by using the multi-layer perceptron layer to obtain the defect semantic segmentation result.

4. The casting appearance defect detection method according to claim 3, characterized in that: The criss-cross attention mechanism operated by the criss-cross attention layer, obtaining a second attention distribution from the third vector sequence and the first context vector, comprises: extracting the first vector sequence and the second vector sequence from the third vector sequence; The first vector sequence is used as a query, the second vector sequence is used as a key, and the first context vector is used as a value to run a cross attention mechanism, and the second attention distribution is obtained from the third vector sequence and the first context vector.

5. The casting appearance defect detection method according to claim 2, characterized in that: The embedding module includes an embedding layer and a position encoding layer; The step of inputting the image block sequence into the embedding module and outputting a first vector sequence corresponding to the image block sequence comprises: Inputting the image block sequence into the embedding layer, and outputting an embedding vector sequence corresponding to the image block sequence, wherein the embedding vector sequence includes a plurality of embedding vectors, each of which corresponds to an image block in the image block sequence; Inputting the image block sequence into the position coding layer, and outputting an offset vector sequence corresponding to the image block sequence, wherein the offset vector sequence includes a plurality of offset vectors, each of which corresponds to an image block in the image block sequence; According to the arrangement order of each image block in the image block sequence, the embedding vector and the offset vector of each image block are concatenated in sequence to obtain the first vector sequence.

6. The casting appearance defect detection method according to claim 2, characterized in that: The multi-scale convolution module includes a plurality of convolution layers and expansion layers connected in series, wherein the convolution kernel size of each convolution layer in the plurality of convolution layers is different; The step of inputting the image block sequence into the multi-scale convolution module and outputting a second vector sequence corresponding to the image block sequence comprises: Inputting the image block sequence into the multiple convolution layers, and outputting a multi-dimensional feature map corresponding to the image block sequence, wherein each dimensional feature map in the multi-dimensional feature map represents a convolution result after performing multiple convolution operations on each image block in the image block sequence; The multidimensional feature map is input into the expansion layer, and the second vector sequence is output, wherein the second vector sequence is a one-dimensional vector sequence obtained by converting the multidimensional feature map.

7. The casting appearance defect detection method according to claim 1, characterized in that: After inputting the image block sequence into a pre-trained defect semantic segmentation model and outputting the defect semantic segmentation result of the casting to be inspected, the method further includes: Based on the defect semantic segmentation result, calculating the actual defect index of the casting to be inspected; Comparing the actual defect index of the casting to be detected with the preset defect index of the casting to be detected; According to the comparison result, determining whether the casting to be detected is a scrap casting; When it is determined that the casting to be inspected is a scrapped casting, the production line track is controlled to transfer the casting to be inspected to a scrapped casting storage area.

8. A casting appearance defect detection device, characterized in that: The device comprises: An acquisition module, used to acquire an appearance image and a design reference image corresponding to the casting to be inspected, wherein the appearance image is a three-channel image and the design reference image is a single-channel image; A fusion module, used for aligning the appearance image and the design reference image and performing channel fusion to obtain a fused image with four channels; A block division module, used for dividing the fused image into blocks to obtain an image block sequence, wherein the image block sequence includes a plurality of image blocks, and each image block in the plurality of image blocks has the same size and number of channels; An output module is used to input the image block sequence into a pre-trained defect semantic segmentation model, and output the defect semantic segmentation result of the casting to be inspected, wherein the defect semantic segmentation result is used to characterize the defect type and defect position of the casting to be inspected, and the defect semantic segmentation model is used to obtain the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image, and determine the defect semantic segmentation result based on the potential semantic association between the three-channel data corresponding to the appearance image and the single-channel data corresponding to the design reference image.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; The processor is used to implement the casting appearance defect detection method described in any one of claims 1 to 7 when executing the program stored in the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for detecting appearance defects of castings according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Express outer package defect detection method and device based on deep learning

    CN111862092A

  • Transform-based defect detection method and electronic equipment

    CN114359283A

  • Surface defect detection method and system based on external semantics and high-frequency information

    CN118396976A

  • Chip image defect segmentation method based on improved SegFormer

    CN119904636A

  • Method for detecting abnormal defect on steel surface based on semi-supervised contrastive learning

    US20240210329A1