Defect Image Generation Method, Device, and Storage Medium

By performing multi-level feature extraction and feature vector fusion on the original defect image, detailed and real defect images are generated, which solves the problem of insufficient details and authenticity in the prior art, and improves the performance of the defect detection system.

CN119477672BActive Publication Date: 2025-06-10JIHUA LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510075012.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-06-10
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The prior art ignores the details in the image when generating defective images, resulting in insufficient details and authenticity of the generated image. Especially when dealing with complex or minor defects, the generated defective images may be different from the actual situation.

Method used

By performing multi-level image feature extraction on the original defect image, feature maps at different levels are obtained, feature vectors corresponding to the feature maps are determined, and each feature vector is fused after traversing each feature map to generate a derived defect image.

Benefits of technology

It realizes the generation of defect images with high realism and rich details, solves the problem of insufficient details and authenticity in traditional methods, and improves the training efficiency and detection accuracy of defect detection systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119477672B_ABST
    Figure CN119477672B_ABST
Patent Text Reader

Abstract

The present application discloses a defective image generation method, device, and storage medium, relating to the technical field of image generation, including: extracting multi-level image features from an original defective image to obtain feature maps of the original defective image at different levels; for any one of the feature maps, determining a feature vector corresponding to the feature map; after traversing each feature map, fusing the feature vectors to obtain a fused feature map; and generating a derivative defective image corresponding to the original defective image based on the fused feature map. By extracting multi-level image features, the present application avoids the problems of insufficient authenticity, missing details, and inaccurate representation of complex or minute defects in the generated images caused by ignoring image details and feature diversity in traditional image generation methods, and realizes the generation of defective images with high realism and rich details under the condition of limited data samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of image generation, and particularly to a method, device, and storage medium for generating defective images. Background Art

[0002] With the emergence of technologies such as Generative Adversarial Network (GAN), Variational Auto-Encoder (VAE), Diffusion Models, and Transformer, the quality and diversity of image generation have reached unprecedented heights. These technologies have been widely applied to the generation of industrial defective images to solve the problem of lack of data samples in the production process, thereby improving the accuracy and stability of defective detection systems. However, despite the many advancements made by these methods, existing technologies still have some problems in actual industrial applications. For example, many existing image generation models often ignore the details in the image when generating defective images, such as the edges of defects or some small features, resulting in insufficient details and authenticity of the generated images. Especially when dealing with complex or tiny defects, the generated defective images may deviate from the actual situation.

[0003] Therefore, how to improve the generation quality of defective images to generate more detailed and realistic defective images is an urgent problem to be solved currently. Summary of the Invention

[0004] The main objective of this application is to provide a method, device, and storage medium for generating defective images, aiming to solve the technical problem of how to improve the generation quality of defective images to generate more detailed and realistic defective images.

[0005] To achieve the above objective, this application proposes a method for generating defective images, which includes:

[0006] Performing multi-level image feature extraction on the original defective image to obtain feature maps of the original defective image at different levels;

[0007] For any one of the feature maps, determining the feature vector corresponding to the feature map;

[0008] After traversing each feature map, fusing each feature vector to obtain a fused feature map;

[0009] Generating a derivative defective image corresponding to the original defective image based on the fused feature map.

[0010] In one embodiment, the step of performing multi-level image feature extraction on the original defect image to obtain feature maps of the original defect image at different levels includes:

[0011] Taking the original defect image as the feature map to be encoded, and performing multiple convolution operations on the feature map to be encoded to obtain the feature map to be pooled at the current level of the feature map to be encoded;

[0012] Performing a pooling operation on the feature map to be pooled to obtain the feature map to be encoded at the next level, and returning to execute the step of performing multiple convolution operations on the feature map to be encoded;

[0013] After traversing each level, based on the feature maps to be pooled at each level, obtain the feature maps of the original defect image at different levels.

[0014] In one embodiment, the step of obtaining the feature maps of the original defect image at different levels based on the feature maps to be pooled at each level includes;

[0015] For the feature map to be pooled at any level, perform weighted averaging on the feature sub-maps at each depth in the feature map to be pooled to obtain an average feature map;

[0016] After traversing each feature map to be pooled, use each average feature map as the feature map of the original defect image at different levels.

[0017] In one embodiment, the step of determining the feature vector corresponding to the feature map includes:

[0018] Flatten the feature map to obtain a multi-dimensional vector corresponding to the feature map, where the dimension of the multi-dimensional vector is determined based on the size of the feature map;

[0019] Perform sampling on the multi-dimensional vector at the corresponding highest dimension in each level to obtain the feature vector corresponding to the feature map.

[0020] In one embodiment, the step of performing sampling on the multi-dimensional vector at the corresponding highest dimension in each level to obtain the feature vector corresponding to the feature map includes:

[0021] Input the multi-dimensional vector into a preset fully connected layer to generate candidate vectors corresponding to the highest dimension in each level through the preset fully connected layer;

[0022] Obtain the generated random noise, and determine the feature vector corresponding to the feature map based on the candidate vector and the random noise.

[0023] In one embodiment, the defect image generation method further includes:

[0024] Construct a loss function based on the candidate vectors, and train the model parameters in the preset convolutional layer and the preset fully connected layer based on the loss function, where the preset convolutional layer is used to perform convolutional operations.

[0025] In one embodiment, the step of fusing the feature vectors to obtain a fused feature map includes:

[0026] For any one feature vector, convert the feature vector into a candidate feature map;

[0027] After traversing all the feature vectors, splice the candidate feature maps to obtain a fused feature map.

[0028] In one embodiment, the step of generating a derivative defect image corresponding to the original defect image based on the fused feature map includes:

[0029] Take the fused feature map as the feature map to be decoded, and perform multiple convolutional operations on the feature map to be decoded to obtain the feature map to be interpolated at the current level of the feature map to be decoded;

[0030] Perform a linear interpolation operation on the feature map to be interpolated to obtain the feature map to be decoded at the previous level of the feature map to be interpolated, and return to execute the step of performing multiple convolutional operations on the feature map to be decoded;

[0031] After the feature map to be decoded undergoes convolutional operations and linear interpolation operations at each level, obtain a target feature map, and generate a derivative defect image corresponding to the original defect image based on the target feature map.

[0032] In addition, to achieve the above object, the present application also proposes a defect image generation system, and the defect image generation system includes:

[0033] An encoder, configured to perform multi-level image feature extraction on an original defect image to obtain feature maps of the original defect image at different levels;

[0034] A sampler, configured to determine, for any one feature map, a feature vector corresponding to the feature map;

[0035] A feature generator, configured to fuse the feature vectors after traversing all the feature maps to obtain a fused feature map;

[0036] A decoder, configured to generate a derivative defect image corresponding to the original defect image based on the fused feature map

[0037] In addition, to achieve the above object, the present application further provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the defective image generation method as described above.

[0038] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the defective image generation method as described above.

[0039] In addition, to achieve the above object, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the defective image generation method as described above.

[0040] One or more technical solutions proposed by the present application have at least the following technical effects:

[0041] The present application first performs multi-level image feature extraction on the original defective image to obtain feature maps of the original defective image at different levels, so as to extract multi-level features in the data by introducing a hierarchical structure, realize feature extraction of the original defective image at different scales, thereby capturing local features and global structure information of the image, ensuring that when generating a defective image subsequently, the hierarchy and structure of the original image can be retained, and improving the authenticity of the generated image; for any one of the feature maps, determine the feature vector corresponding to the feature map, so as to simplify the representation of the feature map by converting the feature map into a representative feature vector, facilitate subsequent feature fusion and defective image generation, and at the same time retain key feature information; after traversing each feature map, fuse the feature vectors to obtain a fused feature map, so as to form a fused feature map containing rich information by integrating feature information at different levels, ensuring that the generated defective image has diversity and authenticity in details and structure; generate a derivative defective image corresponding to the original defective image based on the fused feature map, so as to generate a defective image with rich details and high realism through the reverse generation process from the fused feature map to the defective image, and solve the problem of insufficient details and realism in the traditional method when generating a defective image.

[0042] In summary, the present application extracts multi-level image features from the original defect image, performs feature vector conversion and splicing, and generates a defect image based on the spliced fusion feature map, avoiding the problems of insufficient authenticity, missing details, and inaccurate representation of complex or tiny defects in the generated image caused by ignoring image details and feature diversity in traditional image generation methods. It realizes the generation of defect images with high authenticity and rich details under the condition of limited data samples, enabling the defect detection system to be trained based on high-quality defect samples, thereby improving the training efficiency and detection accuracy of the defect detection system and ensuring the stability and reliability of defect recognition in industrial production. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.

[0044] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the defect image generation method of the present application;

[0046] Figure 2 It is a schematic fusion flowchart of the defect image generation method provided for Embodiment 1 of the present application;

[0047] Figure 3 It is a schematic flowchart provided for Embodiment 2 of the defect image generation method of the present application;

[0048] Figure 4 It is a schematic sampling flowchart of the defect image generation method provided for Embodiment 2 of the present application;

[0049] Figure 5 It is a schematic model framework diagram of the defect image generation method provided for Embodiment 2 of the present application;

[0050] Figure 6 It is a schematic module structure diagram of the defect image generation system according to the embodiment of the present application;

[0051] Figure 7 It is a schematic device structure diagram of the hardware operating environment involved in the defect image generation method according to the embodiment of the present application.

[0052] The implementation, functional features, and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0054] To better understand the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0055] The main solution of the embodiments of the present application is: performing multi-level image feature extraction on the original defect image to obtain feature maps of the original defect image at different levels; for any one of the feature maps, determining the feature vector corresponding to the feature map; after traversing each feature map, fusing the feature vectors to obtain a fused feature map; and generating a derivative defect image corresponding to the original defect image based on the fused feature map.

[0056] Since the prior art often ignores the details in the image when generating defect images, such as the edges of defects or some fine features, resulting in insufficient details and authenticity of the generated images. Especially when dealing with complex or tiny defects, the generated defect images may deviate from the actual situation. Therefore, how to improve the generation quality of defect images to generate more detailed and realistic defect images is an urgent problem to be solved at present.

[0057] The present application provides a solution. By performing multi-level image feature extraction, feature vector conversion and splicing on the original defect image, and generating a defect image based on the fused feature map after splicing, it avoids the problems of insufficient authenticity, missing details, and inaccurate representation of complex or tiny defects in the generated images caused by ignoring image details and feature diversity in traditional image generation methods. It realizes the generation of defect images with high realism and rich details under the condition of limited data samples, enables the defect detection system to be trained based on high-quality defect samples, thereby improving the training efficiency and detection accuracy of the defect detection system, and ensuring the stability and reliability of defect recognition in industrial production.

[0058] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of implementing the above functions. Hereinafter, taking an electronic device as an example, this embodiment and the following embodiments will be described.

[0059] Based on this, the embodiments of the present application provide a method for generating a defect image, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the method for generating a defect image of the present application.

[0060] In this embodiment, the defective image generation method includes steps S10 to S40:

[0061] Step S10, perform multi-level image feature extraction on the original defective image to obtain feature maps of the original defective image at different levels;

[0062] It should be noted that the original defective image refers to the original image of the defective image to be generated, usually a defective product image from an industrial production line. The types of defects in this defective product image can include cracks, depressions, holes, corrosion, scratches, etc. Each type of defect has different visual characteristics, and these characteristics are gradually extracted and abstracted in different levels of the deep learning network. Among them, the defect types can be divided into three categories: micro defects, macro defects, and composite defects. Micro defects are usually small in volume and limited in influence range, such as cracks, scratches, and bubbles. Their feature manifestations are local subtle changes, usually including mutations in edges, corners, and textures. Macro defects are usually more obvious and have a larger influence range, such as holes, depressions, and large-area corrosion. Compared with micro defects, macro defects involve larger-scale structural or surface changes, usually manifested as large-scale shape changes in the image. Composite defects usually involve the superposition of multiple local defects or more complex structural changes, such as large-scale surface corrosion and multi-point surface damage. Micro defects can extract simple local features through the low-level convolutional layers of the deep learning network, focusing on details such as edges and textures. Macro defects can extract more complex local features through the middle-level convolutional layers of the deep learning network, focusing on large-scale shape changes and texture distortions. Composite defects can extract global features through the high-level convolutional layers, integrating local information at multiple levels, and extracting more complex and global irregular damages. These defect types gradually improve the expression of defect features by layer-by-layer extraction and combination of features at different levels in the deep learning network, enabling the system to effectively extract various industrial defect features from small to complex.

[0063] In addition, it should be noted that multi-level image feature extraction refers to using a convolutional neural network (CNN, Convolutional Neural Networks) to process the original defective image layer by layer to extract feature information of the image at different abstraction levels. A feature map refers to the two-dimensional data output by each layer in the convolutional neural network, and this data represents the feature response of the original defective image at that layer.

[0064] It can be understood that since the original defective image has different feature representations at different levels, and traditional image generation models often ignore the details in the image when generating defective images. For example, methods such as generative adversarial networks (GANs) may experience mode collapse during training, that is, the model can only generate a limited type of defective images and ignore other possible defective types. Therefore, in step S10, by extracting the features of the image at different levels, the loss of image information caused by only using single-level features can be avoided, the mode collapse phenomenon is reduced, and thus the comprehensive features of the image are extracted to enhance the model's ability to express different types of defects.

[0065] Exemplarily, first, the original defective image is input into a pre-trained convolutional neural network, which includes multiple convolutional layers and pooling layers. Each layer performs a convolutional operation on the feature map output by the previous layer to extract more abstract and complex features. In this way, the features of the original defective image are extracted layer by layer from the first convolutional layer to the last convolutional layer of the network, obtaining a series of feature maps at different levels, and each feature map represents the feature response of the original defective image at the corresponding level.

[0066] In a feasible implementation manner, step S10 may include steps S11 to S13:

[0067] Step S11, taking the original defective image as the feature map to be encoded, and performing multiple convolutional operations on the feature map to be encoded to obtain the feature map to be pooled at the current level of the feature map to be encoded;

[0068] It should be noted that the feature map to be encoded refers to the original image or the feature map of the previous level as the input before the convolutional operation, and this feature map is the starting point of the convolutional operation at the current level; the current level refers to a specific level in the convolutional neural network, that is, the combination of the current convolutional layer and pooling layer being processed; the feature map to be pooled refers to the feature map that has undergone convolutional operations but has not yet undergone pooling operations, and this feature map contains the feature information of the current level and awaits further processing.

[0069] It can be understood that in order to extract higher-level feature representations from the original defective image, step S11 is performed. By performing multiple convolutional operations on the image to expand the receptive field of the feature map, the problem of too small receptive field of the feature map can be avoided, and thus the image is processed into a feature map that is conducive to extracting higher-level information.

[0070] Exemplarily, first, the original defect image is used as the initial feature map to be encoded. Then, this feature map is fed into the first layer of the convolutional neural network, which contains multiple convolutional kernels. Each convolutional kernel performs a sliding convolutional operation on the feature map to be encoded. Through processes such as 3x3 convolution, batch normalization, and an activation function (such as ReLU), a new feature map is generated, which is the feature map to be pooled at the current level. This process is repeated in each layer to extract more complex features.

[0071] Step S12: Perform a pooling operation on the feature map to be pooled to obtain the feature map to be encoded at the next level, and return to the step of performing multiple convolutional operations on the feature map to be encoded.

[0072] It can be understood that since it is necessary to extract features at different levels and reduce the size of the feature map to reduce the computational complexity, performing step S12 can avoid the problem of waste of computational resources caused by the overly large size of the feature map, thereby achieving the effect of feature dimensionality reduction while maintaining feature saliency.

[0073] Exemplarily, after obtaining the feature map to be pooled at the current level, a pooling operation (such as max pooling or average pooling) is performed on these feature maps. The pooling operation selects a local area on the feature map and outputs a representative value (such as the maximum value or the average value) of this area, thereby reducing the size of the feature map. In this way, the original feature map to be interpolated generates a feature map to be encoded at the next level with a smaller size and more concentrated features after pooling. After completion of pooling, the network will use this new feature map to be encoded as the input and return to the step of performing multiple convolutional operations to continue extracting features at the next level.

[0074] Step S13: After traversing each level, obtain the feature maps of the original defect image at different levels based on the feature maps to be pooled at each level.

[0075] In this embodiment, by using a multi-level CNN to perform layer-by-layer convolution and pooling operations on the original defect image, the problems of insufficient feature extraction and waste of computational resources in traditional image processing methods are avoided. Image features are extracted at multiple abstract levels, and the size of the feature map is effectively reduced, thereby achieving the effects of reducing computational complexity, maintaining feature saliency, and obtaining a comprehensive feature representation. Through this hierarchical feature extraction method, not only the richness and representativeness of image features are improved, but also a more effective feature basis is provided for subsequent image analysis and processing tasks.

[0076] Step S20: For any feature map, determine the feature vector corresponding to the feature map.

[0077] It should be noted that the feature vector refers to the result of flattening the feature map into a one-dimensional vector, and this data contains all the information of the feature map.

[0078] It can be understood that in order to uniformly process feature information at different levels, step S20 is performed, which can avoid incompatibility between feature information and thus achieve effective integration of feature information.

[0079] Exemplarily, for each feature map extracted from the convolutional neural network, through global average pooling (Global Average Pooling, GAP) or global max pooling (Global Max Pooling, GMP) operations, the two-dimensional feature map is converted into a one-dimensional feature vector. In this way, the size of each feature map is reduced to a vector of a fixed length, and this vector is the feature representation of the feature map.

[0080] In a feasible implementation manner, the step of determining the feature vector corresponding to the feature map in step S20 may include steps S21 to S22:

[0081] Step S21, flatten the feature map to obtain a multi-dimensional vector corresponding to the feature map, where the dimension of the multi-dimensional vector is determined based on the size of the feature map;

[0082] It should be noted that the multi-dimensional vector refers to converting the feature map into a one-dimensional array, where each element corresponds to a pixel value in the feature map; the flattening process is to convert a two-dimensional or three-dimensional feature map into a one-dimensional vector, and the dimension of this vector is the width of the feature map multiplied by the height (for a two-dimensional feature map) or the width multiplied by the height multiplied by the depth (for a three-dimensional feature map).

[0083] It can be understood that in order to convert the two-dimensional data structure of the feature map into a one-dimensional data structure for subsequent vector operations and processing, step S21 is performed, which can avoid the inability to apply vectorized machine learning algorithms due to data structure mismatch, and realize converting the information of the feature map into a format suitable for processing by the machine learning model, thereby improving the processing efficiency and compatibility of feature information.

[0084] Exemplarily, in the convolutional neural network, when a feature map is processed by a convolutional layer, it is converted into a one-dimensional array. For example, each pixel value of the feature map is arranged in a certain order (such as from left to right, from top to bottom) in sequence to form a one-dimensional vector. If the size of the feature map is HxW, then the dimension of the flattened multi-dimensional vector is HW. This process can be achieved through the reshape operation of the matrix, that is, reshaping the original two-dimensional matrix of HxW into a one-dimensional matrix with a dimension of HW.

[0085] Step S22: Sample the corresponding highest dimension in each level of the multi-dimensional vector to obtain the feature vector corresponding to the feature map.

[0086] It should be noted that the highest dimension refers to the level with the largest dimension (i.e., the size of the feature map) among the feature maps of multiple levels. This dimension usually corresponds to the feature map of the earliest or deepest level in the network. As the network deepens, the size of the feature map usually decreases, and the early feature maps may have higher dimensions because they retain more details.

[0087] It can be understood that since it is necessary to maintain a consistent vector dimension in feature maps of different levels for feature fusion, step S22 can avoid the difficulties in data processing caused by inconsistent feature vector dimensions, extract equal-length feature vectors from feature maps of different levels, thus facilitating feature fusion and achieving the effects of unified feature representation and improved model performance. In addition, sampling based on the corresponding highest dimension in each level can retain as many details as possible in the features, thereby improving the feature representation quality of the feature vector.

[0088] Exemplarily, after obtaining the flattened multi-dimensional vector, since feature maps of different levels may have different sizes, it is necessary to sample these vectors to maintain a consistent dimension. The specific operation can be that for each level of the feature map, multiply the multi-dimensional vector corresponding to this feature map by a pre-defined sampling matrix, and the size of this sampling matrix is the size of the highest-dimension feature map. If the dimension of the current feature map is less than the highest dimension, the dimension is increased to be consistent with the highest dimension by interpolation (such as nearest neighbor interpolation, linear interpolation, etc.) or by selecting some elements in the multi-dimensional vector. Through the above method, the multi-dimensional vectors of each feature map are sampled to the same dimension, forming the corresponding feature vector.

[0089] In this embodiment, by performing feature map flattening and high-dimension sampling, the problem of data processing difficulties caused by inconsistent feature map sizes is avoided, and the effect of converting feature maps of different levels into feature vectors with a unified dimension is achieved. Such processing not only facilitates subsequent feature fusion and model training, but also ensures the effective extraction and utilization of feature information, improving the efficiency and accuracy of the entire image processing system.

[0090] Step S30: After traversing each feature map, fuse each feature vector to obtain a fused feature map;

[0091] It should be noted that the fused feature map refers to combining all feature vectors into a comprehensive feature representation in a certain way (such as concatenation, weighted summation, etc.).

[0092] It can be understood that in order to combine feature information at different levels, step S30 is performed. By effectively fusing features at different levels, the problem that a single-level feature cannot comprehensively reflect the image content can be avoided, thereby achieving a richer and more comprehensive feature representation.

[0093] Exemplarily, all obtained feature vectors can be concatenated to form a long feature vector, which is the fused feature map and fuses feature information at different levels; or weighted summation can be performed on each feature vector, and the weights can be dynamically adjusted according to the importance of the feature map, and finally a fused feature vector is obtained, which is the fused feature map.

[0094] In a feasible implementation manner, the step of fusing each feature vector in step S30 to obtain a fused feature map may include steps S31 to S32:

[0095] Step S31, for any one feature vector, convert the feature vector into a candidate feature map;

[0096] It should be noted that the candidate feature map refers to converting the feature vector back into a two-dimensional image format for further image processing operations in the convolutional neural network. This process usually involves reshaping the one-dimensional feature vector into a feature map with a specific height and width.

[0097] It can be understood that in order to convert the feature vector back into the format of the feature map to utilize the image processing ability of the convolutional layer in the subsequent decoder, step S31 is performed, which can avoid the problem of loss of image structure information that may occur when the subsequent decoder directly processes features in the vector space, and realizes remapping the feature vector back to the image space to retain the local structure and spatial relationship of the image.

[0098] Exemplarily, first, obtain a feature vector, which is a one-dimensional array flattened from the feature map. Then, according to the size information of the original feature map, reshape this one-dimensional array into a two-dimensional candidate feature map. When the dimension of the feature vector is in the case of, convert the feature vector into a -sized candidate feature map, and this process can be implemented by the reshape function in the programming language.

[0099] Step S32, after traversing each feature vector, concatenate each candidate feature map to obtain a fused feature map.

[0100] It can be understood that since it is necessary to integrate feature information at different levels to obtain a more abundant feature representation, step S32 is performed, which can avoid the problem that the feature information may be incomplete or insufficient to represent complex image content when using the features of a single layer alone, and realizes the integration of multi-level feature information to generate a fused feature map containing more comprehensive and abundant features, thereby improving the accuracy and robustness of image analysis and recognition.

[0101] Exemplarily, after all feature vectors are converted into candidate feature maps, these candidate feature maps are concatenated along the channel dimension (C dimension). If the size of each candidate feature map is HxWxC_i, where C_i is the number of channels of the i-th candidate feature map, then the concatenation operation combines the channels of all candidate feature maps together to form a new feature map with a size of HxWx∑C_i. This process can be implemented through the concatenate function in a programming language, which can concatenate multiple arrays along the specified dimension to obtain a fused feature map containing all feature information. Please refer to Figure 2 , in the figure, the 256-dimensional feature vectors in five levels are respectively converted into 16 x 16 candidate feature maps, and then the candidate feature maps are concatenated by Concatenate. Among them, each circle in the feature vector represents each element in the feature vector.

[0102] In this embodiment, through the conversion from feature vectors to feature maps and the concatenation of feature maps, the problems caused by insufficient single-level feature information in image processing are avoided, and the effect of integrating feature information at different levels to form a more abundant and comprehensive feature representation is realized. It not only retains the structure and detail information of each level of features, but also enhances the model's understanding and expression ability of image content through the vertical concatenation of features.

[0103] Step S40, generating a derivative defect image corresponding to the original defect image based on the fused feature map.

[0104] It can be understood that since it is necessary to inversely generate a defective image from the feature representation, step S40 is performed, which can avoid the technical problems of low image generation quality and unrealistic defects caused by directly operating in the pixel space in traditional methods, and realizes the generation of highly realistic defect images, thereby improving the training quality and detection performance of the defect detection system.

[0105] Exemplarily, a decoder network is used, which can be another convolutional neural network, and its function is to reconstruct an image from the fused feature map. The decoder network gradually increases the size of the image through a series of transposed convolutional layers and upsampling layers, while restoring the details of the image. In this process, some noise or defect generation modules can be introduced, such as an adversarial noise generator, to introduce specific defect patterns in the reconstructed image, and finally obtain a derived defect image that has a similar appearance to the original defective image but contains specific defects.

[0106] In a feasible implementation manner, step S40 may include steps S41 to S43:

[0107] Step S41, taking the fused feature map as the feature map to be decoded, and performing multiple convolutional operations on the feature map to be decoded to obtain the feature map to be interpolated at the current level of the feature map to be decoded;

[0108] It should be noted that the feature map to be decoded refers to the original feature vector or the feature map of the next level before the convolutional operation. This feature map contains the feature information of the original defective image at different levels and is the starting point of the convolutional operation at the current level; the feature map to be interpolated refers to the feature map that needs to increase its size through interpolation operation during the decoding process. This feature map is the output of a certain level in the decoder network and is used to restore the feature representation at a higher resolution.

[0109] It can be understood that since it is necessary to gradually restore the details and structure of the image during the decoding process, step S41 is performed. By performing the corresponding convolutional operation, the problem of losing feature information during image reconstruction can be avoided, and the effect of restoring finer image features from the compressed feature representation can be achieved.

[0110] Exemplarily, first, the fused feature map is passed as input to the first layer of the decoding network. In this layer, through the convolutional operation of a series of convolutional kernels, such as 3x3 convolution, batch normalization, activation function (such as ReLU, etc.), the feature map to be decoded is processed. Each convolutional operation aims to extract or enhance specific attributes of the feature map. These convolutional operations can be transposed convolution or inverse convolution, which are used to gradually construct a higher-resolution feature representation, and finally obtain the feature map to be interpolated at the current level.

[0111] Step S42, performing a linear interpolation operation on the feature map to be interpolated to obtain the feature map to be decoded at the previous level of the feature map to be interpolated, and returning to execute the step of performing multiple convolutional operations on the feature map to be decoded;

[0112] It can be understood that since it is necessary to upsample the size of the feature map to a higher resolution, step S42 is performed. By performing linear interpolation on the feature map at each level, the problem of image blurring or distortion caused by direct upsampling can be avoided, thus achieving smooth and detailed feature map reconstruction.

[0113] Exemplarily, after obtaining the feature map to be interpolated at the current level, linear interpolation (such as bilinear interpolation or nearest neighbor interpolation) is applied to increase the size of the feature map to make it close to the size of the previous level. This process increases the resolution of the feature map, thereby obtaining the feature map to be decoded at the previous level. After the interpolation is completed, the network will use this new feature map to be decoded as the input and return to the step of performing multiple convolution operations to continue reconstructing features at a higher level.

[0114] Step S43, after the feature map to be decoded undergoes convolution operations and linear interpolation operations at each level, a target feature map is obtained, and a derivative defect image corresponding to the original defect image is generated based on the target feature map.

[0115] It should be noted that the target feature map refers to the feature map finally obtained after convolution operations and linear interpolation operations at all levels. This feature map has the same size as the original defect image and contains all the necessary information for generating the defect image.

[0116] It can be understood that since it is necessary to convert the feature information back to the pixel space to generate an image with specific defects, step S43 is performed. The problem of the generated defect image being inconsistent with the original image or the defects being unnatural can be avoided, and an image that is highly realistic and has specific defects is achieved.

[0117] Exemplarily, after the convolution and interpolation operations are completed at all levels of the decoding network, a target feature map with the same size as the original defect image is finally obtained. This target feature map contains all the necessary information for generating the defect image. Next, by applying some generative network techniques (such as Generative Adversarial Network GAN or Variational Autoencoder VAE), a defect image with specific defects, that is, the derivative defect image corresponding to the original defect image, is generated based on the target feature map. This process may include converting the feature map to the pixel space and performing some post-processing steps to ensure the naturalness of the defects and the overall quality of the image.

[0118] In this embodiment, through convolutional operations and linear interpolation operations at each level, the problems of feature information loss, blurred image details, and unnatural defects during image reconstruction are avoided. It is possible to gradually recover a high-resolution and detail-rich target feature map from the fused feature map, and finally generate a defect image that matches the original defect image and contains specific defects. This hierarchical decoding process not only retains the key features of the image but also ensures the authenticity and detail quality of the generated defect image.

[0119] This embodiment provides a method for generating a defect image. By performing multi-level image feature extraction, feature vector conversion and splicing on the original defect image, and generating a defect image based on the spliced fused feature map, the problems in traditional image generation methods, such as insufficient authenticity of the generated image, lack of details, and inaccurate representation of complex or minute defects due to neglecting image details and feature diversity, are avoided. It is possible to generate a defect image with high realism and rich details under the condition of limited data samples, enabling the defect detection system to be trained based on high-quality defect samples, thereby improving the training efficiency and detection accuracy of the defect detection system and ensuring the stability and reliability of defect recognition in industrial production.

[0120] In a feasible embodiment, the step of obtaining the feature maps of the original defect image at different levels based on the pooling feature maps at each level in step S13 may include steps S131 to S132:

[0121] Step S131: For any pooling feature map at a level, perform weighted averaging on the feature sub-maps at each depth in the pooling feature map to obtain an average feature map.

[0122] It should be noted that a feature sub-map refers to each individual image slice divided along the depth dimension (channel dimension) in the pooling feature map, and each feature sub-map represents specific type of feature information of the original defect image at this level; the average feature map refers to the feature map obtained by performing weighted averaging on the feature sub-maps at each depth in the pooling feature map, and this feature map represents a comprehensive and abstract representation of the feature map at this level, which contains important information of all feature sub-maps at this level.

[0123] It can be understood that since it is necessary to find a balance between different feature sub-maps to retain the most important feature information, performing step S131 can avoid the problems of feature information loss or imbalance of features in each feature sub-map caused by directly using the pooling feature map as the feature map of the original defect image at different levels, and achieve the effect of retaining key features and reducing redundant information.

[0124] Exemplarily, at a specific layer of the convolutional neural network, a weight is assigned to each channel (i.e., feature sub-map) of the to-be-pooled feature map. These weights can be preset according to the importance of the feature sub-map or learned through learning. Then, each feature sub-map is multiplied by its corresponding weight, and the results are added together. Finally, this sum is divided by the sum of the weights (or the number of channels if the weights are equal weights) to obtain an average feature map, which represents the weighted average information of all feature sub-maps at this layer.

[0125] Step S132, after traversing each to-be-pooled feature map, each average feature map is used as the feature map of the original defective image at different levels.

[0126] Exemplarily, after performing the weighted average operation at each layer of the network, a series of average feature maps are obtained. As the network propagates from the forward pass to the backward pass, these average feature maps are collected to form a set of feature maps. Each average feature map in this set corresponds to the feature representation of the original defective image at different levels. In the subsequent stages of the network, these average feature maps can be used as inputs.

[0127] In this embodiment, by performing weighted summation and averaging on the feature sub-maps of each depth in the to-be-pooled feature map, the problems of possible feature information loss or dilution of important features in traditional pooling operations are avoided, the effect of retaining key feature information and reducing redundancy is achieved, and at the same time, by using each average feature map as the feature map of the original defective image at different levels, the comprehensiveness and hierarchy of the feature representation are ensured, thereby improving the quality and accuracy of image analysis and processing tasks.

[0128] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as in the above-mentioned embodiment 1 can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , step S22 may include steps S221 to S222:

[0129] Step S221, inputting the multi-dimensional vector into a preset fully connected layer to generate candidate vectors corresponding to the highest dimension in each level through the preset fully connected layer;

[0130] It should be noted that the preset fully connected layer refers to a layer in the neural network where each input is connected to each output, and these connections have learnable weights; the candidate vector refers to an intermediate vector generated by processing the multi-dimensional vector through the preset fully connected layer, and this vector contains the feature representation of the input vector after being transformed by the fully connected layer.

[0131] It can be understood that since it is necessary to convert the flattened high-dimensional feature vector into a low-dimensional vector that is more suitable for subsequent processing, step S221 is performed, which can avoid the problems of high computational complexity and overfitting risk caused by too high dimensions, thus achieving the effects of feature dimensionality reduction and key information extraction.

[0132] Exemplarily, a -dimensional multi-dimensional vector is input into a preset fully connected layer to generate candidate vectors corresponding to the highest dimension in each layer through the preset fully connected layer and . For example, when the dimensions of each layer are 256, 128, 64, 32, and 16 respectively, the highest dimension is 256 dimensions, then the dimensions of the candidate vectors and are 256 dimensions.

[0133] Step S222, obtain the generated random noise, and determine the feature vector corresponding to the feature map based on the candidate vector and the random noise.

[0134] It should be noted that random noise refers to a randomly generated numerical vector, which is usually used to introduce uncertainty and increase data diversity.

[0135] It can be understood that since the acquisition cost of defective samples in actual industrial production is high and the quantity is limited, many traditional methods are difficult to generate defective images with diversity in the case of insufficient data. Therefore, step S222 is performed. By introducing varying random noise in the generation process of defective images to enhance the generalization ability and generation diversity ability of the model, it can avoid the generated feature vectors from being too single or patterned, and achieve the effect of generating feature vectors with both diversity and realism.

[0136] Exemplarily, please refer to Figure 4 . First, generate a random noise vector with the same dimension as the candidate vector . This vector is the highest dimension corresponding to each layer (such as 256 dimensions), and this vector can be created by a random number generator according to a specific distribution (such as the normal distribution N(0, 1)). Then, this random noise vector and the candidate vector and obtained through the fully connected layer before are reparameterized and sampled through the following formula to obtain the feature vector corresponding to the feature map :

[0137]

[0138] The feature vector can then be fed into a feature generator for splicing to obtain a fused feature map. Among them, the circles in the figure represent each element in the feature vector (or each pixel in the feature map), and the connecting lines between the circles represent each element in the multi-dimensional vector of dimensions respectively generates a candidate vector and the corresponding elements in

[0139] In this embodiment, by adopting a fully connected layer combined with the injection of random noise, the risks of insufficient model generalization ability and overfitting caused by single-dimensional feature representation, as well as the problem of lack of diversity in feature representation, are avoided. While maintaining the effectiveness of feature representation, the randomness and diversity of the feature vector are increased, thereby improving the generalization ability of the model and the effect of generating a set of feature vectors with rich and diverse features.

[0140] In a feasible implementation manner, after step S221, step S100 may further be included:

[0141] Step S100: Construct a loss function according to the candidate vector, and train the model parameters in the preset convolutional layer and the preset fully connected layer based on the loss function, where the preset convolutional layer is used to perform a convolutional operation.

[0142] It should be noted that the model parameters include the weights and biases in the preset convolutional layer and the preset fully connected layer.

[0143] It can be understood that since the convolutional layer and the fully connected layer set under default conditions often cannot exhibit excellent performance, step S100 is performed. By training the convolutional layer and the fully connected layer, the problem that the model performance cannot be better exerted due to the mismatch of the model parameters configured in the convolutional layer and the fully connected layer during the image analysis process can be avoided, and the effect of real-time optimizing the model parameters to improve the quality of the defective images generated by the model is achieved.

[0144] Exemplarily, according to the candidate vectors corresponding to the highest dimensions in each layer generated by the preset fully connected layer and construct a loss function as follows:

[0145]

[0146] Among them, i represents the index value of each layer, and it is indicated in this formula that there are five layers; represents the KL divergence between the distribution generated by the encoder and the normal distribution, that is, measures the similarity between these two distributions, and its calculation method is as follows:

[0147]

[0148] It can be seen from the formula that when = 0 and = 1, the loss of this term is the smallest. The smaller this term is, the more the feature vector generated by the encoder conforms to the standard normal distribution, so that the generated defect images are more realistic and diverse. represents the reconstruction error, and its calculation method is as follows:

[0149]

[0150] Among them, represents the real sample, represents the generated sample, represents the height and width of the original defect image, for example . This formula is used to minimize the error between the generated sample and the real sample , which means that an autoencoder model is trained. By constructing the loss function , each weight and each bias in the preset convolutional layer and the preset fully connected layer are trained and adjusted to minimize this loss function as the goal.

[0151] Exemplarily, to help understand the implementation process of the defect image generation method obtained by combining the above Embodiment 1, please refer to Figure 5 , Figure 5 provides a schematic diagram of the model framework of a defect image generation method. Specifically:

[0152] The rectangles in the figure are feature maps. The length of the rectangle represents the size of the feature map (corresponding to different levels), and the number marked on the left side of the rectangle is the specific size. For example, 256 means the size of the feature map is 256 x 256, and 128 means the size of the feature map is 128 x 128, etc. The height represents the dimension of the feature map (i.e., the depth of the feature map, which can also be called the number of feature sub - maps in the feature map), and the bold number marked on the right side of the rectangle is the specific dimension. For example, 64 means the dimension of the feature map is 64, and 1 means the dimension of the feature map is 1, etc. The internal framework corresponding to the sampler in the figure can be referred to Figure 4 , and the internal framework corresponding to the feature generator can be referred to Figure 2, in the figure, the upper part of the frame except the sampler and the feature generator corresponds to the encoder's frame, and the lower part of the frame in the figure corresponds to the decoder's frame. In the whole frame, five levels of image feature extraction corresponding to dimensions of 256, 128, 64, 32, and 16 are performed on the original defect image x, including encoder processing of three convolution operations (i.e., as shown by the solid solid arrows in the figure, such as 3x3 convolution, batch normalization, ReLU activation function, etc.), one flattening (i.e., as shown by the dashed hollow arrow in the figure), and one pooling (i.e., as shown by the solid hollow arrow in the figure) at each level. Then, the feature map with a dimension of 1 obtained after each flattening is input into the sampler to determine the feature vector corresponding to the feature map. Then, the feature vectors at each level are converted into 16x16 feature maps in the feature generator and then spliced to obtain a fused feature map. After that, based on this fused feature map, processing is performed at each level in the decoder (such as 16, 32, 64, 128, 256 in the figure), including two convolution operations (i.e., as shown by the solid solid arrows in the figure, such as 3x3 convolution, batch normalization, ReLU activation function, etc.) and one linear interpolation operation (i.e., as shown by the dotted hollow arrow in the figure) at each level. Finally, a target feature map with a size of 256x256 and a dimension of 64 is obtained to generate a derivative defect image X' corresponding to the original defect image x based on this target feature map.

[0153] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the method for generating defect images in this application. Any simple transformation in more forms based on this technical concept is within the protection scope of this application.

[0154] This application also provides a defect image generation system. Please refer to Figure 6 , the defect image generation system includes:

[0155] An encoder 10 for performing multi-level image feature extraction on the original defect image to obtain feature maps of the original defect image at different levels;

[0156] A sampler 20 for determining the feature vector corresponding to any feature map;

[0157] A feature generator 30 for fusing the feature vectors after traversing each feature map to obtain a fused feature map;

[0158] A decoder 40 for generating a derivative defect image corresponding to the original defect image based on the fused feature map.

[0159] Optionally, the encoder 10 is further configured to:

[0160] Take the original defect image as the feature map to be encoded, and perform multiple convolution operations on the feature map to be encoded to obtain the feature map to be pooled at the current level of the feature map to be encoded;

[0161] Perform a pooling operation on the feature map to be pooled to obtain the feature map to be encoded at the next level, and return to execute the step of performing multiple convolution operations on the feature map to be encoded;

[0162] After traversing each level, obtain the feature maps of the original defect image at different levels based on the feature maps to be pooled at each level.

[0163] Optionally, the encoder 10 is further configured to:

[0164] For the feature map to be pooled at any level, perform weighted averaging on the feature sub-maps at each depth in the feature map to be pooled to obtain an average feature map;

[0165] After traversing each feature map to be pooled, use each average feature map as the feature map of the original defect image at different levels.

[0166] Optionally, the sampler 20 is further configured to:

[0167] Flatten the feature map to obtain a multi-dimensional vector corresponding to the feature map, where the dimension of the multi-dimensional vector is determined based on the size of the feature map;

[0168] Perform sampling on the highest corresponding dimension in each level of the multi-dimensional vector to obtain a feature vector corresponding to the feature map.

[0169] Optionally, the sampler 20 is further configured to:

[0170] Input the multi-dimensional vector into a preset fully connected layer to generate candidate vectors corresponding to the highest dimension in each level through the preset fully connected layer;

[0171] Obtain the generated random noise, and determine the feature vector corresponding to the feature map based on the candidate vector and the random noise.

[0172] Optionally, the training module 50 in the defect image generation system is configured to:

[0173] Construct a loss function according to the candidate vector, and train the model parameters in the preset convolutional layer and the preset fully connected layer based on the loss function, where the preset convolutional layer is used to perform convolution operations.

[0174] Optionally, the feature generator 30 is further configured to:

[0175] For any one feature vector, convert the feature vector into a candidate feature map;

[0176] After traversing each eigenvector, each candidate feature map is stitched together to obtain a fused feature map.

[0177] Optionally, the decoder 40 is further configured to:

[0178] Take the fused feature map as the feature map to be decoded, and perform multiple convolution operations on the feature map to be decoded to obtain the interpolated feature map of the feature map to be decoded at the current level;

[0179] Perform a linear interpolation operation on the interpolated feature map to obtain the feature map to be decoded at the previous level of the interpolated feature map, and return to execute the step of performing multiple convolution operations on the feature map to be decoded;

[0180] After the feature map to be decoded undergoes convolution operations and linear interpolation operations at each level, a target feature map is obtained, and a derivative defect image corresponding to the original defect image is generated based on the target feature map.

[0181] The defect image generation system provided by this application adopts the defect image generation method in the above embodiment, and can solve the technical problem of how to improve the generation quality of defect images to generate more detailed and realistic defect images. Compared with the prior art, the beneficial effects of the defect image generation system provided by this application are the same as those of the defect image generation method provided by the above embodiment, and other technical features in the defect image generation system are the same as those disclosed in the method of the above embodiment, which will not be elaborated here.

[0182] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the defect image generation method in the first embodiment above.

[0183] Next, refer to Figure 7 , which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7The electronic device shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.

[0184] As Figure 7 shown, the electronic device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the electronic device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an electronic device with various systems, it should be understood that it is not required to implement or have all the systems shown. Instead, more or fewer systems may be implemented or had.

[0185] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0186] The electronic device provided by the present application, adopting the defective image generation method in the above embodiments, can solve the technical problem of how to improve the generation quality of defective images to generate more detailed and realistic defective images. Compared with the prior art, the beneficial effects of the electronic device provided by the present application are the same as those of the defective image generation method provided in the above embodiments, and the other technical features in this electronic device are the same as those disclosed in the method of the previous embodiment, which will not be elaborated here.

[0187] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0188] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0189] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the defective image generation method in the above embodiments.

[0190] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM: Random Access Memory), read-only memory (ROM: Read Only Memory), erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0191] The above computer-readable storage medium can be included in an electronic device; or it can exist alone without being assembled into the electronic device.

[0192] The above computer-readable storage medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to: perform multi-level image feature extraction on an original defect image to obtain feature maps of the original defect image at different levels; for any one of the feature maps, determine a feature vector corresponding to the feature map; after traversing each feature map, fuse the feature vectors to obtain a fused feature map; and generate a derivative defect image corresponding to the original defect image based on the fused feature map.

[0193] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0194] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0195] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.

[0196] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned defective image generation method, which can solve the technical problem of how to improve the generation quality of defective images to generate more detailed and realistic defective images. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the defective image generation method provided by the above embodiments, and will not be elaborated here.

[0197] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the defective image generation method as described above.

[0198] The computer program product provided by this application can solve the technical problem of how to improve the generation quality of defective images to generate more detailed and realistic defective images. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as those of the defective image generation method provided by the above embodiments, and will not be elaborated here.

[0199] The above are only some embodiments of this application, and thus do not limit the patent scope of this application. Any equivalent structural transformation made by using the content of the specification and drawings of this application under the technical concept of this application, or direct / indirect application in other related technical fields, is included in the patent protection scope of this application.

Claims

1. A defect image generation method, characterized in that: The defect image generating method comprises: Performing multi-level image feature extraction on the original defect image to obtain feature maps of the original defect image at different levels; For any feature map, determine the feature vector corresponding to the feature map; After traversing each feature map, each feature vector is fused to obtain a fused feature map; The fused feature map is used as a feature map to be decoded, and multiple convolution operations are performed on the feature map to be decoded to obtain a feature map to be interpolated at the current level of the feature map to be decoded; Performing a linear interpolation operation on the feature map to be interpolated to obtain a feature map to be decoded at an upper level of the feature map to be interpolated, and returning to the step of performing multiple convolution operations on the feature map to be decoded; After the feature map to be decoded undergoes convolution operations and linear interpolation operations at various levels, a target feature map is obtained, and a derived defect image corresponding to the original defect image is generated based on the target feature map.

2. The defect image generating method according to claim 1, characterized in that: The step of performing multi-level image feature extraction on the original defect image to obtain feature maps of the original defect image at different levels comprises: The original defect image is used as a feature map to be encoded, and multiple convolution operations are performed on the feature map to be encoded to obtain a pooled feature map of the feature map to be encoded at the current level; Performing a pooling operation on the feature map to be pooled to obtain a feature map to be encoded at the next level, and returning to execute the step of performing multiple convolution operations on the feature map to be encoded; After traversing each level, the feature maps of the original defect image at different levels are obtained based on the feature maps to be pooled at each level.

3. The defect image generating method according to claim 2, characterized in that: The step of obtaining the feature maps of the original defect image at different levels based on the feature maps to be pooled at each level includes: For a feature map to be pooled at any level, weighted averaging is performed on the feature submaps at each depth in the feature map to be pooled to obtain an average feature map; After traversing each feature map to be pooled, each average feature map is used as the feature map of the original defect image at different levels.

4. The defect image generating method according to claim 1, characterized in that: The step of determining the feature vector corresponding to the feature map comprises: Flattening the feature map to obtain a multidimensional vector corresponding to the feature map, wherein the dimension of the multidimensional vector is determined based on the size of the feature map; The multidimensional vector is sampled at the highest dimension in each level to obtain a feature vector corresponding to the feature map.

5. The defect image generating method according to claim 4, characterized in that: The step of sampling the multidimensional vector corresponding to the highest dimension in each level to obtain the feature vector corresponding to the feature map comprises: Inputting the multidimensional vector into a preset fully connected layer to generate a candidate vector corresponding to the highest dimension in each layer through the preset fully connected layer; The generated random noise is obtained, and a feature vector corresponding to the feature map is determined based on the candidate vector and the random noise.

6. The defect image generating method according to claim 5, characterized in that: The defect image generating method further comprises: A loss function is constructed according to the candidate vector, and model parameters in a preset convolutional layer and a preset fully connected layer are trained based on the loss function, wherein the preset convolutional layer is used to perform a convolution operation.

7. The defect image generating method according to claim 1, characterized in that: The step of fusing the feature vectors to obtain a fused feature map comprises: For any feature vector, convert the feature vector into a candidate feature map; After traversing each feature vector, each candidate feature map is concatenated to obtain a fused feature map.

8. An electronic device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the defect image generating method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the defect image generation method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Industrial defect image generation method and device, equipment and storage medium

    CN117671431A

  • Image processing method and device based on multi-level feature extraction, equipment and storage medium

    CN118691845A