Defect detection method and device, storage medium and electronic equipment

By acquiring the intermediate layer vector and historical processing strategies of the image to be detected, processing the image area in combination with global and local attention models, and generating feature maps to identify defects, the problem of inaccurate defect detection in the prior art is solved and higher detection accuracy is achieved.

CN120563501AActive Publication Date: 2025-08-29SHENZHEN XINRUN FULIAN DIGITAL TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511050543.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-08-29
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

In the prior art, due to the diverse object types and differences in picture, the recognition effect is inaccurate when using the same model for defect detection.

Method used

By obtaining the intermediate layer vector of the image to be detected, combining the processing strategy of the historical image, the processing strategy of the image to be detected is determined, and a specific area of ​​the image is processed separately or jointly using the global model and the local attention model to generate a feature map to identify defects.

Benefits of technology

It improves the accuracy of defect detection, adapts to images of different types and complexities, and achieves more accurate defect recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120563501A_ABST
    Figure CN120563501A_ABST
Patent Text Reader

Abstract

The invention relates to a defect detection method and device, a storage medium and electronic equipment. The method comprises the steps of obtaining an intermediate layer vector of a to-be-detected image under the condition that the to-be-detected image is obtained; according to the intermediate layer vector and a historical processing strategy of the historical image, determining a processing strategy of the to-be-detected image, the processing strategy comprising a model for processing the to-be-detected image and a processing area of the model; processing a processing area of the to-be-detected image by using a model corresponding to the processing strategy to obtain a feature map; and identifying the image defect of the to-be-detected image according to the feature map. The technical problem of low defect identification accuracy is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of defect detection, and in particular to a defect detection method, device, storage medium, and electronic device. Background Art

[0002] In the process of detecting defects on the surface of an object, the model can be used to identify the image obtained by photographing the surface of the object to detect defects on the object.

[0003] Existing models can be used to detect surface defects, such as methods for detecting minor visual defects on surfaces. These methods feed images into a pre-trained neural network model. The model's attention mechanism retains subtle features, allowing for the detection of minor surface defects.

[0004] However, due to the variety of object types and the differences in the images taken, if the same model is used for unified processing, the recognition effect will not be accurate. Summary of the Invention

[0005] The present application provides a defect detection method, device, storage medium and electronic device to solve the technical problem of low defect recognition accuracy.

[0006] In the first aspect, the present application provides a defect detection method, comprising: when an image to be detected is obtained, obtaining the intermediate layer vector of the image to be detected; determining the processing strategy of the image to be detected based on the intermediate layer vector and the historical processing strategy of the historical image, wherein the processing strategy includes a model for processing the image to be detected and a processing area of ​​the model; using the model corresponding to the processing strategy to process the processing area of ​​the image to be detected to obtain a feature map; and identifying image defects of the image to be detected based on the feature map.

[0007] In the second aspect, the present application provides a defect detection device, including: an acquisition module, used to obtain the intermediate layer vector of the image to be detected when the image to be detected is acquired; a determination module, used to determine the processing strategy of the image to be detected based on the intermediate layer vector and the historical processing strategy of the historical image, wherein the processing strategy includes a model for processing the image to be detected and a processing area of ​​the model; a processing module, used to use the model corresponding to the processing strategy to process the processing area of ​​the image to be detected to obtain a feature map; an identification module, used to identify the image defects of the image to be detected based on the feature map.

[0008] As an optional example, the above-mentioned processing module includes: a first processing unit, used to input the first processing area of ​​the above-mentioned image to be detected into the above-mentioned global model when the above-mentioned processing strategy is to use the global model to process the first processing area of ​​the above-mentioned image to be detected, so as to obtain the above-mentioned feature map, wherein the model weight of the above-mentioned global model is converted from a first floating point number to a first integer.

[0009] As an optional example, the above-mentioned processing module includes: a second processing unit, used to input the second processing area of ​​the above-mentioned image to be detected into the above-mentioned local attention model when the above-mentioned processing strategy is to use the local attention model to process the second processing area of ​​the above-mentioned image to be detected, so as to obtain the above-mentioned feature map, wherein the model weight of the above-mentioned local attention model is converted from a second floating point number to a half-precision floating point number.

[0010] As an optional example, the above-mentioned processing module includes: a third processing unit, which is used to input the third processing area of ​​the above-mentioned image to be detected into the above-mentioned global model and input the fourth processing area of ​​the above-mentioned image to be detected into the above-mentioned local attention model when the above-mentioned processing strategy is to use the global model to process the third processing area of ​​the above-mentioned image to be detected and to use the local attention model to process the fourth processing area of ​​the above-mentioned image to be detected; combine the global features output by the above-mentioned global model with the local features output by the above-mentioned local attention model to obtain the above-mentioned feature map.

[0011] As an optional example, the third processing unit includes: a processing sub-unit, used to perform upsampling and spatial mapping operations on the global features and the local features, and then spatially align them; splicing the aligned global features and the local features into multi-scale combined features; inputting the multi-scale combined features into a deformable convolution layer, and performing pixel misalignment compensation on the multi-scale combined features through the deformable convolution layer; and adaptively sampling the compensated multi-scale combined features to obtain the feature map.

[0012] As an optional example, the above-mentioned determination module includes: a determination unit, used to determine the historical intermediate layer output vector of the above-mentioned historical image based on the above-mentioned historical image; determine the target historical image from the historical images based on the similarity between the intermediate layer vector of the above-mentioned image to be detected and the historical intermediate layer output vector of the above-mentioned historical image; and determine the processing strategy of the above-mentioned target historical image as the processing strategy of the above-mentioned image to be detected.

[0013] As an optional example, the above-mentioned determination module includes: an input unit, used to input the above-mentioned image to be detected into the strategy network when the above-mentioned processing strategy has not been determined, and the above-mentioned strategy network extracts the intermediate layer vector of the above-mentioned image to be detected, and predicts the above-mentioned processing strategy based on the above-mentioned intermediate layer vector.

[0014] In a third aspect, the present application provides an electronic device comprising: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus, wherein the memory stores a computer program, and the processor is configured to implement any one of the above-mentioned defect detection methods when executing the computer program.

[0015] In a fourth aspect, the present application further provides a computer storage medium storing computer executable instructions, wherein the computer executable instructions are used to execute any of the above-mentioned defect detection methods of the present application.

[0016] The above-mentioned technical solution provided by the embodiment of the present application has the following advantages over the prior art: the solution provided by the embodiment of the present application, by obtaining the intermediate layer vector of the above-mentioned image to be detected when the image to be detected is obtained; determining the processing strategy of the above-mentioned image to be detected based on the above-mentioned intermediate layer vector and the historical processing strategy of the historical image, wherein the above-mentioned processing strategy includes a model for processing the above-mentioned image to be detected and a processing area of ​​the model; using the model corresponding to the above-mentioned processing strategy to process the above-mentioned processing area of ​​the above-mentioned image to be detected to obtain a feature map; identifying the image defects of the above-mentioned image to be detected based on the above-mentioned feature map, so that the intermediate layer vector of the image to be detected can be assisted according to the historical processing strategy of the historical image to determine the model suitable for the image to be detected and the processing area of ​​the image to be detected, processing the processing area by the corresponding model to obtain a feature map, and identifying the image defects of the image to be detected based on the feature map, thereby achieving the effect of using a suitable model to process the processing area of ​​the image to be detected, and improving the accuracy of defect identification of the image to be detected. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0019] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0020] Figure 1 A flowchart of a defect detection method provided in an embodiment of the present application; Figure 2 A diagram for determining a processing strategy for a defect detection method provided in an embodiment of the present application; Figure 3 A characteristic combination diagram of a defect detection method provided in an embodiment of the present application; Figure 4 A schematic structural diagram of a defect detection device provided in an embodiment of the present application; Figure 5 A schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0022] The disclosure below provides many different embodiments or examples for implementing different configurations of the present invention. To simplify the disclosure of the present invention, the components and configurations of specific examples are described below. Of course, these are merely examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numerals and / or letters in different examples. Such repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or configurations discussed.

[0023] In order to solve the technical problem of low defect recognition accuracy in the prior art, the present application provides a defect detection method that can improve the accuracy of defect recognition on an image to be detected.

[0024] Figure 1 This is a flow chart of a defect detection method provided in an embodiment of the present application. Figure 1 As shown, the above-mentioned defect detection method includes: S101, when an image to be detected is obtained, obtaining an intermediate layer vector of the image to be detected; S102, determining a processing strategy for the image to be detected based on the intermediate layer vector and the historical processing strategy of the historical image, wherein the processing strategy includes a model for processing the image to be detected and a processing area of ​​the model; S103, using the model corresponding to the processing strategy to process the processing area of ​​the image to be detected to obtain a feature map; S104: Identify image defects of the image to be detected according to the feature map.

[0025] The above-mentioned defect detection method can be applied in the process of detecting defects of objects in images, for example, detecting defects on the surface of a workpiece, detecting cracks on the surface of a road, etc.

[0026] The image to be inspected can be an image obtained by photographing the object to be inspected for defects, for example, photographing a workpiece to obtain an image containing the workpiece, or photographing a road surface to obtain an image containing the road surface. The image is used for defect detection.

[0027] This embodiment includes multiple models. When performing image defect detection, one or more of these models can be used to perform defect detection. The specific one or more models to be used for defect detection is mainly determined based on the image to be detected and the historical images.

[0028] The idea behind this embodiment is to determine the historical processing strategy for historical images, which in turn correspond to historical intermediate-layer vectors. By comparing the intermediate-layer vectors of the image to be detected with the historical intermediate-layer vectors of the historical images, the type of the image to be detected can be determined, or which historical images it is similar to. Furthermore, the historical processing strategy of the historical images can be used as the processing strategy for the image to be detected.

[0029] The processing strategy in this embodiment includes both the model to be used for inspecting the image to be inspected and the processing areas within the image to be inspected. Therefore, after comparing the model to be used and the processing areas within the image to be inspected, the corresponding model can be used. The corresponding processing areas are processed to generate a feature map. This feature map is then used for final defect identification.

[0030] For example, Figure 2 As shown, for historical images, the corresponding processing strategies have been determined and processed, and the image to be detected can be compared with the historical images for intermediate layer features to determine which historical image processing strategy to use to process the image to be detected.

[0031] In this embodiment, the model used to process the image to be detected is divided into a global model and a local attention model. The global model can be used to extract the global features of the image to be detected. The local attention model can focus more on the local features of the image to be detected. The two models can be used separately or simultaneously. In addition, during the process of using the model, it can also be determined which part of the image to be detected is to be processed by the model.

[0032] The solution provided by the embodiment of the present application obtains the intermediate layer vector of the image to be detected when the image to be detected is obtained; determines the processing strategy of the image to be detected based on the intermediate layer vector and the historical processing strategy of the historical image, wherein the processing strategy includes a model for processing the image to be detected and a processing area of ​​the model; uses the model corresponding to the processing strategy to process the processing area of ​​the image to be detected to obtain a feature map; identifies image defects of the image to be detected based on the feature map, so that the intermediate layer vector of the image to be detected can be assisted according to the historical processing strategy of the historical image to determine the model suitable for the image to be detected and the processing area of ​​the image to be detected, processes the processing area with the corresponding model to obtain a feature map, and identifies the image defects of the image to be detected based on the feature map, thereby achieving the effect of using a suitable model to process the processing area of ​​the image to be detected, and improving the accuracy of defect identification of the image to be detected.

[0033] As an optional example, using a model corresponding to the processing strategy to process the processing area of ​​the image to be detected to obtain a feature map includes: when the processing strategy is to use a global model to process the first processing area of ​​the image to be detected, inputting the first processing area of ​​the image to be detected into the global model to obtain a feature map, wherein the model weight of the global model is converted from a first floating point number to a first integer.

[0034] The first processing area in this embodiment may be the entire area of ​​the image to be detected, or a partial area. If it is a partial area, it may be a key area of ​​the image to be detected.

[0035] The global model in this embodiment can be a convolutional neural network model (CNN), which can also be called a large-scale expert model, and the model is used to extract the global context features of the image to be detected or the processing area of ​​the image to be detected. In order to make the global model more focused on the integrity of the model to be detected, the weight of the global model can be adjusted in this embodiment. The weight of the global model is adjusted according to the degree of focus on the integrity of the image to be detected. For example, the weight can be a floating point number, and the number of digits in the decimal part of the floating point number can be adjusted according to the degree of focus on the integrity of the image to be detected. The more focused on the whole, the more digits in the decimal part. If the focus is on the local, the number of digits in the decimal part can be fewer.

[0036] In this embodiment, in order to allow the global model to focus on the entire image to be detected, weights for the global model can be determined. If the weights are floating-point numbers, the floating-point numbers are converted to integers. The number of bits in the converted integer can be determined based on the degree of focus on the entire image. The greater the focus on the entire image, the fewer bits in the integer. For example, a 32-bit single-precision floating-point number can be converted to an 8-bit integer, or a 64-bit double-precision floating-point number can be converted to a 16-bit integer.

[0037] In this embodiment, the range of change of the weight may also be determined according to the image complexity of the image to be detected. In order to retain more overall features, the higher the image complexity, the greater the range of change of the weight.

[0038] As an optional example, using a model corresponding to the processing strategy to process the processing area of ​​the image to be detected to obtain a feature map includes: when the processing strategy is to use a local attention model to process the second processing area of ​​the image to be detected, inputting the second processing area of ​​the image to be detected into the local attention model to obtain a feature map, wherein the model weights of the local attention model are converted from a second floating point number to a half-precision floating point number.

[0039] The second processing area in this embodiment may be the entire area or a partial area of ​​the image to be detected. If it is a partial area, it may be a key area of ​​the image to be detected.

[0040] The local attention model in this embodiment can be a model based on the Transformer architecture, which is used to identify local features of the image to be detected. In order to make the local attention model more focused on the local features of the image to be detected, the weight of the model can be adjusted, such as increasing the number of digits in the model weight. For example, if the weight of the model is a floating point number, the number of digits in the decimal part of the floating point number can be increased.

[0041] In this embodiment, if the number of bits of weights is simply increased, it will impose a large burden on calculation and memory. After calculation, it is found that when the number of bits of the weights of the local attention model is reduced, the extraction of local features can still be satisfied, and the calculation process and memory optimization are accelerated. For example, if the weights of 32-bit single-precision floating-point numbers are converted to 16-bit half-precision floating-point numbers, the model can still extract local features of the image to be detected, with higher computational efficiency and better memory optimization effect.

[0042] In this embodiment, the range of change of the weights may also be determined according to the image complexity of the image to be detected. In order to retain more local features, the more complex the image, the smaller the range of change of the weights.

[0043] As an optional example, using a model corresponding to the processing strategy to process the processing area of ​​the image to be detected to obtain a feature map includes: when the processing strategy is to use a global model to process the third processing area of ​​the image to be detected and to use a local attention model to process the fourth processing area of ​​the image to be detected, the third processing area of ​​the image to be detected is input into the global model and the fourth processing area of ​​the image to be detected is input into the local attention model; the global features output by the global model are combined with the local features output by the local attention model to obtain a feature map.

[0044] In this embodiment, if the processing strategy includes both using a global model to process the image to be detected and using a local attention model to process the image to be detected, the image to be processed is input into both the global model and the local attention model. The third processing area of ​​the image to be detected involved in the global model can be the entire area or a partial area of ​​the image to be detected. The fourth processing area of ​​the image to be detected involved in the local attention model can be the entire area or a partial area of ​​the image to be detected. The third processing area and the fourth processing area can be the same or different.

[0045] The features obtained by processing the third processing region of the image to be inspected by the global model can be global features, and the features obtained by processing the fourth processing region of the image to be inspected by the local attention model can be local features. The global features and the local features are combined to generate a feature map, which is used to determine whether the image to be inspected contains defects.

[0046] For example, Figure 3 As shown in Figure 2, the image to be detected is processed by both the global model and the local attention model. The global features obtained by the global model are combined with the local features obtained by the local attention model to obtain a feature map.

[0047] As an optional example, the global features output by the global model are combined with the local features output by the local attention model to obtain a feature map, which includes: upsampling and spatially mapping the global features and the local features, and then spatially aligning them; splicing the aligned global features and local features into multi-scale combined features; inputting the multi-scale combined features into a deformable convolution layer, and compensating for pixel misalignment of the multi-scale combined features through the deformable convolution layer; and adaptively sampling the compensated multi-scale combined features to obtain a feature map.

[0048] In this embodiment, for the case where the image to be detected is processed jointly by the global model and the local attention model, there are multiple cases when combining features. First, since the regions of the image to be detected processed by the global model and the local attention model may be different, the global features extracted by the global model and the local features extracted by the local attention model are upsampled and mapped to the same space, so as to spatially align the features and splice the aligned features. Since the regions of the image to be detected processed by the global model and the local attention model may be the same, there may be overlapping parts, or they may be different. For the cases of the same or overlapping parts, pixel misalignment compensation is required. The global features and the local features can be spliced ​​first, and then a multi-scale combined feature is obtained. The multi-scale combined feature is input into the deformable convolution layer. The multi-scale combined feature is compensated for pixel misalignment by the deformable convolution layer. The compensated combined feature can be adaptively sampled to obtain a feature map.

[0049] As an optional example, determining the processing strategy of the image to be detected based on the intermediate layer vector and the historical processing strategy of the historical image includes: determining the historical intermediate layer output vector of the historical image based on the historical image; determining the target historical image from the historical images based on the similarity between the intermediate layer vector of the image to be detected and the historical intermediate layer output vector of the historical image; and determining the processing strategy of the target historical image as the processing strategy of the image to be detected.

[0050] In this embodiment, when determining the processing strategy for the image to be detected, since the intermediate layer vector of the image to be detected is obtained, the intermediate layer vector of the image to be detected can be compared with the historical intermediate layer vectors of the historical images. Since the corresponding processing strategy has been determined for the historical image and it has been processed according to the corresponding processing strategy, it can be determined whether the processing strategy of the historical image is appropriate. The historical images with appropriate strategies can be marked. After comparing the intermediate layer vector of the image to be detected with the intermediate layer vector of the historical images, it can be determined which historical images have a high similarity with the image to be detected. The processing strategy of the marked historical image is then used as the processing strategy for the image to be detected. If the corresponding historical image is not marked, it means that no appropriate processing strategy has been found. If no appropriate processing strategy has been found, a strategy judgment can be performed on the image to be detected to determine the appropriate processing strategy for the image to be detected.

[0051] One approach involves analyzing the differences between adjacent features within the extracted features. High differences between adjacent features indicate greater detail differences in the image being detected, and a local attention model can be used in the processing strategy. Low differences between adjacent features can be used, and a global attention model can be used. The degree of difference can be controlled by a threshold.

[0052] As an optional example, determining the processing strategy of the image to be detected based on the intermediate layer vector and the historical processing strategy of historical images includes: when the processing strategy has not been determined, inputting the image to be detected into the strategy network, extracting the intermediate layer vector of the image to be detected by the strategy network, and predicting the processing strategy based on the intermediate layer vector.

[0053] In this example, when determining a policy for an image to be inspected, the image can be input into the policy network, which then extracts the intermediate layer vectors of the image. Based on these intermediate layer vectors, the network predicts a processing policy. The predicted processing policy can then be used to process the image, and the processing results are fed back to the policy network for training.

[0054] In the process of predicting the processing strategy, the prediction can be made based on the intermediate layer vectors of the entire area of ​​the image to be detected, or the prediction area can be determined from the image to be detected and then the intermediate layer vectors of the prediction area can be used for prediction.

[0055] During prediction, a step-by-step approach can be used, such as selecting a prediction area, predicting a first prediction result, then expanding the prediction area in all directions and continuing the prediction to predict a second prediction result. Based on the difference between the second prediction result and the first prediction result, it is determined whether the prediction result after two predictions is used as the final prediction result. If it is not used as the final prediction result, the prediction area can be further expanded until the final prediction result is determined. If the final prediction result is not determined after expanding to the entire image to be detected, the prediction results of each prediction are comprehensively considered to determine the final prediction result.

[0056] This application solves the problem of accurate fusion of features of different scales and sources by providing deformable convolution for feature fusion. The obtained feature map is used for defect detection, thereby improving the accuracy of defect detection.

[0057] Taking the application of the method in the present application to the scenario of workpiece defect detection as an example, the image to be detected can be a photographed image of the workpiece.

[0058] First, build the model.

[0059] Initialize the global model (also known as the large-scale expert model): Build a large-scale expert model based on a convolutional neural network. This model is designed to extract global contextual features from images to identify product integrity and large-scale anomalies (such as deformation and large stains). During model training or fine-tuning, simulated quantization operations (pseudo-quantization nodes) are introduced. Through quantization-aware training (QAT), the model weights are converted from standard 32-bit floating-point numbers (FP32) to 8-bit integers (INT8). This allows the model to learn to adapt to the accuracy loss caused by quantization during training, ensuring accuracy loss of less than 1% during deployment.

[0060] Initialize the local attention model (also known as the small-scale expert model): Build a small-scale expert model based on the Transformer architecture. This model utilizes the local attention mechanism and is designed for high-precision detection and localization of small local defects (such as scratches and pits). Leveraging the TensorRT framework, the Transformer model's weights are converted from FP32 to 16-bit half-precision floating-point (FP16) mode, ensuring detail detection accuracy while achieving computational acceleration and memory optimization.

[0061] When the workpiece image is received, dynamic routing decisions are made.

[0062] After receiving an image of an industrial product to be inspected, the system performs a historical routing decision reuse query: Feature extraction: The workpiece image or key regions (patches) of the image are fed into a large-scale expert model, and the intermediate layer output is used as a visual feature vector. Similarity comparison: The extracted feature vector is compared with the historical feature vectors of historical images stored in a historical decision library, using the cosine similarity metric and comparing against a preset threshold. The historical decision library stores recent routing decisions. Decision reuse: If a historical region whose visual content meets the threshold is found, the associated routing decision is directly reused, using the same processing strategy as the routing decision to process the workpiece image, skipping policy network inference. Reinforcement learning policy routing: If no reusable historical decision is found, or if the system settings require re-evaluation, the workpiece image or key regions of the image are fed as state input into a pre-trained reinforcement learning (RL) policy network. The RL policy network performs inference and outputs an action that specifies the optimal strategy for processing the current region, such as invoking only the large-scale expert, only the small-scale expert, both, or specifying a specific region to be processed. Historical database update: The routing decision and the corresponding regional feature vector obtained from this reasoning are stored in the historical decision database for subsequent reuse.

[0063] Once the processing strategy is determined based on historical routing, the image data can be distributed to the corresponding expert model. If the decision involves large-scale expert model processing, the large-scale expert model infers the input data, extracts global feature maps, and identifies overall anomalies. If the decision involves small-scale expert model processing, the small-scale expert model infers the specified input data and locates minor defects.

[0064] If the processing strategy involves combining large-scale and small-scale expert models, feature combination is required. The outputs of the large-scale expert (low-resolution global features) and the small-scale expert (high-resolution local features) are collected. Initial spatial alignment is performed through upsampling and spatial mapping, and these are concatenated in the channel dimension to form a multi-scale combined feature map. The combined feature map is then input into a deformable convolution layer. The DCN layer, which can learn the spatial offset of convolution kernel sampling points, can adaptively compensate for pixel misalignment, effectively overcoming pixel-level deviations caused by scale differences, model structures, and imperfect initial alignment. It accurately captures information at corresponding locations from feature maps from different sources. Edge representation is optimized, and through data-driven adaptive sampling, it better focuses on information-complex areas such as defect boundaries. Multi-scale evidence is integrated to significantly improve the accuracy of defect edge segmentation and localization. Fine-grained features are generated, and a fine feature map containing rich multi-scale information, effectively fused and optimized by the DCN, is output.

[0065] The feature map is fed into the final prediction head, which then outputs the final, highly accurate industrial defect detection results for subsequent quality control or process use. The results can include the presence and location of defects.

[0066] The images and detection results described above can be used to reverse-train the model. Using the evaluation results as reward signals, the RL policy network can be updated offline or online to continuously improve the efficiency and accuracy of routing decisions. Based on the feedback data, large- and small-scale expert models can be fine-tuned or incrementally trained to enhance their ability to recognize specific defects or scenarios. The accuracy of the INT8 model is also monitored, and QAT is re-performed when necessary to maintain model accuracy.

[0067] Figure 4 This is a schematic diagram of the structure of a defect detection device provided in an embodiment of the present application. Figure 4 As shown, the above-mentioned defect detection device includes: An acquisition module 401 is used to acquire an intermediate layer vector of the image to be detected when the image to be detected is acquired; A determination module 402 is configured to determine a processing strategy for the image to be detected based on the intermediate layer vector and the historical processing strategy of the historical image, wherein the processing strategy includes a model for processing the image to be detected and a processing region of the model; The processing module 403 is used to process the processing area of ​​the image to be detected using the model corresponding to the processing strategy to obtain a feature map; The recognition module 404 is configured to recognize image defects of the image to be detected based on the feature map.

[0068] The above-mentioned defect detection device can be used in the process of detecting defects of objects in images, for example, detecting defects on the surface of a workpiece, detecting cracks on the surface of a road, etc.

[0069] The image to be inspected can be an image obtained by photographing the object to be inspected for defects, for example, photographing a workpiece to obtain an image containing the workpiece, or photographing a road surface to obtain an image containing the road surface. The image is used for defect detection.

[0070] This embodiment includes multiple models. When performing image defect detection, one or more of these models can be used to perform defect detection. The specific one or more models to be used for defect detection is mainly determined based on the image to be detected and the historical images.

[0071] The idea behind this embodiment is to determine the historical processing strategy for historical images, which in turn correspond to historical intermediate-layer vectors. By comparing the intermediate-layer vectors of the image to be detected with the historical intermediate-layer vectors of the historical images, the type of the image to be detected can be determined, or which historical images it is similar to. Furthermore, the historical processing strategy of the historical images can be used as the processing strategy for the image to be detected.

[0072] The processing strategy in this embodiment includes both the model to be used for inspecting the image to be inspected and the processing areas within the image to be inspected. Therefore, after comparing the model to be used and the processing areas within the image to be inspected, the corresponding model can be used. The corresponding processing areas are processed to generate a feature map. This feature map is then used for final defect identification.

[0073] For example, Figure 2 As shown, for historical images, the corresponding processing strategies have been determined and processed, and the image to be detected can be compared with the historical images for intermediate layer features to determine which historical image processing strategy to use to process the image to be detected.

[0074] In this embodiment, the model used to process the image to be detected is divided into a global model and a local attention model. The global model can be used to extract the global features of the image to be detected. The local attention model can focus more on the local features of the image to be detected. The two models can be used separately or simultaneously. In addition, during the process of using the model, it can also be determined which part of the image to be detected is to be processed by the model.

[0075] The solution provided by the embodiment of the present application obtains the intermediate layer vector of the image to be detected when the image to be detected is obtained; determines the processing strategy of the image to be detected based on the intermediate layer vector and the historical processing strategy of the historical image, wherein the processing strategy includes a model for processing the image to be detected and a processing area of ​​the model; uses the model corresponding to the processing strategy to process the processing area of ​​the image to be detected to obtain a feature map; identifies image defects of the image to be detected based on the feature map, so that the intermediate layer vector of the image to be detected can be assisted according to the historical processing strategy of the historical image to determine the model suitable for the image to be detected and the processing area of ​​the image to be detected, processes the processing area with the corresponding model to obtain a feature map, and identifies the image defects of the image to be detected based on the feature map, thereby achieving the effect of using a suitable model to process the processing area of ​​the image to be detected, and improving the accuracy of defect identification of the image to be detected.

[0076] For other examples of this embodiment, please refer to the above examples and will not be repeated here.

[0077] This embodiment also provides a defect detection system, which includes devices or modules, and implements the above-mentioned defect detection method through a combination of devices or modules. Specific examples can be found in the above examples, which will not be repeated here.

[0078] like Figure 5 As shown, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114. Memory 113, for storing computer programs; In one embodiment of the present application, the processor 111 is configured to implement the defect detection method provided by any one of the aforementioned method embodiments when executing a program stored in the memory 113 .

[0079] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the defect detection method provided in any of the aforementioned method embodiments is implemented.

[0080] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0081] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a general hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0082] It should be understood that the terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "comprise", "include", "contain" and "have" are inclusive and therefore specify the presence of stated features, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the specific order described or illustrated, unless the order of execution is clearly indicated. It should also be understood that additional or alternative steps may be used.

[0083] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A defect detection method, characterized in that: include: When an image to be detected is obtained, an intermediate layer vector of the image to be detected is obtained; determining a processing strategy for the image to be detected based on the intermediate layer vector and the historical processing strategy of the historical image, wherein the processing strategy includes a model for processing the image to be detected and a processing area of ​​the model; Processing the processing area of ​​the image to be detected using a model corresponding to the processing strategy to obtain a feature map; According to the feature map, image defects of the image to be detected are identified.

2. The method according to claim 1, characterized in that The step of processing the processing area of ​​the image to be detected using the model corresponding to the processing strategy to obtain a feature map includes: When the processing strategy is to use a global model to process the first processing area of ​​the image to be detected, the first processing area of ​​the image to be detected is input into the global model to obtain the feature map, wherein the model weight of the global model is converted from a first floating point number to a first integer.

3. The method according to claim 1, characterized in that The step of processing the processing area of ​​the image to be detected using the model corresponding to the processing strategy to obtain a feature map includes: When the processing strategy is to use a local attention model to process the second processing area of ​​the image to be detected, the second processing area of ​​the image to be detected is input into the local attention model to obtain the feature map, wherein the model weight of the local attention model is converted from a second floating point number to a half-precision floating point number.

4. The method according to claim 1, wherein The step of processing the processing area of ​​the image to be detected using the model corresponding to the processing strategy to obtain a feature map includes: When the processing strategy is to use a global model to process the third processing region of the image to be detected and to use a local attention model to process the fourth processing region of the image to be detected, the third processing region of the image to be detected is input into the global model and the fourth processing region of the image to be detected is input into the local attention model; The global features output by the global model are combined with the local features output by the local attention model to obtain the feature map.

5. The method according to claim 4, characterized in that Combining the global features output by the global model with the local features output by the local attention model to obtain the feature map includes: Performing upsampling and spatial mapping operations on the global features and the local features and then performing spatial alignment; Splicing the aligned global features and the local features into a multi-scale combined feature; Inputting the multi-scale combined features into a deformable convolution layer, and performing pixel misalignment compensation on the multi-scale combined features through the deformable convolution layer; Adaptively sampling the compensated multi-scale combined features to obtain the feature map.

6. The method according to claim 1, characterized in that Determining the processing strategy of the image to be detected based on the intermediate layer vector and the historical processing strategy of the historical image includes: Determining a historical intermediate layer output vector of the historical image according to the historical image; Determining a target historical image from the historical images based on a similarity between an intermediate layer vector of the image to be detected and a historical intermediate layer output vector of the historical images; The processing strategy of the target historical image is determined as the processing strategy of the image to be detected.

7. The method according to claim 1, characterized in that Determining the processing strategy of the image to be detected based on the intermediate layer vector and the historical processing strategy of the historical image includes: In the case where the processing strategy is not determined, the image to be detected is input into the strategy network, the strategy network extracts the intermediate layer vector of the image to be detected, and predicts the processing strategy according to the intermediate layer vector.

8. A defect detection device, characterized in that: include: An acquisition module, configured to acquire an intermediate layer vector of an image to be detected when an image to be detected is acquired; a determination module, configured to determine a processing strategy for the image to be detected based on the intermediate layer vector and a historical processing strategy of historical images, wherein the processing strategy includes a model for processing the image to be detected and a processing region of the model; a processing module, configured to process the processing area of ​​the image to be detected using a model corresponding to the processing strategy to obtain a feature map; The recognition module is used to recognize image defects of the image to be detected based on the feature map.

9. An electronic device, characterized in that: include: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor coupled to the at least one bus; At least one memory connected to the at least one bus, wherein a computer program is stored in the memory, and when the processor executes the computer program, the defect detection method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The storage medium stores computer-executable instructions, and the computer-executable instructions are used to execute the defect detection method described in any one of claims 1 to 7 of the present application.

Citation Information

Patent Citations

  • Defect identification method and device, computer equipment and computer readable storage medium

    CN114627089A

  • Image defect detection method and device, medium and electronic equipment

    CN118154524A

  • Optical lens defect detection method and system

    CN118587175A

  • Vision-based surface defect detection method and detection system

    CN119151938A

  • Method for detecting display screen quality, apparatus, electronic device and storage medium

    US20200357109A1