Image detection method and device, program product and storage medium

The sliding window technology extracts local image information and performs multi-scale detection and fusion, which solves the problem of increasing computing and storage requirements caused by direct input of high-resolution images, and achieves efficient image defect detection.

CN120031875AActive Publication Date: 2025-05-23SUZHOU INS IMAGE SOFTWARE TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510503583.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-23
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

In the prior art In image defect detection, high-resolution images are directly input into the model, resulting in increased computing burden and storage requirements, making it difficult to run in resource-constrained devices or real-time applications.

Method used

The local image information is extracted through sliding window technology, decomposed into multiple window images, and defect detection is performed through multi-scale detection module and multi-scale fusion module to avoid direct input of high-resolution images into the model.

Benefits of technology

It effectively reduces the memory usage, improves the inference efficiency and accuracy of the image defect detection model, and is suitable for resource-constrained devices and real-time applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031875A_ABST
    Figure CN120031875A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image detection method and device, a program product and a storage medium, and the method comprises the steps: obtaining a to-be-processed image, and determining the information of a sliding window according to the to-be-processed image; establishing each sliding window based on the sliding window information, and performing image extraction on the to-be-processed image according to a preset sliding direction through each sliding window to obtain a window image of each sliding window; inputting the window image into a predetermined target defect detection model, and performing image defect detection on the window image through a multi-scale detection module and a multi-scale fusion module of the target defect detection model to obtain defect detection data corresponding to the window image; and obtaining a defect detection result of the to-be-processed image based on the defect detection data, and displaying the defect detection result to a user. According to the method, large or small target defect detection can be accurately carried out on the to-be-processed image while equipment computing resources are saved, and the accuracy of image defect detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of image processing technology, and in particular to an image detection method, device, program product and storage medium. Background Art

[0002] With the development of industrial automation and intelligence, the requirements for image defect detection are getting higher and higher. In many fields, such as semiconductor manufacturing, power inspection and material testing, high-resolution images are widely used for defect detection.

[0003] The current image detection method for small defects is to first increase the resolution of the image to be processed, thereby retaining more detailed information of the image to be processed, which helps the defect detection model to capture the characteristics of small defects more clearly, and then use the defect detection model to detect defects on the image to be processed with increased resolution. In this way, the direct input of high-resolution images will significantly increase the computational burden and storage requirements of the model, making it difficult for the model to run on resource-constrained embedded devices or real-time applications. In addition, the model needs to process larger input feature maps, and the training time and video memory requirements will also increase significantly. Summary of the invention

[0004] The embodiments of the present invention provide an image detection method, device, program product and storage medium. On the basis of ensuring the accuracy of image defect detection, high-resolution images do not need to be directly input into the defect detection model, which effectively reduces the occupancy of video memory and improves the reasoning efficiency of the image defect detection model.

[0005] In a first aspect, an embodiment of the present invention provides an image detection method, comprising:

[0006] Acquire an image to be processed, and determine sliding window information according to the image to be processed;

[0007] Establishing each sliding window based on the sliding window information, and performing image extraction on the image to be processed through each sliding window according to a preset sliding direction to obtain a window image of each sliding window;

[0008] Inputting the window image into a predetermined target defect detection model, performing image defect detection on the window image through a multi-scale detection module and a multi-scale fusion module of the target defect detection model, and obtaining defect detection data corresponding to the window image;

[0009] A defect detection result of the image to be processed is obtained based on the defect detection data, and the defect detection result is displayed to a user.

[0010] In a second aspect, an embodiment of the present invention provides an image detection device, the device comprising:

[0011] A data acquisition module, used for acquiring an image to be processed and determining sliding window information according to the image to be processed;

[0012] A first processing module, configured to establish each sliding window based on the sliding window information, and extract the image to be processed through each sliding window according to a preset sliding direction to obtain a window image of each sliding window;

[0013] A second processing module is used to input the window image into a predetermined target defect detection model, perform image defect detection on the window image through a multi-scale detection module and a multi-scale fusion module of the target defect detection model, and obtain defect detection data corresponding to the window image;

[0014] The result determination module is used to obtain the defect detection result of the image to be processed based on the defect detection data, and display the defect detection result to the user.

[0015] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, an image detection method as described in any one of the embodiments of the present invention is implemented.

[0016] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image detection method as described in any one of the embodiments of the present invention.

[0017] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program, which, when executed by a processor, implements the image detection method as described in any one of the embodiments of the present invention.

[0018] In the embodiment of the present invention, an image to be processed is obtained, and sliding window information is determined according to the image to be processed; each sliding window is established based on the sliding window information, and image extraction is performed on the image to be processed according to a preset sliding direction through each sliding window to obtain a window image of each sliding window; the window image is input into a predetermined target defect detection model, and image defect detection is performed on the window image through a multi-scale detection module and a multi-scale fusion module of the target defect detection model to obtain defect detection data corresponding to the window image; a defect detection result of the image to be processed is obtained based on the defect detection data, and the defect detection result is displayed to the user. The method of the embodiment of the present invention does not need to directly input a high-resolution image into the target defect detection model, but accurately and comprehensively extracts image information of the image to be processed through multiple sliding windows. By determining the sliding window information according to the image to be processed, the diversity and comprehensiveness of the information input into the target defect detection model are guaranteed. While saving equipment computing resources, through the multi-scale detection module and the multi-scale fusion module, large or small target defect detection can be performed on the image to be processed, thereby improving the accuracy and efficiency of image defect detection, and at the same time improving the reasoning efficiency of the image defect detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 A first flow chart of an image detection method provided by an embodiment of the present invention;

[0021] Figure 2 A schematic diagram of extracting local information of an image by using a sliding window provided in an embodiment of the present invention;

[0022] Figure 3 A second flow chart of an image detection method provided by an embodiment of the present invention;

[0023] Figure 4 A schematic diagram of the structure of a target defect detection model provided by an embodiment of the present invention;

[0024] Figure 5 A schematic diagram of the structure of a multi-scale fusion module provided in an embodiment of the present invention;

[0025] Figure 6 A flowchart of a model training method provided by an embodiment of the present invention;

[0026] Figure 7A schematic diagram of the structure of an image detection device provided by an embodiment of the present invention;

[0027] Figure 8 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.

[0029] Figure 1 The first flow chart of an image detection method provided by an embodiment of the present invention. The method of the embodiment of the present invention can ensure the accuracy of image defect detection without directly inputting high-resolution images into the defect detection model, effectively reducing the usage of video memory and improving the efficiency of image defect detection. The method can be executed by an image detection device provided by an embodiment of the present invention, and the device can be implemented in software and / or hardware. The following embodiments will be described by taking the device integrated in an electronic device as an example. The electronic device can be a server or a computer device, etc. Figure 1 , the method may specifically include the following steps:

[0030] Step 101: Obtain an image to be processed, and determine sliding window information according to the image to be processed.

[0031] The sliding window information includes the number of sliding windows and the size of each sliding window. The image to be processed is an image for which target defect detection is required. The sliding window is used to extract local information of the image to be processed. Different numbers and sizes of sliding windows can be set for different images to be processed.

[0032] Specifically, when it is necessary to perform target defect detection on the image to be processed, the server can obtain the image to be processed and the required information uploaded by the user, and determine the number of sliding windows and the size of each sliding window according to the size of the image to be processed and the required information. In order to capture information of different scales, multiple sliding windows of different sizes can be generated by expansion and equal division. The sliding windows of this scheme are all square. In the embodiment of this scheme, optionally, the sliding window information is determined according to the image to be processed, including: determining the minimum side length of the image to be processed as the side length of the maximum sliding window; determining the side lengths of the remaining sliding windows based on the predetermined side length of the minimum sliding window, the predetermined number of sliding windows, and the side length of the maximum sliding window.

[0033] Among them, multiple sliding windows of different scales include a maximum sliding window of the largest size, a minimum sliding window of the smallest size, and several intermediate candidate windows (the remaining sliding windows). In order to ensure comprehensive coverage of defect information at different scales, the side length of the maximum sliding window can be set to the minimum side length of the image to be processed. The server can pre-determine the side length of the minimum sliding window and the number of sliding windows based on domain big data, etc., and of course, it can also be dynamically determined based on the specific needs of the user and the size of the image to be processed.

[0034] In an optional implementation, after determining the size of the maximum sliding window, the size of the minimum sliding window, and the number of sliding windows, the sizes of the remaining sliding windows may be determined by scaling in proportion. Alternatively, the sizes of the remaining sliding windows may be determined by dividing at equal intervals: , . represents the side length of the maximum sliding window, represents the side length of the minimum sliding window, i represents the i-th window, represents the side length of the i-th window, and N represents the number of sliding windows (including the maximum sliding window and the minimum sliding window). Through the above steps, sliding windows of different sizes can be accurately set according to the image to be processed, so that the information of the image to be processed can be accurately and comprehensively extracted through multiple sliding windows.

[0035] Step 102: Establish each sliding window based on the sliding window information, and use each sliding window to extract the image to be processed according to a preset sliding direction to obtain a window image of each sliding window.

[0036] Among them, the window image is a partial image of the image to be processed extracted by the sliding window. Specifically, after determining the number of sliding windows and the size of each sliding window, arrange each sliding window according to the size of the sliding window. Determine the sliding step length according to the side length of the minimum sliding window, place the center point of the minimum sliding window on the initial sliding point, slide each sliding window according to a preset sliding direction (for example, from left to right and from top to bottom) and a sliding step length, and obtain the window image of each sliding window corresponding to each sliding step length. In the embodiment of this scheme, optionally, extract the image of the image to be processed by each sliding window according to a preset sliding direction to obtain the window image of each sliding window, including the following steps A1-A2:

[0037] Step A1: Determine the center point of the minimum sliding window as the center point of all sliding windows, and arrange the sliding windows according to the center points of all sliding windows.

[0038] Specifically, after determining the size of each sliding window, the center point of the smallest sliding window is determined as the center point of all sliding windows. The sliding windows are arranged in sequence in a manner of expanding outward from the smallest sliding window until the largest sliding window is arranged. For example, Figure 2 A schematic diagram of using a sliding window to extract local information of an image provided by an embodiment of the present invention. Figure 2 As shown, Figure 2 It includes the image to be processed and 4 sliding windows. Figure 2 The black dotted box in the figure represents the sliding window. The sliding window with the smallest size is the minimum sliding window. The center of the minimum sliding window is used as the center of all sliding windows. All sliding windows are arranged so that the side length of the maximum sliding window is the same as the minimum side length of the image to be processed.

[0039] Step A2: The side length of the minimum sliding window is determined as the sliding step length; the sliding initial point and the sliding end point are determined according to the image to be processed and the minimum sliding window, and each sliding window is slid according to the sliding direction and the sliding step length to obtain the window image extracted by each sliding window corresponding to each step length.

[0040] Among them, the sliding step size specifies the distance that the sliding window slides each time. The sliding initial point is the position of the center point of the sliding window when the sliding window starts sliding, and the sliding end point is the position of the center point of the sliding window when the sliding window ends sliding. The sliding direction is generally from left to right and from top to bottom. The sliding initial point is the point that makes the minimum sliding window located in the upper left corner of the image. After determining the sliding initial point, the minimum sliding window is placed in the upper left corner of the image to be processed, and the area with the same length and width as the input high-resolution image is expanded with the minimum sliding window as the center (the remaining sliding windows and the maximum sliding window), and the insufficient area is automatically filled with 0. Slide each sliding window according to the sliding direction and sliding step size to obtain the window image extracted by each sliding window corresponding to each step size. For example Figure 2 As shown, Figure 2 There are 4 sliding windows in it, and the size of the smallest sliding window can be All sliding windows are slid with a sliding step of 640, and each sliding can generate 4 window images of different scales.

[0041] This solution processes high-resolution images to be processed from top to bottom and from left to right through a sliding window, ensuring comprehensive coverage of defect information of the images to be processed at different scales and ensuring the diversity and comprehensiveness of the information input into the target defect detection model.

[0042] Step 103: input the window image into a predetermined target defect detection model, perform image defect detection on the window image through a multi-scale detection module and a multi-scale fusion module of the target defect detection model, and obtain defect detection data corresponding to the window image.

[0043] Among them, the target defect detection model is predetermined and is used to perform target defect detection on the image to be processed according to the image information extracted by the sliding window. The target defect model includes a multi-scale detection module and a multi-scale fusion module. The multi-scale detection module is used to extract features from the window image, and the multi-scale fusion module is used to perform feature fusion and defect detection on the features output by the multi-scale detection module to obtain defect detection data.

[0044] Specifically, after obtaining N (number of windows) window images of different scales generated M (number of slides) times, the N window images of different scales are processed into image features of the same size in space (for example ). Furthermore, the window images can be grouped according to their different sizes (the original size of the window images before being processed), and the window images can be input into the multi-scale detection module according to the groups. The window images of each group can be independently convolved through the multi-scale detection module, and the feature information of different groups does not interfere with each other. In this way, it is avoided that inputs of different scales need to be traversed and trained in sequence, thereby increasing the training time and inference time of the target defect detection model, and improving the real-time performance of the target defect detection model. Through the multi-scale detection module, the convolved features are fused at different scales (in this scheme, the scales are divided according to the size of the image features) to obtain the detection features output by the multi-scale detection module. After obtaining the detection features, the detection features are input into the multi-scale fusion module, and the multi-scale fusion module is used to perform feature fusion and defect detection on the detection features to obtain the defect detection data corresponding to the window image. This scheme uses a multi-scale sliding window so that the target defect detection model only needs to process 4 blocks. The input of the high-resolution image does not need to be directly input into the model, which effectively reduces the memory usage, making it possible to handle the small defect detection problem of high-resolution images on a commonly used computer with 6G video memory. Taking the high-resolution image of as an example, when the training step is 1, directly training the model needs to consume 22930M (megabytes) of video memory, while the method of this solution only needs to use 2550M video memory, reducing 88% of video memory computing resources.

[0045] In an optional implementation, for the window images extracted by each sliding window, all current window images extracted by the current sliding window are determined as the current feature group; the current feature group is convolved by the multi-scale detection module to obtain the feature tensor of the current feature group; the feature tensor of the current feature group is feature fused according to the feature tensors of other feature groups to obtain the fusion feature of the current feature group, and the fusion feature of all feature groups is determined as the detection feature output by the multi-scale detection module. The detection feature is input into the multi-scale fusion module, and the detection feature is spatially transformed based on the predetermined scaling, offset and spatial transformation unit to obtain the transformation feature of each channel group corresponding to the detection feature; the transformation feature of each channel group is feature spliced ​​and convolved to obtain the defect detection data corresponding to the window image.

[0046] Step 104: Obtain a defect detection result of the image to be processed based on the defect detection data, and display the defect detection result to the user.

[0047] Among them, the defect detection results include the image to be processed with defects marked and the defect analysis results (such as defect severity and defect suggestions). Specifically, after obtaining the defect detection data output by the target defect detection model, a rectangular box, a circle or other shapes are drawn at the defect on the image to be processed according to the defect detection data to obtain the image to be processed with marked defects. The defect detection data is analyzed, and the defect type and confidence of the image to be processed are determined according to the pre-defined defect types, and it is sorted with the image to be processed with marked defects to obtain the defect detection results. The defect detection results are displayed to the user so that the user can intuitively see the defects in the image to be processed, and the defect detection results are further processed and analyzed as needed.

[0048] The technical solution of this embodiment is to obtain the image to be processed, and determine the sliding window information according to the image to be processed; establish each sliding window based on the sliding window information, and extract the image to be processed according to the preset sliding direction through each sliding window to obtain the window image of each sliding window; input the window image into a predetermined target defect detection model, and perform image defect detection on the window image through the multi-scale detection module and the multi-scale fusion module of the target defect detection model to obtain the defect detection data corresponding to the window image; obtain the defect detection result of the image to be processed based on the defect detection data, and display the defect detection result to the user. The technical solution of this embodiment does not need to directly input the high-resolution image into the target defect detection model, but accurately and comprehensively extracts the image information of the image to be processed through multiple sliding windows. By determining the sliding window information according to the image to be processed, the diversity and comprehensiveness of the information input into the target defect detection model are guaranteed. While saving the computing resources of the equipment, the multi-scale detection module and the multi-scale fusion module are used to accurately perform large or small target defect detection on the image to be processed, thereby improving the accuracy and efficiency of image defect detection.

[0049] Figure 3 This is a second flow chart of an image detection method provided by an embodiment of the present invention. This embodiment is a refinement based on the above embodiment. The specific method can be as follows Figure 3 As shown, the method may include the following steps:

[0050] Step 301: Acquire an image to be processed, and determine sliding window information according to the image to be processed.

[0051] The sliding window information includes the number of sliding windows and the size of each sliding window; each sliding window is a square window.

[0052] Step 302: Establish each sliding window based on the sliding window information, and use each sliding window to extract the image to be processed according to a preset sliding direction to obtain a window image of each sliding window.

[0053] Step 303: input the window image to the multi-scale detection module, and perform grouping processing on the window image through the multi-scale detection module to obtain detection features output by the multi-scale detection module.

[0054] The target defect detection model is predetermined by the server and is used to perform target defect detection on the image to be processed based on the image information extracted by the sliding window. The target defect model includes a multi-scale detection module and a multi-scale fusion module. The multi-scale detection module is used to extract features from the window image. In this solution, optionally, the window image is grouped and processed by the multi-scale detection module to obtain the detection features output by the multi-scale detection module, including the following steps B1-B2:

[0055] Step B1: for the window images extracted by each sliding window, all current window images extracted by the current sliding window are determined as the current feature group; the current feature group is convolved through the multi-scale detection module to obtain the feature tensor of the current feature group.

[0056] Specifically, after obtaining the window images generated by all sliding windows, all window images are grouped according to different window sizes to obtain feature groups. For example, Figure 2 As shown in FIG. 1 , there are 4 sliding windows of different sizes, and the number of feature groups is 4. The 4 feature groups are input into the multi-scale detection module, and the multi-scale detection module is used to convolve the 4 feature groups to obtain the feature tensor corresponding to each feature group.

[0057] Step B2: Perform feature fusion on the feature tensor of the current feature group according to the feature tensors of other feature groups to obtain the fused features of the current feature group, and determine the fused features of all feature groups as the detection features output by the multi-scale detection module.

[0058] After obtaining the feature tensors of each feature group, feature fusion is performed on each feature tensor and the feature tensors of other groups respectively to obtain the fused features corresponding to each feature group, and the fused features of all feature groups are determined as the detection features output by the multi-scale detection module.

[0059] For example, Figure 4 This is a schematic diagram of the structure of the target defect detection model provided by an embodiment of the present invention. Figure 4 As shown, the size of the current window image input (Input) to the multi-scale detection module of the target detection model is BatchSize (the number of samples simultaneously input into the network during a training process) = 1, Channel (number of channels) = 3N (N = 4), Height (height) = 640, Width (width) = 640. It is convolved through the multi-scale detection module to obtain the feature tensors F3, F4 and F5 of the current window image. Among them, the size of F3 is 1,32N,80,80; the size of F4 is 1,64N,40,40; the size of F5 is 1,128N,20,20. After obtaining F3, F4 and F5, their features are fused to obtain fused features O3, O4 and O5. The size of O3 is 1,32N,80,80; the size of O4 is 1,64N,40,40; the size of O5 is 1,128N,20,20. Among them Figure 4CT in stands for connection layer, which is used to splice two or more feature maps along a specific dimension. C3 represents the cross-stage partial network, which is used to reduce the amount of calculation and improve the learning ability of the model. O3, O4 and O5 are input into MSFF (multi-scale fusion module) to obtain the fused features O3', O4' and O5'. Finally, target defect detection (i.e. Figure 4 in the Detect). Figure 4 CONV in the figure stands for convolution, and UP stands for upsampling. Through the multi-scale detection module, independent convolution processing is performed on inputs of different scales, which reduces the interference of feature information and improves the efficiency of training and inference.

[0060] Step 304: input the detection features into a multi-scale fusion module, and perform multi-scale fusion on the detection features through the multi-scale fusion module to obtain defect detection data corresponding to the window image.

[0061] The multi-scale fusion module is used to perform feature fusion on the features output by the multi-scale detection module to obtain defect detection data. In this solution, the detection features are multi-scale fused by the multi-scale fusion module to obtain defect detection data corresponding to the window image, including: inputting the detection features into the multi-scale fusion module, calculating the scaling scale and offset corresponding to the detection features, and performing spatial transformation on the detection features based on the scaling scale, offset and spatial transformation unit to obtain the transformation features of each channel group corresponding to the detection features; performing feature splicing and convolution on the transformation features of each channel group to obtain the defect detection data corresponding to the window image.

[0062] The scale is a pre-set parameter for enlarging or reducing the detection feature; the offset is a pre-set parameter for translating the detection feature, which is used to adjust the position of the detection feature. Specifically, in the standard convolution operation, the convolution kernel is calculated on all input channels and then the results are accumulated in the output channel. In this scheme, the input channels are divided into multiple groups, and each group is calculated using an independent convolution kernel. Figure 5 Schematic diagram of the structure of the multi-scale fusion module provided in the embodiment of the present invention. Figure 5 As shown in the figure, after the input of the multi-scale fusion module receives the detection features, it calculates the scale and offset of the detection features of different scales (sizes) relative to the minimum sliding window. Taking N=4 as an example, the detection features are input into the spatial transformer network (STN) through the channel according to the scale and offset, the features output by the STN unit are spliced ​​together, and then the spliced ​​features are convolved to finally obtain defect detection data with full fusion of different scales. Figure 5Concat in stands for concatenation, and CONV stands for convolution. By spatially transforming and concatenating channel features of different scales, it is ensured that the final feature map can fully integrate context information and provide comprehensive feature support for image defect detection.

[0063] Step 305: Obtain a defect detection result of the image to be processed based on the defect detection data, and display the defect detection result to the user.

[0064] In the technical solution of this embodiment, an image to be processed is obtained, and sliding window information is determined according to the image to be processed. The sliding window information includes the number of sliding windows and the size of each sliding window; each sliding window is a square window. Each sliding window is established based on the sliding window information, and the image to be processed is extracted by each sliding window according to a preset sliding direction to obtain the window image of each sliding window. The window image is input into the multi-scale detection module, and the window image is grouped and processed by the multi-scale detection module to obtain the detection features output by the multi-scale detection module. The detection features are input into the multi-scale fusion module, and the detection features are multi-scale fused by the multi-scale fusion module to obtain the defect detection data corresponding to the window image. The defect detection result of the image to be processed is obtained based on the defect detection data, and the defect detection result is displayed to the user. The technical solution of this embodiment, through the multi-scale detection module and the multi-scale fusion module, can better learn context information, effectively improve the robustness and accuracy of defect detection in high-resolution images. In addition, for inputs of different scales, it is not necessary to traverse the loop training but only need to be trained once, which improves the training and reasoning efficiency of the model.

[0065] Figure 6 Flow chart of the model training method provided by the embodiment of the present invention. The specific method can be as follows Figure 6 As shown, the method may include the following steps:

[0066] Step 601: When the pre-established initial defect detection model does not meet the preset conditions, a training image set is obtained.

[0067] Among them, the initial defect detection model is a model that is pre-established and has not been trained, and cannot accurately detect target defects in high-resolution images. The preset conditions are pre-determined based on domain big data, etc., and are used to determine whether the initial defect detection model has completed training, such as whether the number of iterations of model training has reached the preset number, whether the loss function has reached convergence, etc. The training image set is used to train the initial defect target model. The training image set includes a labeled image data set, and the image set contains normal images and defective images. When the initial defect model needs to be trained, the server can obtain the training image set sent by the user.

[0068] Step 602: Input the training image set into the initial defect detection model to obtain output data of the initial defect detection model. Optimize the parameters of the initial defect detection model based on the preset loss function and output data to obtain a target defect detection model that meets preset conditions.

[0069] Specifically, after obtaining the training image set, the images in the training image set are input into the initial defect model to obtain the output data of the initial defect model. The difference between the output data and the actual label is calculated using the loss function of the following formula:

[0070] ;

[0071] in, Represent the corresponding loss weights respectively. . is the target box localization loss, is the target confidence loss, is the classification loss. The model parameters are adjusted according to the value of the loss function to obtain the target defect detection model that meets the preset conditions. In this scheme, the initial learning rate is set to 0.001.

[0072] The technical solution of this embodiment is to obtain a training image set when the pre-established initial defect detection model does not meet the preset conditions; input the training image set into the initial defect detection model to obtain the output data of the initial defect detection model; optimize the parameters of the initial defect detection model based on the pre-set loss function and output data to obtain a target defect detection model that meets the preset conditions. The technical solution of this embodiment significantly improves the performance of the defect detection model by guiding the model to perform parameter optimization through the loss function, so that it can more accurately detect target defects in high-resolution images. It not only improves the accuracy and generalization ability of the model, but also provides flexibility and targeted training capabilities.

[0073] Figure 7 FIG. 1 is a schematic diagram of the structure of an image detection device provided in an embodiment of the present invention, and the device is suitable for executing the image detection method provided in an embodiment of the present invention. Figure 7 As shown, the device may specifically include:

[0074] The data acquisition module 701 is used to acquire the image to be processed and determine the sliding window information according to the image to be processed;

[0075] A first processing module 702 is used to establish each sliding window based on the sliding window information, and extract the image to be processed through each sliding window according to a preset sliding direction to obtain a window image of each sliding window;

[0076] The second processing module 703 is used to input the window image into a predetermined target defect detection model, perform image defect detection on the window image through a multi-scale detection module and a multi-scale fusion module of the target defect detection model, and obtain defect detection data corresponding to the window image;

[0077] The result determination module 704 is used to obtain the defect detection result of the image to be processed based on the defect detection data, and display the defect detection result to the user.

[0078] Optionally, the sliding window information includes the number of sliding windows and the size of each sliding window; the sliding window is a square, and the first processing module 702 is specifically used to: determine the minimum side length of the image to be processed as the side length of the maximum sliding window;

[0079] The side lengths of the remaining sliding windows are determined based on the predetermined side length of the minimum sliding window, the predetermined number of sliding windows, and the side length of the maximum sliding window.

[0080] Optionally, the first processing module 702 is further configured to: determine the center point of the minimum sliding window as the center point of all sliding windows, and arrange the sliding windows according to the center points of all sliding windows;

[0081] Determine the side length of the minimum sliding window as the sliding step length;

[0082] The sliding initial point and the sliding end point are determined according to the image to be processed and the minimum sliding window, and the sliding windows are slid according to the sliding direction and the sliding step length to obtain the window images extracted by the sliding windows corresponding to each step length.

[0083] Optionally, the second processing module 703 is specifically used to: input the window image to the multi-scale detection module, perform grouping processing on the window image through the multi-scale detection module, and obtain the detection features output by the multi-scale detection module;

[0084] The detection features are input into the multi-scale fusion module, and the detection features are multi-scale fused by the multi-scale fusion module to obtain defect detection data corresponding to the window image.

[0085] Optionally, the second processing module 703 is further configured to: for the window images extracted by each sliding window, determine all current window images extracted by the current sliding window as a current feature group;

[0086] Convolving the current feature group through the multi-scale detection module to obtain a feature tensor of the current feature group;

[0087] The feature tensors of the current feature group are subjected to feature fusion according to the feature tensors of other feature groups to obtain fused features of the current feature group, and the fused features of all feature groups are determined as the detection features output by the multi-scale detection module.

[0088] Optionally, the second processing module 703 is further used to: input the detection feature into the multi-scale fusion module, calculate the scaling scale and offset corresponding to the detection feature, and perform spatial transformation on the detection feature based on the scaling scale, the offset and the spatial transformation unit to obtain transformation features of each channel group corresponding to the detection feature;

[0089] Feature splicing and convolution are performed on the transformed features of each channel group to obtain defect detection data corresponding to the window image.

[0090] Optionally, before acquiring the image to be processed, the second processing module 703 is further used to: when the pre-established initial defect detection model does not meet the preset conditions, acquire a training image set;

[0091] Inputting the training image set into the initial defect detection model to obtain output data of the initial defect detection model;

[0092] The initial defect detection model is optimized based on a preset loss function and the output data to obtain a target defect detection model that meets the preset conditions.

[0093] The image detection device provided in the embodiment of the present invention can execute the image detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method. For the contents not described in detail in this embodiment, reference can be made to the description in any method embodiment of the present invention.

[0094] An embodiment of the present invention also provides a computer program product.

[0095] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer program products, which can include one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor, which can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0096] Figure 8 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, referring to Figure 8 , Figure 8 The electronic device 12 shown is only an example and should not limit the functions and scope of use of the embodiments of the present application. Figure 8 As shown, the electronic device 12 is in the form of a general purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that connects various system components (including the system memory 28 and the processing unit 16).

[0097] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor or a local bus using any of a variety of bus architectures. By way of example, these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0098] The electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.

[0099] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Figure 8 not shown, usually called a "hard drive"). Although Figure 8 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present application.

[0100] A program / utility 40 having a set (at least one) of program modules 46 may be stored, for example, in the memory 28, such program modules 46 including but not limited to an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. The program modules 46 generally perform the functions and / or methods of the embodiments described herein.

[0101] The electronic device 12 may also communicate with one or more external devices 14 (e.g., keyboards, pointing devices, displays 24, etc.), may communicate with one or more devices that enable a user to interact with the electronic device 12, and / or may communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface 22. Furthermore, the electronic device 12 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with other modules of the electronic device 12 via a bus 18. It should be understood that although Figure 8 Not shown, other hardware and / or software modules may be used in conjunction with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0102] The processing unit 16 executes various functional applications and data processing by running the program stored in the system memory 28, for example, implementing an image detection method provided in an embodiment of the present invention: acquiring an image to be processed, and determining sliding window information according to the image to be processed; establishing each sliding window based on the sliding window information, and performing image extraction on the image to be processed according to a preset sliding direction through each sliding window to obtain a window image of each sliding window; inputting the window image into a predetermined target defect detection model, performing image defect detection on the window image through a multi-scale detection module and a multi-scale fusion module of the target defect detection model, and obtaining defect detection data corresponding to the window image; obtaining a defect detection result of the image to be processed based on the defect detection data, and displaying the defect detection result to a user.

[0103] The embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, an image detection method provided by all the embodiments of the present invention is implemented: obtaining an image to be processed, and determining sliding window information according to the image to be processed; establishing each sliding window based on the sliding window information, and performing image extraction on the image to be processed according to a preset sliding direction through each sliding window to obtain the window image of each sliding window; inputting the window image into a predetermined target defect detection model, performing image defect detection on the window image through a multi-scale detection module and a multi-scale fusion module of the target defect detection model, and obtaining defect detection data corresponding to the window image; obtaining a defect detection result of the image to be processed based on the defect detection data, and displaying the defect detection result to a user. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic device, device or device of electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection with one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction-executing electronic device, apparatus, or device.

[0104] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction-executing electronic device, apparatus, or device.

[0105] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0106] Computer program code for performing the operation of the present invention may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0107] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. An image detection method, characterized in that: The method comprises: Acquire an image to be processed, and determine sliding window information according to the image to be processed; Establishing each sliding window based on the sliding window information, and performing image extraction on the image to be processed through each sliding window according to a preset sliding direction to obtain a window image of each sliding window; Inputting the window image into a predetermined target defect detection model, performing image defect detection on the window image through a multi-scale detection module and a multi-scale fusion module of the target defect detection model, and obtaining defect detection data corresponding to the window image; A defect detection result of the image to be processed is obtained based on the defect detection data, and the defect detection result is displayed to a user.

2. The method according to claim 1, characterized in that The sliding window information includes the number of sliding windows and the size of each sliding window; the sliding window is a square, and the sliding window information is determined according to the image to be processed, including: Determine the minimum side length of the image to be processed as the side length of the maximum sliding window; The side lengths of the remaining sliding windows are determined based on the predetermined side length of the minimum sliding window, the predetermined number of sliding windows, and the side length of the maximum sliding window.

3. The method according to claim 2, characterized in that The image to be processed is extracted by each sliding window according to a preset sliding direction to obtain a window image of each sliding window, including: Determine the center point of the minimum sliding window as the center point of all sliding windows, and arrange the sliding windows according to the center points of all sliding windows; Determine the side length of the minimum sliding window as the sliding step length; The sliding initial point and the sliding end point are determined according to the image to be processed and the minimum sliding window, and the sliding windows are slid according to the sliding direction and the sliding step length to obtain the window images extracted by the sliding windows corresponding to each step length.

4. The method according to claim 1, characterized in that: The window image is input into a predetermined target defect detection model, and image defect detection is performed on the window image through a multi-scale detection module and a multi-scale fusion module of the target defect detection model to obtain defect detection data corresponding to the window image, including: Inputting the window image into the multi-scale detection module, performing grouping processing on the window image through the multi-scale detection module, and obtaining detection features output by the multi-scale detection module; The detection features are input into the multi-scale fusion module, and the detection features are multi-scale fused by the multi-scale fusion module to obtain defect detection data corresponding to the window image.

5. The method according to claim 4, characterized in that Inputting the window image into the multi-scale detection module, performing grouping processing on the window image by the multi-scale detection module, and obtaining detection features output by the multi-scale detection module, including: For the window images extracted by each sliding window, all current window images extracted by the current sliding window are determined as a current feature group; Convolving the current feature group through the multi-scale detection module to obtain a feature tensor of the current feature group; The feature tensors of the current feature group are subjected to feature fusion according to the feature tensors of other feature groups to obtain fused features of the current feature group, and the fused features of all feature groups are determined as the detection features output by the multi-scale detection module.

6. The method according to claim 4, characterized in that The multi-scale fusion module includes a spatial transformation unit, and the detection features are input into the multi-scale fusion module. The multi-scale fusion module performs multi-scale fusion on the detection features to obtain defect detection data corresponding to the window image, including: Inputting the detection feature into the multi-scale fusion module, calculating the scaling scale and offset corresponding to the detection feature, and performing spatial transformation on the detection feature based on the scaling scale, the offset and the spatial transformation unit to obtain transformation features of each channel group corresponding to the detection feature; Feature splicing and convolution are performed on the transformed features of each channel group to obtain defect detection data corresponding to the window image.

7. The method according to claim 1, characterized in that Before acquiring the image to be processed, the method further includes: When the pre-established initial defect detection model does not meet the preset conditions, a training image set is obtained; Inputting the training image set into the initial defect detection model to obtain output data of the initial defect detection model; The initial defect detection model is optimized based on a preset loss function and the output data to obtain a target defect detection model that meets the preset conditions.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements an image detection method according to any one of claims 1 to 7.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the image detection method as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the image detection method as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method and system for online detection of part defects in additive manufacturing process

    CN111007073A

  • Intelligent fan blade defect detection method based on bimodal fusion

    CN114429457A

  • Surface defect detection method and system, electronic equipment and storage medium

    CN115063357A

  • Circuit board defect detection method, device and equipment and storage medium

    CN119540174A

  • Display panel, manufacturing method thereof and display apparatus having the same

    KR1020230088321A