Target detection method and apparatus, and electronic device and medium
By upsampling the original image and sliding window cutting, small-size targets are detected, which solves the problem of target loss caused by downsampling in feature extraction and improves detection accuracy.
Patent Information
- Application Number
- PCT/CN2024/128680
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-10-30
- Publication Date
- 2025-06-05
AI Technical Summary
The prior art is difficult to effectively detect small-size targets in object detection because downsampling during feature extraction will lead to features loss of small-size targets.
By preset upsampling of the original image, the image to be detected is obtained, and the sliding window is used to divide it into multiple sub-images to be detected, each sub-image is detected in an object, and finally the detection result is mapped back to the original image.
Feature expansion of small-size targets is achieved, target loss caused by downsampling during feature extraction, and the accuracy of small-size target detection is improved.
Smart Images

Figure CN2024128680_05062025_PF_FP_ABST
Abstract
Description
Target detection method, device, electronic device and medium Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a target detection method, device, electronic equipment and medium. Background Art
[0002] The detection of small-sized targets is a key and difficult problem in the field of target detection. Target detection locates the target object in the image through feature extraction. However, for small-sized targets in the image, downsampling during the feature extraction process will cause the features of the small-sized targets to be lost, resulting in the inability to detect small-sized targets in the image.
[0003] Summary of the Invention
[0004] The present invention provides a target detection method, device, electronic device and medium to address the deficiencies in the related art.
[0005] According to a first aspect of an embodiment of the present invention, a target detection method is provided, which includes: upsampling an acquired original image according to a preset upsampling multiple to obtain an image to be detected; using a sliding window to divide the image to be detected into multiple sub-images to be detected; performing target detection on each sub-image to be detected to obtain a first target detection result for each sub-image to be detected; and mapping the first target detection result of each sub-image to be detected to the original image to obtain a target detection result of the original image.
[0006] In some embodiments, the use of a sliding window to divide the image to be detected into multiple sub-images to be detected includes: sliding the window with a set step size, determining the image covered by the window after each sliding of the window as the sub-image to be detected, and the number of overlapping pixels in two adjacent sub-images to be detected is greater than or equal to the upsampling multiple.
[0007] In some embodiments, the target detection is performed on each sub-image to be detected to obtain a first target detection result for each sub-image to be detected, including: extracting a first feature map, a second feature map, and a third feature map of the sub-image to be detected, wherein the dimensions of the first feature map, the second feature map, and the third feature map increase sequentially; extracting a heat map of the third feature map; performing feature splicing on the heat map with the first feature map, the second feature map, and the third feature map, respectively, to obtain first fusion features, second fusion features, and third fusion features of different dimensions; performing target detection on the first fusion features, the second fusion features, and the third fusion features, respectively, to obtain a first target detection result for the sub-image to be detected.
[0008] In some embodiments, the feature splicing of the heat map with the first feature map, the second feature map and the third feature map respectively to obtain first fusion features, second fusion features and third fusion features of different dimensions includes: downsampling the heat map multiple times at different initial positions to obtain a sub-heat map set; feature splicing each sub-heat map in the sub-heat map set with the first feature map to obtain the first fusion feature; feature splicing each sub-heat map in the sub-heat map set with the second feature map to obtain the second fusion feature; feature splicing each sub-heat map in the sub-heat map set with the third feature map to obtain the third fusion feature.
[0009] In some embodiments, target detection is performed on the first fusion feature, the second fusion feature, and the third fusion feature respectively to obtain a first target detection result of the sub-image to be detected, including: performing first sub-target detection on the first fusion feature to obtain a first sub-target detection result; performing second sub-target detection on the second fusion feature to obtain a second sub-target detection result; performing third sub-target detection on the third fusion feature to obtain a third sub-target detection result, and the sizes of the first sub-target, the second sub-target, and the third sub-target increase sequentially; and obtaining the first target detection result of the sub-image to be detected based on the first sub-target detection result, the second sub-target detection result, and the third sub-target detection result.
[0010] In some embodiments, mapping the first target detection result of each sub-image to be detected to the original image to obtain the target detection result of the original image includes: mapping the first target detection result of each sub-image to be detected to the image to be detected according to the position of each sub-image to be detected in the image to be detected to obtain the second target detection result of the image to be detected; and restoring the size of the image to be detected indicated by the second target detection result based on the upsampling multiple to obtain the target detection result of the original image.
[0011] In some embodiments, after mapping the detection boxes of each to-be-detected sub-image to the to-be-detected image to obtain a second object detection result of the to-be-detected image, the method further includes:
[0012] Traverse each sub-image to be detected, and if there is a first target detection result in the target boundary area of the sub-image to be detected, determine whether there is a first target detection result in the adjacent position of another sub-image to be detected adjacent to the sub-image to be detected; if so, use a minimum rectangular frame to enclose the second target detection result mapped from the two adjacent first target detection results on the image to be detected to obtain a candidate target area; if not, determine the mapping area of the first target detection result of the sub-image to be detected on the image to be detected as the candidate target area; cut the image to be detected with the candidate target area as the center area to obtain a third image to be detected; perform target detection on the third image to be detected to obtain a third target detection result; if the confidence of the third target detection result meets the set conditions, determine the third target detection result as the target detection result of the candidate target area.
[0013] In some embodiments, the method also includes: obtaining a fourth target detection result based on the second target detection result in the candidate target area; weightedly fusing the fourth target detection result and the third target detection result to obtain a fifth target detection result of the candidate target area; if the confidence of the fifth target detection result meets the set conditions, the fifth target detection result is determined as the target detection result of the candidate target area.
[0014] In some embodiments, the acquired original image is upsampled according to a preset upsampling multiple to obtain an image to be detected, including: extracting shallow features from the original image using a convolutional layer; extracting deep features from the shallow features using multiple residual groups; upsampling the deep features according to a preset upsampling multiple to obtain an upsampled feature map; and reconstructing the upsampled feature map to obtain the image to be detected.
[0015] In some embodiments, the original image is a screen image of a screen to be inspected, and the method further includes: determining whether the screen to be inspected meets acceptance criteria based on a target detection result of the screen image.
[0016] According to a second aspect of an embodiment of the present invention, there is provided a target detection device, the device comprising:
[0017] The super-resolution reconstruction module is used to upsample the acquired original image according to a preset upsampling multiple to obtain the image to be detected;
[0018] A sliding image cutting module is used to cut the image to be detected into multiple sub-images to be detected using a sliding window;
[0019] An object detection module is used to perform object detection on each sub-image to be detected, and obtain a first object detection result for each sub-image to be detected;
[0020] The post-detection processing module is used to map the first target detection result of each sub-image to be detected into the original image to obtain the target detection result of the original image.
[0021] According to a third aspect of an embodiment of the present invention, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor implements any one of the methods described in the above embodiments by running the executable instructions.
[0022] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the method according to any one of the above embodiments is implemented.
[0023] According to the above embodiments, the present invention upsamples the acquired original image according to a preset upsampling multiple to obtain an image to be detected, uses a sliding window to divide the image to be detected into multiple sub-images to be detected, and performs target detection on each sub-image to be detected to obtain a first target detection result for each sub-image to be detected. The first target detection result of each sub-image to be detected is then mapped to the original image to obtain a target detection result for the original image. By upsampling the original image according to a preset upsampling multiple, the present invention completes feature expansion of small-sized targets, avoids target loss caused by downsampling during the feature extraction process, and improves the accuracy of small-sized target detection.
[0024] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0026] FIG1 is a schematic flow chart of a target detection method according to an embodiment of the present invention.
[0027] FIG2 is a schematic diagram showing a framework of a target detection network according to an embodiment of the present invention.
[0028] FIG3 is a schematic diagram of a target detection framework according to an embodiment of the present invention.
[0029] FIG4 is a schematic diagram showing a framework of a heat map feature extraction module according to an embodiment of the present invention.
[0030] FIG5 is a schematic diagram showing downsampling according to an embodiment of the present invention.
[0031] FIG6 is a schematic diagram showing a framework of a target prediction module according to an embodiment of the present invention.
[0032] FIG. 7 is a schematic diagram showing an image to be detected and sub-images to be detected according to an embodiment of the present invention.
[0033] FIG8 is a schematic diagram of an object detection device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0034] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.
[0035] For detecting small objects, the image can be sliced and the detection results of each slice can be combined to complete the task of detecting small objects in the image. Although this method ensures that the features of small objects are not lost due to resizing, the downsampling operation in the feature extraction process will cause target loss. In addition, the small size of the target will result in insufficient feature information of the target itself, resulting in missed detection.
[0036] In view of this, the present invention provides a target detection method that expands the features of small-sized targets by upsampling the original image according to a preset upsampling factor, thereby avoiding target loss caused by downsampling during the feature extraction process. In this embodiment, a small-sized target refers to a target whose size is within a preset size range.
[0037] The present invention can be applied to screen defect detection scenarios, such as tiny hole detection and spot-like foreign object detection. However, those skilled in the art should understand that the target detection method provided by the present invention can be applied not only to screen defect detection scenarios, but also to other scenarios requiring small-sized target detection. This embodiment uses screen defect detection as an example to introduce the target detection method.
[0038] The following embodiments will introduce a target detection method provided by the present invention in conjunction with the accompanying drawings.
[0039] FIG1 is a flow chart of a target detection method according to an embodiment of the present invention. As shown in FIG1 , the target detection method may include steps 101 to 104 .
[0040] In step 101, the acquired original image is upsampled according to an upsampling multiple to obtain an image to be detected.
[0041] The original image may be an image of the object to be detected acquired by an image acquisition device, or may be an image transmitted by other devices.
[0042] In the scenario of screen inspection, the screen to be inspected can be photographed to obtain an original image of the screen to be inspected. For example, the screen to be inspected can be placed on an inspection station, illuminated, and photographed with a camera to capture an original image of the screen to be inspected. The captured original image is then uploaded to a processor, which is used to execute the target detection method.
[0043] In some embodiments, the purpose of upsampling the acquired original image according to a preset upsampling multiple can be achieved by performing super-resolution reconstruction on the acquired original image. The present invention can use a trained super-resolution network to perform super-resolution reconstruction on the original image and perform feature amplification on small targets in the original image. The upsampling multiple of the super-resolution network is a preset upsampling multiple, and the super-resolution network can adopt a very deep residual channel attention network RCAN (the very deep residual channel attention networks), enhanced deep residual network EDRN (Enhanced Deep Residual Networks) and local-global combined network LGCN (Local–Global Combined Network), etc.
[0044] This embodiment uses the RCAN network as an example to illustrate the super-resolution reconstruction process. The RCAN network can include shallow feature extraction, Residual in Residual (RIR) deep feature extraction, upsampling, and reconstruction. The RIR consists of multiple residual groups RG with a long skip connection LSC. Each RG contains a number of channel residual blocks RCAB with short skip connections SSC. Each channel attention residual block RCAB consists of a simple residual block BN and a channel attention mechanism CA.
[0045] In some embodiments, the RCAN network can be used to perform super-resolution reconstruction on the original image to obtain the image to be detected, which may include: using a convolutional layer to extract shallow features from the original image; using multiple residual groups to extract deep features from the shallow features; upsampling the deep features to obtain an upsampled feature map; and reconstructing the upsampled feature map to obtain the image to be detected.
[0046] The RCAN network enlarges the scale through upsampling operations, and then reconstructs the upsampled feature map through the convolution layer, and finally obtains the super-resolution reconstruction result, that is, the image to be detected.
[0047] The present invention magnifies the original image through a super-resolution network, thereby achieving feature amplification of extremely small-sized objects in the original image.
[0048] In step 102, the image to be detected is divided into multiple sub-images to be detected using a sliding window.
[0049] On the image to be detected, a window of a preset size is slid according to a preset step length to divide the image to be detected into a plurality of sub-images to be detected.
[0050] In step 103, target detection is performed on each sub-image to be detected to obtain a first target detection result for each sub-image to be detected.
[0051] In this embodiment, a trained object detection network may be used to perform object detection on each sub-image to be detected, and a first object detection result for each sub-image to be detected may be obtained.
[0052] The first target detection result includes the vertex coordinates of the detection box, or the vertex coordinates of the detection box and the category of the target in the detection box.
[0053] In step 104, the first object detection result of each sub-image to be detected is mapped to the original image to obtain the object detection result of the original image.
[0054] According to the position of each sub-image to be detected in the image to be detected, the first target detection result of the sub-image to be detected is mapped to the image to be detected, and then the image to be detected is restored to its size to obtain the target detection result of the original image.
[0055] The target detection result may include a detection frame. When the detection frame of the target in each sub-image to be detected is obtained, the coordinates of the detection frame in the sub-image to be detected can be transformed according to the position of the sub-image to be detected in the image to be detected to obtain the coordinates of the same detection frame in the image to be detected, thereby realizing the mapping of the first target detection result in each sub-image to be detected to the image to be detected.
[0056] The present invention performs super-resolution reconstruction on an acquired original image to obtain an image to be detected. The image to be detected is then divided into multiple sub-images to be detected using a sliding window. Target detection is performed on each sub-image to obtain a first target detection result for each sub-image to be detected. The first target detection results of each sub-image to be detected are then mapped onto the original image to obtain a target detection result for the original image. By performing super-resolution reconstruction on the original image, the present invention expands the features of small-sized targets, avoiding target loss due to downsampling during the feature extraction process of target detection, thereby improving the accuracy of small-sized target detection.
[0057] In some embodiments, the use of a sliding window to divide the image to be detected into multiple sub-images to be detected may include: sliding the window with a set step size, determining the image covered by the window after each sliding of the window as the sub-image to be detected, the number of overlapping pixels in two adjacent sub-images to be detected is greater than or equal to the upsampling multiple of the super-resolution network, and the super-resolution network is used to perform super-resolution reconstruction on the original image.
[0058] This embodiment uses a sliding window approach for image slicing, that is, pre-set image slicing parameters. To ensure that small-sized objects are as intact as possible, the number of overlapping pixels between two adjacent sub-images to be detected after slicing is greater than or equal to the upsampling factor. The upsampling factor can be determined based on the size of the original image and the size of the image to be detected. In other words, during the image slicing process, the number of overlapping pixels within adjacent sliding windows is no less than the upsampling factor of the super-resolution network.
[0059] In one implementation, the image can be cut using a preset window size and sliding step size, and the sliding step size should satisfy a preset condition s∈[k,nk], where n is the window size, s is the sliding step size, and k is the upsampling multiple of the super-resolution network.
[0060] In another implementation, the image can be sliced using a preset window size and overlap. When sliding the image by setting the window size and overlap, the number of overlapping pixels can be determined by the following formula: (ns) = p*n
[0061] Where (ns) is the number of overlapping pixels of the sliding window during the sliding process; p is the window overlap; and n is the window size.
[0062] After sliding image slicing using the above-mentioned image slicing method, for an image to be detected with an aspect ratio of 1, the following formula can be used to determine the total number of sub-images to be detected that can be obtained after slicing.
[0063] Among them, N s is the total number of sub-images to be detected; L is the size of the image to be detected; The symbol for rounding up.
[0064] In some embodiments, performing target detection on each sub-image to be detected to obtain a first target detection result for each sub-image to be detected may include: extracting a first feature map, a second feature map, and a third feature map of the sub-image to be detected, wherein the dimensions of the first feature map, the second feature map, and the third feature map increase sequentially; extracting a heat map of the third feature map; performing feature splicing on the heat map with the first feature map, the second feature map, and the third feature map, respectively, to obtain first fusion features, second fusion features, and third fusion features of different dimensions; performing target detection on the first fusion features, the second fusion features, and the third fusion features, respectively, to obtain a first target detection result for the sub-image to be detected.
[0065] In an embodiment of the present invention, a target detection network can be used to perform target detection on each sub-image to be detected. FIG2 is a schematic diagram of a framework of a target detection network according to an embodiment of the present invention. As shown in FIG2 , the target detection network may include a shallow feature extraction module, a heat map feature extraction module, a feature splicing module, and a target prediction module. The shallow feature extraction module is used to extract the first feature map, the second feature map, and the third feature map of the sub-image to be detected; the heat map feature extraction module is used to extract the heat map of the third feature map; the feature splicing module is used to perform feature splicing on the heat map with the first feature map, the second feature map, and the third feature map, respectively, to obtain first fusion features, second fusion features, and third fusion features of different dimensions; the target prediction module is used to perform target detection on the first fusion feature, the second fusion feature, and the third fusion feature, respectively, to obtain the first target detection result of the sub-image to be detected.
[0066] That is to say, in the feature extraction part, various feature contents of the sub-image to be detected are extracted based on the shallow feature extraction module and the heat map feature extraction module, and feature prediction is performed in the target prediction module to complete the target detection task.
[0067] The following embodiments will explain the implementation process in the target detection network.
[0068] In this embodiment, shallow feature extraction can be performed through Resnet50. In one implementation, the feature maps extracted from the 2nd, 3rd, and 4th blocks of the Resnet50 network can be selected for small-size, medium-size, and large-size target prediction, respectively. The resolutions of the feature maps are 1 / 8, 1 / 16, and 1 / 32 of the sub-image to be detected, respectively. The maximum downsampling ratio is no greater than the super-resolution reconstruction ratio of the image, so the feature information of small-size targets in the original image can be effectively retained.
[0069] FIG3 is a schematic diagram of a target detection framework according to an embodiment of the present invention. As shown in FIG3 , assuming that the sub-image to be detected is B*3*256*256, which represents a tensor of batch B, each sample has 3 channels, a height of 256 pixels, and a width of 256 pixels. The sub-image to be detected is input into Resnet50 to obtain a first feature map of B*512*32*32, a second feature map of B*1024*16*16, and a third feature map of B*2048*8*8. A heat map feature extraction module is used to extract a heat map of B*N*64*64 of the third feature map of B*2048*8*8.
[0070] The heatmap feature extraction module extracts features from the high-dimensional feature map, the third feature map, to generate a heatmap. This improves the network's ability to represent small-scale objects. Heatmaps measure the target's characteristics based on its center point and scale. They are typically distributed in a Gaussian circle outward from the target's center. Heatmaps corresponding to different categories correspond to different predicted targets.
[0071] In one implementation, the heat map can be downsampled into a heat map according to the feature map resolution of the shallow feature extraction module, and the heat map can be feature spliced with feature maps of different dimensions to obtain fused features of different dimensions.
[0072] In another implementation, the heat map can be downsampled multiple times at different initial positions to obtain a set of sub-heat maps; each sub-heat map in the sub-heat map set is feature-concatenated with the first feature map to obtain the first fusion feature; each sub-heat map in the sub-heat map set is feature-concatenated with the second feature map to obtain the second fusion feature; each sub-heat map in the sub-heat map set is feature-concatenated with the third feature map to obtain the third fusion feature.
[0073] As shown in Figure 3, the heat map B*N*64*64 is spliced with the first feature map B*512*32*32 to obtain the first fused feature B*(512+2*N)*32*32; the heat map B*N*64*64 is spliced with the second feature map B*1024*16*16 to obtain the second fused feature B*(1024+4*N)*16*16; the heat map B*N*64*64 is spliced with the third feature map B*2048*8*8 to obtain the third fused feature B*(2048+8*N)*8*8; it can be seen that the dimensions of the first fused feature, the second fused feature and the third fused feature increase successively.
[0074] It should be noted that when the heat map B*N*64*64 is spliced with the first feature map B*512*32*32, the heat map B*N*64*64 is downsampled, and its downsampling ratio is 2*2=4; similarly, the downsampling ratios when splicing with the second and third feature maps are 16 and 64 respectively.
[0075] FIG4 is a schematic diagram of a heat map feature extraction module according to an embodiment of the present invention. As shown in FIG4 , the third feature map extracted by ResNet50 is obtained, denoted as F1, and F1 is input into the heat map feature extraction module. In the heat map feature extraction module, F2 is obtained by three deconvolution modules (Deconv), that is, three upsamplings. F2 is then input into the prediction layer for prediction to obtain a heat map. In the "1*1 convolution layer with an N-dimensional stride of 1" in FIG4 , N represents the number of target categories.
[0076] This embodiment performs downsampling on the basis of the extracted heat map, and converts the heat map into a sub-heat map set with a depth of X, where X is the square of the ratio of the heat map size to the feature map size. Figure 5 is a schematic diagram of downsampling according to an embodiment of the present invention. When X=4, the downsampling conversion rule is shown in Figure 5. The horizontal downsampling ratio = the vertical downsampling ratio = 2. The original image is divided into a combination of several 2*2 blocks in sequence. The first sub-heat map of 1 / 4 the size of the original image, which is made up of the (1,1) pixels of each block, is the first dimension of the sub-heat map set. Similarly, the (1,2), (2,1), and (2,2) pixel combinations are the second, third, and fourth dimensions of the sub-heat map set, respectively. Finally, a 4-dimensional sub-heat map set is obtained, and the size of each sub-heat map is 1 / 4 of the heat map size. Wherein, (x, y) indicates that the coordinates of the block where the pixel is located are the xth row and the yth column.
[0077] This embodiment provides two methods for downsampling heatmaps. The first implementation method is to downsample the heatmap into a single heatmap, and the second implementation method is to downsample the heatmap into a set of sub-heatmaps with a depth of X. The first implementation method can be regarded as a subset of the second implementation method, that is, the first implementation method only retains a subgraph of a certain dimension of the second implementation method, usually the first dimension. Therefore, the second implementation method can obtain more comprehensive information, thereby avoiding information offset caused by downsampling.
[0078] After obtaining fusion features of different dimensions, target detection is performed on the first fusion feature, the second fusion feature and the third fusion feature respectively to obtain a first target detection result of the sub-image to be detected, which may include: performing first sub-target detection on the first fusion feature to obtain a first sub-target detection result; performing second sub-target detection on the second fusion feature to obtain a second sub-target detection result; performing third sub-target detection on the third fusion feature to obtain a third sub-target detection result, and the sizes of the first sub-target, the second sub-target and the third sub-target increase sequentially; and obtaining the first target detection result of the sub-image to be detected based on the first sub-target detection result, the second sub-target detection result and the third sub-target detection result.
[0079] In the small-sized target detection scenario, the small-sized targets can also be divided into three categories according to their size: large, medium, and small. The first sub-target is a small-sized target, the second sub-target is a medium-sized target, and the third sub-target is a large-sized target. That is to say, based on the first fusion feature, small-sized targets are detected to obtain small-sized target detection results; based on the second fusion feature, medium-sized targets are detected to obtain medium-sized target detection results; based on the third fusion feature, large-sized targets are detected to obtain large-sized target detection results. Different target detection results can be obtained at different scales of predicted targets. Then, the large-sized target detection results, the medium-sized target detection results, and the small-sized target detection results are subjected to non-maximum suppression (NMS) processing and confidence threshold screening to obtain the first target detection result of the sub-image to be detected.
[0080] As shown in Figure 3, small-size targets are detected based on the first fusion feature B*(512+2*N)*32*32, and small-size target detection results are obtained; medium-size targets are detected based on the second fusion feature B*(1024+4*N)*16*16, and medium-size target detection results are obtained; large-size targets are detected based on the third fusion feature B*(2048+8*N)*8*8, and large-size target detection results are obtained.
[0081] In this embodiment, a target prediction module is used to perform target detection on the first fusion feature, the second fusion feature, and the third fusion feature respectively, and each target detection result is processed to obtain a first target detection result of the sub-image to be detected.
[0082] In some embodiments, target prediction can be performed based on the combined feature map using the YOLOX framework, thereby completing the task of detecting targets in the sub-image to be detected. This embodiment can use separate prediction heads to perform regression predictions on the target category and location, etc. In addition, predictions can be made for targets of different scales based on fused features of different scales, improving the network's expressive power.
[0083] FIG6 is a schematic diagram of a framework of a target prediction module according to an embodiment of the present invention. As shown in FIG6 , taking the large-size target prediction of the third fusion feature B*(2048+8*N)*8*8 as an example, a large-size target detection result can be obtained.
[0084] The present invention obtains the first target detection result in each sub-image to be detected through super-resolution reconstruction, sliding cutting and target detection network. In this case, the first target detection results of each sub-image to be detected are spliced, and the scale is restored based on the sampling ratio on the super-resolution network to obtain the target detection result that meets the final output requirements.
[0085] In some embodiments, mapping the first target detection result of each sub-image to be detected to the original image to obtain the target detection result of the original image may include: mapping the first target detection result of each sub-image to be detected to the image to be detected according to the position of each sub-image to be detected in the image to be detected to obtain the second target detection result of the image to be detected; and restoring the size of the image to be detected indicated by the second target detection result based on the upsampling multiple to obtain the target detection result of the original image.
[0086] In this embodiment, object detection is performed based on the results of the sliding cut image. Therefore, after obtaining the first object detection results for each sub-image to be detected, these first object detection results are concatenated to obtain the second object detection result for the image to be detected. Specifically, the position of the first object detection results based on the cut image is restored to obtain the second object detection result for the image to be detected. The image to be detected, which carries the second object detection result, is then resized based on the upsampling factor of the super-resolution network to obtain the object detection result for the original image.
[0087] The second target detection result includes the vertex coordinates of the detection box, or the vertex coordinates of the detection box and the category of the target in the detection box.
[0088] In the process of mapping the first target detection results of each sub-image to be detected to the image to be detected, that is, in the process of splicing and merging the first target detection results, in the case where target detection results exist at the cut boundaries of adjacent sub-images to be detected, in order to improve the accuracy of the target detection results, the area where the target detection results are located in the image to be detected is cut twice and re-inspected, and the re-inspection results are used to improve the accuracy of target detection.
[0089] FIG7 is a schematic diagram of an image to be detected and each sub-image to be detected according to an embodiment of the present invention. As shown in FIG7 , the sub-image to be detected 701 and the sub-image to be detected 702 adjacent to each other on the left and right sides of the image to be detected 700 are magnified and displayed. When a sliding window is used for image cutting, the same target Q may be cut into two parts, a and b, with a located in the sub-image to be detected 701 and b located in the sub-image to be detected 702. The proportion of target Q in different sub-images to be detected is different, and the target detection results of different sub-images to be detected are different. When a sliding window is used for image cutting, the adjacent boundary areas 703 of the sub-image to be detected 701 and the sub-image to be detected 702 partially overlap. Therefore, when the first target detection results of the sub-image to be detected 701 and the sub-image to be detected 702 in the adjacent boundary areas 703 are mapped to the image to be detected 700, the detection result of the target Q will be inaccurate.
[0090] In addition to the above situation, there is also a situation where there is a target detection result on the cut-out boundary of only one sub-image to be detected on the cut-out boundary of adjacent sub-images to be detected. For this situation, it can be judged according to the confidence level during the stitching process. If its confidence level is less than the threshold, it will be screened out. However, it may be because the cut-out is just above the target, resulting in a decrease in feature quality. Therefore, in this embodiment, for situations where there are target detection results on the cut-out boundary, a second cut-out and re-inspection are performed, and the re-inspection results are used to improve the target detection accuracy.
[0091] In order to reduce the errors caused by the splicing of target detection results, the present invention determines the range of the target detection results involving the sliding window boundary area, performs secondary image cutting, and then inputs them into the target detection network for re-inspection.
[0092] That is, after mapping the detection frames of the sub-images to be detected to the image to be detected to obtain the second object detection result of the image to be detected, the method may further include:
[0093] Traversing each sub-image to be detected, if a first target detection result exists within the target boundary area of the sub-image to be detected, determining whether a first target detection result exists in a nearby position of another sub-image to be detected adjacent to the sub-image to be detected;
[0094] If so, a minimum rectangular frame is used on the image to be detected to enclose the second target detection result mapped from the two adjacent first target detection results to obtain a candidate target area; if not, the region mapped from the first target detection result of the sub-image to be detected on the image to be detected is determined as the candidate target area;
[0095] The image to be detected is cut with the candidate target area as the center area to obtain a third image to be detected;
[0096] Performing target detection on the third image to be detected to obtain a third target detection result;
[0097] If the confidence level of the third target detection result meets a set condition, the third target detection result is determined as the target detection result of the candidate target area.
[0098] Those skilled in the art should understand that the term "adjacent" in this embodiment may include "upper and lower adjacent" and "left and right adjacent".
[0099] As shown in Figure 7, taking the left and right adjacent sub-images to be detected 701 and 702 as examples, if there is a first target detection result a in the target boundary area of the sub-image to be detected 701, then it is determined whether there is a first target detection result b in the adjacent position of the other sub-image to be detected 702 adjacent to the sub-image to be detected 701, where the adjacent position can be understood as a position within the set range of the first target detection result a. If the first target detection result a is in the lower right corner of the sub-image to be detected 701, and the first target detection result b is in the upper left corner of the sub-image to be detected 702, it is obvious that the two are not two parts of the same small target.
[0100] If it is determined that the first target detection result b exists in the vicinity of the sub-image to be detected 702 adjacent to the sub-image to be detected 701, the target detection frame at the splicing edge is surrounded by a minimum rectangular frame to obtain a candidate target area 704, and the image to be detected is cut with the candidate target area as the center area to obtain a third detection image 705. The size of the third detection image 705 should meet the input size of the target detection network.
[0101] In this embodiment, the third image to be detected 705 can be input into the object detection network shown in FIG3 for object detection to obtain a third object detection result. If the confidence level of the third object detection result meets the set conditions, the second object detection result in the candidate object region 704 is deleted, and the third object detection result is mapped to the image to be detected to obtain the object detection result of the candidate object region.
[0102] In this embodiment, to further improve the accuracy of target detection, the second and third target detection results within the candidate target area can be combined and screened to obtain a target detection result for the candidate target area. Therefore, the present invention can also obtain a fourth target detection result based on the second target detection result within the candidate target area; perform a weighted fusion of the fourth and third target detection results to obtain a fifth target detection result for the candidate target area; and if the confidence level of the fifth target detection result meets a set condition, determine the fifth target detection result as the target detection result for the candidate target area.
[0103] The second target detection results within the candidate target area are combined. If the candidate target area includes two second target detection results, the average or maximum value of the confidence levels of the two second target detection results is determined as the fourth target detection result; if the candidate target area includes one second target detection result, the second target detection result is determined as the fourth target detection result.
[0104] The fourth target detection result and the third target detection result are weighted and fused according to a preset weight ratio to obtain a fifth target detection result. For example, the weight ratio can be 2:3. The weighted fusion algorithm averages the two target detection frames. Compared to using confidence level screening, it can combine multiple target frames to obtain a relatively average result, thereby increasing the accuracy of the target frame.
[0105] For example, assuming the fourth detection result is (x11, y11, x12, y12, s1), the third detection result is (x21, y21, x22, y22, s2), and the preset weight ratio is w1:w2, then the target detection result after weighted fusion is (x31, y31, x32, y32, s3), and its calculation formula is as follows:
[0106] Where a∈(x,y,s),j∈(1,2).
[0107] The fifth target detection result is screened based on its confidence level. If the confidence level of the fifth target detection result satisfies a set condition, the fifth target detection result is determined as the target detection result for the candidate target area, and the remaining target detection results within the candidate target area are deleted. If the confidence level of the fifth target detection result does not satisfy the set condition, it indicates that no small-sized targets exist within the candidate target area, and therefore all target detection results within the candidate target area are deleted.
[0108] The present invention completes the task of detecting small-sized targets in images through a combination of multiple modules. On the one hand, the feature expansion of small-sized targets is completed through a super-resolution network, avoiding the loss of targets due to downsampling during the feature extraction process, and solving the problem of loss of detail information caused by the deep learning algorithm due to the small size of small targets in the original image. On the other hand, the feature content of the target detection network is expanded through the heat map feature extraction module, and the feature expression ability of the deep learning algorithm is improved. The heat map extracted by the heat map feature extraction module is fused with the features of different dimensions extracted by the shallow feature extraction module, and the fused features are used for target detection, which can reduce the missed detection rate. In addition, the present invention not only saves computing costs through sliding cutting and post-processing of target detection results, but also improves the accuracy of small target detection while reducing missed detection and false detection.
[0109] The above embodiment introduces the target detection method. The following embodiment will introduce the training process of the super-resolution network and the target detection network.
[0110] For the super-resolution network, data augmentation is performed on the high-resolution images: random flipping, random contrast changes, random hue changes, etc. are performed on the high-resolution images. Then, the high-resolution image dataset after data augmentation is used to generate low-resolution image data through downsampling to construct a paired training dataset. The super-resolution network is trained using the cosine annealing learning rate strategy, and its learning rate expression is:
[0111] Among them, adopting this learning rate strategy can avoid converging to the local optimal solution during training.
[0112] For the object detection network, sample images are collected from the application scenario and data augmentation is used to generate a training sample set. Recommended data augmentation combinations include random contrast enhancement, random brightness adjustment, random hue adjustment, random resizing, image edge augmentation, random flipping, random erasing, mosaic, and mixup. This embodiment also uses a cosine decay learning rate strategy to optimize the object detection network.
[0113] For ease of understanding, the following embodiment illustrates the overall process of the testing phase. During the testing phase, the target detection task is completed based on the trained network model.
[0114] Get the input original image. Assume that the size of the original image is 1024*1024.
[0115] The original image is input into the trained RCAN network for super-resolution reconstruction to obtain the image to be detected with an upsampling ratio of 32.
[0116] The image to be detected is input into the sliding cutting module, and the image is cut using a sliding window with a window size of 256 and a step size of 128, obtaining 65,025 sub-images to be detected.
[0117] Each sub-image to be detected is input into the target detection network for target detection to obtain a first target detection result.
[0118] Each first detection result is input into the post-detection processing module to obtain the target detection result of the original image.
[0119] In the field of screen defect detection, the original image is a screen image of a screen to be inspected, and the method further includes: determining whether the screen to be inspected meets acceptance criteria based on a target detection result of the screen image. If the target detection result of the screen image indicates that the number or location of target detection results on the screen to be inspected does not meet the acceptance criteria, then the screen to be inspected does not meet the acceptance criteria; otherwise, the screen to be inspected meets the acceptance criteria.
[0120] Based on the same inventive concept, the present invention further provides a target detection device. FIG8 is a schematic diagram of a target detection device according to an embodiment of the present invention. As shown in FIG8 , the device includes:
[0121] The super-resolution reconstruction module 801 is used to upsample the acquired original image according to a preset upsampling multiple to obtain an image to be detected;
[0122] A sliding image cutting module 802 is used to cut the image to be detected into multiple sub-images to be detected using a sliding window;
[0123] The target detection module 803 is configured to perform target detection on each sub-image to be detected, and obtain a first target detection result for each sub-image to be detected;
[0124] The post-detection processing module 804 is configured to map the first target detection result of each sub-image to be detected into the original image to obtain the target detection result of the original image.
[0125] The present invention uses a super-resolution module to enlarge the original image, with the aim of amplifying the features of extremely small-sized targets. Secondly, the image to be detected after feature amplification is input into a sliding slicing module, and slicing is performed based on a sliding window. In this module, the integrity of small targets is ensured by restricting the sliding window parameters and step size settings. Subsequently, the sub-image to be detected after slicing is input into a target detection module, and shallow high-dimensional features and heat map features are extracted respectively. The target detection task is completed based on the target detection module. Finally, the first target detection result of the target detection module is post-processed to obtain the target detection result of the original image.
[0126] In some embodiments, the sliding image cutting module 802 is specifically used to slide the window with a set step size, and determine the image covered by the window after each sliding window as the sub-image to be detected, and the number of overlapping pixels in two adjacent sub-images to be detected is greater than or equal to the upsampling multiple.
[0127] In some embodiments, the target detection module 803 is specifically used to extract the first feature map, the second feature map and the third feature map of the sub-image to be detected, and the dimensions of the first feature map, the second feature map and the third feature map increase successively; extract the heat map of the third feature map; perform feature splicing on the heat map with the first feature map, the second feature map and the third feature map respectively to obtain first fusion features, second fusion features and third fusion features of different dimensions; perform target detection on the first fusion features, the second fusion features and the third fusion features respectively to obtain the first target detection result of the sub-image to be detected.
[0128] In some embodiments, the target detection module 803 is specifically used to perform multiple downsampling of the heat map at different initial positions to obtain a sub-heat map set; perform feature splicing of each sub-heat map in the sub-heat map set with the first feature map to obtain the first fusion feature; perform feature splicing of each sub-heat map in the sub-heat map set with the second feature map to obtain the second fusion feature; perform feature splicing of each sub-heat map in the sub-heat map set with the third feature map to obtain the third fusion feature.
[0129] In some embodiments, the target detection module 803 is specifically used to perform first sub-target detection on the first fusion feature to obtain a first sub-target detection result; perform second sub-target detection on the second fusion feature to obtain a second sub-target detection result; perform third sub-target detection on the third fusion feature to obtain a third sub-target detection result, and the sizes of the first sub-target, the second sub-target and the third sub-target increase in sequence; based on the first sub-target detection result, the second sub-target detection result and the third sub-target detection result, the first target detection result of the sub-image to be detected is obtained.
[0130] In some embodiments, the post-detection processing module 804 is specifically used to map the first target detection result of each sub-image to be detected to the image to be detected according to the position of each sub-image to be detected in the image to be detected, so as to obtain the second target detection result of the image to be detected; and restore the size of the image to be detected indicated by the second target detection result based on the upsampling multiple to obtain the target detection result of the original image.
[0131] In some embodiments, the post-detection processing module 804 is also used to traverse each sub-image to be detected. If there is a first target detection result in the target boundary area of the sub-image to be detected, it is determined whether there is a first target detection result in the adjacent position of another sub-image to be detected adjacent to the sub-image to be detected; if so, a second target detection result mapped by using a minimum rectangular frame to enclose the two adjacent first target detection results on the image to be detected to obtain a candidate target area; if not, the mapping area of the first target detection result of the sub-image to be detected on the image to be detected is determined as the candidate target area; the image to be detected is cut with the candidate target area as the center area to obtain a third image to be detected; target detection is performed on the third image to be detected to obtain a third target detection result; if the confidence of the third target detection result meets the set conditions, the third target detection result is determined as the target detection result of the candidate target area.
[0132] In some embodiments, the post-detection processing module 804 is also used to obtain a fourth target detection result based on the second target detection result in the candidate target area; perform weighted fusion on the fourth target detection result and the third target detection result to obtain a fifth target detection result of the candidate target area; if the confidence of the fifth target detection result meets the set conditions, the fifth target detection result is determined as the target detection result of the candidate target area.
[0133] This embodiment may also include a display device of the above-mentioned screen. The display device in this embodiment may be: electronic paper, mobile phone, tablet computer, television, laptop computer, digital photo frame, navigator, or any other product or component with a display function.
[0134] It should be noted that in the accompanying drawings, the sizes of layers and regions may be exaggerated for clarity of illustration. It will also be understood that when an element or layer is referred to as being "on" another element or layer, it may be directly on the other element, or there may be an intermediate layer. In addition, it will be understood that when an element or layer is referred to as being "under" another element or layer, it may be directly under the other element, or there may be more than one intermediate layer or element. In addition, it will also be understood that when a layer or element is referred to as being "between" two layers or elements, it may be the only layer between the two layers or elements, or there may also be more than one intermediate layer or element. Similar reference numerals throughout the text indicate similar elements.
[0135] In the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance. The term "plurality" refers to two or more, unless otherwise clearly defined.
[0136] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the disclosure herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the following claims.
[0137] It should be understood that the present invention is not limited to the exact construction described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A target detection method, characterized in that: The method comprises: Upsampling the acquired original image according to a preset upsampling multiple to obtain an image to be detected; Using a sliding window to divide the image to be detected into a plurality of sub-images to be detected; Performing target detection on each sub-image to be detected to obtain a first target detection result for each sub-image to be detected; The first target detection result of each sub-image to be detected is mapped to the original image to obtain the target detection result of the original image.
2. The method according to claim 1, characterized in that The method of dividing the image to be detected into a plurality of sub-images to be detected by using a sliding window includes: The window is slid with a set step size, and the image covered by the window after each sliding of the window is determined as the sub-image to be detected, and the number of overlapping pixels in two adjacent sub-images to be detected is greater than or equal to the upsampling multiple.
3. The method according to claim 1, characterized in that The performing target detection on each sub-image to be detected to obtain a first target detection result of each sub-image to be detected includes: Extracting a first feature map, a second feature map, and a third feature map of the sub-image to be detected, wherein the dimensions of the first feature map, the second feature map, and the third feature map increase sequentially; Extracting a heat map of the third characteristic map; Perform feature splicing on the heat map with the first feature map, the second feature map and the third feature map respectively to obtain first fusion features, second fusion features and third fusion features of different dimensions; Target detection is performed on the first fusion feature, the second fusion feature, and the third fusion feature respectively to obtain a first target detection result of the sub-image to be detected.
4. The method according to claim 3, characterized in that The step of performing feature splicing on the heat map with the first feature map, the second feature map, and the third feature map to obtain first fusion features, second fusion features, and third fusion features of different dimensions includes: Downsampling the heat map at different initial positions multiple times to obtain a set of sub-heat maps; Performing feature concatenation on each sub-heatmap in the sub-heatmap set and the first feature map to obtain the first fusion feature; Performing feature concatenation on each sub-heatmap in the sub-heatmap set and the second feature map to obtain the second fusion feature; Each sub-heat map in the sub-heat map set is feature concatenated with the third feature map to obtain the third fusion feature.
5. The method according to claim 3 or 4, characterized in that: The performing target detection on the first fusion feature, the second fusion feature and the third fusion feature respectively to obtain a first target detection result of the sub-image to be detected includes: Performing first sub-target detection on the first fusion feature to obtain a first sub-target detection result; Performing second sub-target detection on the second fused features to obtain a second sub-target detection result; Performing a third sub-target detection on the third fusion feature to obtain a third sub-target detection result, wherein the sizes of the first sub-target, the second sub-target and the third sub-target are increased in sequence; A first target detection result of the sub-image to be detected is obtained according to the first sub-target detection result, the second sub-target detection result and the third sub-target detection result.
6. The method according to claim 1, characterized in that Mapping the first target detection result of each sub-image to be detected into the original image to obtain the target detection result of the original image includes: According to the position of each sub-image to be detected in the image to be detected, mapping the first target detection result of each sub-image to be detected to the image to be detected, so as to obtain the second target detection result of the image to be detected; The image to be detected indicated by the second target detection result is restored to its original size based on the upsampling multiple to obtain the target detection result of the original image.
7. The method according to claim 6, characterized in that After mapping the detection frame of each to-be-detected sub-image to the to-be-detected image to obtain a second target detection result of the to-be-detected image, the method further includes: Traversing each sub-image to be detected, if there is a first target detection result in the target boundary area of the sub-image to be detected, determining whether there is a first target detection result in the vicinity of another sub-image to be detected adjacent to the sub-image to be detected; If it exists, a minimum rectangular frame is used on the image to be detected to enclose the second target detection result mapped from two adjacent first target detection results, to obtain a candidate target area; if it does not exist, a mapping area of the first target detection result of the sub-image to be detected on the image to be detected is determined as the candidate target area; The image to be detected is cut with the candidate target area as the center area to obtain a third image to be detected; Performing target detection on the third image to be detected to obtain a third target detection result; If the confidence of the third target detection result meets the set condition, the third target detection result is determined as the target detection result of the candidate target area.
8. The method according to claim 7, characterized in that The method further comprises: Obtaining a fourth target detection result according to the second target detection result in the candidate target area; Performing weighted fusion on the fourth target detection result and the third target detection result to obtain a fifth target detection result of the candidate target area; If the confidence level of the fifth target detection result meets a set condition, the fifth target detection result is determined as the target detection result of the candidate target area.
9. The method according to claim 1, characterized in that: The method of upsampling the acquired original image according to a preset upsampling multiple to obtain the image to be detected includes: Use convolutional layers to extract shallow features from the original image; Extracting deep features from the shallow features using multiple residual groups; Upsampling the deep features according to a preset upsampling multiple to obtain an upsampled feature map; The upsampled feature map is reconstructed to obtain an image to be detected.
10. The method according to any one of claims 1 to 9, characterized in that The original image is a screen image of a screen to be detected, and the method further includes: According to the target detection result of the screen image, it is determined whether the screen to be detected meets the acceptance criteria.
11. A target detection device, characterized in that: The device comprises: A super-resolution reconstruction module is used to perform super-resolution reconstruction on the acquired original image to obtain an image to be detected; A sliding image cutting module, used for cutting the image to be detected into multiple sub-images to be detected by using a sliding window; The target detection module is used to perform target detection on each sub-image to be detected, and obtain a first target detection result for each sub-image to be detected; The post-detection processing module is used to map the first target detection result of each sub-image to be detected into the original image to obtain the target detection result of the original image.
12. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 10 by running the executable instructions.
13. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instructions are executed by a processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Aerial image small target detection method based on feature fusion and up-sampling
CN111461217A
Small target detection method and device, equipment and medium
CN113221895A
Target detection method, target detection device and computer readable storage medium
CN114419428A
Method and device for detecting tiny target of image
CN115565049A
Target detection method and device, electronic equipment and medium
CN117593511A
Cited By
Solder paste plate surface foreign matter detection method and system based on AI drive and medium
CN121544871A
Small target detection method based on combination of target super-resolution, background degradation and attention
CN121746692A
Small target detection method based on target super-resolution and background degradation and attention combination
CN121746692B
Method for establishing modified-yolo-based surface defect detection method and defect-detection system
TWI939049B