Method for detecting object in image and related equipment

CN120129923APending Publication Date: 2025-06-10BOE TECHNOLOGY GROUP CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380010541.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-09-08
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing object detection technologies are difficult to detect extremely small and large objects in images at the same time, especially in industrial product images, where resources are wasted and detection time is long.

Method used

Object detection is performed using a trained Faster-RCNN neural network by generating multiple feature maps for detecting objects of different sizes and fusing them into a fusion feature map.

Benefits of technology

Accurate detection of objects of various sizes in the image, especially those of small sizes, improve detection speed and save computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120129923A_ABST
    Figure CN120129923A_ABST
Patent Text Reader

Abstract

A method for detecting an object in an image, a computer readable storage medium, and an electronic device. The method comprises: acquiring an image to be detected (102); generating, based on the image, a plurality of feature maps for detecting objects having different sizes (104); fusing the plurality of feature maps to obtain a fused feature map (106); and detecting an object (108) based on the fused feature map, enabling detection of objects of various sizes, improving the detection speed and saving computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Method and related device for detecting objects in images Technical Field

[0001] The present disclosure relates to the technical field of image processing, and more particularly, to a method, a computer-readable storage medium, and an electronic device for detecting an object in an image. Background Art

[0002] In recent years, object detection technology, as a research topic in image processing, has been widely applied in fields such as industrial production and intelligent manufacturing. Image processing technology uses a camera to capture images of the industrial product to be inspected and transmits them to an image processing system. The image processing system detects the target's features based on the distribution of pixels, brightness, color, and other information within the image. For example, it can detect defects in industrial products such as scratches, holes, and dents.

[0003] Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, a computer-readable storage medium, and an electronic device for detecting an object in an image.

[0005] In a first aspect of the present disclosure, a method for detecting an object in an image is provided. The method comprises: acquiring an image to be detected; generating, based on the image, a plurality of feature maps for detecting objects of different sizes; fusing the plurality of feature maps to obtain a fused feature map; and detecting the object based on the fused feature map.

[0006] In an embodiment of the present disclosure, the multiple feature maps include: a first feature map for detecting an object having a first size, a second feature map for detecting an object having a second size, and a third feature map for detecting an object having a third size, wherein the first size is larger than the second size, and the second size is larger than the third size.

[0007] In an embodiment of the present disclosure, generating the first feature map based on the image includes: downsampling the image to obtain first image data; and performing a first convolution process on the first image data to generate the first feature map.

[0008] In an embodiment of the present disclosure, the downsampling ratio is 6 times, and in the first convolution process, the size of the convolution kernel is 3 and the step size is 2.

[0009] In an embodiment of the present disclosure, generating the second feature map based on the image includes: performing a second convolution process on the image to obtain second image data; and performing a third convolution process on the second image data to generate the second feature map.

[0010] In an embodiment of the present disclosure, in the second convolution processing, the size of the convolution kernel is 3 and the step size is 3, and wherein, in the third convolution processing, the size of the convolution kernel is 4 and the step size is 4.

[0011] In an embodiment of the present disclosure, generating the third feature map based on the image includes: performing a fourth convolution process on the image to obtain third image data; performing a fifth convolution process on the third image data to obtain fourth image data; performing a first pooling process and a second pooling process on the fourth image data, respectively, to obtain fifth image data and sixth image data; splicing the fifth image data and the sixth image data to obtain spliced ​​image data; and performing a sixth convolution process on the spliced ​​image data to generate the third feature map.

[0012] In an embodiment of the present disclosure, the first pooling process includes a first maximum pooling process, and wherein the second pooling process includes a second maximum pooling process.

[0013] In an embodiment of the present disclosure, in the fourth convolution processing, the size of the convolution kernel is 3 and the step size is 3, in the fifth convolution processing, the size of the convolution kernel is 3 and the step size is 1, and wherein, in the first pooling processing, the size of the pooling kernel is 3 and the step size is 2, in the second pooling processing, the size of the pooling kernel is 5 and the step size is 2, and wherein, in the sixth convolution processing, the size of the convolution kernel is 3 and the step size is 3.

[0014] In an embodiment of the present disclosure, a trained neural network is used to detect the object based on the fused feature map.

[0015] In an embodiment of the present disclosure, the trained neural network is a Faster-RCNN network, wherein the Faster RCNN network includes a backbone network, a region proposal network, and a regression classification network.

[0016] In an embodiment of the present disclosure, the SimOTA algorithm is used to train the region proposal network.

[0017] In an embodiment of the present disclosure, the SimOTA algorithm is used to train the region proposal network, including: taking feature points that fall within the true box or near the true box as candidate positive samples; calculating the cost matrix of the predicted box of the candidate positive sample and the true box; for each true box, ranking the intersection-over-union (IoU) values ​​of the true box and the predicted box from large to small, and selecting the top m predicted boxes; summing the IoU values ​​corresponding to the top m predicted boxes, and the sum is n; and ranking the cost matrix from small to large, selecting the top n predicted boxes as positive samples, and the others as negative samples.

[0018] In an embodiment of the present disclosure, when using the SimOTA algorithm, if the size of the true frame is smaller than a predetermined threshold, the size of the true frame is adjusted to the predetermined threshold.

[0019] In an embodiment of the present disclosure, if there are M true boxes and N predicted boxes, the size of the cost matrix is ​​M×N, and each element in the cost matrix is ​​the value of the loss function between the true box and the predicted box.

[0020] In an embodiment of the present disclosure, the loss function includes cross entropy loss for classification and IOU loss for regression.

[0021] In an embodiment of the present disclosure, the cross entropy loss for classification is expressed as:

[0022] Among them, N cls is the number of selected prediction boxes, p i is the probability that the predicted box is the real box, p i * =0 is a positive sample, p i * =0 is a negative sample, L cls (p i , p i * ) is represented as:

[0023] In an embodiment of the present disclosure, the IOU loss for regression is expressed as:

[0024] Among them, t i is the offset predicted by the prediction box, t i * is the offset of the predicted box relative to the real box, where L reg (t i , t i * ) is represented as:

[0025] Among them, R is smooth L1 function, which is expressed as:

[0026] According to a second aspect of the present disclosure, a device for detecting an object in an image is provided. The device includes: an acquisition module configured to acquire an image to be detected; a feature map generation module configured to generate, based on the image, multiple feature maps for detecting objects of different sizes; a fusion module configured to fuse the multiple feature maps to obtain a fused feature map; and an object detection module configured to detect the object based on the fused feature map.

[0027] According to a third aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, wherein when executed by a processor, the computer program instructions cause the processor to perform the method according to the first aspect of the present disclosure.

[0028] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the method according to the first aspect of the present disclosure.

[0029] Further aspects and scope of adaptability become apparent from the description provided herein. It should be understood that various aspects of the present application can be implemented individually or in combination with one or more other aspects. It should also be understood that the description and specific embodiments herein are intended for illustrative purposes and are not intended to limit the scope of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The drawings described herein are for illustrative purposes only of selected embodiments and not all possible implementations, and are not intended to limit the scope of the present application, wherein:

[0031] FIG1 shows a schematic flowchart of a method for detecting an object in an image according to an embodiment of the present disclosure;

[0032] FIG2 shows an exemplary STEM network according to an embodiment of the present disclosure;

[0033] FIG3 shows a network architecture in which the method shown in FIG1 is implemented according to an embodiment of the present disclosure, the network architecture including the stem network and the Faster-RCNN network shown in FIG2 ;

[0034] FIG4 shows an example of a result detected by a method according to an embodiment of the present disclosure;

[0035] FIG5 shows an example of a result detected by a method according to an embodiment of the present disclosure;

[0036] FIG6 is a block diagram showing a schematic structure of an electronic device according to an embodiment of the present disclosure; and

[0037] FIG7 shows an exemplary structural block diagram of a device for detecting an object in an image according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure. The embodiments of the present disclosure will be described in detail below with reference to the drawings and in conjunction with the embodiments. It should be noted that the features in the embodiments of the present disclosure can be combined with each other in the absence of conflict.

[0039] Faster-RCNN can detect targets (also referred to as objects in this article). Faster-RCNN-based methods can generate a bounding box (bbox), which can define the target or object. When performing object detection on industrial products, images of the products to be detected may contain extremely small objects (e.g., objects of 2 to 3 pixels) as well as larger objects (e.g., objects of 300 to 8000 pixels). Directly using Faster-RCNN for object detection requires a large amount of video memory due to the high image resolution, making object detection impossible on existing hardware. Using Faster-RCNN for detection after downsampling the image may result in failure to detect small objects due to the loss of information about them. Typically, if these objects need to be detected simultaneously, an image pyramid structure is used to extract multi-scale features, and a sliding window method is used to process the image. However, this method is complex and time-consuming.

[0040] Traditional Faster-RCNN primarily processes natural scenes. To handle complex natural scenes, each layer of the Faster-RCNN backbone network is relatively heavyweight. Because natural scenes are complex and the objects to be detected are not very small, the resolution of natural scene images is typically within 1024×1024. However, images of industrial products to be inspected have a higher resolution, typically around 8000×5000. Therefore, directly using this heavyweight backbone network to process these images would result in a significant waste of resources and be time-consuming.

[0041] Figure 1 illustrates a schematic flow chart of a method for detecting objects in an image according to an embodiment of the present disclosure. As shown in Figure 1 , in block 102, an image to be detected is acquired. In embodiments of the present disclosure, the image may be an image of an industrial product to be detected, and the object may refer to a defect present in the industrial product, such as a burn, puncture, or crease. The image may be acquired in real time by an image acquisition device or pre-stored in a storage device, without limitation.

[0042] Continuing with reference to FIG. 1 , in block 104 , multiple feature maps for detecting objects of different sizes are generated based on the acquired image. In an embodiment of the present disclosure, the multiple feature maps may include a first feature map, a second feature map, and a third feature map. The first feature map may be used to detect objects of a first size, the second feature map may be used to detect objects of a second size, and the third feature map may be used to detect objects of a third size, wherein the first size is larger than the second size, and the second size is larger than the third size. In this embodiment, objects of the first size may be larger defects in industrial products, objects of the second size may be medium-sized defects, and objects of the third size may be smaller defects. In one embodiment, for an image with an image resolution of 8000, the first size may be in the range of 1 to 32 pixels, the second size may be in the range of 32 to 256 pixels, and the third size may be greater than 256 pixels. In another embodiment, for an image with an image resolution of 15000, the first size may be in the range of 1 to 64 pixels, the second size may be in the range of 64 to 512 pixels, and the third size may be greater than 512 pixels.

[0043] It should be noted that the embodiments of this disclosure use three dimensions to segment objects in an image. Those skilled in the art may use more or fewer dimensions to segment objects based on actual needs. Furthermore, the following examples illustrate the use of grayscale images as an example. It should be understood that color images may also be used.

[0044] FIG2 shows an exemplary stem network according to an embodiment of the present disclosure. The following will describe in detail how to generate the first, second, and third feature maps in conjunction with the three branches of the stem network of FIG2. It should be noted that the image data and feature maps mentioned below are represented in the form of tensor shapes [B, C, H, W], where B represents the batch, C represents the channel (for example, the channel of the grayscale image is 1 and the channel of the RGB image is 3), H represents the height of the image, and W represents the width of the image.

[0045] As shown in Figure 2, the first branch (labeled as "branch 1") is used to generate a first feature map. In this branch, first, the acquired image is downsampled to obtain first image data. In one embodiment, the downsampling process can use a bilinear interpolation function, the downsampling ratio can be 6 times, and the tensor shape of the obtained first image data is [B, 1, H / 6, W / 6]. Then, a first convolution process is performed on the first image data [B, 1, H / 6, W / 6] to generate a first feature map. In one embodiment, the convolution kernel size used in the first convolution process is 3 and the step size is 2. The tensor shape of the generated first feature map is [B, 16, H / 12, W / 12]. In the processing of the first branch, a large amount of computing resources can be saved by downsampling. The generated first feature map retains the image data of large-sized objects, but loses the image data of small-sized objects, and is mainly used to detect large-sized objects.

[0046] As shown in Figure 2, the second branch (labeled as "branch 2") is used to generate a second feature map. In this branch, first, a second convolution process is performed on the acquired image to obtain second image data. In one embodiment, the size of the convolution kernel used in the second convolution process is 3, and the step size is 3. The tensor shape of the obtained second image data is [B, 8, H / 3, W / 3]. Then, a third convolution process is performed on the second image data [B, 8, H / 3, W / 3] to generate a second feature map. In one embodiment, the size of the convolution kernel used in the third convolution process is 4, and the step size is 4. The tensor shape of the generated second feature map is [B, 16, H / 12, W / 12]. The second feature map generated by the processing of the second branch retains the image data of medium-sized objects and is mainly used to detect medium-sized objects.

[0047] Continuing with reference to Figure 2, the third branch (labeled as "branch 3") is used to generate a third feature map. In this branch, a fourth convolution process is performed on the acquired image to obtain third image data. In one embodiment, the size of the convolution kernel used in the fourth convolution process is 3, and the stride is 3. The tensor shape of the obtained third image data is [B, 8, H / 2, W / 2]. Next, a fifth convolution process is performed on the third image data [B, 8, H / 2, W / 2] to obtain fourth image data. In one embodiment, the size of the convolution kernel used in the fifth convolution process is 3, and the stride is 1. The tensor shape of the obtained fourth image data is [B, 16, H / 2, W / 2]. Then, a first pooling process, for example, a first maximum pooling process, is performed on the fourth image data [B, 16, H / 2, W / 2] to obtain fifth image data. In one embodiment, the size of the pooling kernel used in the first pooling process is 3, and the stride is 2. The tensor shape of the obtained fifth image data is [B, 8, H / 4, W / 4]. A second pooling process, for example, a second maximum pooling process, is performed on the fourth image data [B, 16, H / 2, W / 2] to obtain sixth image data. In one embodiment, the size of the pooling kernel used in the second pooling process is 5, and the stride is 2. The tensor shape of the obtained sixth image data is [B, 8, H / 4, W / 4]. Next, the fifth image data [B, 8, H / 4, W / 4] and the sixth image data [B, 8, H / 4, W / 4] are spliced ​​to obtain spliced ​​image data [B, 32, H / 4, W / 4]. In one embodiment, a fully connected layer can be used to splice the fifth image data and the sixth image data. Finally, a sixth convolution process is performed on the spliced ​​image data [B, 32, H / 4, W / 4] to generate a third feature map. In one embodiment, the size of the convolution kernel used in the sixth convolution process is 3, and the stride is 3. The tensor shape of the generated third feature map is [B, 32, H / 12, W / 12]. The third feature map generated by the processing of the third branch retains image data of large-sized objects and is mainly used for detecting large-sized objects.

[0048] It should be noted that, in the embodiment of the present disclosure, the above three branches may be executed in parallel, and the first pooling process and the second pooling process may be executed in parallel.

[0049] In an embodiment of the present disclosure, each of the above-mentioned convolution processes may sequentially use a two-dimensional convolution layer, a normalization layer (for example, a batch normalization layer or a group normalization layer), and an activation function.

[0050] Continuing with FIG1 , in block 106 , the first, second, and third feature maps are fused. In embodiments of the present disclosure, a fully connected layer may be used to fuse the first, second, and third feature maps to obtain a fused feature map having a tensor shape of [B, 64, H / 12, W / 12]. In one embodiment, the fusion process may fuse image data channel by channel.

[0051] In block 108, objects are detected based on the fused feature map. In an embodiment of the present disclosure, a trained neural network may be used to detect objects. In one embodiment, the trained neural network may be a Faster-RCNN network, which will be described in detail below.

[0052] It should be noted that the flowchart shown in FIG1 is only for example, and those skilled in the art will appreciate that various modifications may be made to the flowchart shown or the steps described therein.

[0053] FIG3 shows a network architecture in which the method shown in FIG1 is implemented according to an embodiment of the present disclosure. As shown in FIG3 , the network architecture includes the stem network and the Faster-RCNN network shown in FIG2 . In an embodiment of the present disclosure, the Faster-RCNN network may include a backbone network, a region proposal network, and a regression classification network. It should be noted that although FIG3 shows the stem network and the Faster-RCNN network separately, it is understood that the stem network can also be embedded between appropriate layers of the backbone network of the Faster-RCNN network.

[0054] In an embodiment of the present disclosure, the backbone network may be a residual network ResNet50. ResNet50 receives the fused feature map output by the stem network. The fused feature map is processed by the five layers of ResNet50. In an embodiment of the present disclosure, the tensor shape of the image data output by the 0th layer is [B, 256, H / 24, W / 24], the tensor shape of the image data output by the 1st layer is [B, 512, H / 48, W / 48], the tensor shape of the image data output by the 2nd layer is [B, 1024, H / 96, W / 96], the tensor shape of the image data output by the 3rd layer is [B, 2048, H / 192, W / 192], and the tensor shape of the image data output by the 4th layer is [B, 2048, H / 384, W / 384]. Then, the image data processed by the five layers is provided to the feature pyramid network FPN to generate unified image data.

[0055] In the embodiment of the present disclosure, the FPN processing process is as follows: the 4th layer features are processed by a layer of convolutional network, the convolutional network kernel size is 1, the number of output channels is 256, and the generated tensor shape is [B, 256, H / 384, W / 384]; the 3rd layer features are processed by a layer of convolutional network, the convolutional network kernel size is 1, the number of output channels is 256, and the generated tensor shape is [B, 256, H / 192, W / 192]; the 4th layer features are double-upsampled and added to the 3rd layer features as the third layer candidate features, and the third layer candidate features are processed by a layer of 3x3 convolutional network to become the new third layer features; the 2nd layer features are processed by a layer of convolutional network, the convolutional network kernel size is 1, the number of output channels is 256, and the generated tensor shape is [B, 256, H / 96, W / 96]; the 3rd layer candidate features are double-upsampled and added to the 2nd layer features As the second-layer candidate features, the second-layer candidate features are processed by a 3x3 convolutional network to become the new second-layer features; the first-layer features are processed by a convolutional network with a kernel size of 1 and an output channel number of 256. The generated tensor shape is [B, 256, H / 48, W / 48]; the second-layer candidate features are upsampled twice and added to the first-layer features as the first-layer candidate features. The first-layer candidate features are processed by a 3x3 convolutional network to become the new first-layer features; the 0th-layer features are processed by a convolutional network with a kernel size of 1 and an output channel number of 256. The generated tensor shape is [B, 256, H / 24, W / 24]; the first-layer candidate features are upsampled twice and added to the 0th-layer features as the 0th-layer candidate features. The 0th-layer candidate features are processed by a 3x3 convolutional network to become the new 0th-layer features; finally, The new multi-layer feature tensor data is provided to the region proposal network (RPN) and regression classification network (RCNN) for subsequent processing.

[0056] In an embodiment of the present disclosure, multiple layers of feature tensor data are input into a region proposal network. In the region proposal network, first, based on the feature tensor data of each layer, different multi-layer convolutional networks are used for processing to generate a proposed feature tensor, which is then processed using two independent branches. In an embodiment of the present disclosure, the first branch uses convolution to generate a proposal box; the other branch uses a convolutional network for processing and uses sigmoid for activation to classify the corresponding proposal box as foreground or background. The corresponding foreground proposal box (also referred to as a preliminary bounding box in this application) will serve as the input of the next stage (that is, the RCNN stage).

[0057] Next, the preliminary bounding box and the multi-layer feature tensor data are provided to the regression classification network. In the regression classification network, based on the multi-layer feature tensor data and the preliminary bounding box, the region of interest pooling layer cuts the multi-layer feature tensor data based on the preliminary bounding box, and the cut features are the corresponding candidate target features. The candidate target features are then processed using two independent branches. In the embodiment of the present disclosure, the first branch is processed using a 4-layer convolutional network to ultimately generate the corresponding regression parameters from the preliminary bounding box to the final target; the other branch is processed using a two-layer fully connected network to generate the final classification result; the final bounding box is obtained through the cooperation of the classification branch and the regression branch, that is, the category of each object and the position of the final bounding box of the object are obtained.

[0058] In an embodiment of the present disclosure, non-maximum suppression (NMS) can also be used to process the preliminary bounding box output by the region proposal network, and the top 512 preliminary bounding boxes with the highest confidence are selected as candidate bounding boxes; then ROIAlignPool is used to clip the multi-layer feature tensor based on the candidate box to obtain clipped data, and a 4-layer 3x3 convolutional network and a one-layer full connection are used to process the clipped data to obtain the final regression parameters; and two layers of full connection are used to obtain the final classification result.

[0059] It's important to note that in the field of object recognition, the location of an object is determined by drawing a bounding box around one or more objects. Classification is the process of assigning a label to an object (i.e., determining whether the object belongs to the background or a defect target such as a burn, puncture, or tear).

[0060] Figures 4 and 5 illustrate examples of detection results using methods according to embodiments of the present disclosure. The circular boxes in Figure 4 depict detected small objects, such as small defects in industrial products. The square boxes in Figure 5 depict detected medium-sized objects, such as medium-sized defects in industrial products. This demonstrates that the methods of the present disclosure can accurately detect objects of various sizes in an image, particularly small objects.

[0061] In addition, before using the network architecture shown in Figure 3 to detect objects, the region proposal network in Faster-RCNN needs to be trained. Traditionally, in the process of training the region proposal network, a binary classification label (e.g., 0 or 1) is assigned to each prediction box, where 0 represents a negative sample and 1 represents a positive sample. The full name of IoU is Intersection over Union, which is a concept used in target detection. In the traditional region proposal network, it is necessary to predefine prediction boxes of different sizes and proportions in advance. By calculating the overlap rate between the prediction box and the "true box", that is, the ratio of their intersection and union, it is determined whether the corresponding prediction box is a positive sample or a negative sample during training. In an embodiment of the present disclosure, if the intersection over union (IoU) of the prediction box and the true box is greater than 0.7, it is called a positive sample. If the intersection over union (IoU) of the prediction box and the true box is less than 0.3, it is called a negative sample. Other prediction boxes are neither positive samples nor negative samples and are not used for final training. After the region proposal network is trained using this method, Faster-RCNN is used to detect objects in industrial products. Because object sizes vary greatly—for example, from 2 pixels to 5000 pixels, a size difference of over 2000 times, far exceeding the range of object sizes in natural scenes—it's difficult to define appropriate prediction boxes so that each target has a sufficient number of prediction boxes. Consequently, the sensitivity of the Intersection over Union (IoU) for objects of different sizes varies significantly, especially for small objects. This results in most prediction boxes becoming negative samples due to minor positional deviations, leading to inaccurate label assignments. Therefore, the traditional Faster-RCNN neural network is insufficient for detecting defects in industrial products.

[0062] To this end, in an embodiment of the present disclosure, the SimOTA algorithm can be used to train the region proposal network in Faster-RCNN. First, the feature points that fall within the true box or near the true box are taken as candidate positive samples. Next, the cost matrix of the predicted box and the true box of the candidate positive sample is calculated. If there are M true boxes and N predicted boxes, the size of the cost matrix is ​​M×N, and each element in the cost matrix is ​​the value of the loss function of the true box and the predicted box, where the loss function includes the cross entropy loss for classification and the IOU loss for regression. Then, for each true box, the IoU value is ranked from large to small, the top m predicted boxes are selected, and the IoU values ​​corresponding to the m predicted boxes are summed (the sum value is n). Finally, the cost corresponding to each true box in the cost matrix is ​​ranked from small to large, and the top n predicted boxes are selected as positive samples, and the others are used as negative samples. In this process, in order to better handle small targets, a small target threshold τ (typical value is 64) is set. If the scale of the true box is smaller than τ, its size is expanded to τ for the above processing.

[0063] In an embodiment of the present disclosure, the cross entropy loss for classification can be expressed as:

[0064] Among them, N cls is the number of selected prediction boxes, p i is the probability that the predicted box is the real box, p i * =0 is a positive sample, p i * =0 is a negative sample, L cls (p i , p i * ) is the logarithmic loss of the two classes:

[0065] In an embodiment of the present disclosure, the IOU loss for regression can be expressed as:

[0066] Among them, t i is the offset predicted by the prediction box, t i * is the offset of the predicted box relative to the real box, where L reg (t i , t i * ) is expressed as:

[0067] Among them, R is smooth L1 Function, which is expressed as:

[0068] In an embodiment of the present disclosure, when the SimOTA algorithm is used to train a region proposal network, if the size of the true box is smaller than a predetermined threshold, the size of the true box is adjusted to the predetermined threshold.

[0069] Table 1 shows a comparison of object detection in an image using a method according to an embodiment of the present disclosure and a sliding window method based on Faster-RCNN. As shown in Table 1, the detected objects include burns, punctures, and creases on products. Larger values ​​indicate a better method.

[0070] Table 1

[0071] As can be seen from Table 1, the time taken by the method disclosed herein is only about 1 / 100 of that of the conventional sliding window method. In addition, the method disclosed herein also exhibits good performance.

[0072] It can be seen from the above description that by adopting the method according to the embodiments of the present disclosure, by generating multiple feature maps for detecting objects of different sizes, objects of various sizes in the image, especially small-sized objects, can be detected, thereby improving the detection speed and saving computing resources.

[0073] Figure 6 shows a schematic block diagram of an electronic device 60 for detecting objects in an image, according to an embodiment of the present disclosure. Electronic device 60 includes one or more processors 602 and a memory 604. Memory 604 is coupled to processor 602 via a bus and an I / O interface 606, and stores instructions executable by processor 602. When the instructions are executed by processor 602, electronic device 60 can perform the steps of the method for detecting objects in an image described in any of the above embodiments.

[0074] The memory 604 may include readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory, and may further include read-only memory (ROM).

[0075] Memory 604 may also include a program / utility having a set (at least one) of program modules, such program modules including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0076] The bus may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0077] The electronic device 60 can also communicate with one or more external devices (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 60, and / or any device that enables the electronic device 60 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can occur via an input / output (I / O) interface 606. Furthermore, the electronic device 60 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter. The network adapter can communicate with other modules of the electronic device 60 via a bus. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device 60, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0078] 7 shows an exemplary structural block diagram of a device 700 for detecting objects in an image according to an embodiment of the present disclosure. The device 700 may include an acquisition module 710 , a feature map generation module 720 , a fusion module 730 , and an object detection module 740 .

[0079] In an embodiment of the present disclosure, the acquisition module 710 is configured to acquire an image to be inspected, which may be an image of an industrial product to be inspected, and the object may refer to defects existing in the industrial product, such as burns, punctures, creases, etc.

[0080] In an embodiment of the present disclosure, the feature map generation module 720 is configured to generate a plurality of feature maps for detecting objects of different sizes based on the acquired image. The plurality of feature maps may include a first feature map, a second feature map, and a third feature map. The first feature map may be used to detect an object of a first size, the second feature map may be used to detect an object of a second size, and the third feature map may be used to detect an object of a third size, wherein the first size is larger than the second size, and the second size is larger than the third size.

[0081] In an embodiment of the present disclosure, the fusion module 730 is configured to fuse multiple feature maps to obtain a fused feature map. The fusion module 730 will further explain in detail how to generate the first, second and third feature maps in conjunction with the three branches of the stem network introduced in Figure 2. In the processing of the first branch shown in Figure 2, the image is downsampled to obtain first image data; and a first convolution process is performed on the first image data to generate a first feature map. In the processing of the second branch shown in Figure 2, a second convolution process is performed on the image to obtain second image data; and a third convolution process is performed on the second image data to generate a second feature map. In the processing of the third branch shown in Figure 2, a fourth convolution process is performed on the image to obtain third image data; a fifth convolution process is performed on the third image data to obtain fourth image data; a first pooling process and a second pooling process are performed on the fourth image data respectively to obtain fifth image data and sixth image data; the fifth image data and the sixth image data are spliced ​​to obtain spliced ​​image data; and a sixth convolution process is performed on the spliced ​​image data to generate a third feature map.

[0082] In an embodiment of the present disclosure, the object detection module 740 is configured to detect objects based on the fused feature map. In an embodiment of the present disclosure, a trained neural network may be used to detect objects.

[0083] Each unit of the device 700 can further implement the functions and methods described in Figures 1-3, which will not be repeated here.

[0084] In an embodiment of the present disclosure, a computer-readable storage medium is further provided, on which computer program instructions are stored. When executed by, for example, a processor, the computer program instructions can implement the steps of the method for detecting an object in an image described in any of the above embodiments. In some possible implementations, various aspects of the present application can also be implemented in the form of a program product, which includes program code. When the program product is executed on a terminal device, the program code is used to cause the terminal device to perform the steps described in the method for detecting an object in an image according to various exemplary embodiments of the present application.

[0085] According to an embodiment of the present application, a program product for implementing the above method can be a portable compact disc read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present application is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0086] The program product may be implemented in any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0087] The computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, wherein the readable program code is carried. The data signal propagated may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or component. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0088] The program code for performing the operations of the present application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0089] Those skilled in the art will appreciate that various aspects of the present application can be implemented as systems, methods, or program products. Therefore, various aspects of the present application can be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation that combines hardware and software aspects, which may be collectively referred to herein as a "circuit," "module," or "system."

[0090] The foregoing description is merely a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Those skilled in the art will readily appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present disclosure shall be included within the scope of protection of the present disclosure.

Claims

1. A method for detecting an object in an image, comprising: Acquire an image to be detected; Based on the image, generating a plurality of feature maps for detecting objects having different sizes; Fusing the multiple feature maps to obtain a fused feature map; as well as Based on the fused feature map, the object is detected.

2. The method according to claim 1, wherein: The multiple feature maps include: a first feature map for detecting an object having a first size, a second feature map for detecting an object having a second size, and a third feature map for detecting an object having a third size, wherein the first size is larger than the second size, and the second size is larger than the third size.

3. The method according to claim 2, wherein: Generating the first feature map based on the image includes: downsampling the image to obtain first image data; and A first convolution process is performed on the first image data to generate the first feature map.

4. The method according to claim 3, wherein: The downsampling ratio is 6 times, and in the first convolution process, the size of the convolution kernel is 3 and the step size is 2.

5. The method according to claim 2, wherein: Generating the second feature map based on the image includes: performing a second convolution process on the image to obtain second image data; and A third convolution process is performed on the second image data to generate the second feature map.

6. The method according to claim 5, wherein: In the second convolution process, the size of the convolution kernel is 3 and the step size is 3, and wherein, in the third convolution process, the size of the convolution kernel is 4 and the step size is 4.

7. The method according to claim 2, wherein: Generating the third feature map based on the image includes: performing a fourth convolution process on the image to obtain third image data; performing a fifth convolution process on the third image data to obtain fourth image data; Performing a first pooling process and a second pooling process on the fourth image data respectively to obtain fifth image data and sixth image data; splicing the fifth image data and the sixth image data to obtain spliced ​​image data; and A sixth convolution process is performed on the spliced ​​image data to generate the third feature map.

8. The method according to claim 7, wherein: The first pooling process includes a first maximum pooling process, and wherein the second pooling process includes a second maximum pooling process.

9. The method according to claim 7 or 8, wherein: In the fourth convolution processing, the size of the convolution kernel is 3 and the step size is 3, in the fifth convolution processing, the size of the convolution kernel is 3 and the step size is 1, and wherein, in the first pooling processing, the size of the pooling kernel is 3 and the step size is 2, in the second pooling processing, the size of the pooling kernel is 5 and the step size is 2, and wherein, in the sixth convolution processing, the size of the convolution kernel is 3 and the step size is 3.

10. The method according to any one of claims 1 to 9, wherein: Based on the fused feature map, the object is detected using a trained neural network.

11. The method according to claim 10, wherein: The trained neural network is a Faster-RCNN network, wherein the Faster RCNN network includes a backbone network, a region proposal network, and a regression classification network.

12. The method according to claim 11, wherein: The SimOTA algorithm is used to train the region proposal network.

13. The method according to claim 11, wherein: Using the SimOTA algorithm to train the region proposal network includes: The feature points that fall within the true frame or near the true frame are taken as candidate positive samples; Calculate the cost matrix of the predicted box of the candidate positive sample and the true box; For each true box, rank the intersection over union (IoU) values ​​of the true box and the predicted box from large to small, and select the top m predicted boxes; Sum the IoU values ​​corresponding to the first m prediction boxes, where the sum is n; and The cost matrix is ​​ranked from small to large, and the first n prediction boxes are selected as positive samples, and the others are selected as negative samples.

14. The method according to claim 13, wherein: When using the SimOTA algorithm, if the size of the true frame is smaller than a predetermined threshold, the size of the true frame is adjusted to the predetermined threshold.

15. The method according to claim 13 or 14, wherein: If there are M real boxes and N predicted boxes, the size of the cost matrix is ​​M×N, and each element in the cost matrix is ​​the value of the loss function of the real box and the predicted box.

16. The method according to claim 15, wherein: The loss function includes cross entropy loss for classification and IOU loss for regression.

17. The method according to claim 16, wherein: The cross entropy loss for classification is expressed as: Among them, N cls is the number of selected prediction boxes, p i is the probability that the predicted box is the true box, p i * =0 is a positive sample, p i * = 0 is a negative sample, L cls (p i , p i * ) is represented as:

18. The method according to claim 16, wherein: The IOU loss for regression is expressed as: Among them, t i is the offset predicted by the prediction box, t i * is the offset of the predicted box relative to the real box, where L reg (t i , t i * ) is represented as: Among them, R is smooth L1 function, which is expressed as:

19. An apparatus for detecting an object in an image, comprising: An acquisition module is configured to acquire an image to be detected; a feature map generation module, configured to generate a plurality of feature maps for detecting objects with different sizes based on the image; A fusion module, configured to fuse the multiple feature maps to obtain a fused feature map; as well as The object detection module is configured to detect the object based on the fused feature map.

20. A computer-readable storage medium having computer program instructions stored thereon, wherein: The computer program instructions, when executed by a processor, cause the processor to perform a method according to any one of claims 1 to 18.

21. An electronic device comprising: processor; as well as a memory storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method according to any one of claims 1 to 18.