Method and device for detecting target object in image, storage medium and equipment

By generating heat maps and initializing the location of query points, combined with the Transformer module for detection, the problem of slow R-CNN inference speed is solved, faster target object detection is achieved, and computing power consumption is reduced.

CN119963650AActive Publication Date: 2025-05-09NEOLITHIC HUITONG TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510452154.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-09
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The inference speed of R-CNN is slower, resulting in the detection speed of target objects being slower.

Method used

By obtaining the feature map of the image, generating a heat map, initializing the location of the query point, and using the Transformer module to detect the feature map and query point, the detection of the target object is achieved.

Benefits of technology

This method reduces the number of invalid query points, improves the inference speed of the model, and can accelerate the inference speed of the model without losing model performance, and consumes less computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963650A_ABST
    Figure CN119963650A_ABST
Patent Text Reader

Abstract

The invention discloses a target object detection method and device in an image, a storage medium and equipment, and belongs to the technical field of image processing. The method comprises the steps of obtaining a to-be-recognized image and a trained target detection model, wherein the target detection model at least comprises a feature extraction network and a Transform module; performing feature extraction on the image by using the feature extraction network to obtain a feature map; generating a thermodynamic diagram according to the feature map, wherein the thermodynamic diagram comprises all potential target objects in the image; initializing each query point according to the thermodynamic diagram, wherein each query point represents the position of a potential target object in the image; and detecting the feature map and each initialized query point by using the Transform module to obtain a target object in the image. According to the method, the query point can be placed at the position of the potential target object shown in the thermodynamic diagram, and the reasoning speed is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method, device, storage medium and equipment for detecting a target object in an image. Background Art

[0002] The object detection method based on deep learning refers to identifying the category, location, and size of the target object of interest in a given image based on a deep learning algorithm. Figure 1 As shown, through the target object detection method, the two dogs, the person, the bag and the car in the image can be identified.

[0003] The Region-Convolutional Neural Networks (R-CNN) in the deep learning algorithm adopts a two-stage approach to achieve the target object detection task. The first stage uses a heuristic method to find the area in the image where the target object may exist. The second stage uses CNN to classify all potential areas to determine whether the target object actually exists.

[0004] However, R-CNN suffers from slow inference speed, which results in slower detection of target objects. Summary of the invention

[0005] The present application provides a method, apparatus, storage medium and device for detecting target objects in an image, which are used to solve the problem that the reasoning speed of R-CNN is slow, resulting in a slow detection speed of target objects. The technical solution is as follows: According to a first aspect of the present application, a method for detecting a target object in an image is provided, the method comprising: Obtaining an image to be recognized and a trained target detection model, wherein the target detection model includes at least a feature extraction network and a Transformer module; Using the feature extraction network to extract features from the image to obtain a feature map; Generating a heat map according to the feature map, wherein the heat map includes all potential target objects in the image; Initializing each query point according to the heat map, the query point representing a potential location of a target object in the image; The Transformer module is used to detect the feature map and each initialized query point to obtain the target object in the image.

[0006] In a possible implementation, initializing each query point according to the heat map includes: Get the preset allocation threshold; The positions in the heat map whose values ​​are greater than the allocation threshold are allocated to each query point, and different query points are separated by n positions, where n≥2.

[0007] In a possible implementation, allocating the positions in the heat map whose values ​​are greater than the allocation threshold to each query point includes: When traversing the values ​​of each position in the heat map, if the value of the i-th row and j-th column is greater than the allocation threshold, the position of the i-th row and j-th column is allocated to a query point, where i and j are positive integers; Continue traversing the values ​​in the heat map from a new position that is n positions away from the position of the i-th row and j-th column.

[0008] In a possible implementation, allocating positions in the heat map with values ​​greater than the allocation threshold to each query point further includes: If the value in the i-th row and j-th column is less than or equal to the allocation threshold, then continue to traverse the values ​​in the heat map from the new position in the i-th row and j+1-th column.

[0009] In a possible implementation, the detecting the feature map and each initialized query point by using the Transformer module to obtain the target object in the image includes: The feature map and each initialized query point are processed using multiple Transformer layers in the Transformer module, so that after being superimposed by multiple Transformer layers, each query point gradually approaches the actual position of the target object it represents until each target object is detected.

[0010] According to a second aspect of the present application, a device for detecting a target object in an image is provided, the device comprising: An acquisition module, used to acquire an image to be identified and a trained target detection model, wherein the target detection model includes at least a feature extraction network and a Transformer module; An extraction module, used to extract features from the image using the feature extraction network to obtain a feature map; A generating module, configured to generate a heat map according to the feature map, wherein the heat map includes all potential target objects in the image; An initialization module, configured to initialize each query point according to the heat map, wherein the query point represents a potential target object in the image; The detection module is used to detect the feature map and each initialized query point using the Transformer module to obtain the target object in the image.

[0011] In a possible implementation, the initialization module is further used to: Get the preset allocation threshold; The positions in the heat map whose values ​​are greater than the allocation threshold are allocated to each query point, and different query points are separated by n positions, where n≥2.

[0012] In a possible implementation, the initialization module is further used to: When traversing the values ​​of each position in the heat map, if the value of the i-th row and j-th column is greater than the allocation threshold, the position of the i-th row and j-th column is allocated to a query point, where i and j are positive integers; Continue traversing the values ​​in the heat map from a new position that is n positions away from the position of the i-th row and j-th column.

[0013] According to a third aspect of the present application, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the target object detection method in an image as described above.

[0014] According to a fourth aspect of the present application, a computer device is provided, the computer device comprising the above-mentioned target object detection device in the image.

[0015] The beneficial effects of the technical solution provided by this application include at least: By generating a heat map of potential target objects based on the feature map of the image, and then initializing the positions of each query point (Query) based on the heat map, query points are placed at the positions of these potential target objects instead of randomly placing query points as in the classic Deform-DETR. This approach not only reduces the number of invalid query points, but also can filter out query points in areas without target objects. Each initialized query point is closer to the real target object, and fewer Transformer layers and fewer query points can be used to achieve the same effect. Ultimately, the inference speed of the model is accelerated without losing model performance. It converges much faster than the classic Deform-DETR and consumes less computing power. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1It is a schematic diagram of the recognition result of a target object in an image; Figure 2 It is a comparison chart of the image recognition result and the heat map; Figure 3 This is a schematic diagram of how the query gradually approaches the target object after being processed by multiple Transformer layers. Figure 4 is a flow chart of a method for detecting a target object in an image provided by an embodiment of the present application; Figure 5 is a flow chart of a method for detecting a target object in an image provided by an embodiment of the present application; Figure 6 It is a schematic diagram of the processing flow of a target detection model provided by an embodiment of the present application; Figure 7 It is a structural block diagram of a target object detection device in an image provided by an embodiment of the present application. DETAILED DESCRIPTION

[0018] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below in conjunction with the accompanying drawings.

[0019] In order to speed up the reasoning speed of R-CNN, Yolo and SSD proposed a one-stage target object detection scheme. Specifically, Yolo and SSD place many anchor boxes in the image space, and then classify each anchor box (negative samples indicate that the anchor box is not associated with any target object, and positive samples indicate that the anchor box is associated with certain target objects). For anchor boxes predicted as positive samples, the error between the anchor box and the most matching target box is also predicted. Although Yolo and SSD have made great improvements in performance indicators and reasoning speed, this prior-based scheme is more dependent on the placement of anchor boxes and its performance indicators and reasoning speed still have much room for improvement.

[0020] In order to reduce the model performance's over-reliance on the prior box, CenterNet proposes a new one-stage target object detection scheme. Specifically, CenterNet reformulates the target object detection task into identifying the center position of the target object and the size of the regression box. To achieve this goal, CenterNet uses a heat map to encode the center position of the target object, then takes the features of the center position and decodes them to obtain the size of the target object. Figure 2 As shown in the figure, the left figure shows the image that needs to be detected, and the right figure shows the heat map. The position of each local maximum in the heat map represents the center position of an object. Compared with previous solutions, CenterNet is easier to train and its reasoning speed is significantly better than previous methods.

[0021] DETR (Data-Efficient Image Transformer) directly predicts the top left and top right corners of the target object with the help of the powerful representation ability of Transformer, and uses bipartite graph matching to calculate the training loss to train the model. Based on innovative improvements, DETR has once again refreshed the results on the COCO dataset. Although the architecture is elegant and the indicators are off the charts, the training speed of DETR is extremely slow, and the training resources consumed are more than 10 times that of previous methods.

[0022] Deform-DETR proposes a deformable attention mechanism and solves the problem of slow convergence of DETR training. Specifically, Deform-DETR will randomly initialize some query points (Query), each Query represents a potential target object in the image, and by stacking Transformer layers, each Query will continue to move closer to the target object it represents, as shown in the schematic diagram. Figure 3 shown. Figure 3 In the figure, star 0 indicates the position of the query output by the 0th layer Transformer, star 1 indicates the position of the query output by the 1st layer Transformer, star 2 indicates the position of the query output by the 2nd layer Transformer, and the star without a number indicates the actual position of the target object. At the 0th layer, the position of the query is randomly initialized, and after each layer of Transformer, the query gradually approaches the actual position of the target object.

[0023] Although Deform-DETR is significantly improved over the original DETR, the performance of Deform-DETR is greatly affected by the initial position of each query. That is, if the initial position of a query is far away from the actual target object, more Transformer layers need to be stacked to enable the query to successfully predict the target object. In addition, for areas in the image without target objects, placing queries in these areas is a waste of resources.

[0024] In order to further improve the performance of Deform-DETR, this application uses heat maps to predict potential target objects in images and places queries at the locations of these potential target objects instead of randomly placing queries as in Deform-DETR. This approach not only reduces the number of invalid queries but also reduces the difficulty of model training, making it possible to use fewer Transformer layers and fewer queries to achieve the same effect, ultimately accelerating the model's inference speed without losing model performance.

[0025] like Figure 4 As shown, it shows a method flow chart of a method for detecting a target object in an image provided by an embodiment of the present application, and the method for detecting a target object in an image can be applied to a computer device. The method for detecting a target object in an image may include: Step 401, obtaining an image to be recognized and a trained target detection model, where the target detection model at least includes a feature extraction network and a Transformer module.

[0026] The image to be identified may be an image captured by a perception module in the vehicle, or may be an image acquired from other devices. The source of the image is not limited in this embodiment.

[0027] The target detection model can be a model improved based on Deform-DETR, which at least includes a feature extraction network and a Transformer module. The feature extraction network can be CNN or Transformer or other networks.

[0028] Step 402: extract features from the image using a feature extraction network to obtain a feature map.

[0029] Step 403: Generate a heat map based on the feature map, and the heat map contains all potential target objects in the image.

[0030] In this embodiment, the idea of ​​CenterNet can be adopted to generate a heat map according to the feature map. The values ​​of each position in the heat map are in the interval [0, 1]. If there is no target object at a position in the heat map, the value of the position is 0. If there is a target object at a position in the heat map, the value of the position is greater than 0 and less than or equal to 1.

[0031] Step 404 , initializing each query point according to the heat map, where the query point represents the location of a potential target object in the image.

[0032] The purpose of initializing query points according to the heat map is to place query points at locations with potential target objects. This approach not only reduces the number of invalid query points, but also makes each initialized query point closer to the real target object.

[0033] The query point initialization method is described in detail below.

[0034] Step 405: Use the Transformer module to detect the feature map and each initialized query point to obtain the target object in the image.

[0035] The Transformer module includes multiple Transformer layers, and the feature map and each initialized query point can be detected through the superimposed Transformer layers to obtain the target object in the image.

[0036] In summary, the target object detection method in the image provided by the embodiment of the present application generates a heat map of potential target objects according to the feature map of the image, and then initializes the positions of each query point (Query) based on the heat map, and places query points at the positions of these potential target objects instead of randomly placing query points as in the classic Deform-DETR. This approach not only reduces the number of invalid query points, but also can filter out query points in areas without target objects. Moreover, each initialized query point is closer to the real target object, and fewer Transformer layers and fewer query points can be used to achieve the same effect. Ultimately, the inference speed of the model is accelerated without losing model performance, and the convergence speed of the model is faster than that of the classic Deform-DETR, and the computing power consumption is smaller.

[0037] like Figure 5 As shown, it shows a flow chart of a method for detecting a target object in an image provided by an embodiment of the present application, and the method for detecting a target object in an image can be applied to a computer device. The method for detecting a target object in an image may include: Step 501, obtaining an image to be identified and a trained target detection model, where the target detection model at least includes a feature extraction network and a Transformer module.

[0038] The image to be identified may be an image captured by a perception module in the vehicle, or may be an image acquired from other devices. The source of the image is not limited in this embodiment.

[0039] The target detection model can be a model improved based on Deform-DETR, which at least includes a feature extraction network and a Transformer module. The feature extraction network can be a CNN or a Transformer or other networks.

[0040] Step 502: extract features from the image using a feature extraction network to obtain a feature map.

[0041] Step 503: Generate a heat map based on the feature map, and the heat map contains all potential target objects in the image.

[0042] In this embodiment, the idea of ​​CenterNet can be adopted to generate a heat map according to the feature map. The values ​​of each position in the heat map are in the interval [0, 1]. If there is no target object at a position in the heat map, the value of the position is 0. If there is a target object at a position in the heat map, the value of the position is greater than 0 and less than or equal to 1.

[0043] Step 504, obtain a preset allocation threshold; allocate positions in the heat map with values ​​greater than the allocation threshold to each query point, and different query points are separated by n positions, n≥2, and the query point represents the position of a potential target object in the image.

[0044] The purpose of initializing query points according to the heat map is to place query points at locations with potential target objects. This approach not only reduces the number of invalid query points, but also makes each initialized query point closer to the real target object.

[0045] Specifically, allocating locations in the heat map where the values ​​are greater than the allocation threshold to each query point may include: (1) When traversing the values ​​of each position in the heat map, if the value of the i-th row and j-th column is greater than the allocation threshold, the position of the i-th row and j-th column is assigned to a query point, where i and j are positive integers; (2) Continue to traverse the values ​​in the heat map from the new position that is n positions away from the position in the i-th row and j-th column; (3) If the value in the i-th row and j-th column is less than or equal to the allocation threshold, continue to traverse the values ​​in the heat map from the new position in the i-th row and j+1-th column.

[0046] Simply put, if a position is assigned to a query point, the traversal will continue at a new position that is n (n ≥ 2) positions away from the position; if a position is not assigned to a query point, the traversal will continue at a new position that is 1 position away from the position.

[0047] Step 505: Use the Transformer module to detect the feature map and each initialized query point to obtain the target object in the image.

[0048] The Transformer module includes multiple Transformer layers, and the feature map and each initialized query point can be detected through the superimposed Transformer layers to obtain the target object in the image.

[0049] Specifically, the feature map and each initialized query point are detected using a Transformer module to obtain a target object in the image, including: using multiple Transformer layers in the Transformer module to process the feature map and each initialized query point, so that after multiple Transformer layers are superimposed, each query point gradually approaches the actual position of the target object it represents until each target object is detected.

[0050] like Figure 6 As shown in the figure, after an image is input into the target detection model, the feature extraction network CNN is used to extract its features to obtain a feature map; then a heat map is generated based on the feature map, and each query point is initialized using the heat map; finally, the Transformer module is used to process the feature map and each query point to obtain each target object in the image.

[0051] In summary, the target object detection method in the image provided by the embodiment of the present application generates a heat map of potential target objects according to the feature map of the image, and then initializes the positions of each query point (Query) based on the heat map, and places query points at the positions of these potential target objects instead of randomly placing query points as in the classic Deform-DETR. This approach not only reduces the number of invalid query points, but also can filter out query points in areas without target objects. Moreover, each initialized query point is closer to the real target object, and fewer Transformer layers and fewer query points can be used to achieve the same effect. Ultimately, the inference speed of the model is accelerated without losing model performance, and the convergence speed of the model is faster than that of the classic Deform-DETR, and the computing power consumption is smaller.

[0052] like Figure 7 As shown, it shows a structural block diagram of a target object detection device in an image provided by an embodiment of the present application, and the target object detection device in an image can be applied to a computer device. The target object detection device in an image may include: An acquisition module 710 is used to acquire an image to be recognized and a trained target detection model, where the target detection model includes at least a feature extraction network and a Transformer module; An extraction module 720 is used to extract features from an image using a feature extraction network to obtain a feature map; A generating module 730 is used to generate a heat map according to the feature map, wherein the heat map contains all potential target objects in the image; Initialization module 740, used to initialize each query point according to the heat map, the query point represents a potential target object in the image; The detection module 750 is used to detect the feature map and each initialized query point using the Transformer module to obtain the target object in the image.

[0053] In an optional embodiment, the initialization module 740 is further used to: Get the preset allocation threshold; The positions in the heat map with values ​​greater than the allocation threshold are allocated to each query point, and there are n positions between different query points, where n ≥ 2.

[0054] In an optional embodiment, the initialization module 740 is further used to: When traversing the values ​​of each position in the heat map, if the value of the i-th row and j-th column is greater than the allocation threshold, the position of the i-th row and j-th column is assigned to a query point, where i and j are positive integers; Continue traversing the values ​​in the heat map from the new position that is n positions away from the position in the i-th row and j-th column.

[0055] In an optional embodiment, the initialization module 740 is further used to: If the value in the i-th row and j-th column is less than or equal to the allocation threshold, continue to traverse the values ​​in the heat map from the new position in the i-th row and j+1-th column.

[0056] In an optional embodiment, the detection module 750 is further configured to: The feature map and each initialized query point are processed using multiple Transformer layers in the Transformer module, so that after multiple Transformer layers are superimposed, each query point gradually approaches the actual position of the target object it represents until each target object is detected.

[0057] In summary, the target object detection device in the image provided by the embodiment of the present application generates a heat map of potential target objects according to the feature map of the image, and then initializes the positions of each query point (Query) based on the heat map, and places query points at the positions of these potential target objects instead of randomly placing query points as in the classic Deform-DETR. This approach not only reduces the number of invalid query points, but also can filter out query points in areas without target objects. Moreover, each initialized query point is closer to the real target object, and fewer Transformer layers and fewer query points can be used to achieve the same effect. Ultimately, the inference speed of the model is accelerated without losing model performance, and the convergence speed of the model is faster than that of the classic Deform-DETR, and the computing power consumption is smaller.

[0058] An embodiment of the present application provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the target object detection method in an image as described above.

[0059] An embodiment of the present application provides a computer device, wherein the computer device includes the above-mentioned device for detecting a target object in any image.

[0060] It should be noted that: the target object detection device in the image provided by the above embodiment only uses the division of the above functional modules as an example when performing target object detection in the image. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the target object detection device in the image is divided into different functional modules to complete all or part of the functions described above. In addition, the target object detection device in the image provided by the above embodiment and the target object detection method embodiment in the image belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0061] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0062] The above description is not intended to limit the embodiments of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the protection scope of the embodiments of the present application.

Claims

1. A method for detecting a target object in an image, characterized in that: The method comprises: Obtaining an image to be recognized and a trained target detection model, wherein the target detection model includes at least a feature extraction network and a Transformer module; Using the feature extraction network to extract features from the image to obtain a feature map; Generating a heat map according to the feature map, wherein the heat map includes all potential target objects in the image; Initializing each query point according to the heat map, the query point representing a potential location of a target object in the image; The Transformer module is used to detect the feature map and each initialized query point to obtain the target object in the image.

2. The method for detecting a target object in an image according to claim 1, characterized in that: Initializing each query point according to the heat map includes: Get the preset allocation threshold; The positions in the heat map whose values ​​are greater than the allocation threshold are allocated to each query point, and different query points are separated by n positions, where n≥2.

3. The method for detecting a target object in an image according to claim 2, characterized in that: The step of allocating the positions in the heat map where the values ​​are greater than the allocation threshold to each query point comprises: When traversing the values ​​of each position in the heat map, if the value of the i-th row and j-th column is greater than the allocation threshold, the position of the i-th row and j-th column is allocated to a query point, where i and j are positive integers; Continue traversing the values ​​in the heat map from a new position that is n positions away from the position of the i-th row and j-th column.

4. The method for detecting a target object in an image according to claim 2, wherein: The step of allocating the positions in the heat map where the values ​​are greater than the allocation threshold to each query point further includes: If the value in the i-th row and j-th column is less than or equal to the allocation threshold, then continue to traverse the values ​​in the heat map from the new position in the i-th row and j+1-th column.

5. The method for detecting a target object in an image according to any one of claims 1 to 4, characterized in that: The using the Transformer module to detect the feature map and each initialized query point to obtain the target object in the image includes: The feature map and each initialized query point are processed using multiple Transformer layers in the Transformer module, so that after being superimposed by multiple Transformer layers, each query point gradually approaches the actual position of the target object it represents until each target object is detected.

6. A device for detecting a target object in an image, characterized in that: The device comprises: An acquisition module, used to acquire an image to be recognized and a trained target detection model, wherein the target detection model includes at least a feature extraction network and a Transformer module; An extraction module, used to extract features from the image using the feature extraction network to obtain a feature map; A generating module, configured to generate a heat map according to the feature map, wherein the heat map includes all potential target objects in the image; An initialization module, configured to initialize each query point according to the heat map, wherein the query point represents a potential target object in the image; The detection module is used to detect the feature map and each initialized query point using the Transformer module to obtain the target object in the image.

7. The target object detection device in an image according to claim 6, characterized in that: The initialization module is also used for: Get the preset allocation threshold; The positions in the heat map whose values ​​are greater than the allocation threshold are allocated to each query point, and different query points are separated by n positions, where n≥2.

8. The target object detection device in an image according to claim 7, characterized in that: The initialization module is also used for: When traversing the values ​​of each position in the heat map, if the value of the i-th row and j-th column is greater than the allocation threshold, the position of the i-th row and j-th column is allocated to a query point, where i and j are positive integers; Continue traversing the values ​​in the heat map from a new position that is n positions away from the position of the i-th row and j-th column.

9. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the method for detecting a target object in an image as described in any one of claims 1 to 5.

10. A computer device, characterized in that: The computer device comprises the device for detecting a target object in an image as claimed in any one of claims 6 to 8.

Citation Information

Patent Citations

  • Image detection method, device, equipment, storage medium and computer program product

    CN112597837A

  • Image detection method and device

    CN117593738A

  • DETR structure-based three-dimensional target detection method, device and system

    CN118823735A

  • Object detection method and apparatus, device, and storage medium

    WO2024183181A1