Feature map self-supervised dense crowd small target detection method and electronic equipment

Through the self-supervision method of feature maps, two-stage training combined with self-supervision loss function is adopted to force constraints on the deep feature map to retain shallow detail information, solving the problem of missed detection and high false detection rates for small target detection in dense populations and improving detection performance.

CN120339596AActive Publication Date: 2025-07-18SHENZHEN MAXVISION TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510828129.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

In dense crowd environments, small object detection is easily affected by factors such as occlusion, lighting, motion blur and low image resolution, resulting in high missed detection and false detection rates and insufficient detection capabilities of the existing technology.

Method used

The self-supervision method of feature maps is adopted to generate feature maps at different levels through multi-scale feature extraction, and a self-supervision loss function is constructed in combination with the structural similarity index SSIM, which forces and constrains the deep feature map to retain the detailed information of the shallow feature map, and optimize network parameters using dual-stage training.

Benefits of technology

It reduces the probability of missed and misdetection of small targets in dense populations, improves detection performance, and solves the problem of information attenuation and gradient imbalance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339596A_ABST
    Figure CN120339596A_ABST
Patent Text Reader

Abstract

The invention provides a feature map self-supervised dense crowd small target detection method and electronic equipment, and the method comprises the steps: S1, carrying out the multi-scale feature extraction of a training image through a target detection network, and generating feature maps of different levels; s2, in the first stage, pre-training a target detection network by using a main loss function of a detection task; s3, adding a self-supervision loss function in the second stage to perform joint training on the target detection network, wherein the second stage comprises the following steps: S3.1, calculating a structural similarity index SSIM for feature maps of adjacent hierarchies to serve as a self-supervision signal; and S3.2, constructing a self-supervised loss function based on the structural similarity index SSIM, forcibly constraining the deep feature map, retaining detail information of the shallow feature map, and jointly optimizing network parameters with a main loss function of a detection task. According to the feature map self-supervised dense crowd small target detection method and the electronic equipment, the detection performance of small targets in dense crowds can be comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of small target detection in dense crowds. More specifically, it relates to a method for detecting small targets in dense crowds with self-supervised feature maps and an electronic device. Background Art

[0002] Object detection is a basic algorithm in computer vision. The purpose of object detection is to determine the position coordinate information of the target of interest in the image, providing strong support for subsequent target analysis. Therefore, the accuracy of detection is directly related to the feasibility of subsequent analysis results.

[0003] However, in the real situation of a dense crowd environment, the human body target is not only small, but also limited by occlusion, lighting, motion blur, low image resolution, etc. This directly leads to the small human body target being prone to missed detection and false detection due to the small amount of information carried in its region of interest. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a method for detecting small targets in dense crowds with self-supervised feature maps and an electronic device, so as to solve the technical problem of insufficient detection ability in the process of detecting small targets in dense crowds in the prior art.

[0005] To achieve the above purpose, the technical solution adopted in this application is: providing a method for detecting small targets in dense crowds with self-supervised feature maps, the method for detecting small targets in dense crowds with self-supervised feature maps includes the steps of: S1, using an object detection network to perform multi-scale feature extraction on training images to generate feature maps of different levels; S2, pre-training the object detection network with the main loss function of the detection task in the first stage; S3, adding a self-supervised loss function in the second stage to jointly train the object detection network, which includes the steps of: S3.1, calculating the structural similarity index SSIM for adjacent-level feature maps as a self-supervised signal; S3.2, constructing a self-supervised loss function based on the structural similarity index SSIM to force the deep feature maps to retain the detailed information of the shallow feature maps, and jointly optimizing the network parameters with the main loss function of the detection task.

[0006] In a preferred embodiment, the formula for calculating the structural similarity index SSIM for adjacent-level feature maps is: SSIM(F1,F2) = (l(F1,F2) × c(F1,F2) × s(F1,F2)) α Among them, F1 and F2 represent two feature maps of adjacent levels, SSIM(F1, F2) represents the structural similarity index SSIM of two feature maps of adjacent levels, l(F1, F2) represents the luminance similarity between two images, c(F1, F2) represents the contrast similarity between two images, s(F1, F2) represents the structural similarity between two images, α is the dependent variable, and α is directly proportional to the density of the target population.

[0007] In a preferred embodiment, the method for calculating the structural similarity index SSIM for feature maps of adjacent levels includes the steps of: S3.11, calculating the mean value of the feature maps of each level along the channel dimension to generate a single-channel mean feature map F mean (x, y), and the calculation formula is: , where C is the number of feature maps of the corresponding level, and F i (x, y) is the value of the pixel coordinates (x, y) in the i-th feature map; S3.12, spatially aligning the mean feature maps of adjacent levels.

[0008] In a preferred embodiment, the method for calculating the structural similarity index SSIM for feature maps of adjacent levels further includes the steps of: S3.13, dividing the spatially aligned mean feature maps into s×s grids, applying different weights to each grid respectively, and calculating the structural similarity index SSIM by weighted calculation.

[0009] In a preferred embodiment, the grid weights are dynamically adjusted according to the center point position of the true annotation box, and the weight of the grid where the center of the annotation box falls is increased.

[0010] In a preferred embodiment, the calculation formula for the weight coefficient k of each grid is:

[0011] where n is the number of true annotation boxes in the mean feature map, and m is the number of center points of the true annotation boxes in the corresponding grid.

[0012] In a preferred embodiment, the calculation formula for the weight coefficient k of each grid is:

[0013] where n is the number of true annotation boxes in the current image, m is the number of center points of the true annotation boxes in the corresponding grid, a is the basic weight of each grid, and a is a positive integer.

[0014] In a preferred embodiment, s is taken as 4 and a is taken as 1.

[0015] In a preferred embodiment, the method for obtaining the self-supervised loss function L includes the steps of: S3.21, uniformly adjust the structural similarity index SSIM to [-1, 1]. The closer to 1, the higher the structural similarity between the two images; the closer to -1, the lower the structural similarity. S3.22, the formula for calculating the self-supervised loss function L is: , where X = 1 - SSIM(F1, F2).

[0016] In a preferred embodiment, after constructing the self-supervised loss function, the method further includes the steps of: Let the feature extraction layers of the object detection network be T1 to T n ; Based on the analysis of the structural similarity index SSIM of the feature maps of adjacent levels, determine that the feature extraction layer where the boundary structure of small targets in dense crowds undergoes a qualitative change is T m , where 1 < m < n; Establish the total self-supervised loss function L total = L 1&2 + L 2&3 + …… + L (m-1)&m , where L 1&2 represents the self-supervised loss between the feature map of the feature extraction layer T1 and the feature map of the feature extraction layer T2, and L 2&3 represents the self-supervised loss between the feature map of the feature extraction layer T2 and the feature map of the feature extraction layer T3, and L (m-1)&m represents the self-supervised loss between the feature map of the feature extraction layer T m-1 and the feature map of the feature extraction layer T m .

[0017] The present application also provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the above-mentioned method is implemented.

[0018] The beneficial effects of the feature map self-supervised dense crowd small target detection method and the electronic device provided by the present application are as follows: By adopting a two-stage training method, the mid-low-level consistency constrained by the self-supervised loss function and the high-level features constrained by the main loss function form a clear division of labor, cooperate with each other, and complement each other's advantages and disadvantages. The method forcibly constrains the explicit retention of shallow detail information in the deep network to solve the problems of information attenuation and gradient imbalance in small target detection, reduce the probability of missed detection and false detection, and comprehensively improve the detection performance of small targets in dense crowds. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a flowchart of the feature map self-supervised dense crowd small object detection method provided by the embodiment of the present application; Figure 2 It is a schematic diagram of Mosaic data augmentation provided by the embodiment of the present application; Figure 3 It is a flowchart of step S3 provided by the embodiment of the present application; Figure 4 It is a schematic diagram of dividing the spatially aligned mean feature map into an s×s grid provided by the embodiment of the present application; Figure 5 It is a flowchart of establishing the total self-supervised loss function provided by the embodiment of the present application; Figure 6 It is a schematic diagram of the object detection network provided by the embodiment of the present application; Figure 7 It is Figure 6 a schematic diagram after visualizing the hierarchical feature maps output by the object detection network in the first stage in Figure 8 It is Figure 6 a schematic diagram after visualizing the hierarchical feature maps output by the object detection network in the second stage in Detailed implementation manners

[0021] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application more clearly understood, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0022] It should be noted that when an element is referred to as "fixed to" or "disposed on" another element, it can be directly on the other element or indirectly on the other element. When an element is referred to as "connected to" another element, it can be directly connected to the other element or indirectly connected to the other element.

[0023] It should be understood that the orientation or positional relationship indicated by terms such as "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application.

[0024] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, the meaning of "a plurality" is two or more, unless otherwise specifically defined.

[0025] Please refer to Figure 1 together. Now, a feature map self-supervised dense crowd small object detection method provided by an embodiment of the present application will be described. The feature map self-supervised dense crowd small object detection method includes the steps: S1. Use an object detection network to perform multi-scale feature extraction on training images to generate feature maps of different levels.

[0026] In S1, in this embodiment, CNN, YOLOX, and SSD are preferably used as the basic detection networks. In another embodiment, it can also be adaptively replaced with the YOLOv8 network architecture in real-time dense detection scenarios, the DETR network architecture in complex occlusion scenarios, etc.

[0027] The size of the training images in this embodiment is 640×640. In another embodiment, the size of the training images being 640×640 can also be adaptively replaced with 320×320, 1280×1280, etc.

[0028] When performing multi-scale feature extraction on 640×640 training images, feature maps of 320×320, 160×160, 80×80, 40×40, and 20×20 can be obtained step by step.

[0029] S2. In the first stage, use the main loss function of the detection task to pre-train the object detection network.

[0030] In S2, it can be understood that the training using the main loss function of the detection task in the first stage essentially belongs to conventional supervised training, without adding self-supervised loss, and only making the network parameters converge preliminarily through conventional supervised training.

[0031] Please refer to Figure 2, for example, pre-training using the COCO dataset, adopting Mosaic data augmentation, which includes operations such as random scaling, cropping, and flipping. The number of training epochs can be determined according to the actual situation and is usually set to 10 to 30 epochs. Preferably, the number of training epochs is set to 15 epochs, the optimizer is SGD, the initial learning rate is 0.01, and the cosine annealing strategy is adopted to adjust the learning rate. The main loss function includes object detection losses, such as Obj Loss, classification loss, IoU Loss, etc.

[0032] Thus, the essence of the first stage is to utilize the advantages of supervised learning to directly optimize the task objective, enabling the object detection network to directly learn the mapping relationship from input to output through labeled data (such as classification labels and detection boxes), clarifying the task objective, and first allowing the object detection network to initially and quickly learn the basic human detection ability. Although the object detection network after pre-training has limitations in generalization ability (if the training data distribution is significantly different from the real scenario, such as annotation bias, the model performance may decline), it is difficult to handle unknown categories (unable to automatically identify categories not present in the training set), and the feature level is single (both classification labels and detection boxes are abstract semantic information, that is, only high-level features are available).

[0033] It is worth supplementing that the classification of the feature level includes: (1) Low-level features (edges / textures): The pixel information of small targets (such as pedestrians below 20×20 pixels) is extremely scarce, and their identifiability highly depends on edge sharpness (such as limb contours) and local textures (such as clothing folds). These information belong to low-level visual features.

[0034] (2) Intermediate features (geometric structures): The local geometric relationship of small targets (such as the relative position of the head and torso) is the key to distinguishing targets from noise and belongs to intermediate features.

[0035] (3) High-level features (semantics): The semantic features extracted by the deep network (such as the "person" and "car" categories annotated by the detection box) are often ineffective for small targets because small targets may only have a few pixels left in the deep feature map, and the semantic information has been severely lost.

[0036] S3. In the second stage, a self-supervised loss function is added for joint training of the object detection network, which includes the steps: S3.1. Calculate the structural similarity index SSIM for adjacent-level feature maps as the self-supervised signal; S3.2. Based on the structural similarity index SSIM, construct a self-supervised loss function to force the deep feature map to retain the detailed information of the shallow feature map, and jointly optimize the network parameters with the main loss function of the detection task.

[0037] In S3, in the second stage, a self-supervised loss function is also added for joint training of the object detection network. That is, during the training process of the second stage, there are both the supervision constraint of the main loss function and the self-supervision constraint of the self-supervised loss function.

[0038] Specifically, in S3.1, the structural similarity index SSIM (Structural Similarity) is calculated for adjacent-level feature maps as the self-supervised signal. The structural similarity index SSIM is an index used to evaluate the visual similarity of two images. It considers not only the luminance similarity and contrast similarity of the images, but also the structural similarity of the overall images.

[0039] It is worth adding that the luminance similarity aims to ensure the consistency of the overall intensity distribution of the target, and the contrast similarity aims to retain the edge and texture differences of the target. Among them, basic visual elements such as the overall intensity distribution, edges, and textures belong to low-level features. The structural similarity aims to ensure that key geometric structures (such as the contours of small targets) are not damaged. Among them, combined features such as geometric structures and local target contours belong to intermediate-level features. That is, the structural similarity index SSIM can force adjacent-level feature maps to align in the following dimensions: edge consistency (low-level), texture consistency (low-level), and geometric consistency (intermediate-level).

[0040] In this way, when jointly optimizing the network parameters with the main loss function of the detection task in S3.2, the low- and intermediate-level consistency constrained by the self-supervised loss function and the high-level features constrained by the main loss function form the following cooperative relationship: (1) Clear division of labor: The self-supervised loss function uses the self-supervised loss function to force the deep feature maps to retain the detailed information of the shallow feature maps, and forces the consistency of low-level and intermediate-level features, responsible for retaining the "physical presence" (edges / shapes) of the target and solving the problem of "where is the target" (localization). The semantic features under the constraint of the main loss function are responsible for judging the "identity" (category) of the target and solving the problem of "what is it" (classification).

[0041] (2) Synergistic effect: During the object detection process, the structure-enhanced feature maps constrained by the self-supervised loss function provide accurate position information, while the deep semantic features constrained by the main loss function provide class confidence. The combination of the two can reduce both missed detections (relying on structure) and false detections (relying on semantics) simultaneously.

[0042] (3) Complementary advantages and disadvantages: First, use the main loss function of the detection task to pre-train the object detection network, enabling the network to initially learn basic detection capabilities (such as the localization and classification of large targets), and initially narrow the gap between adjacent feature maps. On the basis of relatively stable network parameters, then introduce the self-supervised loss function to prevent situations such as difficult training convergence or "gradient explosion", and avoid gradient conflicts caused by early introduction of self-supervision.

[0043] Thus, compared with the prior art, the method for self-supervised dense crowd small object detection provided by this application, by adopting a two-stage training method, enables the middle and low-level consistency constrained by the self-supervised loss function and the high-level features constrained by the main loss function to form a cooperative relationship with clear division of labor, synergistic effect, and complementary advantages and disadvantages, and forcibly constrains the explicit retention of shallow detail information in the deep network to solve the problems of information attenuation and gradient imbalance in small object detection, reduce the probability of missed detection and false detection, and comprehensively improve the detection performance of small objects in dense crowds.

[0044] In a sparse population scenario, usually the whole human body can be seen; while in a dense crowd, the problem of human body occlusion will be infinitely magnified, and most human body targets only have one head, which will lead to the omission of the detection of the human body target. To solve the above problems, the following embodiments are adopted in this application: The formula for calculating the structural similarity index SSIM for adjacent-level feature maps is: SSIM(F1,F2) = (l(F1,F2) × c(F1,F2) × s(F1,F2)) α Wherein, F1 and F2 represent two adjacent-level feature maps, SSIM(F1,F2) represents the structural similarity index SSIM of two adjacent-level feature maps, l(F1,F2) represents the luminance similarity between two images, c(F1,F2) represents the contrast similarity between two images, s(F1,F2) represents the structural similarity between two images, α is a dependent variable, and α is directly proportional to the density of the target population.

[0045] It can be understood that the calculation method of the structural similarity index SSIM in this embodiment is different from the conventional structural similarity index SSIM. l(F1,F2) × c(F1,F2) × s(F1,F2) ∈ [0,1]. The closer to 1, the higher the structural similarity of two adjacent-level feature maps, and the closer to 0, the lower the structural similarity of two adjacent-level feature maps. Since α is directly proportional to the density of the target population, when the crowd is dense and α increases, SSIM(F1,F2) decreases accordingly. After the decrease of SSIM(F1,F2) is fed back to the self-supervised loss function, it strengthens the constraint for the deep feature map to retain the detail information of the shallow feature map; similarly, when the crowd is sparse, it weakens the constraint for the deep feature map to retain the detail information of the shallow feature map. In this way, it achieves an adaptive scenario, dynamically adjusts the degree of constraint, and takes into account the improvement of small object detection performance and training efficiency.

[0046] Since the number of feature maps at each level is inconsistent, and the scales of feature maps at different levels are also inconsistent. However, the structural similarity index SSIM generally measures the structural similarity between two images. It can be seen that there are technical barrier problems of large computational complexity and inaccuracy when calculating the structural similarity index SSIM for feature maps of adjacent levels. To solve the above problems, the following embodiments are adopted in this application: Please refer to Figure 3 simultaneously. A method for calculating the structural similarity index SSIM for feature maps of adjacent levels includes the steps: S3.11, calculating the mean value of the feature maps of each level along the channel dimension to generate a single-channel mean feature map F mean (x, y). The calculation formula is: , where C is the number of feature maps at the corresponding level, and F i (x, y) is the value of the pixel coordinate (x, y) in the i-th feature map; S3.12, spatially aligning the mean feature maps of adjacent levels.

[0047] It can be understood that in this embodiment, the single-channel mean feature map F mean (x, y) of each level is calculated through the above formula, so as to unify the feature maps at each scale. Then, the mean feature maps of adjacent levels are spatially aligned. Thus, two images with the same size can be obtained, and each image can represent the feature structure of its corresponding level. In this way, the technical barrier problems of large computational complexity and inaccuracy when calculating the structural similarity index SSIM for feature maps of adjacent levels are solved.

[0048] Please refer to Figure 3 and Figure 4 simultaneously. Further, a method for calculating the structural similarity index SSIM for feature maps of adjacent levels further includes the steps: S3.13, dividing the spatially aligned mean feature map into s×s grids, applying different weights to each grid respectively, and calculating the structural similarity index SSIM by weighted calculation.

[0049] It can be understood that the formation of small target population images is usually formed by long-distance shooting. Thus, when applying different weights to each grid respectively, the weights of the grids in the upper region of the image can be increased, and / or the weights of the grids in the lower region of the image can be reduced to enhance the attention of small targets at a long distance. Of course, according to empirical values, the weights of the grids in the central region of the image can also be increased, and / or the weights of the grids in the edge region of the image can be reduced to enhance the attention of small targets in the central region of interest and reduce the attention of the edge background region.

[0050] Thus, when calculating the structural similarity index SSIM using the two mean feature maps after spatial alignment, the factor of large variance in pixel values of the overall pictures of the two images can be overcome, the detail supervision ability of the feature maps for small targets in the crowd can be improved, and further the detection performance of small targets in dense crowds can be enhanced.

[0051] In a best embodiment, the grid weights are dynamically adjusted according to the center point position of the true annotation box, and the grid weights where the center of the annotation box falls are increased.

[0052] Since annotation boxes are set for the training images during the pre-training process, in this embodiment, the grid weights can be adaptively and dynamically adjusted directly according to the center points of the annotation boxes of the true small human targets, increasing the grid weights where the crowd is dense and decreasing the grid weights where the crowd is sparse. In this way, it can accurately adapt to images in any scenario, making the detail supervision ability of the feature maps for small targets in the crowd reach the best, and further enhancing the detection performance of small targets in dense crowds.

[0053] In a specific embodiment, the calculation formula for the weight coefficient k of each grid is:

[0054] where n is the number of true annotation boxes in the mean feature map, and m is the number of center points of the true annotation boxes in the corresponding grid.

[0055] It can be understood that based on the calculation formula of the weight coefficient k in this embodiment, only the grids containing the center points of the annotation boxes will obtain weights, and the weights of the remaining grids are 0. Therefore, this embodiment has the following effects: (1) The self-supervised loss completely ignores the background area and only forces the feature maps around the target center to maintain structural consistency, enhancing the binding force of the deep feature maps to retain the detail information of the shallow feature maps.

[0056] (2) Since the grid weights of the center points of the true annotation boxes are 0, only the SSIM of some grids needs to be calculated, reducing the amount of computation, improving the training efficiency, and accelerating the convergence of the self-supervised loss function.

[0057] (3) The gradient is completely driven by the target area, enhancing the anti-interference ability.

[0058] However, at the same time, since only the grids containing the targets are calculated for the loss in the above embodiment, there are also the following disadvantages: the model will completely ignore the feature expression of the background area, fail for the slightly offset target edges, and the unconstrained background area may be misjudged as unknown targets, that is, the robustness decreases.

[0059] In an improved embodiment, the calculation formula for the weight coefficient k of each grid is:

[0060] Where n is the number of true annotation boxes of the current image, m is the number of center points of the true annotation boxes in the corresponding grid, a is the basic weight of each grid, and a is a positive integer.

[0061] It can be understood that by setting the basic weight for each grid, it is ensured that the target-free area (such as a pure background) can still participate in self-supervised learning, avoiding excessive bias, maintaining uniform initialization, ensuring that the network continuously learns the global feature consistency, and thus improving the robustness. For example, there may be a 1-2 pixel offset in the center of the annotation box (especially in a dense scene), or some small target annotations may be missed (such as a heavily occluded pedestrian).

[0062] In order to balance the detection performance and robustness of extremely dense crowd small targets, in an optimal specific embodiment, the value of s is 4, and the value of a is 1.

[0063] In a preferred embodiment, the method for obtaining the self-supervised loss function L includes the steps of: S3.21, uniformly adjust the structural similarity index SSIM to [-1, 1]. The closer it is to 1, the higher the structural similarity between the two images. The closer the value is to -1, the lower the structural similarity. S3.22, the formula for calculating the self-supervised loss function L is: , where X = 1 - SSIM(F1, F2).

[0064] It can be understood that the calculation formula of the self-supervised loss function L in this embodiment combines the advantages of both L1 loss and L2 loss. When the SSIM(F1, F2) value of the two feature maps F1 and F2 is 1, that is, when they are very similar, the loss reaches the minimum.

[0065] In a preferred embodiment, please refer to Figure 5 together. After constructing the self-supervised loss function, it further includes the steps of: Let the feature extraction layers of the object detection network be T1 to T n ; Based on the analysis of the structural similarity index SSIM of the feature maps of adjacent levels, determine that the feature extraction layer where the boundary structure of the dense crowd small target undergoes a qualitative change is T m , where 1 < m < n; Establish the total self-supervised loss function L total = L 1&2 + L 2&3 + …… + L (m-1)&m where L 1&2 represents the self-supervised loss between the feature map of the feature extraction layer T1 and the feature map of the feature extraction layer T2, L2&3 Represents the self-supervised loss, L, between the feature map of the feature extraction layer T2 and the feature map of the feature extraction layer T3 (m-1)&m Represents the feature extraction layer T m-1 's feature map and the self-supervised loss of the feature map of the feature extraction layer T m 's feature map.

[0066] It can be understood that, taking the object detection network shown in Figure 6 as an example, when pre-training the object detection network with the main loss function of the detection task in the first stage, an input training image of 640×640, the visualization of the feature maps of each level output is as shown in Figure 7 As shown, it can be seen that as the level gradually deepens, the resolution of the feature map gradually becomes smaller, and the details of the small objects (human bodies) in the picture gradually disappear. By the dark3 layer (T3), the size of the feature map is 80x80, and the information of small objects can still be clearly distinguished. However, by the dark4 layer (T4) and dark5 layer (T5), the boundary information of small objects can no longer be distinguished at all. It should be noted that in the 80x80 feature map of the PAFPN (feature fusion module), the boundary information of small objects is not as obvious as that in the dark3 feature map. After analyzing a large number of different training images, it is found that there are nodes where the boundary structure of small objects in dense crowds undergoes a qualitative change during the continuous feature extraction process for each training image, and the level where the node of qualitative change occurs may also be different according to the size of the detection box and the degree of crowd density.

[0067] Thus, in this embodiment, based on the analysis of the structural similarity index SSIM of adjacent-level feature maps, the feature extraction layer where the boundary structure of small objects in each training image undergoes a qualitative change can be determined. And in the object detection network, small objects are mainly detected by large-resolution feature maps. Therefore, in order to improve the detection ability of the model for small objects in dense crowds, the detailed information of large-resolution feature maps needs to be improved. Therefore, according to the method of this embodiment, the total self-supervised loss function L total is introduced. During the training process of the second stage, the visualization of the feature maps of each level output is as shown in Figure 8 As shown, it can be seen that as the level gradually deepens, the change in the boundary structure of small objects in the feature maps from the stem layer (T1) to the dark3 layer (T3) is not obvious compared with the corresponding feature maps in the first stage. However, when it reaches the dark4 layer (T4), the clarity of the boundary structure of small objects still maintains a relative gradient loss, and there is an obvious improvement compared with the change in the boundary structure of small objects in the corresponding feature maps in the first stage.

[0068] It can be seen that after establishing the total self-supervised loss function L according to the above formula total there are two beneficial effects: (1) Determine that the feature extraction layer where the small target boundary structure undergoes qualitative change is T m After that, only perform forced constraints on its upper-level feature map, and discard the ineffective forced constraints on the lower-level feature map, thereby improving the self-supervised effect.

[0069] (2) According to the above formula for the total self-supervised loss function L total associate and constrain the self-supervised losses between multiple consecutive levels, which is beneficial to the loss gradient between levels and further prevents situations such as difficult training convergence or "gradient explosion".

[0070] This application also provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the above-mentioned method is implemented.

[0071] The electronic device can be an intelligent device such as an intelligent camera, a mobile phone, a computer, or an intelligent vehicle.

[0072] The above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of this application shall be included in the protection scope of this application.

Claims

1. A self-supervised small target detection method for dense crowds in feature maps, characterized in that Including the steps: S1. Use a target detection network to perform multi-scale feature extraction on training images to generate feature maps of different levels. S2. In the first stage, pre-train the target detection network using the main loss function of the detection task. S3. In the second stage, add a self-supervised loss function to jointly train the target detection network, which includes the steps: S3.

1. Calculate the structural similarity index SSIM for adjacent-level feature maps as a self-supervised signal. S3.

2. Based on the structural similarity index SSIM, construct a self-supervised loss function to force the deep feature maps to retain the detailed information of the shallow feature maps, and jointly optimize the network parameters with the main loss function of the detection task.

2. The feature map self-supervised dense crowd small object detection method according to claim 1, characterized in that The formula for calculating the structural similarity index SSIM for adjacent-level feature maps is: SSIM(F1,F2) = (l(F1,F2) × c(F1,F2) × s(F1,F2)) α where F1 and F2 represent two adjacent-level feature maps, SSIM(F1,F2) represents the structural similarity index SSIM of the two adjacent-level feature maps, l(F1,F2) represents the luminance similarity between the two images, c(F1,F2) represents the contrast similarity between the two images, s(F1,F2) represents the structural similarity between the two images, α is a dependent variable, and α is directly proportional to the density of the target population.

3. The feature map self-supervised dense crowd small object detection method according to claim 2, characterized in that, The method for calculating the structural similarity index SSIM for adjacent-level feature maps includes the steps: S3.

11. Calculate the mean value of the feature maps at each level along the channel dimension to generate a single-channel mean feature map F mean (x, y), and the calculation formula is: , where C is the number of feature maps at the corresponding level, and F i (x, y) is the value of the pixel coordinates (x, y) in the i-th feature map; S3.

12. Align the mean feature maps of adjacent levels spatially.

4. The feature map self-supervised dense crowd small object detection method according to claim 3, characterized in that The method for calculating the structural similarity index SSIM for adjacent-level feature maps further includes the steps: S3.

13. Divide the spatially aligned mean feature maps into s×s grids, apply different weights to each grid respectively, and calculate the structural similarity index SSIM by weighted calculation.

5. The feature map self-supervised dense crowd small object detection method according to claim 4, wherein Dynamically adjust the grid weights according to the center point position of the true annotation box, and increase the weight of the grid where the center of the annotation box falls.

6. The feature map self-supervised dense crowd small target detection method according to claim 5, wherein The calculation formula for the weight coefficient k of each grid is: , where n is the number of true annotation boxes in the mean feature map, and m is the number of center points of the true annotation boxes in the corresponding grid.

7. The feature map self-supervised dense crowd small object detection method according to claim 5, characterized in that, The calculation formula for the weight coefficient k of each grid is: , where n is the number of true annotation boxes in the current image, m is the number of center points of the true annotation boxes in the corresponding grid, and a is the base weight of each grid, and a is a positive integer.

8. The feature map self-supervised dense crowd small object detection method according to claim 1, wherein The method for obtaining the self-supervised loss function L includes the steps: S3.

21. Uniformly adjust the structural similarity index SSIM to [-1,1]. The closer it is to 1, the higher the structural similarity between the two images, and the closer the value is to -1, the lower the structural similarity. S3.

22. The formula for calculating the self-supervised loss function L is: , where X = 1 - SSIM(F1,F2).

9. The feature map self-supervised dense crowd small object detection method according to claim 8, characterized in that, After constructing the self-supervised loss function, it further includes the steps: Let the feature extraction layers of the object detection network be T1 to T in sequence n ; Based on the analysis of the structural similarity index (SSIM) of the feature maps of adjacent levels, it is determined that the feature extraction layer where the boundary structure of small targets in dense crowds undergoes a qualitative change is T m , where 1 < m < n; Establish the total self-supervised loss function L total =L 1&2 +L 2&3 +……+L (m-1)&m , where L 1&2 represents the self-supervised loss between the feature map of the feature extraction layer T1 and the feature map of the feature extraction layer T2, and L 2&3 represents the self-supervised loss between the feature map of the feature extraction layer T2 and the feature map of the feature extraction layer T3, and L (m-1)&m represents the self-supervised loss between the feature map of the feature extraction layer T m-1 and the feature map of the feature extraction layer T m .

10. An electronic device includes a memory, a processor, and a computer program stored on the memory, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Small-scale target detection method based on weak edge

    CN110852317A

  • Progressive cascaded face detection method based on feature enhancement in unconstrained scene

    CN111553230A

  • Unsupervised defect detection method based on quantization auto-encoder

    CN115375604A

  • Training method of inter-class occlusion target detection network model based on weak supervision semantic segmentation and feature compensation

    CN116503603A

  • Seismic image processing method based on Swin Transform

    CN118710523A