Selective Inheritance Learning Method, Device, and Electronic Device for Target Counting

Through selective inheritance learning methods, extracting and fusing features with different resolutions, the problem of scale changes in target counting is solved, and efficient target counting and good generalization ability are achieved.

CN116740481BActive Publication Date: 2025-06-27SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310765059.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-27
Publication Date
2025-06-27
Estimated Expiration
2043-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of scale changes in target counting, resulting in increased complexity of the overall counting method, large training overhead and long inference time.

Method used

The selective inheritance learning method is adopted to extract the target image input feature feature network, extract multiple first features of different resolutions, and extract and pass the second feature from the low-resolution fusion features through a preset selective inheritance adapter, fuse to generate high-resolution fusion features, and finally perform target counting at the highest resolution.

Benefits of technology

It effectively solves the problem of scale changes, improves prediction accuracy and counting performance, reduces inference time, and maintains good generalization ability on different scales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116740481B_ABST
    Figure CN116740481B_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer vision technology, and discloses a selective inheritance learning method, device and electronic device for object counting. The method inputs a target image into a feature extraction network to extract a plurality of first features with different resolutions; in the order of increasing resolution, uses a preset selective inheritance adapter to extract and transfer a second feature from the low-resolution fused feature, and fuses the second feature with the high-resolution first feature to generate a high-resolution fused feature; wherein, the second feature is the most discriminative feature region in the low-resolution fused feature; uses the fused feature at the highest resolution as the target feature, and performs object counting based on the target feature. The embodiments of the present application can select the most discriminative feature regions for feature learning at each resolution, and gradually inherit the discriminative features from low resolution to high resolution, which can effectively solve the problem of scale change.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and particularly to a selective inheritance learning method, device, and electronic device for object counting. Background Art

[0002] Object counting using computer vision technology has a wide range of applications. For example, counting people in a dense scene can estimate the population density and prevent stampede accidents. Counting vehicles in a traffic scene can assist traffic management. Cell counting in a medical scene can assist in disease diagnosis. Animal counting in the wild can help protect rare species, and crop counting can accurately estimate crop yields.

[0003] Deep learning is a commonly used paradigm for object counting currently. The constructed deep model first extracts effective features of the input image regarding the object to be counted, then outputs a single-channel density map. Finally, it is supervised by the Gaussian density map generated by the labeled object center points, thereby continuously optimizing the model to make the distance between the predicted density map and the true density map continuously shrink.

[0004] In the existing technical solutions, the method of scaling image patches is a new paradigm for solving scale changes in object counting. In the first stage, it first estimates an appropriate scaling factor for a given image patch or feature patch, and then scales the image or feature to a suitable scale according to the scale prediction factor before performing the density prediction in the second stage. This method increases the implementation complexity of the overall counting method and the training cost. Secondly, it is difficult to predict the density prediction map of the entire image at one time, increasing the time required for inference.

[0005] The multi-scale fusion method is also widely used to construct a scale-aware object counting method. Generally speaking, the multi-scale fusion method first constructs a multi-scale feature representation of the image, and then uses an attention mechanism to weight and fuse features or density maps of different resolutions. However, since the multi-scale fusion method does not independently optimize the target scale range that each scale is good at, this type of method cannot effectively solve the problem of scale changes. Summary of the Invention

[0006] The embodiments of this application provide a selective inheritance learning method for object counting to solve the problem in the prior art that the problem of scale changes cannot be effectively solved.

[0007] Correspondingly, the embodiments of this application also provide a selective inheritance learning device for object counting, an electronic device, and a computer-readable storage medium to ensure the implementation and application of the above method.

[0008] To solve the above technical problems, an embodiment of the present application discloses a selective inheritance learning method for object counting, and the method includes:

[0009] Input the target image into the feature extraction network to extract multiple first features with different resolutions;

[0010] In the order from low to high resolution, use a preset selective inheritance adapter to extract and transfer the second feature from the low-resolution fused feature, and fuse the second feature with the high-resolution first feature to generate a high-resolution fused feature;

[0011] Use the fused feature at the highest resolution as the target feature, and perform object counting based on the target feature;

[0012] Wherein, the fused feature at the lowest resolution is the first feature at the lowest resolution; the selective inheritance adapter is arranged between adjacent first features; the second feature is the most discriminative feature region in the low-resolution fused feature.

[0013] An embodiment of the present application also discloses a selective inheritance learning device for object counting, and the device includes:

[0014] A feature extraction module, configured to input the target image into the feature extraction network to extract multiple first features with different resolutions;

[0015] A feature transfer module, configured to, in the order from low to high resolution, use a preset selective inheritance adapter to extract and transfer the second feature from the low-resolution fused feature, and fuse the second feature with the high-resolution first feature to generate a high-resolution fused feature;

[0016] An object counting module, configured to use the fused feature at the highest resolution as the target feature, and perform object counting based on the target feature;

[0017] Wherein, the fused feature at the lowest resolution is the first feature at the lowest resolution; the selective inheritance adapter is arranged between adjacent first features; the second feature is the most discriminative feature region in the low-resolution fused feature.

[0018] An embodiment of the present application also discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor implements one or more of the methods in the embodiments of the present application when executing the program.

[0019] The embodiments of the present application also disclose a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the methods described in one or more of the embodiments of the present application are implemented.

[0020] In the embodiments of the present application, a target image is input into a feature extraction network to extract multiple first features with different resolutions. In the order from the lowest resolution to the highest resolution, a second feature is extracted and transferred from the low-resolution fusion feature by using a preset selective inheritance adapter, and the second feature is fused with the high-resolution first feature to generate a high-resolution fusion feature. Among them, the fusion feature with the lowest resolution is the first feature with the lowest resolution. The selective inheritance adapter is arranged between adjacent first features. The second feature is the most discriminative feature region in the low-resolution fusion feature. Then, the fusion feature at the highest resolution is used as the target feature, and target counting is performed based on the target feature. The embodiments of the present application can select the most discriminative feature regions for feature learning at each resolution and gradually inherit the discriminative features from the low resolution to the high resolution, which can effectively solve the problem of scale change.

[0021] Additional aspects and advantages of the embodiments of the present application will be given in the following description section, which will become apparent from the following description or be understood through the practice of the present application. Description of the Drawings

[0022] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0023] Figure 1 is a flowchart of the selective inheritance learning method for target counting provided by the embodiments of the present application;

[0024] Figure 2 is a schematic structural diagram of the first example provided by the embodiments of the present application;

[0025] Figure 3 is a schematic structural diagram of the second example provided by the embodiments of the present application;

[0026] Figure 4 is a comparison diagram of effects provided by the embodiments of the present application;

[0027] Figure 5 is a schematic structural diagram of the selective inheritance learning method for target counting provided by the embodiments of the present application;

[0028] Figure 6 is a schematic structural diagram of the electronic device provided by the embodiments of the present application. Detailed Embodiments

[0029] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and should not be construed as a limitation of the present application.

[0030] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more of the associated listed items.

[0031] Those skilled in the art can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0032] The solution provided by the embodiments of the present application can be executed by any electronic device, such as a terminal device or a server. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this. Regarding the technical problems existing in the prior art, the selective inheritance learning method, device, and electronic device provided by the present application for target counting are intended to solve at least one of the technical problems in the prior art.

[0033] The following uses specific embodiments to elaborate in detail on the technical solution of the present application and how the technical solution of the present application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0034] An embodiment of the present application provides a possible implementation manner. As Figure 1 shown, a flowchart of a selective inheritance learning method for target counting is provided. This solution can be executed by any electronic device. Optionally, it can be executed on the server side or the terminal device.

[0035] As Figure 1 shown in, this method may include the following steps:

[0036] Step 101, input the target image into the feature extraction network to extract multiple first features with different resolutions.

[0037] In the embodiment of the present application, the target image may be an RGB image to be processed. Generally, low resolution is good at capturing large-size objects, while high resolution is sensitive to small-size objects. Therefore, it is crucial to extract multiple corresponding first features at multiple resolutions from the target image.

[0038] Step 102, in the order from low to high resolution, use a preset selective inheritance adapter to extract and transfer the second feature from the low-resolution fused feature, and fuse the second feature with the high-resolution first feature to generate a high-resolution fused feature.

[0039] Among them, the fused feature with the lowest resolution is the first feature with the lowest resolution; the selective inheritance adapter is arranged between adjacent first features; the second feature is the most discriminative feature region in the low-resolution fused feature.

[0040] When fusing low-resolution features into high-resolution features, directly upsampling the first feature has two drawbacks: First, multiple upsamplings will damage the already converged features at each resolution, resulting in a decrease in the confidence of the density map of large-size objects; Second, the weak discriminative features of large-size objects at high resolution are noise, and directly upsampling and fusing will introduce this part of the noise into the high resolution.

[0041] In the embodiments of the present application, multiple first features are processed in the order of increasing resolution. Specifically, a selective inheritance adapter can be used between adjacent first features. The low resolution and high resolution mentioned in the context refer to two adjacent first features. The most discriminative features in the low-resolution first features are passed to the high resolution and then fused with the high-resolution first features to achieve inheritance learning from low resolution to high resolution, which can improve the prediction accuracy.

[0042] Step 103: Use the transferred features at the highest resolution as target features and perform target counting based on the target features.

[0043] In the embodiments of the present application, all targets are finally counted at the highest resolution, which can improve the counting performance.

[0044] In the embodiments of the present application, a target image is input into a feature extraction network to extract multiple first features with different resolutions; in the order of increasing resolution, a preset selective inheritance adapter is used to extract and transfer second features from the low-resolution fused features, and the second features are fused with the high-resolution first features to generate high-resolution fused features; among them, the fused features at the lowest resolution are the first features at the lowest resolution; the selective inheritance adapter is arranged between adjacent first features; the second features are the most discriminative feature regions in the low-resolution fused features; then, the fused features at the highest resolution are used as target features and target counting is performed based on the target features. The embodiments of the present application can select the most discriminative feature regions for feature learning at each resolution and gradually inherit the discriminative features from low resolution to high resolution, which can effectively solve the problem of scale variation.

[0045] As a first example, as Figure 2 shown, for the input RGB image, first use a multi-scale feature extraction network (Feature Extractor) to extract four first features with different resolutions. The multi-scale feature extraction network in this embodiment can use existing convolutional backbone networks, such as HRNet-W48, VGG-19, etc. The lowest resolution is 1 / 32 of the input RGB image. The spatial resolution of each scale space increases by 2 times from bottom to top, that is, in the order of increasing resolution (in the order from bottom to top in Figure 2 ), the first features at the four different resolutions are 1 / 32 of the input RGB image (hereinafter referred to as 1 / 32 resolution), 1 / 16 of the input RGB image (hereinafter referred to as 1 / 16 resolution), 1 / 8 of the input RGB image (hereinafter referred to as 1 / 8 resolution), and 1 / 4 of the input RGB image (hereinafter referred to as 1 / 4 resolution).

[0046] A selective inheritance adapter FSIA is disposed between adjacent first features, that is, between the first feature R3 with a resolution of 1 / 32 and the first feature R2 with a resolution of 1 / 16, between the first feature R2 with a resolution of 1 / 16 and the first feature R1 with a resolution of 1 / 8, and between the first feature R1 with a resolution of 1 / 8 and the first feature R0 with a resolution of 1 / 4. In this embodiment, starting from the first feature with the lowest resolution, that is, using the selective adapter between the first feature with a resolution of 1 / 32 and the first feature with a resolution of 1 / 16 to process the first feature R3 with a resolution of 1 / 32, the first feature R3 with a resolution of 1 / 32 can be directly used as the fused feature with a resolution of 1 / 32. From the fused feature extract the second feature with a resolution of 1 / 32 and transfer the second feature with a resolution of 1 / 32 to the resolution of 1 / 16; at the resolution of 1 / 16, fuse the second feature with a resolution of 1 / 32 and the first feature R2 with a resolution of 1 / 16 to generate a fused feature with a resolution of 1 / 16. And transfer the fused feature down to the next 1 / 8 resolution to be processed. And so on, until a fused feature with a resolution of 1 / 4 is obtained. Use the fused feature as the target feature and input it into a preset counting head (that is, the Counting Head in Figure 2 ) for target counting.

[0047] In an alternative embodiment, the selective inheritance adapter is provided between each set of adjacent first features, and the selective inheritance adapter is the same one.

[0048] Combined with the above first example, as Figure 2 shown, selective inheritance adapters are provided between the first feature with a resolution of 1 / 8 and the first feature with a resolution of 1 / 4, between the first feature with a resolution of 1 / 4 and the first feature with a resolution of 1 / 2, and between the first feature with a resolution of 1 / 2 and the first feature with a resolution of 1 / 1. The selective inheritance adapters set at these three places are the same one. By sharing the selective inheritance adapter, it can be ensured that the discriminative features selected at different resolutions have the same distribution.

[0049] In an alternative embodiment, the selective inheritance adapter includes a scale non-custom feature forward propagation network and a soft mask generator. Using the preset selective inheritance adapter to extract and transfer the second feature from the fused feature with a lower resolution includes:

[0050] Process the fused feature with a lower resolution using the soft mask generator to obtain an unaggregated feature, and generate a first attention map using the unaggregated feature.

[0051] Use the first attention map to disentangle the low-resolution fused features to obtain the second feature.

[0052] As a second example, as Figure 3 shown, UFN is a scale non-custom feature forward propagation network; SMG is a soft mask generator; R j is the low-resolution fused feature. The scale non-custom feature forward propagation network UFN continuously and independently propagates the scale non-custom features forward to high resolution and plays a role in the next feature selection. Specifically, the soft mask generator SMG makes a discrimination on the low-resolution fused feature to determine the aggregated features and unaggregated features, and generates the first attention map using the unaggregated features. Then, the scale non-custom feature forward propagation network UFN uses the first attention map to disentangle the low-resolution fused feature to obtain the most discriminative feature region of the low-resolution fused feature as the second feature. The scale non-custom feature forward propagation network UFN obtains the most discriminative second feature, which can suppress the noise components in the second feature.

[0053] In an optional embodiment, before using the fused feature at the highest resolution as the target feature and performing target counting based on the target feature, the method further includes:

[0054] Extract prediction features from the first feature using the selective inheritance adapter;

[0055] Obtain the prediction result at the corresponding resolution using the prediction features, and calculate the loss value.

[0056] In the embodiments of the present application, each resolution's processing is regarded as a network branch. Then, for the first feature of each network branch, prediction features can be extracted, and then the prediction density map (i.e., the prediction result) corresponding to this resolution can be obtained based on the prediction features. By comparing the prediction density map with the preset true density map, the loss value of the prediction density map is calculated, and the network branch corresponding to this resolution is trained using the loss value to improve the network accuracy.

[0057] Combined with the above first example, as Figure 2 shown, taking the branch corresponding to the 1 / 8 resolution as an example, use the selective inheritance adapter FSIA to extract the prediction feature O1 from the first feature R1, input the prediction feature O1 into the counting head Counting Head (Copy), and obtain the prediction density map at the 1 / 2 resolution after 2 upsampling processes in the counting head Use the prediction density map and the preset true density map Calculate the loss value l1 corresponding to this branch. Similarly, for the predicted feature O3 obtained by processing the branch corresponding to the 1 / 32 resolution, after 2 times of upsampling processing in the counting head Counting Head (Copy), a predicted density map with a resolution of 1 / 8 is obtained, and the loss value l3 corresponding to this branch is calculated by using the predicted density map and the preset true density map; for the predicted feature O2 obtained by processing the branch corresponding to the 1 / 16 resolution, after 2 times of upsampling processing in the counting head Counting Head (Copy), a predicted density map with a resolution of 1 / 4 is obtained, and the loss value l2 corresponding to this branch is calculated by using the predicted density map and the preset true density map; for the predicted feature O0 obtained by processing the branch corresponding to the 1 / 4 resolution, after 2 times of upsampling processing in the counting head Counting Head, a predicted density map with a resolution of 1 / 1 is obtained, and the loss value l0 corresponding to this branch is calculated by using the predicted density map and the preset true density map.

[0058] In an alternative embodiment, the selective inheritance adapter further includes a scale-customized feature forward propagation network. The extracting of the predicted feature from the first feature by using the selective inheritance adapter includes:

[0059] Process the low-resolution fused feature by using a soft mask generator to obtain the aggregated feature, and generate a second attention map by using the aggregated feature;

[0060] Disentangle the low-resolution fused feature by using the second attention map to obtain the predicted feature.

[0061] Combined with the above second example, as Figure 3 shown, CFN is the scale-customized feature forward propagation network; SMG is the soft mask generator; is the low-resolution fused feature. The soft mask generator SMG makes a discrimination on the low-resolution fused feature to determine the aggregated feature and the unaggregated feature, and generate a second attention map by using the aggregated feature. Then, the scale-customized feature forward propagation network CFN uses the second attention map to disentangle the low-resolution fused feature to obtain a third feature, and fuses the third feature with the high-resolution first feature to generate the predicted feature

[0062] In an alternative embodiment, the determining of processing the low-resolution fused feature by using the soft mask generator to obtain the unaggregated feature and generating a first attention map by using the unaggregated feature includes:

[0063] Determine the target region in the fused features of the low resolution;

[0064] Generate the first attention map using the unaggregated features in the target region.

[0065] In an optional embodiment, the determining the target region in the fused features of the low resolution includes:

[0066] Calculate the loss value for the prediction results of the same feature region at different resolutions;

[0067] Take the feature region as the target feature corresponding to the resolution with the minimum loss value.

[0068] Combined with the above first example, as Figure 3 shown, for the same feature region, process it at 1 / 32 resolution, 1 / 16 resolution, 1 / 8 resolution, and 1 / 4 resolution respectively to obtain the prediction results of this region. If the loss value of the prediction result of this region at 1 / 32 resolution is the smallest, then take this region as the target region in the first feature at 1 / 32 resolution. During the training and prediction process, at 1 / 32 resolution, mainly focus on the features in the target feature region.

[0069] In this embodiment, taking Figure 2 the first feature R2 at 1 / 16 resolution in Figure 3 as an example, specifically: find the feature region in the first feature R2 that is most suitable for prediction at the current stage, and then set the corresponding region to 1 in the mask map that is all 0 to generate the first attention map. By applying the first attention map to the predicted density map and the true density map, the first feature R2 will focus on the target region closest to the true distribution, realizing the training of the selective inheritance adapter FSIA. In addition to training the feature selection ability of the selective inheritance adapter FSIA, the first feature R2 can also inherit the most discriminative second feature in the first feature R3. That is, re-optimize the most discriminative region in the first feature R3 in the optimization target at the stage of processing the first feature R2. Therefore, the feature selection and inheritance learning in this step are supervised by two loss functions respectively. Among them, the inheritance loss will activate the selective inheritance adapter FSIA, enabling the soft mask generator SMG to decouple the first feature R3 into aggregated features and unaggregated features. At the same time, it also ensures that the features after fusing the second feature in the first feature R2 and the first feature R3 can still give accurate predictions like the first feature R3. Through this hierarchical selection and inheritance learning, finally, all the targets are collected at the highest resolution for the final prediction.

[0070] In an optional embodiment, both the scale non-custom feature forward propagation network and the scale custom feature forward propagation network include two convolutional layers;

[0071] The soft mask generator includes three convolutional kernels.

[0072] The embodiments of the present application have been verified through multiple experiments. Its effectiveness has been verified on nine large-scale object counting or localization datasets, namely NWPU-Crowd, Shanghai Tech Part A, Shanghai Tech Part B, UCF-QNRF, FDST, JHU-CROWD++, MTC, URBAN_TREE, and TRANCOS, and relatively advanced performance has been achieved in each dataset. In the most challenging and largest-scale-changing NWPU-Crowd dataset, the embodiments of the present application are superior to other methods in terms of the overall mean absolute error (MAE) and mean square error (MSE) counting metrics and most subset partitions. The embodiments of the present application reduce the MAE on the NWPU_Crowd test set to 63.7, a 12.3% reduction compared to the second-ranked P2PNet (the method proposed by Qingyu Song, Changan Wang, Zhengkai Jiang, etc. in the paper "Rethinking Counting and Localization in Crowds: A Purely Point-Based Framework"). When directly applying the embodiments of the present application to vehicle counting (TRANCOS) and corn counting (MTC) tasks, the counting effects of existing methods are also significantly improved. For vehicle counting, the embodiments of the present application further reduce the MAE of the estimated value. In terms of plant counting, the model of the embodiments of the present application is superior to the state-of-the-art methods, with the MAE and MSE reduced by 12.9% and 14.0% respectively.

[0073] In particular, in the task of locating trees in urban environments using remote sensing images. The embodiment of the present application only uses the RGB format of URBAN_TREE as input. With fewer input channels, the embodiment of the present application still achieved a detection accuracy of 75.0% and a tree recall rate of 74.9%, which is a significant improvement compared to ITC (the method proposed by Jonathan Ventura, Camille Pawlak, Milo Honsberger, et al. in the paper “Individual Tree Detection in Large-Scale Urban Environments using High-Resolution Multispectral Imagery”) which also uses the RGB format.

[0074] Figure 4 This is a comparison of the multi-scale fusion density prediction map predicted by the general multi-scale feature fusion method and the density prediction map predicted by the method in the embodiment of the present application; the first column is the input image and the true density map; the second column is the multi-scale fusion density prediction map; the third column is the density prediction map in the embodiment of the present application (i.e., the density prediction map of this method). Figure 4 As shown, the method in the embodiment of the present application obtains a better density prediction map. Therefore, compared with the scaling image block and the general multi-scale feature fusion method, the embodiment of the present application first significantly improves the generalization ability at different scales (resolutions). Secondly, no additional learning tasks are introduced, and no additional labels are required. The embodiment of the present application can achieve similar counting performance when the target size changes from 0 to 1 million pixels;

[0075] Since all features are gathered at the highest resolution for prediction in the embodiment of the present application, the final result is the same as the original Figure 1 Therefore, while achieving the most advanced counting results in various counting tasks, the predicted density map obtained by the embodiment of the present application can also be directly used for target positioning tasks, and the positioning performance is relatively excellent.

[0076] Based on the same principle as the method provided in the embodiment of the present application, the embodiment of the present application also provides a selective inheritance learning device based on target counting, such as Figure 5 As shown, the device comprises:

[0077] The feature extraction module 501 is used to input the target image into the feature extraction network to extract a plurality of first features with different resolutions.

[0078] In the embodiments of the present application, the target image may be an RGB image to be processed. Generally, it is good at capturing large-sized objects at low resolution, and is sensitive to small-sized objects at high resolution. Therefore, it is crucial to extract multiple corresponding first features at multiple resolutions from the target image.

[0079] The feature transfer module 502 is configured to extract and transfer the second feature from the low-resolution fusion feature in the order of increasing resolution by using a preset selective inheritance adapter, and fuse the second feature with the high-resolution first feature to generate a high-resolution fusion feature.

[0080] Among them, the fusion feature with the lowest resolution is the first feature with the lowest resolution; the selective inheritance adapter is arranged between adjacent first features; the second feature is the most discriminative feature region in the low-resolution fusion feature.

[0081] When fusing low-resolution features into high-resolution features, directly upsampling the first feature has two drawbacks: First, multiple upsamplings will destroy the converged features at each resolution, resulting in a decrease in the confidence of the density map of large-sized objects; Second, the weakly discriminative features of large-sized objects at high resolution are noise, and directly upsampling and fusing will introduce this part of noise into the high resolution.

[0082] In the embodiments of the present application, the multiple first features are processed in the order of increasing resolution. Specifically, a selective inheritance adapter can be used between adjacent first features. The low resolution and high resolution mentioned in the context refer to two adjacent first features. The most discriminative feature in the low-resolution first feature is transferred to the high resolution, and then fused with the high-resolution first feature to achieve inheritance learning from low resolution to high resolution, which can improve the prediction accuracy.

[0083] The target counting module 503 is configured to use the fusion feature at the highest resolution as the target feature and perform target counting based on the target feature.

[0084] In the embodiments of the present application, all targets are finally counted at the highest resolution, which can improve the counting performance.

[0085] In the embodiment of the present application, the feature extraction module inputs the target image into the feature extraction network to extract multiple first features with different resolutions; the feature transfer module extracts and transfers the second feature from the low-resolution fusion feature in the order of increasing resolution by using a preset selective inheritance adapter, and fuses the second feature with the high-resolution first feature to generate a high-resolution fusion feature; wherein, the fusion feature with the lowest resolution is the first feature with the lowest resolution; the selective inheritance adapter is arranged between adjacent first features; the second feature is the most discriminative feature region in the low-resolution fusion feature; then, the target counting module uses the fusion feature at the highest resolution as the target feature and performs target counting according to the target feature. The embodiment of the present application can select the most discriminative feature region for feature learning at each resolution and gradually inherit the discriminative features from low resolution to high resolution, which can effectively solve the problem of scale change.

[0086] In an alternative embodiment, the selective inheritance adapter includes a scale non-custom feature forward propagation network and a soft mask generator, and the feature transfer module 502 includes:

[0087] The first feature inheritance sub-module is used to process the low-resolution fusion feature by using the soft mask generator to obtain the unaggregated feature, and generate the first attention map by using the unaggregated feature;

[0088] The second feature inheritance sub-module uses the first attention map to disentangle the low-resolution fusion feature to obtain the second feature.

[0089] In an alternative embodiment, the device further includes:

[0090] The prediction feature extraction module is used to extract the prediction feature from the first feature by using the selective inheritance adapter;

[0091] The result prediction module is used to obtain the prediction result at the corresponding resolution by using the prediction feature and calculate the loss value.

[0092] In an alternative embodiment, the selective inheritance adapter further includes a scale custom feature forward propagation network, and the prediction feature extraction module includes:

[0093] The first prediction feature extraction sub-module is used to process the low-resolution fusion feature by using the soft mask generator to obtain the aggregated feature, and generate the second attention map by using the aggregated feature;

[0094] The second prediction feature extraction sub-module is used to disentangle the low-resolution fusion feature by using the second attention map to obtain the prediction feature.

[0095] In an optional embodiment, the first feature inheritance sub-module includes:

[0096] An area segmentation unit, configured to determine a target area in the low-resolution fusion feature;

[0097] A mapping graph generation unit, configured to generate the first attention mapping graph by using the unaggregated features in the target area.

[0098] In an optional embodiment, the area segmentation unit includes:

[0099] A loss calculation unit, configured to calculate a loss value for prediction results of the same feature area at different resolutions;

[0100] An area confirmation unit, configured to use the feature area as the target feature corresponding to the resolution with the minimum loss value.

[0101] In an optional embodiment, the selective inheritance adapter is provided between each group of adjacent first features, and the selective inheritance adapter is the same one.

[0102] In an optional embodiment, both the scale non-custom feature forward propagation network and the scale custom feature forward propagation network include two convolutional layers;

[0103] The soft mask generator includes three convolutional kernels.

[0104] The selective inheritance learning device for object counting provided by the embodiments of the present application can implement Figures 1 to 4 each process implemented in the method embodiments. To avoid repetition, details are not described herein again.

[0105] The selective inheritance learning device for object counting in the embodiments of the present application can execute the selective inheritance learning method for object counting provided by the embodiments of the present application. Their implementation principles are similar. The actions performed by each module, sub-module, and unit in the selective inheritance learning device for object counting in each embodiment of the present application correspond to the steps in the selective inheritance learning method for object counting in each embodiment of the present application. For the detailed function descriptions of the modules of the selective inheritance learning device for object counting, reference can specifically be made to the descriptions in the corresponding selective inheritance learning method for object counting shown above. Details are not described herein again.

[0106] Based on the same principle as the method shown in the embodiments of the present application, the embodiments of the present application also provide an electronic device, which may include but is not limited to: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the selective inheritance learning method for target counting shown in any optional embodiment of the present application by calling the computer program. Compared with the prior art, the selective inheritance learning method for target counting provided by the present application inputs a target image into a feature extraction network to extract a plurality of first features with different resolutions; in the order from low to high resolution, a preset selective inheritance adapter is used to extract and transfer a second feature from the low-resolution fused feature, and the second feature is fused with the high-resolution first feature to generate a high-resolution fused feature; wherein, the fused feature with the lowest resolution is the first feature with the lowest resolution; the selective inheritance adapter is arranged between adjacent first features; the second feature is the most discriminative feature region in the low-resolution fused feature; then, the fused feature at the highest resolution is used as the target feature, and target counting is performed according to the target feature. The embodiments of the present application can select the most discriminative feature region for feature learning at each resolution and gradually inherit the discriminative features from low resolution to high resolution, which can effectively solve the problem of scale change.

[0107] In an optional embodiment, an electronic device is also provided, as Figure 6 shown Figure 6 The electronic device 600 shown may be a server, including: a processor 601 and a memory 603. Among them, the processor 601 and the memory 603 are connected, such as connected through a bus 602. Optionally, the electronic device 600 may further include a transceiver 604. It should be noted that in actual applications, the transceiver 604 is not limited to one, and the structure of the electronic device 600 does not constitute a limitation to the embodiments of the present application.

[0108] The processor 601 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logical blocks, modules and circuits described in combination with the disclosure of the present application. The processor 601 may also be a combination for implementing computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0109] The bus 602 may include a path for transmitting information between the above components. The bus 602 can be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 602 can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0110] The memory 603 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0111] The memory 603 is used to store the application program code for implementing the solution of this application, and is controlled by the processor 601 for execution. The processor 601 is used to execute the application program code stored in the memory 603 to implement the content shown in the foregoing method embodiments.

[0112] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0113] The server provided by this application can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication means, and this application does not make any restrictions here.

[0114] An embodiment of this application provides a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiment.

[0115] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. Their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0116] It should be noted that the above computer-readable storage medium of the present application can also be a computer-readable signal medium or a combination of a computer-readable storage medium and a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. And in the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0117] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.

[0118] The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.

[0119] According to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the selective inheritance learning method, apparatus, and electronic device for target counting provided in the above various alternative implementation manners.

[0120] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0122] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the module itself in some cases. For example, the feature extraction module can also be described as "a feature extraction module for inputting a target image into a feature extraction network and extracting a plurality of first features with different resolutions".

[0123] The above description is only for the preferred embodiments of this application and the explanation of the technical principles applied. Those skilled in the art should understand that the scope of disclosure involved in this application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A selective inheritance learning method for target counting, characterized in that The method includes: Inputting a target image into a feature extraction network to extract multiple first features with different resolutions; In the order from low to high resolution, using a preset selective inheritance adapter to extract and transfer a second feature from the low-resolution fused feature, and fusing the second feature with the high-resolution first feature to generate a high-resolution fused feature; Taking the fused feature at the highest resolution as the target feature, and performing target counting based on the target feature; Wherein, the fused feature at the lowest resolution is the first feature at the lowest resolution; the selective inheritance adapter is arranged between adjacent first features; the second feature is the most discriminative feature region in the low-resolution fused feature; The selective inheritance adapter includes a scale non-custom feature forward propagation network and a soft mask generator. The step of using the preset selective inheritance adapter to extract and transfer the second feature from the low-resolution fused feature includes: Processing the low-resolution fused feature with the soft mask generator to obtain an unaggregated feature, and generating a first attention map using the unaggregated feature; Using the first attention map to disentangle the low-resolution fused feature to obtain the second feature.

2. The selective inheritance learning method for target counting according to claim 1, wherein Before taking the fused feature at the highest resolution as the target feature and performing target counting based on the target feature, the method further includes: Using the selective inheritance adapter to extract a prediction feature from the first feature; Using the prediction feature to obtain a prediction result at the corresponding resolution and calculating a loss value.

3. The selective inheritance learning method for target-oriented counting according to claim 2, wherein The selective inheritance adapter further includes a scale custom feature forward propagation network. The step of using the selective inheritance adapter to extract the prediction feature from the first feature includes: Processing the low-resolution fused feature with the soft mask generator to obtain an aggregated feature, and generating a second attention map using the aggregated feature; Using the second attention map to disentangle the low-resolution fused feature to obtain the prediction feature.

4. The selective inheritance learning method for target-oriented counting according to any one of claims 1, characterized in that The step of processing the low-resolution fused feature with the soft mask generator to obtain an unaggregated feature and generating a first attention map using the unaggregated feature includes: Determining a target region in the low-resolution fused feature; Generating the first attention map using the unaggregated feature in the target region.

5. The selective inheritance learning method for target-oriented counting according to claim 4, wherein The step of determining the target region in the low-resolution fused feature includes: Calculating a loss value for the prediction results of the same feature region at different resolutions; Taking the feature region as the target feature corresponding to the resolution with the minimum loss value.

6. The selective inheritance learning method for target counting according to claim 1, characterized in that A selective inheritance adapter is arranged between each group of adjacent first features, and the selective inheritance adapter is the same one.

7. The selective inheritance learning method for target-oriented counting according to claim 3, wherein Both the scale non-custom feature forward propagation network and the scale custom feature forward propagation network include two convolutional layers; The soft mask generator includes three convolutional kernels.

8. A selective inheritance learning device for target counting, characterized in that, The device includes: A feature extraction module for inputting a target image into a feature extraction network to extract multiple first features with different resolutions; A feature transfer module, configured to extract and transfer a second feature from low-resolution fused features in ascending order of resolution by using a preset selective inheritance adapter, and fuse the second feature with a high-resolution first feature to generate a high-resolution fused feature; A target counting module, configured to use the fused feature at the highest resolution as a target feature and perform target counting based on the target feature; wherein, the fused feature at the lowest resolution is the first feature at the lowest resolution; the selective inheritance adapter is disposed between adjacent first features; the second feature is the most discriminative feature region in the low-resolution fused features; The selective inheritance adapter includes a scale non-custom feature forward propagation network and a soft mask generator. The extracting and transferring the second feature from the low-resolution fused features by using the preset selective inheritance adapter includes: processing the low-resolution fused features by using the soft mask generator to obtain unaggregated features, and generating a first attention map by using the unaggregated features; disentangling the low-resolution fused features by using the first attention map to obtain the second feature.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • SAR image target detection method and device, electronic equipment and storage medium

    CN113657196A

  • Visual learning method based on multi-knowledge fusion

    CN115115918A