Multi-class dense target counting and positioning method based on dynamic threshold and classification positioning score fusion

By introducing dynamic threshold and classified positioning score fusion technology in dense group counting and positioning methods, combined with Transformer technology, the accuracy of counting and positioning in multi-category target scenarios is solved, and efficient counting and positioning effects in dense multi-category scenarios are achieved.

CN120071352APending Publication Date: 2025-05-30陈楷锦 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510033679.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing dense group counting and positioning methods are difficult to achieve accurate counting and positioning in multi-category target scenarios, especially in dense multi-category scenarios, where data set labeling is high cost and ineffective.

Method used

A classification and positioning score fusion method based on dynamic threshold is adopted, and a pseudo-enclosed frame is generated from point-noted target objects using Transformer technology. Combining the visual Transformer framework and dynamic threshold module, classification scores and positioning scores are fused, and the model is trained through an optimization algorithm to achieve accurate counting and positioning of multi-category intensive targets.

Benefits of technology

Accurate counting and positioning of multi-category dense group scenarios is realized, and density imbalance problem is adapted to the problem of dense multi-category scenarios, and point positioning screening is optimized, which is highly generalized and robust.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071352A_ABST
    Figure CN120071352A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-class dense target counting and positioning method based on dynamic threshold and classification positioning score fusion. The method comprises the following steps: acquiring point labeling data, and preprocessing image data; the method comprises the following steps: establishing a visual Transform framework, embedding a threshold regression head behind a backbone network, embedding a positioning score regression head in a Transform decoder, completing network construction, and initializing network parameters; in the training stage, the actual positioning score is obtained by matching the prediction point with the marking point, and the prediction positioning score is supervised; the classification score and the positioning score are fused to obtain a joint score, an optimal threshold is searched, and a prediction threshold is supervised; training the whole model by using an optimization algorithm; and in the test stage, model parameters after training are reserved, a predicted dynamic threshold value is used for screening the joint score to obtain a prediction result, and a test is carried out on a data set. According to the invention, by introducing dynamic threshold prediction and classification positioning score fusion, accurate multi-class dense group counting and positioning are realized, and the current optimal effect is achieved in counting and positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer vision and image processing, and more specifically, relates to a multi-class dense target counting and positioning method based on dynamic threshold and classification score fusion. Background Art

[0002] Dense crowd counting and positioning is one of the important research topics in the fields of dense crowd analysis and remote sensing positioning, aiming to count and locate dense target objects (such as people or vehicles) appearing in image data, and output the number of targets in the current scene and the position of each target. Among them, positioning refers to outputting the two-dimensional coordinates of each target key point (such as the head in crowd positioning and the center of the car in car positioning). As a basic problem in the fields of dense crowd analysis and remote sensing positioning, the dense crowd counting and positioning technology has broad and practical application prospects.

[0003] The existing dense crowd counting and positioning methods focus on single-class objects and scenes. For complex scenes containing multiple-class objects, the existing methods cannot give accurate and effective counting and positioning results. However, in the real tasks of dense crowd analysis and remote sensing positioning, the targets often contain more than two classes. Therefore, it has very important application value to study the dense crowd counting and positioning method for multi-class targets. Summary of the Invention

[0004] In view of the above defects or improvement requirements of the prior art, the present invention provides a multi-class dense target counting and positioning method based on dynamic threshold and classification and positioning score fusion, which can use the Transformer technology to generate pseudo bounding boxes from point-annotated target objects, and apply them to common detection methods for existing outdoor point cloud scenes and indoor point cloud scenes, so as to alleviate the technical problem of excessive dataset annotation cost.

[0005] To achieve the above object, an embodiment provides a multi-class dense target counting and positioning method based on dynamic threshold and classification and positioning score fusion, and the method includes:

[0006] Step 10: Obtain the point annotation data of the dataset and preprocess the image data;

[0007] Step 20: Build a vision Transformer framework, embed a threshold regression head after the backbone network, and embed a positioning score regression head in the Transformer decoder to complete the network construction and initialize the network parameters;

[0008] In step 30, the training phase, the actual localization score is obtained by matching the predicted points with the annotated points, and the predicted localization score is supervised; the classification score and the localization score are fused to obtain the joint score, the optimal threshold is searched, and the predicted threshold is supervised; the entire model is trained using the optimization algorithm;

[0009] Step 40: In the testing phase, the model parameters after training are retained, the predicted dynamic threshold is used to filter the joint score to obtain the prediction result, and the test is carried out on the dataset;

[0010] In one embodiment of the present invention, in the said step 10:

[0011] Obtain image data and the point annotation coordinates in the scene, and the total number M of objects.

[0012] In one embodiment of the present invention, in the said step 20 specifically includes:

[0013] Read the image data Send it into the image encoder to obtain the image features

[0014] Input the above image feature f into a dynamic threshold module to obtain the predicted threshold of the scene At the same time, after adding the position encoding to f, send it into a Vision Transformer Encoder to obtain the encoded feature f e , then, the model will use the encoded feature f e Input it into the Transformer Decoder and the object queries defined during the model initialization Calculate to obtain the decoded feature f d .

[0015] The decoded feature f d is respectively input into a classification regression head to obtain the classification scores of each query where c is the number of pre-given scene categories; input it into a coordinate regression head to obtain the predicted positions of each query Input it into a localization score prediction head to obtain the predicted localization scores of each query The joint score is obtained by weighted fusion of the classification score and the localization score The calculation method is:

[0016]

[0017] Finally, use the joint score Combined with the predicted dynamic threshold The predicted queries are screened, and the combined score corresponding to each query is the predicted score, and the corresponding predicted coordinates are the positions of the predicted points.

[0018] In one embodiment of the present invention, step 30 specifically includes:

[0019] The classification scores of each query The coordinates of the predicted point Match with the GT points (actual points) through the Matcher. If there is no successful match with the GT points, the actual positioning score of this point is 0; if the match is successful, the distance between this point and the GT points is mapped to the actual positioning score, and the smaller the distance, the higher the actual positioning score. The mapping method is:

[0020]

[0021] Where D is the L2 distance and m is a pre-given normalization constant (m is a positive integer).

[0022] Use L1 loss to calculate the predicted positioning score and the actual positioning score to obtain the positioning score loss The calculation formula is as follows:

[0023]

[0024] Where M is the number of actual objects in the scene.

[0025] At the same time, take the maximum value of the predicted combined score in the category dimension to obtain the sorted score After sorting in descending order, select the Mth score as the best threshold Where M is the number of actual objects in the scene, and use L2 loss to calculate the predicted threshold and the best threshold to obtain the threshold loss The calculation formula is as follows:

[0026]

[0027] Finally, count all the loss functions and train until the total loss function reaches a predetermined value. The formula for the total loss function is as follows:

[0028]

[0029] respectively represent the regression loss, classification loss, positioning score loss, and threshold loss.

[0030] In one embodiment of the present invention, in step 40, the model predicts the combined score and the dynamic threshold After taking the top-k of the combined scores and comparing with the threshold, the points greater than or equal to the prediction threshold are selected as the final prediction points.

[0031] Generally speaking, compared with the prior art by the above technical solutions conceived by the present invention, the following beneficial effects are achieved:

[0032] (1) The embodiments of the present invention can achieve accurate counting and positioning in multi-class dense crowd scenarios;

[0033] (2) Further, the present invention uses a dynamic threshold module and a classification and localization score fusion module, which can better adapt to the density imbalance problem in dense multi-class scenarios, achieve good counting effects, and at the same time optimize the point localization and screening to achieve good localization effects. Moreover, the two modules are lightweight and simple, and can be easily migrated to other models, with strong generalization and robustness;

[0034] (3) Further, the present invention uses a vision Transformer as the backbone, which has strong feature representation capabilities. The Transformer structure can be easily extended, and pre-trained weights in the two-dimensional image field are used. At the same time, further improvements on the Transformer framework are applicable to this method. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a schematic flowchart of a multi-class dense target counting and positioning method based on dynamic threshold and classification and localization score fusion proposed by an embodiment of the present invention;

[0036] Figure 2 is a schematic diagram of the overall framework of a multi-class dense target counting and positioning method based on dynamic threshold and classification and localization score fusion proposed by an embodiment of the present invention;

[0037] Figure 3 is an example of the results of a multi-class dense target counting and positioning method based on dynamic threshold and classification and localization score fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0039] The following further describes a multi-class dense target counting and localization method based on dynamic threshold and classification localization score fusion of the present invention in conjunction with the accompanying drawings and specific embodiments. It should be particularly noted that the following description is exemplary and aims to further illustrate the present invention rather than limit the embodiments.

[0040] Example 1:

[0041] The method of this embodiment uses a deep neural network to extract features, uses point annotations to match classification scores and predicted coordinates to obtain actual localization score supervision for predicting localization scores, uses the actual number of objects in the scene to screen the sorted joint scores to obtain the best threshold for supervising the predicted threshold, calculates the localization score loss, threshold loss, classification loss, and regression loss for training, and then uses the trained model for counting and localization. This embodiment mainly elaborates on how to train to obtain the required model. Combining Figure 1 and Figure 2 , this embodiment provides a training method for a multi-class dense target counting and localization method based on dynamic threshold and classification localization score fusion, including:

[0042] Step 10: Obtain the point annotation data of the dataset and preprocess the image data.

[0043] In this embodiment, the size of the image data is constrained, and after normalization, data augmentation methods such as random cropping, random scaling, and random flipping are performed. At the same time, the point annotation data is read, and the number of objects in each category is counted.

[0044] Step 20: Build a Vision Transformer framework, embed a threshold regression head after the backbone network, and embed a localization score regression head in the Transformer decoder to complete the network construction and initialize the network parameters;

[0045] Specifically, step 20 specifically includes:

[0046] Read the image data Send it into the ResNet50 network to extract the intermediate layer output to obtain the image features

[0047] Input the above image feature f into a dynamic threshold module to obtain the predicted threshold of this scene At the same time, add positional encoding to f and then send it into a Vision Transformer Encoder to obtain the encoded feature f e , after that, the model will input the encoded feature f e into the Transformer Decoder and the object queries defined during model initialization (In this example, the number of queries is set to 500) Calculate the decoded feature f d 。

[0048] The decoded feature f d Input into a classification regression head respectively to obtain the classification scores of each query where c is the number of pre - given scene categories; input into a coordinate regression head to obtain the predicted positions of each query Input into a localization score prediction head to obtain the predicted localization scores of each query Obtain the combined score by weighted fusion of the classification score and the localization score The calculation method is as follows:

[0049]

[0050] Step 30: By matching the predicted points with the labeled points, obtain the actual localization score to supervise the predicted localization score; fuse the classification score and the localization score to obtain the combined score, search for the optimal threshold to supervise the predicted threshold; use the optimization algorithm to train the entire model;

[0051] This step specifically includes:

[0052] The model predicts the classification scores of each query The coordinates of the predicted point Match with the GT points (actual points) through the Matcher. If it fails to match successfully with the GT points, the actual localization score of this point is 0; if it matches successfully, map the distance between this point and the GT points to the actual localization score. In this example, the mapping method is:

[0053]

[0054] where D is the L2 distance between the predicted point and the GT point. Then use the L1 loss to calculate the predicted localization score and the actual localization score to obtain the localization score loss The calculation formula is as follows:

[0055]

[0056] where M is the number of objects in the scene. At the same time, the predicted combined score Take the maximum value in the category dimension to obtain the sorted score After sorting in descending order, select the M - th score as the optimal threshold Use the L2 loss to calculate the predicted threshold And the optimal threshold Obtain the threshold loss The calculation formula is as follows:

[0057]

[0058] Finally, sum up all the loss functions and train until the total loss function reaches a predetermined value. The formula for the total loss function is as follows:

[0059]

[0060] represent the regression loss, classification loss, localization score loss, and threshold loss respectively. In this example, Adam is used as the optimization algorithm for the model. The model is trained for a total of 300 rounds. At the same time, to avoid unstable threshold supervision caused by insufficient convergence of the joint score, the threshold loss does not participate in the calculation of the total loss in the first 50 rounds.

[0061] Step 50: Use the trained model to perform multi-class dense crowd counting and localization on the test set of the SimMICL-10k dataset.

[0062] Table 1 Multi-class counting evaluation results

[0063]

[0064]

[0065] Table 2 Multi-class localization evaluation results

[0066]

[0067] It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi-category dense target counting and positioning method based on dynamic threshold and classification positioning score fusion, characterized in that: The method comprises: Step 10: Obtain the point annotation data of the data set and preprocess the image data; Step 20: Build a visual Transformer framework, embed a threshold regression head after the backbone network, embed a positioning score regression head in the Transformer decoder, complete the network construction, and initialize the network parameters; Step 30: In the training phase, the actual positioning score is obtained by matching the predicted points with the annotated points, and the predicted positioning score is supervised; the classification score and the positioning score are integrated to obtain a joint score, and the optimal threshold is searched to supervise the predicted threshold; the optimization algorithm is used to train the entire model; Step 40: In the testing phase, retain the trained model parameters, use the predicted dynamic threshold to filter the joint score to obtain the predicted result, and test it on the dataset.

2. The multi-category dense target counting and positioning method based on dynamic threshold and classification positioning score fusion according to claim 1 is characterized in that: The step 20 specifically includes: Reading image data Send to the image encoder to get image features The above image feature f is input into a dynamic threshold module to obtain the prediction threshold of the scene At the same time, f is added to the position code and sent to a visual Transformer Encoder to obtain the encoded feature f e , after which the model will encode the feature f e Input the object defined in Transformer Decoder during model initialization Calculate the decoded feature f d . Decoded feature f d Input a classification regression head respectively to get the classification score of each query Where c is the number of scene categories given in advance; input a coordinate regression head to get the predicted position of each query Input a positioning score prediction head to get the positioning score predicted by each query The joint score is obtained by weighted fusion of the classification score and the positioning score The calculation method is: Finally, using the joint score Combined prediction dynamic threshold The predicted query is obtained by filtering. The joint score corresponding to each query is the predicted score, and the corresponding predicted coordinates are the predicted point positions.

3. According to the multi-category dense target counting and positioning method based on dynamic threshold and classification positioning score fusion according to claim 1, the dynamic threshold module in step 20 refers to: The feature f after the backbone network is passed through two layers of convolutional networks and ReLU activation function. After the feature is flattened, it passes through a linear layer to obtain the prediction threshold.

4. According to the multi-category dense target counting and positioning method based on dynamic threshold and classification positioning score fusion according to claim 1, in the step 30, the supervised prediction positioning score specifically includes: Score of each query category Prediction point coordinates Matching is performed with the GT point (actual point) through Matcher. If the match with the GT point is not successful, the actual positioning score of the point is 0; if the match is successful, the distance between the point and the GT point is mapped to the actual positioning score. The smaller the distance, the higher the actual positioning score. The mapping method is: Wherein, D is the L2 distance, and m is a pre-given normalization constant (m is a positive integer). Use L1loss to calculate the predicted positioning score and the actual positioning score to get the positioning score loss The calculation formula is as follows:

5. According to the multi-category dense target counting and positioning method based on dynamic threshold and classification positioning score fusion according to claim 1, in the step 30, the supervised prediction threshold specifically includes: The predicted joint score After taking the maximum value on the category dimension, the ranking score is obtained After sorting in descending order, select the Mth score as the optimal threshold Where M is the actual number of objects in the scene, and L2loss is used to calculate the prediction threshold. With the optimal threshold Get the threshold loss The calculation formula is as follows:

6. The multi-category dense target counting and positioning method based on dynamic threshold and classification positioning score fusion according to claim 1 is characterized in that: In step 30, the total loss function of the training is calculated according to the following formula: Until the total loss function Reach the predetermined value. in, They represent regression loss, classification loss, localization score loss, and threshold loss respectively.