A method for constructing a multi-modal crowd counting model

By constructing a multimodal crowd counting model and using multi-head self-attention blocks and multilayer perceptrons for feature fusion and prediction, the problem of inaccurate counting in complex scenarios by traditional methods is solved, and higher accuracy crowd counting is achieved.

CN115359428BActive Publication Date: 2026-07-24ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ANHUI UNIV
Filing Date
2022-08-29
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Traditional methods of crowd counting based on color images are difficult to accurately count people under lighting conditions and in complex scenes. In particular, the false recognition rate is high in dark scenes and with background interference. A multimodal crowd counting model is needed to improve detection performance.

Method used

A multimodal crowd counting model is constructed by extracting features from multimodal source signals using an encoder, fusing and enhancing features through multi-head self-attention blocks, and combining a prediction head and a multilayer perceptron for density map prediction and supervision, thus forming a multimodal crowd counting model.

Benefits of technology

It improves the accuracy of crowd counting, especially in complex scenarios where crowd counting can be performed more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359428B_ABST
    Figure CN115359428B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal crowd counting model construction methods, comprising: extracting multi-modal feature from multi-modal source signal;Set learnable counting feature;Cascade multi-modal feature and counting feature, form the fusion feature of counting guidance;Through multi-head self-attention block, the fusion feature of counting guidance is enhanced, and enhanced feature is formed;Split enhanced feature, form enhanced multi-modal feature and enhanced counting feature;Enhanced multi-modal feature channel cascade is used Prediction head carries out the prediction of density map;Using multilayer perception, enhanced counting feature is reduced channel, and forms count value;Using density map true value supervises density map, using count value supervises the count value of density map statistics, using count value supervises;Through training set training forms multi-modal crowd counting model.The model constructed by the application can improve crowd counting precision by the guidance of counting information, multi-modal fusion is implemented by multi-head self-attention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision, and in particular to a method for constructing a multimodal crowd counting model. Background Technology

[0002] Crowd counting can predict crowd distribution by estimating the number of people, which is beneficial for controlling traffic and social distancing, preventing stampedes, and suppressing the spread of viruses. Traditional crowd counting methods relying on color images are inevitably affected by lighting conditions and complex scenes; even the most advanced methods cannot obtain accurate counting information. For example, in dark scenes, color images are almost impossible to identify; in scenes with a lot of background interference, the false recognition rate is also very high. Therefore, there is an urgent need for a multimodal crowd counting model and its construction method that utilizes modal information other than color images, such as depth information, thermal infrared information, and sound information, to assist in crowd counting. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method for constructing a multimodal crowd counting model. Under the guidance of learnable counting information, the constructed model uses multi-head self-attention to promote the fusion between different modalities, so as to improve detection performance.

[0004] The specific technical solution adopted in this invention is as follows: A method for constructing a multimodal population counting model includes the following steps: S1. Extract multimodal features from multimodal source signals using an encoder; S2. Set a learnable counting feature; S3, cascaded multimodal features and counting features, form counting-guided fusion features; S4. Enhance the counting-guided fusion features by using a multi-head self-attention block to form enhanced features; S5. Split the enhanced features to form enhanced multimodal features and enhanced counting features; S6. Concatenate the enhanced multimodal features into channels and use the prediction head to predict the density map; S7. Use a multilayer perceptron to perform a channel reduction operation on the enhanced counting features to form a count value; S8. Supervise the density plot using the density plot truth value, supervise the count values ​​of the density plot statistics using the count truth value, and supervise the count values ​​using the count truth value; S9. A multimodal population counting model is formed through training on the training set.

[0005] Furthermore, in step S1, an encoder is used From multimodal source signals Extracting multimodal features ; (1) in, Represents the modal number, from 1 to The multimodal source signal This refers to color mode, depth mode, thermal infrared mode, and sound mode; the encoder This refers to an image encoder and a sound encoder; the image encoder can be ResNet, Swin Transformer, PVT, etc.; the sound encoder can be wav2vec, data2vec, etc. Furthermore, in step S2, a learnable counting feature is set. Random initialization; Furthermore, in step S3, cascaded multimodal features and counting features To form a fusion feature guided by counting ; (2) in This indicates a cascading operation along the block; Furthermore, in step S4, the counting-guided fusion features are processed through a multi-head self-attention block. Enhancement is performed to form enhanced features. ; (3) in This represents a multi-head self-attention block, which consists of two multi-head self-attention layers; Furthermore, in step S5, the enhanced features are split. To form enhanced multimodal features and enhanced counting features ; (4) Furthermore, in step S6, the enhanced multimodal features are... Channel cascading is performed, and density mapping is generated using a prediction head. The prediction; (5) in This indicates a cascading operation along the channel direction. This indicates the prediction of the density map of the cascaded features, specifically consisting of two... convolution and a The convolutional structure; Furthermore, in step S7, a multilayer perceptron is used to enhance the counting features. Perform a down-channel operation to generate a count value. ; (6) in Represents a multilayer perceptron; Furthermore, in step S8, the density map truth value is used. density map To conduct supervision, using the true value of the count. Supervise the count values ​​of the density plot statistics and use the true count values. For count value To supervise; (7) in, This represents the count value calculated based on the density map. This refers to the distribution matching loss proposed in the paper "Distribution Matching for Crowd Counting". It means Paradigm; Furthermore, in step S9, a multimodal crowd counting model is formed through training the training set, which includes an RGB-T crowd counting training set, an RGB-D crowd counting training set, and an RGB-Audio crowd counting training set.

[0013] Compared with existing technologies, the beneficial effects of this invention are reflected in: This invention proposes a method for constructing a multimodal crowd counting model. Guided by counting information, multimodal fusion is implemented by multi-head self-attention to improve the accuracy of crowd counting. Attached Figure Description

[0014] Figure 1 This is a diagram of a multimodal population counting model according to the present invention.

[0015] The present invention will be further described below through specific embodiments and in conjunction with the accompanying drawings, but the embodiments of the present invention are not limited thereto. Detailed Implementation

[0016] The embodiments of the present invention are described in detail below. These embodiments are implemented based on the technical solution of the present invention, and provide detailed implementation methods and specific operation processes. However, the scope of protection of the present invention is not limited to the following embodiments.

[0017] This invention provides a method for constructing an RGB-Thermal multimodal crowd counting model, as follows: Figure 1 As shown, it includes the following steps: S1. Extract multimodal features from color modal and thermal infrared modal source signals using an encoder; S2. Set a learnable counting feature; S3, cascaded multimodal features and counting features, form counting-guided fusion features; S4. Enhance the counting-guided fusion features by using a multi-head self-attention block to form enhanced features; S5. Split the enhanced features to form enhanced multimodal features and enhanced counting features; S6. Concatenate the enhanced multimodal features into channels and use the prediction head to predict the density map; S7. Use a multilayer perceptron to perform a channel reduction operation on the enhanced counting features to form a count value; S8. Supervise the density plot using the density plot truth value, supervise the count values ​​of the density plot statistics using the count truth value, and supervise the count values ​​using the count truth value; S9. An RGB-Thermal multimodal crowd counting model is formed by training the RGB-Thermal crowd counting training set.

[0018] Furthermore, in step S1, an encoder is used From color mode and thermal infrared mode source signals Extracting multimodal features ; (8) in, Represents the modal number, from 1 to The multimodal source signal This refers to color mode and thermal infrared mode; the encoder This refers to an image encoder; the image encoder is a PVT. Furthermore, in step S2, a learnable counting feature is set. Random initialization; Furthermore, in step S3, cascaded multimodal features and counting features To form a fusion feature guided by counting ; (9) in This indicates a cascading operation along the block; Furthermore, in step S4, the counting-guided fusion features are processed through a multi-head self-attention block. Enhancement is performed to form enhanced features. ; (10) in This represents a multi-head self-attention block, which consists of two multi-head self-attention layers; Furthermore, in step S5, the enhanced features are split. To form enhanced multimodal features and enhanced counting features ; (11) Furthermore, in step S6, the enhanced multimodal features are... Channel cascading is performed, and density mapping is generated using a prediction head. The prediction; (12) in This indicates a cascading operation along the channel direction. This indicates the prediction of the density map of the cascaded features, specifically consisting of two... convolution and a The convolutional structure; Furthermore, in step S7, a multilayer perceptron is used to enhance the counting features. Perform a down-channel operation to generate a count value. ; (13) in Represents a multilayer perceptron; Furthermore, in step S8, the density map truth value is used. density map To conduct supervision, using the true value of the count. Supervise the count values ​​of the density plot statistics and use the true count values. For count value To supervise; (14) in, This represents the count value calculated based on the density map. This refers to the distribution matching loss proposed in the paper "Distribution Matching for Crowd Counting". It means Paradigm; Furthermore, in step S9, an RGB-Thermal multimodal crowd counting model is formed by training the RGB-Thermal crowd counting training set.

[0026] The RGB-Thermal multimodal population counting model was compared with three other models, and the results are shown in Table 1.

[0027] Table 1 Experimental Results

[0028] As shown in Table 1, the model method of this invention achieves optimal results in both GAME and RMSE evaluation metrics.

[0029] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for constructing a multimodal population counting model, characterized in that, Includes the following steps: S1. Extract multimodal features from multimodal source signals using an encoder; S2. Set a learnable counting feature; S3, cascaded multimodal features and counting features, form counting-guided fusion features; S4. Enhance the counting-guided fusion features by using a multi-head self-attention block to form enhanced features; S5. Split the enhanced features to form enhanced multimodal features and enhanced counting features; S6. Concatenate the enhanced multimodal features into channels and use the prediction head to predict the density map; S7. Use a multilayer perceptron to perform a channel reduction operation on the enhanced counting features to form a count value; S8. Supervise the density plot using the density plot truth value, supervise the count values ​​of the density plot statistics using the count truth value, and supervise the count values ​​using the count truth value; S9. A multimodal population counting model is formed through training on the training set.