Foggy day crowd counting method based on Transform code cross attention
Through the local enhancement and cross-attention mechanism based on the Transformer encoder, the difficulty of population counting in haze weather is solved, and accurate population count estimation in haze environment is achieved, and the annotation cost is reduced.
Patent Information
- Application Number
- CN202510334831.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-08
AI Technical Summary
In smog weather, traditional methods are difficult to effectively process low-quality images, resulting in difficulties in population counting and behavioral analysis, especially the reduction in clarity, occlusion and density changes between people need to capture long-range dependency context information.
Using a Transformer encoder-based method, features are enhanced through local enhancement modules, and location information tokens are introduced, feature interaction and global average pooling are performed in combination with cross attention mechanisms to generate population numbers.
It improves the accuracy of population counting in smog environments, reduces errors, reduces dependence on precise labeling data, saves labeling costs, and captures population structure dependence in heavy fog scenes.
Smart Images

Figure CN120279480A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a foggy crowd counting method based on Transformer encoding cross-attention. Background Art
[0002] Crowd Counting is an important task in the field of computer vision, aiming to estimate or predict the number of people in a specific area from images or videos; this technology is widely used in many fields such as smart cities, public safety, traffic monitoring, and social media analysis. The Transformer model has achieved remarkable results in the field of computer vision. The Transformer utilizes powerful global information modeling capabilities to enhance the ability to identify and count people in complex scenes. Nevertheless, there are still many challenges to be solved, especially in application scenarios under harsh environments such as haze and low light; Hazy weather refers to a large amount of tiny water droplets or particles suspended in the air, resulting in low visibility and having a great impact on image quality. In hazy weather, the quality of images drops significantly. Traditional feature extraction methods based on color, texture, or shape cannot effectively process these low-quality images, making the research on crowd counting and behavior analysis in foggy environments particularly important. Although convolutional neural networks can capture the connections between short-range features, for haze weather images, the clarity between people decreases, and the movement, occlusion, and density changes of the crowd all require capturing long-range dependent context information to obtain better predictions of the number of people. Summary of the Invention
[0003] Object of the Invention: In order to overcome the deficiencies in the prior art, the present invention provides a foggy crowd counting method based on Transformer encoding cross-attention. The features extracted by the encoder are output through a local enhancement module, and a designed position information token encoding is introduced and incorporated into the decoder input. In the decoder, cross-attention is used to interact the locally enhanced features with the original input features, and a global average pooling operation is performed on the output sequence, and the number of people is obtained through a counting regression head.
[0004] Technical Solution: To achieve the above object, a foggy crowd counting method based on Transformer encoding cross-attention of the present invention includes the following steps:
[0005] Step 1, obtain a preprocessed image, and use the encoder network of the Transformer to extract features of the preprocessed image to obtain the image feature Fr;
[0006] Step 2, input the image feature Fr into a local enhancement module to update and output a locally enhanced feature F'r;
[0007] Step 3: Output the feature sequence composed of the locally enhanced features F′r by the local enhancement module, incorporate the position tokens into the feature sequence and input it into the decoder module to output the final decoded feature Wr;
[0008] Step 4: Compose the final decoded feature Wr output by the decoder module into the final feature sequence, perform global average pooling on the final feature sequence, and feed it to the regression head to generate the predicted crowd count.
[0009] Further, in Step 1, the preprocessed image is transformed from a two-dimensional space into a one-dimensional vector sequence and input into the encoder of the Transformer, which includes the following steps:
[0010] Step 1-1: Uniformly segment the preprocessed image to obtain a number of image patches with the same fixed size;
[0011] Step 1-2: Flatten each image patch into a one-dimensional vector, and perform a linear transformation on the one-dimensional vector to obtain the embedded vector representation of each image patch.
[0012] Step 1-3: Input the embedded vector into the Transformer encoder, and the self-attention mechanism and the feed-forward neural network extract the global features of the image to obtain the output feature Fr.
[0013] Further, the encoder of the Transformer is stacked by multiple identical sub-layers, and the output of each layer will be used as the input of the next layer. The formula for the multi-layer transformation in the encoder of the Transformer is:
[0014] W′ r-1 = MSA(LN(W r-1 )) + W r-1
[0015] W r = MLP(LN(W′ r-1 )) + W′ r-1
[0016] In the formula, MSA is the self-attention mechanism operation, MLP is the feed-forward neural network operation, and LN is the layer normalization.
[0017] Furthermore, the local enhancement module includes a residual block and a Linear layer; the residual block includes a Conv layer, a ReLU function, and a DeformConv layer; the output features of the Transformer encoder are reorganized into spatial two-dimensional image features as the input of the local enhancement module; the spatial two-dimensional image features pass through the Conv layer, the ReLU function, the DeformConv layer, and the ReLU function in sequence to output features, and the spatial two-dimensional image is added and fused with the output features through a skip connection to obtain the output of the residual block; the output of the residual block passes through the Linear layer and then outputs the local enhancement feature F'r; the calculation formula of the residual block is as follows:
[0018] y = DeformConv(ReLU(Conv(Norm(x)))) + x
[0019] In the formula, DeformConv is a deformable convolution operation, ReLU is an activation function, Conv is a convolution operation, and Norm is a normalization operation.
[0020] Furthermore, the third step includes the following steps:
[0021] Step 2-1: Introduce the local enhancement feature F'r output by the local enhancement module and add an additional positional information token T to form the input of the decoder module;
[0022] Step 2-2: Perform self-attention mechanism calculation on the input of the decoder module to obtain the self-attention mechanism output;
[0023] Step 2-3: Perform cross-attention mechanism calculation on the self-attention mechanism output to obtain the cross-attention mechanism output, where the K key and V value come from the encoder output, and the Q query comes from the output of the previous layer of the decoder;
[0024] Step 2-4: After passing through the fully connected layer and the normalization layer, the cross-attention mechanism output outputs the final decoded feature Wr.
[0025] Furthermore, the cross-attention mechanism is obtained by transforming the self-attention mechanism, and the calculation formula of the cross-attention mechanism is expressed as follows:
[0026] CrossAttention = concat[Attention(Q, K, V)]
[0027]
[0028] In the formula, concat is a fusion operation; Q, K, and V are the learning matrices of the query, key, and value respectively.
[0029] Further, the steps of training and testing the foggy crowd counting model constructed by the Transformer encoder, local enhancement module, and Transformer decoder based on Transformer encoding cross-attention include:
[0030] Step 3-1: Randomly obtain several image data from the image dataset and perform preprocessing;
[0031] Step 3-2: Input the preprocessed image data into the foggy crowd counting model, train the model using a weakly supervised training method, and update the parameters through backpropagation;
[0032] Step 3-3: Determine whether the number of training times has reached the test round. If not, return to Step 3-2; if so, perform testing;
[0033] Step 3-4: After the testing process is completed, output the test results. When the test results reach the set threshold, determine whether the training round of the test has reached the maximum round. If not, return to Step 3-2; if so, end the training.
[0034] Further, the testing process includes the following steps:
[0035] Step 4-1: Randomly select images and ground truths from the test dataset;
[0036] Step 4-2: Load the weights of the crowd counting model obtained by weakly supervised training, input the images of the dataset into the trained model for feature extraction and update, and output the predicted number of people by the model;
[0037] Step 4-4: Calculate the absolute error and squared error between the predicted number of people and the ground truth annotation;
[0038] Step 4-5: Determine whether the absolute error and squared error of the test results reach the corresponding set values. If not, retrain the model; if so, end the test.
[0039] Further, the L1 loss is used as the loss function in the testing process, and the formula is as follows:
[0040]
[0041] In the formula, Pi is the predicted crowd quantity of the i-th image, Gi is the corresponding ground truth of the i-th image, and M is the batch size of the training images.
[0042] Beneficial effects: A foggy-day crowd counting method based on Transformer encoding cross-attention of the present invention uses a local enhancement module to strengthen the features output by the Transformer encoder, avoiding insufficient extraction of local information by self-attention. The position information tokens are added in parallel to the Transformer decoder of the multi-head cross-attention, effectively solving the problem of blurred crowd positions in low visibility. Finally, a global average pooling operation is performed on the output sequence and fed into the counting regression head to generate an accurate number of people. The present invention adopts a weakly supervised training method, and the weakly supervised crowd counting method can reduce the dependence on accurately labeled data, effectively saving the labeling costs of manpower and material resources. Using Transformer as the backbone network can better capture the crowd structure dependencies at different positions, making the model more accurate in estimating the number of people in a foggy scenario and reducing errors. Description of the Drawings
[0043] Figure 1 It is a block diagram of a foggy-day crowd counting network based on Transformer encoding cross-attention;
[0044] Figure 2 It is a flowchart of a foggy-day crowd counting model based on Transformer encoding cross-attention;
[0045] Figure 3 It is a network framework diagram of the encoder module;
[0046] Figure 4 It is a network framework diagram of the local enhancement module;
[0047] Figure 5 It is a network framework diagram of the decoder module;
[0048] Figure 6 It is a training flowchart of a foggy-day crowd counting model based on Transformer encoding cross-attention. Detailed Embodiments
[0049] The present invention will be further described below with reference to the accompanying drawings.
[0050] As Figure 1-2 shown, a foggy-day crowd counting method based on Transformer encoding cross-attention. The features extracted by the encoder of the Transformer form local enhanced features through the local enhancement module. The position information token encoding is designed and incorporated into the decoder input. In the decoder, cross-attention is used to interact the locally enhanced features with the original input features. A global average pooling operation is performed on the output sequence, and the number of people is obtained through the counting regression head. It includes the following steps:
[0051] Step 1: Obtain the preprocessed image, and use the encoder network of Transformer to extract features from the preprocessed image to obtain the image feature Fr; that is, obtain the encoded feature tokens.
[0052] Step 2: Input the image feature Fr into the local enhancement module to update and output the local enhanced feature F′r.
[0053] Step 3: Form a feature sequence from the local enhanced feature F′r output by the local enhancement module, and incorporate the position tokens into the feature sequence and input it into the decoder module to output the final decoded feature Wr; the encoded feature tokens and position tokens after the operation of the local enhancement module are incorporated and fused and input into the decoder module.
[0054] Step 4: Form the final feature sequence from the final decoded feature Wr output by the decoder module, perform global average pooling on the final feature sequence to obtain the pooling tokens, and feed the pooling tokens to the regression head to generate the predicted crowd count.
[0055] As Figure 3 shown, the Transformer backbone network consists of a multi-head self-attention mechanism MSA and a multi-layer perceptron MLP or a feed-forward neural network; each layer contains a residual connection and layer normalization, and the weight distribution of each layer's processing and attention mechanism; in the above Step 1, the preprocessed image is transformed from a two-dimensional space into a one-dimensional vector sequence and input into the Transformer encoder; it includes the following steps:
[0056] Step 1-1: Uniformly divide the preprocessed image to obtain a number of image patches with the same fixed size.
[0057] Step 1-2: Flatten each image patch into a one-dimensional vector, and perform a linear transformation on the one-dimensional vector to obtain the embedded vector representation of each image patch.
[0058] Step 1-3: Input the embedded vector into the Transformer encoder, and use the self-attention mechanism and the feed-forward neural network to extract the global features of the image to obtain the output feature Fr.
[0059] The encoder of the Transformer is stacked by multiple identical sub-layers, and the output of each layer will be used as the input of the next layer. The multi-layer transformation representation formula in the Transformer encoder is:
[0060] W′ r-1 =MSA(LN(W r-1 ))+W r-1
[0061] W r =MLP(LN(W′r-1 )) + W' r-1
[0062] In the formula, MSA is the self-attention mechanism operation, MLP is the feed-forward neural network operation, and LN is the layer normalization;
[0063] As Figure 3 shown, for example, when r = 1, the input W of the first sub-layer r-1 is the embedding vector. After passing through the layer normalization operation and the self-attention mechanism operation in sequence, the result of the output is added and fused with the embedding vector to obtain the output W' of the multi-head self-attention mechanism r-1 ; the output W' of the multi-head self-attention mechanism r-1 After passing through the layer normalization operation and the multi-layer perceptron operation in the feed-forward neural network in sequence, the result of the output is added and fused with the output W' of the multi-head self-attention mechanism r-1 to obtain the output W of the feed-forward neural network r ; the output W of the feed-forward neural network r is used as the output of the first layer. r represents the r-th layer, and W r represents the query of the r-th layer.
[0064] The output W of the feed-forward neural network r is used as the input of the second sub-layer. Then, after passing through the layer normalization operation and the self-attention mechanism operation in sequence, the result of the output is added and fused with the embedding vector to obtain the output W' of the multi-head self-attention mechanism r ; the output W' of the multi-head self-attention mechanism r After passing through the layer normalization operation and the multi-layer perceptron operation in the feed-forward neural network in sequence, the result of the output is added and fused with the output W' of the multi-head self-attention mechanism r to obtain the output W of the feed-forward neural network r+1 ; the output W of the feed-forward neural network r+1 is used as the input of the third sub-layer, and so on until the final output feature Fr is obtained after all sub-layers in the encoder of the Transformer are calculated.
[0065] The traditional global attention mechanism may ignore local important information or fail to effectively process local features lost due to occlusion. After local information enhancement, the model can more accurately capture the crowd density distribution that can still be identified even under foggy conditions, avoiding ignoring important local features under the suppression of global information. Since the crowd targets overlap due to factors such as light, climate, or occlusion in foggy days, the local enhancement module is used to enhance the learning of multi-scale features.
[0066] As Figure 4As shown, the local enhancement module includes a residual block and a Linear layer; the residual block includes a Conv layer, a ReLU function, and a DeformConv layer; the output features of the Transformer encoder are reorganized into spatial two-dimensional image features as the input of the local enhancement module; let the input module feature be X ∈ R B×N×C , (B represents the batch size, N represents the length of the query sequence, C represents the feature dimension of each image patch), the input feature sequence is reshaped into B×C×width×width, where width represents the number of patches in each row of the image block; the spatial two-dimensional image features pass through a Conv layer, a ReLU function, a DeformConv layer, and a ReLU function in sequence to output features, and the spatial two-dimensional image is added and fused with the output features through a skip connection to obtain the output of the residual block; the output of the residual block passes through a Linear layer to output the local enhancement feature F′r; the residual block introduces a skip connection to directly add the input to the output, avoiding the problem of gradient disappearance and enhancing information flow; through the operation of the deformable convolutional layer DeformConv layer, the model's perception of local details is enhanced; the formula of the residual block is as follows:
[0067] y = DeformConv(ReLU(Conv(Norm(x)))) + x
[0068] In the formula, DeformConv is a deformable convolution operation, which is used to replace the traditional convolution to enhance the perception ability of the convolution operation for local regions; ReLU is an activation function; Conv is a convolution operation; Norm is a normalization operation, and x is the input feature.
[0069] As Figure 5 shown, in step 3, position information tokens are introduced in the decoder to provide global context position information and guide image spatial positioning. Cross-attention is used in the decoder. Even when local features are damaged, the decoder can still generate accurate output feature sequences based on global information; it includes the following steps:
[0070] Step 2-1: Introduce the local enhancement feature F′r output by the local enhancement module and add an additional position information token T to form the input of the decoder module;
[0071] Step 2-2: Perform self-attention mechanism calculation on the input of the decoder module to obtain the output of the self-attention mechanism;
[0072] Step 2-3: Perform cross-attention mechanism calculation on the output of the self-attention mechanism to obtain the output of the cross-attention mechanism, where the K key and V value come from the encoder output, and the Q query comes from the output of the previous layer of the decoder;
[0073] Step 2-4: After the output of the cross-attention mechanism passes through the fully connected layer and the normalization layer, the final decoded feature Wr is output.
[0074] Introducing position information tokens can provide global position perception, enabling the model to understand the relative relationships of different positions in the image through spatial position information even without clear features, which is beneficial for the model to further learn the spatial distribution pattern of people; in addition, in the decoder module design, additional layers are arranged to perform the multi-head cross-attention mechanism, which is beneficial for distinguishing elements at different positions and capturing global dependencies. The decoder module sequentially passes the input containing query Q, key K, and value V through layer normalization and the self-attention mechanism to obtain the first output. The first output is added and fused with the output containing query Q, key K, and value V output by the encoder to obtain the second input containing query Q, key K, and value V; the second input sequentially passes through the cross-attention mechanism and layer normalization to obtain the second output; the second output and the fused feature containing query Q, key K, and value V are added and fused to obtain the third input; the third input sequentially passes through the fully connected layer and layer normalization to obtain the third output, and the third output and the third input are added and fused to obtain the output of the decoder module, and the final decoded feature Wr is output.
[0075] The cross-attention mechanism is transformed based on the self-attention mechanism. The multi-head self-attention mechanism includes multiple self-attention mechanisms, and the cross-attention mechanism performs a cross-fusion operation on the outputs of multiple self-attention mechanisms; the calculation formula of the cross-attention mechanism is expressed as follows:
[0076] CrossAttention = concat[Attention(Q, K, V)]
[0077] The self-attention mechanism allows each patch to interact with the information of other patches and generates a new representation through weighted averaging; the embedding representation of each patch is linearly transformed to obtain the query Q, key K, and value V vectors, and the attention weights are generated by calculating the similarity (dot product) between them; then through Softmax normalization, the weighted representation of each patch is obtained; the formula is as follows:
[0078]
[0079] In the formula, concat is the fusion operation; Q, K, and V are the learning matrices of the query, key, and value respectively.
[0080] In Step 4, the sequence output by the decoder module is regressed to generate the predicted count. Selecting the regression head based on global average pooling to reduce the sequence length can generate a richer discriminative semantic crowd pattern and achieve better counting performance compared to using additional regression tokens.
[0081] Such asFigure 6 As shown, the steps for training and testing the foggy crowd counting model based on Transformer-encoded cross-attention constructed by the Transformer encoder, local enhancement module, and Transformer decoder include:
[0082] Step 3-1: Randomly obtain several image data from the image dataset and perform preprocessing.
[0083] Step 3-2: Input the preprocessed image data into the foggy crowd counting model, and train the model using a weakly supervised training method. Train the model flexibly and attentively with limited labeled data, and update the parameters through backpropagation.
[0084] Step 3-3: Determine whether the number of training times has reached the test rounds. Otherwise, return to Step 3-2. If yes, perform testing.
[0085] Step 3-4: After the testing process is completed, output the test results. At this time, it is necessary to determine whether the obtained test results reach the set threshold during the testing process. When not reached, retrain, that is, return to Step 3-2. When the test results reach the set threshold, determine whether the training rounds of the test reach the maximum rounds. Otherwise, return to Step 3-2. If yes, end the training.
[0086] The model is trained using a weakly supervised training method. Specifically, the weakly supervised training method trains the model with limited labeled data to achieve efficient and accurate crowd counting. Compared with traditional supervised learning methods, weakly supervised learning is more flexible in obtaining labeled data, can significantly reduce the labeling cost, and still maintain high counting performance. During the training process of this model, different from the previous method of generating density maps to calculate the loss function and update parameters to guide the model to learn the crowd distribution, the total number of people in the image is directly used as the training label, significantly reducing the label annotation cost and improving the generalization and robustness of the model. At the same time, the model training uses the mini-batch training method, randomly selecting foggy images from the dataset for training each time. Before model training, the dataset is first subjected to necessary preprocessing, and then training hyperparameters are configured, including learning rate, batch size, optimization algorithm, and loss function, etc. In addition, a validation mechanism needs to be set in advance and monitoring logs are configured to track the training progress and performance in real time.
[0087] The testing process includes the following steps:
[0088] Step 4-1: Randomly select images and ground truths from the test dataset.
[0089] Step 4-2: Load the weights of the crowd counting model obtained through weak supervision training, input the images of the dataset into the trained model for feature extraction and update, and output the predicted number of people in the model.
[0090] Step 4-4: Calculate the absolute error and squared error between the predicted number of people and the true value annotation.
[0091] Step 4-5: Determine whether the absolute error and squared error of the test result reach the corresponding set values, that is, the set thresholds. Otherwise, retrain the model. If so, the test ends. Determine whether the absolute error of the test result is lower than or equal to the set absolute error threshold and whether the squared error is lower than or equal to the set average error threshold. The test only ends when both the absolute error and the average error are lower than or equal to the corresponding thresholds. After the test process is completed, it is judged during the training process whether the maximum number of training rounds is reached. Only when the maximum number of training rounds is reached does the training end. At this time, the training of the foggy-day crowd counting model based on Transformer-coded cross-attention is completed. Only then can the trained foggy-day crowd counting model based on Transformer-coded cross-attention be used to count the number of people in the images under foggy conditions.
[0092] In the above test process, the L1 loss is used as the loss function, which is suitable for scenarios that are not sensitive to outliers. Its advantage is that it can produce sparse solutions and is helpful for feature selection. During the training process of the model, the L1 loss is also used as the loss function. The calculation formula of the L1 loss is as follows:
[0093]
[0094] In the formula, Pi is the predicted crowd quantity of the i-th image, Gi is the corresponding ground truth of the i-th image, and M is the batch size of the training images.
[0095] The design and application of the loss function are the key links in realizing model training. Model training usually relies on limited or incomplete annotation information. Therefore, the loss function needs to be able to guide the model to learn effective feature representations from this limited information. The loss function can guide the model to learn. By minimizing the loss function, the model can learn the mapping relationship from the input image to the crowd count or density map, and at the same time can effectively avoid the overfitting problem. The above L1 loss function is a commonly used loss function, especially suitable for regression tasks. The performance of the model is measured by calculating the average of the absolute differences between the predicted value and the true value. The error of the predicted value is calculated through the loss function, and then the update of the model parameters is guided until the optimal solution is reached.
[0096] Embodiment
[0097] Image data is selected from the Hazy-ShanghaiTechRGBD dataset, which is a synthetic haze crowd counting dataset. The sunny RGB-D crowd counting benchmark is applied to this dataset according to the haze simulation algorithm. In particular, the dataset contains corresponding depth information, which greatly helps the authenticity of haze generation. The dataset covers diverse scenarios from the city center to the park, and a total of 1193 training images and test images are divided, containing 144,512 crowd annotations. Before training, the dataset is preprocessed. Since the input of the Transformer encoder is different from that of the convolutional neural network, the input image is segmented into image patches of 384*384, and then mapped to the embedding features using the linear projection layer of the positional embedding. The processed image is input into the foggy day crowd counting model based on the Transformer encoded cross-attention to obtain the number of people in the generated image.
[0098] The evaluation metrics used are the mean absolute error (MAE) and the mean squared error (MSE). The smaller the values of these two metrics, the higher the accuracy. Through experiments, the MAE of the TransCrowd method is 8.67 and the MSE is 13.1. The MAE of the DAANet method is 8.41 and the MSE is 12.46. The MAE of the present invention is 8.33 and the MSE is 11. Compared with other conventional counting and foggy day targeted counting methods, the present invention can improve the accuracy.
[0099] The above is only a description of the preferred embodiments of the present invention. Those of ordinary skill in the art can make several modifications and optimizations based on the above disclosure without departing from the basic principle content. These improvements and optimizations should be regarded as the protection scope understood by the present invention.
Claims
1. A foggy day crowd counting method based on Transformer encoded cross-attention, characterized in that: It includes the following steps: Step 1: Obtain the preprocessed image, and use the encoder network of Transformer to extract the features of the preprocessed image to obtain the image feature Fr; Step 2: Input the image feature Fr into the local enhancement module to update and output the local enhanced feature F′r; Step 3: Compose the feature sequence composed of the local enhanced feature F′r output by the local enhancement module, and incorporate the position token into the feature sequence and input it into the decoder module to output the final decoded feature Wr; Step 4: Compose the final decoded feature Wr output by the decoder module into the final feature sequence, perform global average pooling on the final feature sequence, and feed it to the regression head to generate the predicted crowd count.
2. The method for foggy crowd counting based on Transformer encoding cross-attention according to claim 1, wherein: In step 1, the preprocessed image is transformed from a two-dimensional space into a one-dimensional vector sequence and input into the encoder of Transformer; it includes the following steps: Step 1-1: Uniformly segment the preprocessed image to obtain a number of image patches with the same fixed size; Step 1-2: Flatten each image patch into a one-dimensional vector, and perform a linear transformation on the one-dimensional vector to obtain the embedded vector representation of each image patch, Step 1-3: Input the embedded vector into the Transformer encoder, and use the self-attention mechanism and the feed-forward neural network to extract the global features of the image to obtain the output feature Fr.
3. A foggy crowd counting method based on Transformer-encoded cross-attention according to claim 2, characterized in that: The encoder of the Transformer is stacked by multiple identical sub-layers, and the output of each layer will be used as the input of the next layer. The formula for the multi-layer change in the encoder of the Transformer is: W′ r-1 = MSA(LN(W r-1 )) + W r-1 W r = MLP(LN(W' r-1 )) + W' r-1 In the formula, MSA is the self-attention mechanism operation, MLP is the feed-forward neural network operation, and LN is the layer normalization.
4. A foggy crowd counting method based on Transformer-encoded cross-attention according to claim 1, characterized in that: The local enhancement module includes a residual block and a Linear layer; the residual block includes a Conv layer, a ReLU function, and a DeformConv layer; the output feature of the Transformer encoder is reorganized into a spatial two-dimensional image feature as the input of the local enhancement module; the spatial two-dimensional image feature passes through the Conv layer, the ReLU function, the DeformConv layer, and the ReLU function in sequence to output the feature, and the spatial two-dimensional image is added and fused with the output feature through a skip connection to obtain the output of the residual block; the output of the residual block passes through the Linear layer to output the local enhanced feature F′r; the calculation formula of the residual block is as follows: y = DeformConv(ReLU(Conv(Norm(x)))) + x In the formula, DeformConv is the deformable convolution operation, ReLU is the activation function, Conv is the convolution operation, and Norm is the normalization operation.
5. A method for foggy crowd counting based on Transformer encoded cross-attention according to claim 1, characterized in that: Step 3 includes the following steps: Step 2-1: Output the local enhanced feature F′r of the local enhancement module, and introduce an additional position information token T to form the input of the decoder module; Step 2-2: Perform self-attention mechanism calculation on the input of the decoder module to obtain the output of the self-attention mechanism; Step 2-3: Perform cross-attention mechanism calculation on the output of the self-attention mechanism to obtain the output of the cross-attention mechanism, where the key K and value V come from the encoder output, and the query Q comes from the output of the previous decoder layer; Step 2-4: After the output of the cross-attention mechanism passes through the fully connected layer and the normalization layer, the final decoded feature Wr is output.
6. The method for counting foggy-day crowds based on Transformer-encoded cross-attention according to claim 5, characterized in that: The cross-attention mechanism is obtained by transforming the self-attention mechanism, and the calculation formula of the cross-attention mechanism is as follows: CrossAttention = concat[Attention(Q, K, V)] In the formula, concat is the fusion operation; Q, K, and V are the learning matrices of the query, key, and value respectively.
7. A foggy crowd counting method based on Transformer-encoded cross-attention according to claim 1, characterized in that: The steps for training and testing the foggy crowd counting model based on Transformer encoding cross-attention constructed by the Transformer encoder, local enhancement module, and Transformer decoder include: Step 3-1: Randomly obtain a number of image data from the image dataset and perform preprocessing; Step 3-2: Input the preprocessed image data into the foggy crowd counting model, and use the weakly supervised training method to train the model and update the parameters by backpropagation; Step 3-3: Determine whether the number of training times has reached the test round. Otherwise, return to Step 3-2. If yes, perform testing; Step 3-4: After the testing process is completed, output the test result. When the test result reaches the set threshold, determine whether the training round of the test has reached the maximum round. Otherwise, return to Step 3-2. If yes, end the training.
8. A foggy day crowd counting method based on Transformer encoding cross-attention according to claim 1, characterized in that: The testing process includes the following steps: Step 4-1: Randomly select images and ground truths from the test dataset; Step 4-2: Load the weights of the crowd counting model obtained by weakly supervised training, and input the images of the dataset into the trained model for feature extraction and update, and output the predicted number of people in the model; Step 4-4: Calculate the absolute error and squared error between the predicted number of people and the ground truth annotation; Step 4-5: Determine whether the absolute error and squared error of the test result reach the corresponding set values. Otherwise, retrain the model. If yes, the testing ends.
9. A foggy crowd counting method based on Transformer-encoded cross-attention according to claim 8, characterized in that: The L1 loss is used as the loss function in the testing process, and the formula is as follows: In the formula, Pi is the predicted number of people in the i-th image, Gi is the corresponding ground truth of the i-th image, and M is the batch size of the training images.
Citation Information
Patent Citations
River and lake remote sensing image segmentation method based on deformable convolution and self-attention model
CN115601549A
Remote sensing image change detection network and detection method based on double twinborn branches
CN116524361A
Weakly supervised crowd counting method based on cross window self-attention mechanism
CN119206612A
Cross-token guided Transform weak supervision positioning method
CN119478333A