Counting Diagram-Assisted Cross-Modal Pedestrian Flow Monitoring Method and System

By using RGBT datasets and cross-modal crowd counting network models, the problem of lighting changes and scale changes in the prior art affecting the accuracy of crowd counting is solved, and more accurate density map generation and crowd counting are achieved.

CN114724081BActive Publication Date: 2025-05-27ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210348731.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-01
Publication Date
2025-05-27
Estimated Expiration
2042-04-01

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the problems of light changes and scale changes in high-crowding scenarios, resulting in the impact of the accuracy of population counting.

Method used

The RGBT data set is used to generate a counting map and a cross-modal population counting network model is designed. The model can be trained across three modalities: RGB, thermal imaging map, and counting map to generate a more accurate density map.

Benefits of technology

By taking into account the changes in lighting and scale, more accurate density maps are generated, which improves the accuracy and reliability of population counting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114724081B_ABST
    Figure CN114724081B_ABST
Patent Text Reader

Abstract

Counting Map-Assisted Cross-Modal Pedestrian Flow Monitoring Method, including: Step 1, generating a counting map, using the model LibraNet based on deep reinforcement learning, with the VGG16 convolutional neural network as the front end and the deep reinforcement learning network as the back end, to output the counting map; Step 2, performing crowd counting based on a cross-modal network model, designing a network model, taking RGB, thermal imaging map, and counting map as three-modal inputs, outputting a density map, summing the values in the density map, and obtaining the total number of people in the RGB image used as the input. The present invention also includes a counting map-assisted cross-modal pedestrian flow monitoring system. The model provided by the present invention can take into account both illumination changes and scale changes and generate a more accurate density map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for crowd monitoring based on a cross-modal convolutional neural network. Background Art

[0002] As one of the most basic tasks in crowded scene analysis, crowd counting has great application prospects in aspects such as crowd monitoring and security prevention, and has received extensive attention in the field of computer vision in recent years. However, due to factors such as the complex distribution of people, scale changes of images, illumination effects, occlusion, and background clutter, crowd counting based on a single image remains an active and challenging topic.

[0003] Early crowd counting used the method of pedestrian detection, but this method was greatly affected by occlusion and had poor performance in high-crowded scenes. Due to the inherent defects of pedestrian detection, researchers turned their attention to regression methods, learning the relationships between features extracted from cropped image patches, and then calculating the number of specific objects. However, this method cannot show the spatial distribution of the crowd. In recent years, with the proposal of density maps and the development of convolutional neural networks, CNN-based counting methods have been continuously proposed, and this method takes the density map as the output, which can not only reflect the spatial distribution of people but also accurately estimate the number of people. Among the CNN-based methods, the mainstream ones include multi-level feature map fusion methods, perspective-assisted methods, methods for improving the traditional L2 loss function, methods for optimizing density maps, and so on. These methods more or less pay attention to the scale changes of single images, but rarely pay attention to illumination changes. And methods that pay attention to illumination changes, such as cross-modal RGBT counting methods, rarely pay attention to scale changes. Summary of the Invention

[0004] The present invention aims to overcome the above-mentioned drawbacks of the prior art and provides an image-based method and system for crowd monitoring that can take into account both illumination changes and scale changes.

[0005] The present invention mainly uses the RGBT dataset, which provides RGB images, thermal images, and annotation data, and the annotation data selected is the annotation data of thermal images. This dataset provides bright and dim scenes. The present invention first generates a counting map. Then a new cross-modal crowd counting network model is constructed to be able to cross three modalities: RGB, thermal images, and counting maps. Then the convolutional neural network is trained.

[0006] The counting-map-assisted cross-modal crowd monitoring method of the present invention includes the following steps: 1. Generating a counting map;

[0007] The counting map is for density Figure 1The integral of a fixed area can reflect the scale change of the image. The counting map generation method uses a model based on deep reinforcement learning, LibraNet. This model LibraNet takes the VGG16 convolutional neural network as the front end and the deep reinforcement learning network as the back end to output the counting map. Specifically, it includes:

[0008] 11) Train the deep reinforcement model; Select the RGBT-CC dataset as the training dataset. This dataset is divided into a training set, a test set, and a validation set. Each set has RGB, thermal images, and annotation data. The annotation data shows the pixel positions of the head centers of each person in the thermal image. Use the thermal images and annotation data of the training set as training data, and use the thermal images and annotation data of the test set as test data to train the deep reinforcement model LibraNet. Use {-10, -5, -2, -1, 1, 2, 5, 10, 999} as the action set, where 999 is the termination action, and the remaining actions represent the change magnitude of the count value in the pixel area. For example, 5 means adding 5 to the count value, and -5 means subtracting 5 from the count value.

[0009] 12) Generate the counting map; Fix the trained LibraNet model and input all the thermal images in the RGBT-CC dataset to generate the counting map. Each value in the counting map can reflect the approximate quantity information of the corresponding thermal image pixel area.

[0010] 13) Optimize the counting map; Because the length and width of the generated counting map are only where n represents the number of pooling layers of VGG16 used as the front end of LibraNet, upsampling is required to make the size the same as the original Figure 1 Consistent; Use the nearest neighbor interpolation method for upsampling and divide by 2 at each pixel position 2*n , so that the total count value remains unchanged and the image size is the same as the original Figure 1 Consistent.

[0011] 2 Perform crowd counting based on a cross-modal network model;

[0012] Design a cross-modal network model. This model has four branches, including three input branches and one output branch. The three input branches respectively input RGB, thermal imaging maps, and counting maps as three modalities; the output branch is a shared branch, which is initialized to 0, receives and refines the information of the three modalities, and its output is a density map. Except for the shared branch, the remaining branches are all composed of VGG-19. Since VGG-19 can be divided into 5-layer modules, except for the last layer module, each layer module is finally equipped with a 2×2 pooling layer, so the three input branches can all be divided into 5-layer modules. The shared branch is composed of VGG-19 after removing the first two layers and can be divided into 4-layer modules. After designing the model, use the RGBT-CC dataset for training. After training the model, perform crowd counting based on this model; specifically including:

[0013] 21) Generate context information I; use the L-level pyramid pooling layer to extract the context information I from the feature maps F generated by each layer module of each branch in the network; specifically, for the l-th layer (l = 1, 2,..., L), use a 2 l-1 ×2 l-1 max pooling layer, with the h×w feature map F as the input, output a feature, and then upsample it to h×w using the nearest neighbor interpolation method to form the context feature F l ; finally, the context information I can be calculated by Equation (1):

[0014]

[0015] where represents the feature concatenation operation, and Conv 1×1 represents a 1×1 convolutional layer.

[0016] 22) Refine the feature map of the shared branch; the feature maps generated by each layer module of the RGB branch, thermal imaging branch, counting map branch, and shared branch are transformed into context information through Equation (1), and this context information is used as the input to output the refined feature map of the shared branch; the specific formula is as follows:

[0017]

[0018] where I r 、I t 、I c 、I s are the context information calculated by Equation (1) from the feature maps F r 、F t 、F c 、F s generated by each layer module of the RGB branch, thermal imaging branch, counting map branch, and shared branch, is the refined feature map of the shared branch, Ir2s , I t2s , I c2s are respectively I r , I t , I c for the residual information of I s , ω r2s , ω t2s , ω c2s are the weight parameters obtained by inputting the corresponding context information I r2s , I t2s , I c2s through a 1×1 convolutional layer; ⊙ is the element-wise multiplication operation;

[0019] 23) Refine the feature maps of the RGB branch, thermal imaging branch, and counting map branch; Generate the context features from the refined shared branch feature map through formula (1) Then, taking as the core, refine the feature maps generated by each layer module of the RGB branch, thermal imaging branch, and counting map branch; The specific formula is as follows:

[0020]

[0021]

[0022]

[0023] In the formula, I s2r , I s2t , I s2c are respectively the refined context information of the shared branch for the context information I r , I t , I c of each layer module of the RGB branch, thermal imaging branch, and counting map branch, ω s2r , ω s2t , ω s2c are the weight parameters obtained by inputting the corresponding context information I s2r , I s2t , I s2c ; F r , F t , F c are respectively the feature maps generated by each layer module of the RGB branch, thermal imaging branch, and counting map branch, are respectively the refined feature maps of the feature maps generated by each layer module of the RGB branch, thermal imaging branch, and counting map branch.

[0024] 24) Data preprocessing; first calculate the mean and variance values of the RGB images and thermal images in the RGBT-CC dataset, and normalize them; during training, the random sampling method is adopted to sample blocks of size 256×256.

[0025] 25) Network training; after setting the path of the dataset, training can be carried out; during training, only the cross-modal part is trained, and the counting map generation module is fixed; after 20 training epochs, validation is performed once every training epoch. When the validation effect is the best, the model parameters are recorded, and a test is performed to obtain the evaluation results of the test set; the model parameters, optimizer parameters, and training epochs are recorded for each training epoch.

[0026] 26) Crowd counting; after the network model is trained, crowd counting can be carried out; RGB images and thermal images are required for crowd counting, and counting maps are generated through thermal images and the LibraNet model; taking RGB images, thermal images, and counting maps as inputs, density maps are output; as a two-dimensional array, the density map can reflect the counting results. Summing the values in the density map can obtain the total number of people in the input RGB image.

[0027] The present invention also includes a cross-modal human flow monitoring system assisted by a counting map, characterized by: a counting map generation module and a crowd counting module.

[0028] The advantages of the present invention are:

[0029] The model provided by the present invention can take into account both illumination changes and scale changes and generate more accurate density maps. Brief Description of the Drawings

[0030] Figure 1 is a flowchart of the method of the present invention.

[0031] Figure 2 is a cross-modal network structure diagram. Convi_j indicates that the number of channels of this module is i and there are j convolutional layers. The convolutional kernel size of each convolutional layer is 3×3. Except for the last module, a 2×2 pooling layer is equipped at the end of each layer. In the four-line branch, the first line is the RGB branch, the second line is the shared branch, the third line is the thermal imaging branch, and the fourth line is the counting map branch. Detailed Embodiments

[0032] The technical solution of the present invention will be further described below with reference to the drawings.

[0033] The cross-modal human flow monitoring method assisted by a counting map of the present invention includes the following steps: 1 Generating a counting map;

[0034] The counting map is for density Figure 1The integral of a fixed area can reflect the scale change of the image. The counting map generation method uses a model based on deep reinforcement learning, LibraNet. This model LibraNet takes the VGG16 convolutional neural network as the front end and the deep reinforcement learning network as the back end, and outputs a counting map. Specifically, it includes:

[0035] 11) Train the deep reinforcement model; Select the RGBT-CC dataset as the training dataset. This dataset is divided into a training set, a test set, and a validation set. Each set has RGB images, thermal images, and annotation data, and the annotation data shows the pixel positions of the head centers of each person in the thermal image. Use the thermal images and annotation data of the training set as training data, and use the thermal images and annotation data of the test set as test data to train the deep reinforcement model LibraNet. Use {-10, -5, -2, -1, 1, 2, 5, 10, 999} as the action set, where 999 is the termination action, and the remaining actions represent the change in the count value in the pixel area. For example, 5 means adding 5 to the count value, and -5 means subtracting 5 from the count value.

[0036] 12) Generate a counting map; Fix the trained LibraNet model and input all the thermal images in the RGBT-CC dataset to generate a counting map. Each value in the counting map can reflect the approximate quantity information of the corresponding pixel area in the thermal image.

[0037] 13) Optimize the counting map; Because the length and width of the generated counting map are only where n represents the number of pooling layers of VGG16 used as the front end of LibraNet, upsampling is required to make the size the same as the original Figure 1 Use the nearest neighbor interpolation method for upsampling and divide by 2 at each pixel position 2*n to keep the total count value unchanged and the image size the same as the original Figure 1 consistent.

[0038] 2 Perform crowd counting based on a cross-modal network model;

[0039] Design a cross-modal network model. This model has four branches, including three input branches and one output branch. The three input branches respectively input RGB, thermal images, and counting maps as three modalities. The output branch is a shared branch, which is initialized to 0, receives and refines the information of the three modalities, and its output is a density map. Except for the shared branch, the other branches are all composed of VGG-19. Since VGG-19 can be divided into 5-layer modules, except for the last layer module, each layer module is finally equipped with a 2×2 pooling layer. Therefore, the three input branches can all be divided into 5-layer modules. The shared branch is composed of VGG-19 after removing the first two layers and can be divided into 4-layer modules. After designing the model, use the RGBT-CC dataset for training. After training the model, perform crowd counting based on this model. Specifically, it includes:

[0040] 21) Generate context information I; use the L-level pyramid pooling layer to extract the context information I from the feature maps F generated by each layer module of each branch in the network. Specifically, for the l-th layer (l = 1, 2,..., L), use a 2 l-1 ×2 l-1 max pooling layer, with the h×w feature map F as the input, output a feature, and then upsample it to h×w using the nearest neighbor interpolation method to form the context feature F l ; Finally, the context information I can be calculated by Equation (1):

[0041]

[0042] where represents the feature concatenation operation, and Conv 1×1 represents a 1×1 convolutional layer.

[0043] 22) Refine the feature map of the shared branch; the feature maps generated by each layer module of the RGB branch, thermal imaging branch, counting map branch, and shared branch are transformed into context information through Equation (1), and this context information is used as the input to output the refined feature map of the shared branch. The specific formula is as follows:

[0044]

[0045]

[0046] where I r , I t , I c , I s are the feature maps F r , F t , F c , F sThe context information calculated by formula (1) is the refined shared branch feature map, I r2s 、I t2s 、I c2s are respectively the residual information of I r 、I t 、I c with respect to I s ; ω r2s 、ω t2s 、ω c2s are the weight parameters obtained by inputting the corresponding context information I r2s 、I t2s 、I c2s into the 1×1 convolutional layer; ⊙ is the element-wise multiplication operation;

[0047] 23) Refine the feature maps of the RGB branch, thermal imaging branch, and counting map branch; Generate the context features from the refined shared branch feature map through formula (1) Then, taking as the core, refine the feature maps generated by each layer module of the RGB branch, thermal imaging branch, and counting map branch; The specific formula is as follows:

[0048]

[0049]

[0050]

[0051] In the formula, I s2r 、I s2t 、I s2c are respectively the residual information of the refined shared branch context information with respect to the context information I r 、I t 、I c of each layer module of the RGB branch, thermal imaging branch, and counting map branch, ω s2r 、ω s2t 、ω s2c are the weight parameters obtained by inputting the corresponding context information I s2r 、I s2t 、I s2c into the 1×1 convolutional layer, F r 、F t 、F c are respectively the feature maps generated by each layer module of the RGB branch, thermal imaging branch, and counting map branch, are respectively the refined feature maps of the feature maps generated by each layer module of the RGB branch, thermal imaging branch, and counting map branch.

[0052] 24) Data preprocessing: First, calculate the mean and variance values of the RGB images and thermal images in the RGBT-CC dataset and normalize them. During training, the random sampling method is adopted to sample blocks of size 256×256.

[0053] 25) Network training: After setting the path of the dataset, training can be carried out. During training, only the cross-modal part is trained, and the counting map generation module is fixed. After 20 training epochs, validation is performed once every training epoch. When the validation effect is the best, the model parameters are recorded, and a test is conducted to obtain the evaluation results of the test set. The model parameters, optimizer parameters, and training epochs are recorded for each training epoch.

[0054] 26) Crowd counting: After the network model is trained, crowd counting can be carried out. RGB images and thermal images are required for crowd counting. The counting map is generated through the thermal image and the LibraNet model. Taking the RGB image, thermal image, and counting map as inputs, a density map is output. As a two-dimensional array, the density map can reflect the counting result. Summing the values in the density map can obtain the total number of people in the input RGB image.

[0055] The present invention also includes a cross-modal human flow monitoring system assisted by a counting map, which is characterized by: a counting map generation module and a crowd counting module, where:

[0056] Counting map generation module. The counting map is the integral of the density Figure 1 over a fixed area and can reflect the scale change of the image. The counting map generation method uses the model LibraNet based on deep reinforcement learning. This model LibraNet uses the VGG16 convolutional neural network as the front end and the deep reinforcement learning network as the back end to output the counting map. Specifically, it includes:

[0057] 11) Training the deep reinforcement model. Select the RGBT-CC dataset as the training dataset. This dataset is divided into a training set, a test set, and a validation set. Each set has RGB images, thermal images, and annotation data. The annotation data shows the pixel positions of the head centers of each person in the thermal image. Use the thermal images and annotation data of the training set as training data, and use the thermal images and annotation data of the test set as test data to train the deep reinforcement model LibraNet. Use {-10, -5, -2, -1, 1, 2, 5, 10, 999} as the action set, where 999 is the termination action, and the remaining actions represent the change magnitude of the count value in the pixel area. For example, 5 means adding 5 to the count value, and -5 means subtracting 5 from the count value.

[0058] 12) Generate a counting map. Fix the trained LibraNet model and input all the thermal images in the RGBT-CC dataset to generate a counting map. Each value in the counting map can reflect the approximate quantity information of the corresponding thermal image pixel area.

[0059] 13) Counting map optimization. Since the length and width of the generated counting map are only where n represents the number of pooling layers of VGG16 used as the front end of LibraNet, upsampling is required to make the size consistent with the original Figure 1 one. Use the nearest neighbor interpolation method for upsampling and divide by 2 at each pixel position 2*n to keep the total count value unchanged and the image size consistent with the original Figure 1 one.

[0060] Crowd counting module. This module uses a trained cross-modal network model, inputs RGB, thermal images, and the counting map as three modalities, and outputs a density map. The model has four branches, including three input branches and one output branch. The three input branches respectively input RGB, thermal images, and the counting map; the output branch is a shared branch, which is initialized to 0, receives and refines the information of the three modalities, and its output is the density map. Except for the shared branch, the remaining branches are all composed of VGG-19. Since VGG-19 can be divided into 5-layer modules, except for the last layer, each layer module is finally equipped with a 2×2 pooling layer, so the three input branches can all be divided into 5-layer modules. The shared branch is composed of VGG-19 with the first two layers removed and can be divided into 4-layer modules. It is trained through the RGBT-CC dataset. After training the model, crowd counting is performed based on this model. Specifically, it includes:

[0061] 21) Generate context information I; use an L-level pyramid pooling layer to extract context information I from the feature maps F generated by each layer module of each branch in the network; specifically, for the l-th layer (l = 1, 2,..., L), use a 2 l-1 ×2 l-1 max pooling layer, take the h×w feature map F as the input, and output a feature, and then upsample it to h×w using the nearest neighbor interpolation method to form the context feature F l ; finally, the context information I can be calculated by Equation (1):

[0062]

[0063] In the formula represents the feature concatenation operation, and Conv 1×1 represents a 1×1 convolutional layer.

[0064] 22) Refine the feature map of the shared branch; the feature maps generated by each layer module of the RGB branch, thermal imaging branch, counting map branch, and shared branch are converted into context information through formula (1), and this context information is used as the input to output the refined feature map of the shared branch; the specific formula is as follows:

[0065]

[0066] In the formula, I r , I t , I c , I s are the feature maps F r , F t , F c , F s generated by each layer module of the RGB branch, thermal imaging branch, counting map branch, and shared branch, the context information calculated through formula (1), is the refined feature map of the shared branch, I r2s , I t2s , I c2s are respectively the residual information of I r , I t , I c with respect to I s , ω r2s , ω t2s , ω c2s are the weight parameters obtained from the context information I r2s , I t2s , I c2s input by the 1×1 convolutional layer; ⊙ is the element-wise multiplication operation;

[0067] 23) Refine the feature maps of the RGB branch, thermal imaging branch, and counting map branch; generate the context features from the refined feature map of the shared branch through formula (1), and then, with as the core, refine the feature maps generated by each layer module of the RGB branch, thermal imaging branch, and counting map branch; the specific formula is as follows:

[0068]

[0069]

[0070]

[0071] In the formula, I s2r , I s2t , I s2c are respectively the refined context information of the shared branch The context information I of each layer module of the RGB branch, thermal imaging branch, and counting map branch r 、I t 、I c The residual information of, ω s2r 、ω s2t 、ω s2c is the weight parameter obtained by inputting the corresponding context information I into the 1×1 convolutional layer s2r 、I s2t 、I s2c The weight parameters obtained, F r 、F t 、F c are the feature maps generated by each layer module of the RGB branch, thermal imaging branch, and counting map branch respectively, are the feature maps refined from the feature maps generated by each layer module of the RGB branch, thermal imaging branch, and counting map branch respectively.

[0072] 24) Data preprocessing; First, calculate the mean and variance values of the RGB images and thermal images in the RGBT-CC dataset and normalize them; During training, the random sampling method is adopted to sample blocks of size 256×256.

[0073] 25) Network training; After setting the path of the dataset, training can be carried out; During training, only the cross-modal part is trained, and the counting map generation module is fixed; After training for 20 epochs, validation is performed once every training epoch. When the validation effect is the best, the model parameters are recorded, and a test is conducted to obtain the evaluation results of the test set; The model parameters, optimizer parameters, and training epochs are recorded for each training epoch.

[0074] 26) Crowd counting; After the network model is trained, crowd counting can be carried out; Crowd counting requires the use of RGB images and thermal images, and a counting map is generated through the thermal images and the LibraNet model; Taking the RGB images, thermal images, and counting map as inputs, a density map is output; As a two-dimensional array, the density map can reflect the counting result. Summing the values in the density map can obtain the total number of people in the input RGB image.

Claims

1. Counting graph-assisted cross-modal pedestrian flow monitoring method, including the following steps: Step 1, generate a counting graph; The counting graph is the integral of a certain area of the density graph and can reflect the scale change of the image. The counting graph generation method uses a model based on deep reinforcement learning, LibraNet. This model LibraNet uses the VGG16 convolutional neural network as the front end and the deep reinforcement learning network as the back end to output the counting graph. Specifically including: 11) Train the deep reinforcement model; Select the RGBT-CC dataset as the training dataset. This dataset is divided into a training set, a test set, and a validation set. Each set has RGB, thermal images, and annotation data. The annotation data shows the pixel positions of the head centers of each person in the thermal image. Use the thermal images and annotation data of the training set as training data, and use the thermal images and annotation data of the test set as test data to train the deep reinforcement model LibraNet. Use {-10, -5, -2, -1, 1, 2, 5, 10, 999} as the action set, where 999 is the termination action, and the remaining actions represent the change size of the count value in the pixel area; 12) Generate a counting graph; Fix the trained LibraNet model and input all the thermal images in the RGBT-CC dataset to generate a counting graph. Each value in the counting graph can reflect the quantity information of the corresponding thermal image pixel area; 13) Counting graph optimization; since the length and width of the generated counting graph are only where n represents the number of pooling layers of VGG16 as the front end of LibraNet, and upsampling is required to make the size consistent with the original image; nearest neighbor interpolation is used for upsampling, and each pixel position is divided by 2 2*n , so that the total count value remains unchanged and the image size is consistent with the original image; Step 2, perform crowd counting based on the cross-modal network model; Design a cross-modal network model. This model has four branches, including three input branches and one output branch. The three input branches respectively input RGB, thermal images, and counting graphs as three modalities. The output branch is a shared branch, which is initialized to 0, receives and refines the information of the three modalities, and its output is the density graph. Except for the shared branch, the other branches are all composed of VGG-19. Since VGG-19 can be divided into 5-layer modules, except for the last layer module, each layer module is finally equipped with a 2×2 pooling layer, so the three input branches can all be divided into 5-layer modules. The shared branch is composed of VGG-19 after removing the first two layers and can be divided into 4-layer modules; After designing the model, use the RGBT-CC dataset for training; After training the model, perform crowd counting based on this model. Specifically including: (21) Generate context information I; use an L-level pyramid pooling layer to extract context information I from the feature maps F generated by each module in each layer of each branch in the network; specifically, for the l-th layer, l = 1, 2,..., L, use a l-1 2 l-1 × 2 maximum pooling layer, with the h × w feature map F as the input, output a feature, and then upsample it to h × w using the nearest neighbor interpolation method to form the context feature F l ; finally, the context information I can be calculated by Equation (1): wherein represents a feature concatenation operation, and Conv 1×1 represents a 1×1 convolutional layer; 22) Refine the feature map of the shared branch; The feature maps generated by each layer module of the RGB branch, thermal image branch, counting graph branch, and shared branch are transformed into context information through formula (1), and this context information is used as the input to output the refined feature map of the shared branch; The specific formula is as follows: I r2s = I r - I s , ω r2s = Conv 1×1 (I r2s ) I t2s = I t - I s , ω t2s = Conv 1×1 (I t2s ), (2) I c2s = I c - I s , ω c2s = Conv 1×1 (I c2s ), where I r 、I t 、I c 、I s are the feature maps F r 、F t 、F c 、F s calculated by formula (1), is the refined feature map of the shared branch, I r2s 、I t2s 、I c2s are respectively the residual information of I r 、I t 、I c with respect to I s , ω r2s 、ω t2s 、ω c2s are the weight parameters obtained from the corresponding context information I r2s 、I t2s 、I c2s input by the 1×1 convolutional layer; ⊙ is the element-wise multiplication operation; 23) Refine the feature maps of the RGB branch, thermal imaging branch, and counting map branch; refine the refined shared branch feature maps Generate context features through formula (1) Then, taking as the core, refine the feature maps generated by each layer module of the RGB branch, thermal imaging branch, and counting map branch; the specific formula is as follows: where I s2r , I s2t , I s2c are respectively the refined shared branch context information for the context information I r , I t , I c of each layer module of the RGB branch, thermal imaging branch, and counting map branch. ω s2r , ω s2t , ω s2c are the weight parameters obtained by inputting the corresponding context information I s2r , I s2t , I s2c into the 1×1 convolutional layer. F r , F t , F c are respectively the feature maps generated by each layer module of the RGB branch, thermal imaging branch, and counting map branch, and are respectively the refined feature maps of the feature maps generated by each layer module of the RGB branch, thermal imaging branch, and counting map branch; 24) Data preprocessing; First calculate the mean and variance values of the RGB images and thermal images in the RGBT-CC dataset and normalize them. During training, the random sampling method is adopted to sample blocks of size 256×256; 25) Network training; After setting the path of the dataset, start the training; During training, only train the cross-modal part and fix the counting graph generation module; After 20 training epochs, perform validation once every training epoch. When the validation effect is the best, record the model parameters and conduct a test to obtain the evaluation results of the test set; Record the model parameters, optimizer parameters, and training epochs for each training epoch. 26) Crowd counting; After the network model is trained, perform crowd counting; Crowd counting requires the use of RGB images and thermal images to generate a counting graph through the thermal image and the LibraNet model; Take the RGB image, thermal image, and counting graph as inputs and output a density map; The density map, as a two-dimensional array, can reflect the counting result; Summing the values in the density map can obtain the total number of people in the input RGB image.

2. Cross-modal pedestrian flow monitoring system assisted by a counting graph Characterized in that: A counting graph generation module and a crowd counting module, where: The counting graph generation module. The counting graph is the integral of a certain area of the density map and can reflect the scale change of the image; The counting graph generation method uses a model based on deep reinforcement learning, LibraNet. This model LibraNet uses the VGG16 convolutional neural network as the front end and the deep reinforcement learning network as the back end to output the counting graph; Specifically includes: 11) Training the deep reinforcement model; Select the RGBT-CC dataset as the training dataset. This dataset is divided into a training set, a test set, and a validation set. Each set has RGB images, thermal images, and annotation data. The annotation data shows the pixel positions of the head centers of each person in the thermal image; Use the thermal images and annotation data in the training set as training data, and use the thermal images and annotation data in the test set as test data to train the deep reinforcement model LibraNet; Use {-10, -5, -2, -1, 1, 2, 5, 10, 999} as the action set, where 999 is the termination action, and the remaining actions represent the change magnitude of the count value in the pixel area. 12) Generating the counting graph; Fix the trained LibraNet model and input all the thermal images in the RGBT-CC dataset to generate the counting graph; Each value in the counting graph can reflect the quantity information of the corresponding thermal image pixel area. 13) Counting graph optimization; since the length and width of the generated counting graph are only where n represents the number of pooling layers of VGG16 as the front end of LibraNet, and upsampling is required to make the size consistent with the original image; nearest neighbor interpolation is used for upsampling, and each pixel position is divided by 2 2*n so that the total count value remains unchanged and the image size is consistent with the original image; Crowd counting module. This module inputs RGB images, thermal images, and counting maps as three modalities into a trained cross-modal network model and outputs a density map. The model has four branches, including three input branches and one output branch. The three input branches respectively input RGB images, thermal images, and counting maps. The output branch is a shared branch, which is initialized to 0, receives and refines the information of the three modalities, and its output is the density map. Except for the shared branch, the other branches are all composed of VGG-19. Since VGG-19 can be divided into 5-layer modules, and except for the last layer, each layer module is finally equipped with a 2×2 pooling layer, so the three input branches can all be divided into 5-layer modules. The shared branch is composed of VGG-19 with the first two layers removed and can be divided into 4-layer modules. It is trained using the RGBT-CC dataset. After training the model, crowd counting is performed based on this model. Specifically, it includes: (21) Generate context information I; use an L-level pyramid pooling layer to extract context information I from the feature maps F generated by each module in each layer of each branch in the network; specifically, for the l-th layer, l = 1, 2,..., L, use a l-1 ×2 l-1 max pooling layer, take the h×w feature map F as the input, output a feature, and then upsample it to h×w using the nearest neighbor interpolation method to form the context feature F l ; finally, the context information I can be calculated by Equation (1): In the formula represents the feature concatenation operation, and Conv 1×1 represents a 1×1 convolutional layer; 22) Refine the feature map of the shared branch. The feature maps generated by each layer module of the RGB branch, thermal imaging branch, counting map branch, and shared branch are transformed into context information through formula (1), and this context information is used as the input to output the refined feature map of the shared branch. The specific formula is as follows: where I r 、I t 、I c 、I s are the feature maps F r 、F t 、F c 、F s calculated by formula (1), is the refined feature map of the shared branch, I r2s 、I t2s 、I c2s are respectively the residual information of I r 、I t 、I c with respect to I s , ω r2s 、ω t2s 、ω c2s are the weight parameters obtained by inputting the corresponding context information I r2s 、I t2s 、I c2s into the 1×1 convolutional layer; ⊙ is the element-wise multiplication operation; 23) Refine the feature maps of the RGB branch, the thermal imaging branch, and the counting map branch; the refined shared branch feature maps Generate context features through formula (1) Then, As the core, refine the feature maps generated by each layer module of the RGB branch, the thermal imaging branch, and the counting map branch; the specific formula is as follows: where I s2r , I s2t , I s2c are the refined shared branch context information for the context information I of each layer module of the RGB branch, thermal imaging branch, and counting map branch r , I t , I c residual information, ω s2r , ω s2t , ω s2c are the weight parameters obtained by inputting the corresponding context information I s2r , I s2t , I s2c through a 1×1 convolutional layer, F r , F t , F c are the feature maps generated by each layer module of the RGB branch, thermal imaging branch, and counting map branch respectively, are the refined feature maps of the feature maps generated by each layer module of the RGB branch, thermal imaging branch, and counting map branch respectively; 24) Data preprocessing. First, calculate the mean and variance values of the RGB images and thermal images in the RGBT-CC dataset and normalize them. During training, the random sampling method is adopted to sample blocks of size 256×256. 25) Network training. After setting the path of the dataset, training is carried out. During training, only the cross-modal part is trained, and the counting map generation module is fixed. After training for 20 epochs, validation is performed once every 1 training epoch. When the validation effect is the best, the model parameters are recorded, and a test is performed to obtain the evaluation results of the test set. The model parameters, optimizer parameters, and training epochs are recorded for each training epoch. 26) Crowd counting. After the network model is trained, crowd counting is performed. Crowd counting requires the use of RGB images and thermal images to generate a counting map through the thermal image and the LibraNet model. Using the RGB image, thermal image, and counting map as inputs, a density map is output. The density map, as a two-dimensional array, can reflect the counting result. By summing the values in the density map, the total number of people in the input RGB image can be obtained.

Citation Information

Patent Citations

  • A human-certificate integrated verification terminal, system and method based on multi-mode face recognition

    CN109902780A

  • Lightweight feature fusion crowd counting method and system

    CN112861718A