Fine-grained crowd counting method and system based on double-branch cooperation
By introducing a dual-branch collaboration method in the population counting technology, combining the technical means of density branches and semantic branches, the problem of insufficient ability to distinguish groups in the existing technology is solved, and a more efficient and accurate fine-grained population counting and classification is achieved.
Patent Information
- Application Number
- CN202411945567.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has shortcomings in population counting accuracy, fine analysis of population categories and adaptability to complex scenarios, especially in scenarios where crowds are dense and severely obstructed, counting accuracy is poor.
A fine-grained population counting method based on dual-branch collaboration is proposed. By constructing density branches and semantic branches, multi-scale fusion of population density distribution and mining of human and environmental context information, the fine-grained population counting model is generated after the fusion.
It improves the accuracy and category distinction ability of population counting, can estimate the number of people more accurately, and deeply analyze the relationship between the details of people's behavior and the environment, adapt to complex scenarios, and reduces counting errors.
Smart Images

Figure CN120014534A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision and deep learning technology, and more specifically, to a fine-grained crowd counting method and system based on dual-branch collaboration. Background Art
[0002] In the field of crowd analysis, crowd counting and classification are important research topics, which are widely used in many aspects such as public safety and business intelligence. Traditional crowd counting methods, such as those based on manual counting and simple sensors, are inefficient and difficult to meet the needs of large-scale complex scenarios.
[0003] With the development of computer vision technology, image-based crowd counting methods have emerged. Existing technical methods mostly focus on estimating the overall crowd density, using simple feature extraction and regression models. Although they improve counting efficiency to a certain extent, their accuracy is poor in complex scenarios. For example, in crowded street scenes, the crowd is severely obscured, and individual features are difficult to accurately extract, resulting in large counting errors; at large-scale events, the crowd is unevenly distributed and the scale varies significantly. These methods are difficult to adapt and cannot accurately reflect the number of people.
[0004] In terms of crowd classification, traditional methods mainly rely on manually set rules to distinguish crowd categories based on appearance characteristics (such as clothing and body shape). However, this method has obvious limitations. On the one hand, in complex environments, such as when the light changes strongly or the background interference is serious, the appearance characteristics are easily affected, resulting in classification errors; on the other hand, it is difficult to effectively distinguish people with similar behavioral characteristics but different actual categories, such as people with similar walking speeds but different purposes, and it is impossible to deeply analyze the relationship between crowd behavior details and the environment, making it difficult to meet the needs of refined crowd analysis.
[0005] In summary, the existing technology has many shortcomings in terms of crowd counting accuracy, detailed analysis of crowd categories and adaptability to complex scenarios. New methods are urgently needed to overcome these limitations and achieve more accurate and efficient fine-grained crowd counting and classification.
[0006] Prior art, such as a Chinese patent application with publication number "CN118823685A", discloses a crowd positioning method based on a hybrid expert network, which includes the following steps: Step S1: Divide the crowd data set into a training set, a validation set and a test set, and perform preprocessing on the crowd images and corresponding crowd position labels in the data set; Step S2: Construct a convolutional neural network encoder, encode the training image input and extract deep features; Step S3: Construct a hybrid expert Transformer module, combine its self-attention mechanism and the hybrid expert network, perform special processing on different areas of the input features, and further refine the crowd position features; Step S4: Construct a KMO matcher to match the real crowd coordinates with the predicted crowd coordinates and supervise the model training; Step S5: Construct and train a crowd positioning model based on a hybrid expert network; Step S6: Input the crowd image into the crowd positioning model and output the predicted crowd position information.
[0007] The problem with the above-mentioned prior art is that the method focuses on crowd positioning and only outputs crowd position information. When dealing with scenes with dense crowds and severe occlusion, the crowd counting accuracy may be poor. Summary of the invention
[0008] In order to solve the above technical problems, the present invention proposes a fine-grained crowd counting method and system based on dual-branch collaboration.
[0009] The technical solution of the present invention is as follows:
[0010] The present invention proposes a fine-grained crowd counting method based on dual-branch collaboration, comprising the following steps:
[0011] Step S1, collect crowd images and perform annotation processing, construct a fine-grained crowd counting data set, divide the fine-grained crowd counting data set into a training set and a test set, and perform data preprocessing on the fine-grained crowd counting data set; crowd image annotation processing includes: crowd classification information, individual behavior attributes, and crowd location and distribution relationship;
[0012] Step S2, constructing a crowd feature encoder to extract features from crowd images in the training set;
[0013] Step S3, constructing a density branch, performing multi-scale fusion on the features extracted by the crowd feature encoder, and obtaining the total number of people in the scene through density regression of the multi-scale fused features;
[0014] Step S4, constructing a semantic branch, mining the contextual information of people and the environment in the scene based on the features extracted by the crowd feature encoder, and generating semantic graphs of different categories;
[0015] Step S5, fusing the density branch and the semantic branch to construct a fine-grained crowd counting model based on dual-branch collaboration, using the training set images to train the fine-grained crowd counting model based on dual-branch collaboration, and obtaining a trained fine-grained crowd counting model based on dual-branch collaboration;
[0016] Step S6, input the crowd images of the test set into the trained fine-grained crowd counting model based on dual-branch collaboration, output the corresponding crowd density map and semantic map, and complete the fine-grained crowd counting and classification tasks.
[0017] As a preferred implementation, the crowd feature encoder adopts a VGG16 neural network model, wherein:
[0018] The feature map resizing function is:
[0019]
[0020] Where: w i and b i are the weights and biases of the convolutional layer; Feature maps of different dimensions for input; is the adjusted feature map; Resize(·) is the size adjustment function;
[0021] The feature map fusion function is:
[0022]
[0023] Where: is the fused feature map output by the crowd feature encoder; ConvGroup(·) is multiple convolution group operations; Concat(·) is a concatenation operation; i∈M.
[0024] As a preferred implementation, the density branch is constructed to perform multi-scale fusion on the extracted features; the specific formula for multi-scale fusion is as follows:
[0025]
[0026] Where: is the feature after multi-scale fusion of density branch; Upsample(·) is the upsampling function; ConvModule1(·) is the density branch convolution module operation; Concat(·) is the feature concatenation operation in the new dimension.
[0027] As a preferred implementation, the semantic branch performs feature fusion through a dual feature propagation module, extracts key feature indexes, and enhances key features using a cross-attention mechanism; wherein:
[0028] The feature fusion formula is as follows:
[0029]
[0030] Where: X fuse is the fused feature; is the feature of the input feature propagation module; DownSample(·) is the downsampling operation;
[0031] The index selection function is as follows:
[0032] index SFP ,w=Select SFP (X fuse );
[0033] Where: index SFP is the feature index point; w is the weight; Select SFP (·) Index selection function;
[0034] The specific expression of the cross-attention mechanism to enhance the key features is as follows:
[0035] CrossAttention(Q,K,V)=softmax(Q,K T )V;
[0036] Where: Q is the significant feature composition of the query vector filtered by the index function; K is the key feature; V is the value feature; K and V are obtained by weighting the feature vectors at the corresponding positions in the significant features through weights.
[0037] As a preferred implementation, in the feature processing step of the semantic branch, the feature distinguishing ability of the semantic branch is enhanced by a category loss function, and the category loss function is specifically as follows:
[0038] L CD =max(Sim(v m ,v n ),0);
[0039] Where: L CD is the category loss function; v m and v n is the representative feature embedding of this category; Sim(v m ,v n ) is the similarity between two feature vectors.
[0040] As a preferred implementation, in the process of using the training set images to train the fine-grained crowd counting model based on dual-branch collaboration, the stochastic gradient descent method is used to update the model parameters, and the overall objective function is:
[0041] L=L MAE +L MSE+γL CD ;
[0042] Where: L is the overall objective function; L MAE is the mean absolute error between the predicted density map and the true density map; L MSE The absolute deviation between the predicted value and the true value; γ is the balance factor.
[0043] On the other hand, the present invention also provides a fine-grained crowd counting system based on dual-branch collaboration, comprising:
[0044] The data preparation module collects crowd images and performs annotation processing, constructs a fine-grained crowd counting dataset, divides the fine-grained crowd counting dataset into a training set and a test set, and performs data preprocessing on the fine-grained crowd counting dataset; crowd image annotation processing includes: crowd classification information, individual behavior attributes, and crowd location and distribution relationship;
[0045] Feature extraction module, builds a crowd feature encoder and extracts features from crowd images in the training set;
[0046] Density analysis module: builds density branches, performs multi-scale fusion on the features extracted by the crowd feature encoder, and uses density regression to obtain the total number of people in the scene.
[0047] The semantic analysis module builds semantic branches and deeply mines the contextual information of people and the environment in the scene based on the features extracted by the crowd feature encoder to generate semantic graphs of different categories.
[0048] The model training module integrates the density branch and the semantic branch to build a fine-grained crowd counting model based on dual-branch collaboration. The training set images are used to train the fine-grained crowd counting model based on dual-branch collaboration to obtain a trained fine-grained crowd counting model based on dual-branch collaboration.
[0049] The model prediction module inputs the crowd images of the test set into the trained fine-grained crowd counting model based on dual-branch collaboration, outputs the corresponding crowd density map and semantic map, and completes the fine-grained crowd counting and classification tasks.
[0050] As a preferred implementation, the feature extraction module and the crowd feature encoder adopt a VGG16 neural network model, wherein:
[0051] The feature map resizing function is:
[0052]
[0053] Where: w i and b i are the weights and biases of the convolutional layer; Feature maps of different dimensions for input; is the adjusted feature map; Resize(·) is the size adjustment function;
[0054] The feature map fusion function is:
[0055]
[0056] Where: is the fused feature map output by the crowd feature encoder; ConvGroup(·) is multiple convolution group operations; Concat(·) is a concatenation operation; i∈M.
[0057] On the other hand, the present invention further provides an electronic device having a computer program stored thereon, wherein when the computer program is executed by a processor, the computer program implements a fine-grained crowd counting method based on dual-branch collaboration as described in any embodiment of the present invention.
[0058] On the other hand, the present invention also provides a computer-readable medium for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement a fine-grained crowd counting method based on dual-branch collaboration as described in any embodiment of the present invention.
[0059] The present invention has the following beneficial effects:
[0060] 1. Improved accuracy of crowd counting: By integrating multi-scale features with density branches to regress crowd density distribution, the number of people can be estimated more accurately. In complex scenes such as crowded streets and large-scale events, the scale change and occlusion problems can be effectively dealt with. Compared with existing technologies, the mean absolute error and the mean absolute error of the category average are significantly reduced, making the crowd count estimation closer to reality.
[0061] 2. Deepen the understanding of crowd categories: The semantic branch's two-stage salient feature propagation module 2-SFP mines crowd and environmental context information, the foreground module extracts crowd features, and the background module mines environmental clues. The synergy enables the model to accurately infer crowd categories. It can not only distinguish between standing and sitting people, but also determine walking directions, breaking through the limitations of existing technologies, deepening the details of individual behavior and environmental associations, and providing richer and more accurate semantic information.
[0062] 3. Enhanced model differentiation ability: The category difference loss function increases the feature distance of different crowd categories and strengthens the ability to distinguish semantic branches. In the case of similar appearance of different crowd categories, such as approaching and leaving people in grayscale images, it can still be accurately identified, reducing the probability of misclassification, improving the recognition of people in complex scenes, and making the classification results more reliable.
[0063] 4. Improved flexibility and versatility: In view of the limitations of fully supervised fine-grained crowd counting tasks, this invention adopts dual-branch design and innovative loss function to better adapt to new crowd categories or scene changes. When encountering new behavior patterns or layout changes in actual applications, effective counting and classification can be achieved without large-scale retraining, reducing manpower and time costs and enhancing the applicability of the model in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0065] Figure 1 It is a schematic diagram of the process of the present invention;
[0066] Figure 2 is a diagram of a network model structure in an embodiment of the present invention;
[0067] Figure 3 is a structural diagram of a salient feature propagation module in an embodiment of the present invention;
[0068] Figure 4 This is the calculation process of the category difference loss in the embodiment of the present invention. DETAILED DESCRIPTION
[0069] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0070] It should be understood that the step numbers used in this document are only for convenience of description and are not intended to limit the order in which the steps are executed.
[0071] It should be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0072] The terms “include” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0073] The term "and / or" means and includes any and all possible combinations of one or more of the associated listed items.
[0074] Embodiment 1:
[0075] In order to make the purpose, technical solutions and advantages of the present invention more clear, the following will be combined with the specific embodiments of the present application and refer to the attached Figure 1 , clearly and completely describe the technical solution of the present invention.
[0076] To solve the problems in the prior art, the present invention provides a fine-grained crowd counting method based on dual-branch collaboration, comprising the following steps:
[0077] Step S1, collect crowd images and perform annotation processing, construct a fine-grained crowd counting data set, divide the fine-grained crowd counting data set into a training set and a test set, and perform data preprocessing on the fine-grained crowd counting data set; crowd image annotation processing includes: crowd classification information, individual behavior attributes, and crowd location and distribution relationship;
[0078] Step S11: Collect crowd images and perform annotation processing to construct a fine-grained crowd counting dataset. For the images in the dataset, the dataset is divided into a training set and a test set according to a certain ratio. The training set includes 400 images and the test set includes 1200 images. The fine-grained crowd counting dataset includes images with standing and sitting labels.
[0079] Step S12: Perform data augmentation on the images in the training set to increase the number of samples in the data set, including randomly flipping images, randomly cropping images, photometric distortion, etc.;
[0080] Step S13: Preprocess the image after data enhancement. First, the image size is cropped to pixels, and then the cropped image is normalized to convert the image data into a standard normal distribution. In order to ensure that the size and position of the crowd area in the label correspond to the crowd image, the same operation is performed on the label at each step of image data enhancement and image preprocessing.
[0081] Step S2, constructing a crowd feature encoder to extract features from crowd images in the training set;
[0082] The crowd feature encoder adopts the VGG16 neural network model, where:
[0083] The feature map resizing function is:
[0084]
[0085] Where: w i and b i are the weights and biases of the convolutional layer; Feature maps of different dimensions for input; is the adjusted feature map; Resize(·) is the size adjustment function;
[0086] The feature map fusion function is:
[0087]
[0088] Where: is the fused feature map output by the crowd feature encoder; ConvGroup(·) is multiple convolution group operations; Concat(·) is a concatenation operation; i∈M.
[0089] Step S21: Input the crowd image into the first 10 convolutional layers of the VGG16 network to obtain dimensions of Four-layer feature map Among them, h, w and c represent the height, width and number of channels of the feature map respectively, c=32, h=512, w=512, and the encoded features from the last three stages of the feature encoder are retained. In the process of decoding features in subsequent modules, these three feature maps will serve as auxiliary information to help the decoder make full use of features at different levels and realize the effective fusion of detail information and context information, thereby improving the prediction accuracy and detail retention ability of the model.
[0090] Step S22: The feature map obtained from step S21 They are input into the 1×1 convolutional layer respectively, the number of channels of the corresponding features is adjusted, and then the size of the feature map is adjusted by the Resize(·) function. The specific expression is:
[0091]
[0092] Among them, w i ,b i are the weights and biases of the 1×1 convolutional layer. Through the 1×1 convolutional layer, The dimensions are adjusted to
[0093] Step S23: Get the data from step S22 Perform fusion processing to obtain the converted features It is used as the input of the subsequent density branch and semantic branch. The specific expression is:
[0094]
[0095] Among them, ConvGroup(·) represents multiple 3×3 convolution groups, and Concat(·) means that the features are concatenated in a new dimension;
[0096] Step S3, constructing a density branch, performing multi-scale fusion on the features extracted by the crowd feature encoder, and obtaining the total number of people in the scene through density regression of the multi-scale fused features;
[0097] The density branch is constructed to perform multi-scale fusion of the extracted features; the specific formula for multi-scale fusion is as follows:
[0098]
[0099] Where: is the feature after multi-scale fusion of density branch; Upsample(·) is the upsampling function; ConvModule1(·) is the density branch convolution module operation; Concat(·) is the feature concatenation operation in the new dimension.
[0100] Step S31: Given the input obtained by the previous module Perform the first stage of decoding and compare it with the encoder features The concatenation is performed to achieve effective fusion of information from different sources and provide rich context information for the subsequent decoding process. After convolution and upsampling, the first stage of decoding is completed, and the first stage of decoding features are obtained. In this process, the model can initially integrate feature information at different levels and begin to build a preliminary understanding of the population density distribution. The specific expression is as follows:
[0101]
[0102] Among them, ConvModule1(·) represents a 3×3 convolution module, Upsample(·) represents an upsampling function, and Concat(·) represents concatenation of features in a new dimension;
[0103] Step S32: The first stage decoding features obtained from step S31 Perform the second stage of decoding on it and compare it with the third layer feature of the encoder After splicing again, the same convolution and upsampling steps as step S31 are performed to complete the second stage of decoding. This step aims to enable the model to comprehensively utilize rich contextual information and accurately capture the global semantic and structural features of the image through such a fusion method.
[0104] Step S33: The fused features obtained in step S32 are subjected to density regression to obtain a category-independent total density map, which reflects the density distribution of all people in the image. By accumulating the values in the total density map, the model can obtain the number of all people in the scene.
[0105] Step S4, constructing a semantic branch, mining the contextual information of people and the environment in the scene based on the features extracted by the crowd feature encoder, and generating semantic graphs of different categories;
[0106] The semantic branch performs feature fusion through a dual feature propagation module, extracts key feature indexes, and enhances key features using a cross-attention mechanism; wherein:
[0107] The feature fusion formula is as follows:
[0108]
[0109] Where: X fuse is the fused feature; is the feature of the input feature propagation module; DownSample(·) is the downsampling operation;
[0110] The index selection function is as follows:
[0111] index SFP ,w=Select SFP (X fuse );
[0112] Where: index SFP is the feature index point; w is the weight; Select SFP (·) Index selection function;
[0113] The specific expression of the cross-attention mechanism to enhance the key features is as follows:
[0114] CrossAttention(Q,K,V)=softmax(Q,K T )V;
[0115] Where: Q is the significant feature composition of the query vector filtered by the index function; K is the key feature; V is the value feature; K and V are obtained by weighting the feature vectors at the corresponding positions in the significant features through weights.
[0116] In the feature processing step of the semantic branch, the feature distinguishing ability of the semantic branch is enhanced by a category loss function, and the category loss function is specifically as follows:
[0117] L CD =max(Sim(v m ,v n),0);
[0118] Where: L CD is the category loss function; v m and v n is the representative feature embedding of this category; Sim(v m ,v n ) is the similarity between two feature vectors.
[0119] Step S41: Figure 3 As shown, in order to enhance the crowd feature extraction, that is, to enhance the foreground feature, we select It is worth noting that although It is also rich in foreground information, but due to the low feature resolution, some details will be lost, resulting in suboptimal results, so we choose The characteristics and features from the encoder The input foreground salient feature propagation module P-SFP is effectively fused. Specifically, the two features are first input into the P-SFP module as low-resolution deep features. They ignore low-level information such as background texture and have certain abstractness and high semantics. Inside the P-SFP module, the input encoding features and decoding features are first integrated together, and then a unique index selection function Select P-SFP (·) is processed, and the specific formula is as follows:
[0120]
[0121] index P-SFP ,w=Select P-SFP (X fuse )
[0122] Select P-SFP (·) The structure consists of a 3×3 convolutional layer, a Sigmoid activation layer, and an adaptive maximum pooling layer. The convolutional layer performs a preliminary transformation on the input features and extracts preliminary feature information; the Sigmoid activation layer limits the feature value to the range of 0 to 1, obtains the relative weight of each feature, highlights the key features and suppresses irrelevant information; the adaptive maximum pooling layer selects several feature index points with the largest weight in the entire feature map. P-SFP , and output the corresponding weight w. These key feature indexes extracted by the index selection structure will guide the model from the encoding feature The significant feature vectors are selected for effective enhancement.
[0123] The model uses the selected salient features as the query vector Query of the cross-attention module, and decodes the features according to the weight w The feature vectors at the corresponding positions in the image are weighted to form key features Key and value features Value, and the relationship between the salient features and the coding features is modeled through calculation to infer their spatial relationship and interaction mode. Finally, the salient region features strengthened by the cross attention module are fused with the original coding features, and the index index is used to represent the salient region features. P-SFP Add back to the corresponding position in the decoded feature as the output feature of the P-SFP module The specific formula is as follows:
[0124] CrossAttention(Q,K,V)=softmax(Q,K T )V
[0125] This step aims to accurately extract and enhance key features that are closely related to the population itself, ensuring that the model can deeply explore and identify the unique attributes of the population;
[0126] Step S42: In order to further mine the environmental context information related to the crowd, select features rich in environmental background details The output features of the foreground salient feature propagation module obtained in step S41 are fused. and features from the encoder Input the background salient feature propagation module B-SFP. Among them, As a shallow feature with a resolution of 1 / 4 of the input image, it contains rich detail information and location information. The processing flow of the B-SFP module is similar to that of the P-SFP module. It also first extracts key feature indexes through the index selection structure, and then uses the cross attention mechanism to enhance the features. The output features of the B-SFP module are recorded as
[0127] Step S43: The output features of the foreground salient feature propagation module and the background salient feature propagation module are and Perform feature fusion to obtain feature X s ,This fusion method is the same as the density branch, ensuring that the model can integrate information from different sources and complement each other's details. The fused features are processed through a series of convolution and upsampling operations to generate semantic maps of different categories. In the semantic map, each pixel value is in the interval, which represents the probability that the pixel in the image belongs to a specific category;
[0128] Step S44: Obtain feature X from step S43 sPerform a region-based average pooling (RAP) operation. When performing this operation, you need to use category-specific crowd region masks, which provide accurate location information for each category of people. By extracting the features corresponding to the mask position of each crowd category, and then calculating the average of these features through the regional average operation, we can get the representative feature embedding v of the category. k (k=1,2). This step can capture the average feature value from each region and use it as the feature representation of the region, providing a basis for the subsequent calculation of the category difference loss;
[0129] Step S45: Figure 4 As shown, the category difference loss function L is applied to the representative features v1 and v2 obtained in step S44. CD To supervise. The category difference loss function is defined as:
[0130] L CD =max(Sim(v1,v2),0)
[0131] Among them, Sim(v1,v2) measures the similarity between two feature vectors by calculating the cosine similarity, and its calculation formula is:
[0132]
[0133] The value range of the calculation result is defined as [-1,1]. If the calculation result is close to 1, it indicates that there is a high similarity between the two feature vectors; if the result is close to -1, it means that the similarity between the two vectors is extremely low; when the result is 0, it means that the two vectors are orthogonal, that is, there is no similarity between them. By introducing the category difference loss function, the model will be supervised during the training process, prompting it to create the largest possible distance for the feature representations of different categories in the embedding space, thereby encouraging the model to learn and capture the individual characteristics of a specific category, enhance the sensitivity to the characteristics of different categories of people, and improve the accuracy of the model's inference of population categories, even when the characteristics of the population are highly similar. Excellent performance can be maintained.
[0134] Step S5, such as Figure 2 As shown, the density branch and the semantic branch are fused to construct a fine-grained crowd counting model based on dual-branch collaboration, and the training set images are used to train the fine-grained crowd counting model based on dual-branch collaboration to obtain a trained fine-grained crowd counting model based on dual-branch collaboration;
[0135] In the process of using the training set images to train the fine-grained crowd counting model based on dual-branch collaboration, the stochastic gradient descent method is used to update the model parameters, and the overall objective function is:
[0136] L=L MAE +L MSE +γL CD ;
[0137] Where: L is the overall objective function; L MAE is the mean absolute error between the predicted density map and the true density map; L MSE The absolute deviation between the predicted value and the true value; γ is the balance factor.
[0138] Step S51: Use the first ten layers of VGG-16 as the basic architecture for feature extraction to build a feature encoder, and retain the encoding features of the last three stages during the encoding process It is used to provide different levels of information supplement to the decoder in the density branch and decoding branch. At the same time, the density branch and semantic branch are constructed. The density branch is designed according to the process in step S3 above, and is used to fuse multi-scale features to regress the total density map of the entire scene; the semantic branch is constructed according to the process in step S4, and uses the two-stage significant feature propagation module to deeply explore the key information between people and social environment and generate a category semantic map;
[0139] Step S52: Input the training set image into the feature encoder for encoding, and input the obtained encoded features into the density branch and the decoding branch respectively. In the density branch, input the feature With encoder characteristics After concatenation, it undergoes convolution, upsampling, and After splicing, convolution, upsampling, and multi-scale feature fusion, the class-independent total density map is finally obtained through density regression. and After being processed by the foreground salient feature propagation module P-SFP and Then it is processed by the background salient feature propagation module B-SFP to obtain Finally and Perform feature fusion and generate category semantic graph through a series of operations;
[0140] Step S53: Perform a dot multiplication operation on the total density map obtained by the density branch regression and the category semantic map generated by the semantic branch to obtain the density map of each category. These density maps, as the core of the model output, directly reflect the quantity distribution of each category of people in the scene;
[0141] Step S54: Calculate the mean absolute error (MAE), mean square error (MSE) and category difference loss (L according to the population density maps and true labels output by the model. CD). MAE and MSE are used to measure the error between the predicted density map and the true density map, and the category difference loss is obtained by processing the semantic branch features according to the calculation method of step S4. Specifically, MAE and MSE are calculated as follows:
[0142]
[0143] Where k (k = 1, 2) represents the population category, Represents the density map of the k-th group of people predicted by the model, Y k Represents the density map of the kth class generated by the true label;
[0144] Step S55: Back propagate the calculated loss value and update the model parameters using the stochastic gradient descent method. During the back propagation process, according to the overall objective function:
[0145] L=L MAE +L MSE +γL CD
[0146] (γ is the category difference loss L CD The gradient of each parameter is calculated and the parameters are adjusted to gradually reduce the loss value.
[0147] Step S56: Repeat the above steps S52 to S55 in batches until the calculated loss value converges and stabilizes. During the training process, the model parameters are continuously adjusted so that the model can better fit the training data and improve the accuracy of fine-grained crowd counting tasks. Through multiple iterative training, the model gradually learns the complex relationship between crowd density distribution and crowd categories, so that it can accurately estimate the number of various types of people in the scene and classify them. When the loss value stabilizes, save the trained model parameters, complete the training process of the fine-grained crowd counting model based on dual-branch collaboration, and obtain a model that can accurately perform fine-grained crowd counting tasks in practical applications;
[0148] Step S6, input the crowd images of the test set into the trained fine-grained crowd counting model based on dual-branch collaboration, output the corresponding crowd density map and semantic map, and complete the fine-grained crowd counting and classification tasks.
[0149] Step S61: Input the test crowd image into the trained fine-grained crowd counting model based on dual-branch collaboration to generate a total density map of the test image, which reflects the overall density distribution of all people in the image, and generates an accurate category semantic map, in which each pixel value represents the probability that the pixel belongs to a specific category.
[0150] Step S62: Perform a dot multiplication operation on the total density map obtained by the density branch and the category semantic map generated by the semantic branch to obtain the density map of each category. These category density maps clearly show the density distribution of people of different categories in the image, and provide detailed information for further analysis of the number, distribution and behavior patterns of the crowd. At the same time, the output category semantic map intuitively reflects the crowd category to which each pixel in the image belongs, assisting in understanding the classification of the crowd. Finally, the model successfully outputs the crowd density map and semantic map corresponding to the test crowd image, completes the fine-grained crowd counting and classification tasks for the test image, and provides valuable crowd analysis results for practical applications.
[0151] Embodiment 2:
[0152] This embodiment provides a fine-grained crowd counting system based on dual-branch collaboration, including:
[0153] The data preparation module collects crowd images and performs annotation processing, constructs a fine-grained crowd counting dataset, divides the fine-grained crowd counting dataset into a training set and a test set, and performs data preprocessing on the fine-grained crowd counting dataset; crowd image annotation processing includes: crowd classification information, individual behavior attributes, and crowd location and distribution relationship;
[0154] Feature extraction module, builds a crowd feature encoder and extracts features from crowd images in the training set;
[0155] Density analysis module: builds density branches, performs multi-scale fusion on the features extracted by the crowd feature encoder, and uses density regression to obtain the total number of people in the scene.
[0156] The semantic analysis module builds semantic branches and deeply mines the contextual information of people and the environment in the scene based on the features extracted by the crowd feature encoder to generate semantic graphs of different categories.
[0157] The model training module integrates the density branch and the semantic branch to build a fine-grained crowd counting model based on dual-branch collaboration. The training set images are used to train the fine-grained crowd counting model based on dual-branch collaboration to obtain a trained fine-grained crowd counting model based on dual-branch collaboration.
[0158] The model prediction module inputs the crowd images of the test set into the trained fine-grained crowd counting model based on dual-branch collaboration, outputs the corresponding crowd density map and semantic map, and completes the fine-grained crowd counting and classification tasks.
[0159] Embodiment three:
[0160] This embodiment provides an electronic device having a computer program stored thereon. When the computer program is executed by a processor, the method for fine-grained crowd counting based on dual-branch collaboration as described in any embodiment of the present invention is implemented.
[0161] Embodiment 4:
[0162] This embodiment provides a computer-readable medium for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement a fine-grained crowd counting method based on dual-branch collaboration as described in any embodiment of the present invention.
[0163] In the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can be represented by: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, c can be single or multiple.
[0164] Those of ordinary skill in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented in a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0165] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0166] In several embodiments provided in the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory; hereinafter referred to as: ROM), random access memory (Random Access Memory; hereinafter referred to as: RAM), disk or optical disk, and other media that can store program codes.
[0167] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A fine-grained crowd counting method based on dual-branch collaboration, characterized in that: The following steps are involved: Step S1, collecting crowd images and performing annotation processing, constructing a fine-grained crowd counting dataset, dividing the fine-grained crowd counting dataset into a training set and a test set, and performing data preprocessing on the fine-grained crowd counting dataset; Crowd image annotation processing includes: crowd classification information, individual behavior attributes, and crowd location and distribution relationship; Step S2, constructing a crowd feature encoder to extract features from crowd images in the training set; Step S3, constructing a density branch, performing multi-scale fusion on the features extracted by the crowd feature encoder, and obtaining the total number of people in the scene through density regression of the multi-scale fused features; Step S4, constructing a semantic branch, mining the contextual information of people and the environment in the scene based on the features extracted by the crowd feature encoder, and generating semantic graphs of different categories; Step S5, fusing the density branch and the semantic branch to construct a fine-grained crowd counting model based on dual-branch collaboration, using the training set images to train the fine-grained crowd counting model based on dual-branch collaboration, and obtaining a trained fine-grained crowd counting model based on dual-branch collaboration; Step S6, input the crowd images of the test set into the trained fine-grained crowd counting model based on dual-branch collaboration, output the corresponding crowd density map and semantic map, and complete the fine-grained crowd counting and classification tasks.
2. According to claim 1, a fine-grained crowd counting method based on dual-branch collaboration is characterized in that: The crowd feature encoder adopts the VGG16 neural network model, where: The feature map resizing function is: Where: w i and b i are the weights and biases of the convolutional layer; Feature maps of different dimensions are input; is the adjusted feature map; Resize(·) is the size adjustment function; The feature map fusion function is: Where: is the fused feature map output by the crowd feature encoder; ConvGroup(·) is multiple convolution group operations; Concat(·) is a concatenation operation; i∈M.
3. According to claim 1, a fine-grained crowd counting method based on dual-branch collaboration is characterized in that: The density branch is constructed to perform multi-scale fusion of the extracted features; the specific formula for multi-scale fusion is as follows: Where: It is the feature after multi-scale fusion of density branch; Upsample(·) is the upsampling function; ConvModule1(·) is the density branch convolution module operation; Concat(·) is the feature concatenation operation in the new dimension.
4. According to claim 1, a fine-grained crowd counting method based on dual-branch collaboration is characterized in that: The semantic branch performs feature fusion through a dual feature propagation module, extracts key feature indexes, and enhances key features using a cross-attention mechanism; wherein: The feature fusion formula is as follows: Where: X fuse is the fused feature; is the feature of the input feature propagation module; DownSample(·) is the downsampling operation; The index selection function is as follows: index SFP ,w=Select SFP (X fuse ); Where: index SFP is the feature index point; w is the weight; Select SFP (·) Index selection function; The specific expression of the cross-attention mechanism to enhance the key features is as follows: CrossAttention(Q,K,V)=softmax(Q,K T )V; Where: Q is the significant feature composition of the query vector filtered by the index function; K is the key feature; V is the value feature; K and V are obtained by weighting the feature vectors at the corresponding positions in the significant features through weights.
5. According to claim 4, a fine-grained crowd counting method based on dual-branch collaboration is characterized in that: In the feature processing step of the semantic branch, the feature distinguishing ability of the semantic branch is enhanced by a category loss function, and the category loss function is specifically as follows: L CD =max(Sim(v m ,v n ),0); Where: L CD is the category loss function; v m and v n is the representative feature embedding of this category; Sim(v m ,v n ) is the similarity between two feature vectors.
6. The fine-grained crowd counting method based on dual-branch collaboration according to claim 1 is characterized in that: In the process of using the training set images to train the fine-grained crowd counting model based on dual-branch collaboration, the stochastic gradient descent method is used to update the model parameters, and the overall objective function is: L=L MAE +L MSE +γL CD ; Where: L is the overall objective function; L MAE is the mean absolute error between the predicted density map and the true density map; L MSE The absolute deviation between the predicted value and the true value; γ is the balance factor.
7. A fine-grained crowd counting system based on dual-branch collaboration, characterized in that: include: The data preparation module collects crowd images and performs annotation processing, builds a fine-grained crowd counting dataset, divides the fine-grained crowd counting dataset into a training set and a test set, and performs data preprocessing on the fine-grained crowd counting dataset; Crowd image annotation processing includes: crowd classification information, individual behavior attributes, and crowd location and distribution relationship; Feature extraction module, builds a crowd feature encoder and extracts features from crowd images in the training set; Density analysis module: builds density branches, performs multi-scale fusion on the features extracted by the crowd feature encoder, and uses density regression to obtain the total number of people in the scene. The semantic analysis module builds semantic branches and deeply mines the contextual information of people and the environment in the scene based on the features extracted by the crowd feature encoder to generate semantic graphs of different categories. The model training module integrates the density branch and the semantic branch to build a fine-grained crowd counting model based on dual-branch collaboration. The training set images are used to train the fine-grained crowd counting model based on dual-branch collaboration to obtain a trained fine-grained crowd counting model based on dual-branch collaboration. The model prediction module inputs the crowd images of the test set into the trained dual-branch collaborative In the fine-grained crowd counting model, the corresponding crowd density map and semantic map are output to complete the fine-grained crowd counting and classification tasks.
8. The fine-grained crowd counting system based on dual-branch collaboration according to claim 7 is characterized in that: The feature extraction module and the crowd feature encoder adopt the VGG16 neural network model, where: The feature map resizing function is: Where: w i and b i are the weights and biases of the convolutional layer; Feature maps of different dimensions are input; is the adjusted feature map; Resize(·) is the size adjustment function; The feature map fusion function is: Where: is the fused feature map output by the crowd feature encoder; ConvGroup(·) is multiple convolution group operations; Concat(·) is a concatenation operation; i∈M.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the fine-grained crowd counting method based on dual-branch collaboration as described in any one of claims 1 to 6 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a fine-grained crowd counting method based on dual-branch collaboration as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Crowd positioning method and system based on hybrid expert network, and storage medium
CN118823685A