Video pedestrian flow estimation method based on one-to-many matching strategy
Through a one-to-many matching strategy, a convolutional neural network and a self-attention mechanism are used to calculate the similarity of pedestrian features, which solves the problems of parameter sensitivity and occlusion sensitivity in existing technologies and achieves high-precision video pedestrian flow estimation.
Patent Information
- Application Number
- CN202510725650.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-03
AI Technical Summary
Among existing video pedestrian flow estimation methods, bipartite graph matching relies on empirical thresholds, resulting in high parameter sensitivity, while implicit matching strategies are sensitive to occlusion, motion, and posture changes, resulting in insufficient counting accuracy.
A one-to-many matching strategy is adopted to extract pedestrian features through a convolutional neural network, calculate feature similarity using a decoupled multi-head self-attention mechanism, and predict the matching probability through a discriminator, allowing pedestrians to match other pedestrians with adjacent spatial positions, reducing dependence on parameters.
The accuracy and robustness of pedestrian flow counting are improved, and it can accurately count different crowd densities in different scenarios, output intuitive visual results, and reduce the model's sensitivity to parameters.
Smart Images

Figure CN120635981A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pedestrian flow detection, and in particular relates to a video pedestrian flow estimation method based on a one-to-many matching strategy. Background Art
[0002] With the continuous development of counting technologies such as artificial intelligence and computer vision, in order to reduce the risks brought by crowding, video crowd counting (VCC) has been widely adopted based on video surveillance and computer vision technology to monitor crowd density.
[0003] Existing technologies for counting pedestrian traffic in videos use bipartite graphs or privacy matching, but these methods suffer from the following issues: Existing bipartite graph matching often requires an empirical threshold to filter out false matches, increasing the model's sensitivity to parameters. While implicit matching strategies eliminate the drawbacks of threshold parameters, their core concept, particularly in training, still adheres to one-to-one matching. This highly relies on discriminative individual feature representations, making it highly sensitive to appearance changes caused by occlusion, motion, and posture changes.
[0004] Therefore, a video pedestrian flow estimation method based on one-to-many matching using “group context” is proposed. Summary of the Invention
[0005] The purpose of this invention is to provide a method for estimating pedestrian flow from video based on a one-to-many matching strategy. By making full use of the group characteristics of small groups of people through one-to-many matching, the accuracy of pedestrian flow counting is improved. The technical solution is as follows: A video pedestrian flow estimation method based on a one-to-many matching strategy includes the following steps: S1. Periodically sample the video to obtain video frames , two consecutive frames of images in the video frame Perform pedestrian detection and positioning to obtain the pedestrian's head positioning coordinates; S2, according to the head positioning coordinates, from two consecutive frames of image Cropped pedestrian image patches and , using convolutional neural networks to extract image patches The appearance features of the pedestrians are obtained by using the head positioning coordinates for position encoding to obtain the position features, and the appearance features and position features are combined to obtain the pedestrian features. and ; S3, pedestrian features of two consecutive frames of images and Splicing into a matrix , and then use the decoupled multi-head self-attention mechanism to calculate the matrix The similarity between the elements of the self-attention map A is obtained. The pedestrian feature similarity part is extracted according to the self-attention map A, and the mean of the pedestrian feature similarity part is spliced with the semantic matching feature obtained according to the self-attention map A to obtain the pedestrian feature with matching semantic enhancement. ; S4. Separate pedestrian features , get the pedestrian features of the two frames of images , calculate the pedestrian features of the two frames of images The Hadamard product between them is used to obtain a three-dimensional similarity matrix, which is input into the discriminator to predict the pedestrian matching probability. The number of people flowing in or out is distinguished according to the matching probability, and the number of people flowing in each frame is accumulated to obtain the total flow.
[0006] Furthermore, in step S1, a universal head locator PET is used to detect and locate pedestrians in the video frame image.
[0007] Furthermore, the step S2 specifically includes the following steps: S201, using the pedestrian head positioning coordinates as a reference point, from two consecutive frames of image The pixel size of the pedestrians cut out is The rectangular area is obtained by bilinear interpolation Pedestrian image patches of size and ; S202, using convolutional neural network ConvNeXt-s to extract image blocks Pedestrian appearance features at a medium resolution of 1 / 32; S203: Using sinusoidal position coding to obtain a position code with 256 channels based on the coordinates of the pedestrian's head positioning point, inputting a one-dimensional convolution layer with a convolution kernel size of 1 to obtain a position code with 1152 channels, thereby obtaining a position feature; S204: Combine the pedestrian appearance features and position features to obtain pedestrian features. and .
[0008] Furthermore, step S3 includes: S301, pedestrian features of two consecutive frames of images and Spliced into a four-dimensional pedestrian feature matrix , the four-dimensional pedestrian feature matrix is transformed into The three-dimensional pedestrian feature matrix is obtained by flattening along the spatial dimension, and the number of channels is reduced to 256 through a two-dimensional convolution layer with a convolution kernel of 1, resulting in a 256-dimensional pedestrian feature matrix. As input to the self-attention layer; S302, using the decoupled 8-head self-attention mechanism to transform the 256-dimensional pedestrian feature matrix Perform implicit feature matching; in the self-attention layer, the 256-dimensional pedestrian feature matrix Projection is performed to obtain query Q, key K, and value V. The self-attention graph is obtained by calculating the similarity between query Q and key K using the self-attention mechanism. , and multiplied by the value V to obtain a feature with semantic matching ; S303, from the self-attention map Extract the pedestrian feature similarity part of the two frames , and then the pedestrian feature similarity part Averaging along spatial dimensions , the average value and features Splicing to obtain matching semantically enhanced pedestrian features ,in Where, is the decoupled self-attention layer, For splicing operation, express Similarity to oneself, express and The similarity of express and The similarity of express Similarity to oneself.
[0009] Furthermore, step S4 specifically includes the following steps: S401, pedestrian features Separate the pedestrian features of the two frames , the pedestrian features Transformed into m and Two-dimensional feature descriptor, c is the number of feature channels, m and n are the number of pedestrians in the previous and next frames respectively; , , and each pedestrian feature and They are expanded into one-dimensional feature vectors, and the Hadamard product of each pedestrian feature is performed to obtain m The three-dimensional similarity matrix is used as the input of the discriminator; S402: The similarity matrix is passed through the multi-layer perceptron of the discriminator to obtain a logits vector graph, and the probability value of pedestrian matching is obtained through the softmax function. The number of people flowing in or out is distinguished based on the probability value, and the number of people flowing in is accumulated to obtain the total flow.
[0010] Furthermore, the step S402 specifically includes: Input the similarity matrix into the multi-layer perceptron of the fully connected network of the discriminator layer, and output m The logits vector graph is used to obtain the probability value of each pair of pedestrians matching through the Softmax function. : in, is the sigmoid function, represents a multilayer perceptron, represents the Hadamard product, is the matching probability of the output, for exist The index number in for exist The index number in ; If the matching probability If the probability is greater than 0.5, the two behaviors are considered to be the same pedestrians in the two frames before and after, and are not counted; if the matching probability is less than or equal to 0.5, the matching is considered to have failed, and the pedestrians not matched in the next frame are regarded as new pedestrians and counted once to get the number of pedestrians flowing in. and the number of people who left ; in, is the rounding symbol; The flow counting result is the accumulation of the number of people flowing in. ,in, hour, That is, the total number of people detected in the first frame of the video stream.
[0011] Furthermore, in step S403, the multilayer perceptron is trained using a cross entropy loss function and a one-to-many matching strategy is introduced, in which pedestrians are allowed to match themselves in adjacent frames and other pedestrians with a normalized distance less than 0.2.
[0012] Furthermore, the one-to-many matching strategy is as follows: The discriminator regards several adjacent pedestrians as small groups of people and performs a one-to-many match between the pedestrians in the previous frame and the small groups of people in the next frame. This matching strategy uses spatial position information to allow a pedestrian to be matched with other pedestrians adjacent to its spatial position.
[0013] The beneficial effects of the present invention are: (1) The proposed method adopts a one-to-many matching approach, avoiding the problem of sensitivity to appearance features and detection results caused by the one-to-one matching strategy. Through one-to-many matching, the movement characteristics of pedestrians walking in groups in public places are deeply explored, the movement patterns of small groups of people are fully explored, the counting accuracy of the model is improved, and the model is more robust to different crowd densities in different scenarios, while also outputting intuitive visualization results.
[0014] (2) The present invention uses an implicit matching method to perform feature matching. The overall network architecture is simple and efficient, avoiding the manual setting and adjustment of relevant parameters, eliminating the model's sensitivity to such parameters, and improving the matching accuracy with this learnable matching method, thereby improving the counting accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a flow chart of the video pedestrian flow estimation method provided by the present invention; Figure 2 This is a system framework diagram of the video pedestrian flow estimation method provided by the present invention; Figure 3 It is a comparison chart of related tasks and related work provided by the present invention; Figure 4 This is a comparison chart of one-to-one matching and one-to-many matching provided by the present invention; Figure 5 This is a diagram of a small group of people walking together, as provided by the present invention; Figure 6 A pair of original images of the preceding and following frames of a video containing pedestrians provided by the present invention; Figure 7 It is a one-to-many matching result visualization diagram provided by the present invention. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0017] The application principle of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0018] The technical solution process of the present invention is roughly divided into the following steps: The first part locates and extracts features from the video to be detected, using PET to locate pedestrian heads and extract their characteristic expressions. The second part generates matching context, mining the similarities between the characteristic expressions of pedestrians in the previous and next frames to generate new pedestrian features with implicit matching relationships. The third part implements one-to-many matching, using a three-layer perceptron to determine the similarity vectors of each pair of pedestrian features, achieving a one-to-many matching effect for people and ultimately outputting a traffic count. The traffic counting task targeted by this invention is a completely new task, distinct from multi-target tracking, line crossing counting, and person re-identification.
[0019] The detailed steps of the present invention are described below with reference to the accompanying drawings: like Figure 1 As shown in FIG, a video pedestrian flow estimation method based on a one-to-many matching strategy is described. The specific steps are as follows: S1. Periodically sample the video to obtain video frames , the universal head locator PET will be used to align two consecutive frames of video frames. Perform pedestrian detection and positioning to obtain the coordinates of the pedestrian's head positioning point; S2, according to the coordinates of the head positioning point, from two consecutive frames of image Cropped pedestrian image patches and , using convolutional neural networks to extract image patches The appearance features of the pedestrians are obtained by position encoding using the coordinates of the head positioning points, and the appearance features and position features are combined to obtain the pedestrian features. and ; The specific technical solution is: S201, using the pedestrian head positioning coordinates as a reference point, from two consecutive frames of image The pixel size of the pedestrians cut out is The rectangular area is obtained by bilinear interpolation (an interpolation method that performs linear interpolation in two directions, which is an existing technology) Pedestrian image patches of size and ; S202, using convolutional neural network ConvNeXt-s to extract image blocks Pedestrian appearance features at a medium resolution of 1 / 32; S203. Sine position coding is used to obtain a position code with 256 channels based on the coordinates of the pedestrian's head positioning point. A one-dimensional convolution layer with a convolution kernel size of 1 is input to obtain a position code with 1152 channels to obtain a position feature. The sinusoidal position coding formula is: Among them, pos represents the position, i represents the dimension index, d represents the total dimension of the code, and PE is the position code; S204: Combine the pedestrian appearance features and position features to obtain pedestrian features. and .
[0020] S3, pedestrian features of two consecutive frames of images and Splicing into a matrix , and then use the decoupled multi-head self-attention mechanism to calculate the matrix The similarity between the elements of the self-attention map is calculated, and the matching semantic part of the self-attention map is averaged and concatenated with the pedestrian features after the decoupled self-attention layer to obtain the pedestrian features enhanced by matching semantics. ; The specific technical solution is: S301, pedestrian features of two consecutive frames of images and Splice to The four-dimensional pedestrian feature matrix , the four-dimensional matrix is transformed into Flatten along the spatial dimension to obtain The three-dimensional pedestrian feature matrix is obtained by reducing the number of channels to 256 through a two-dimensional convolution layer with a convolution kernel of 1, and a 256-dimensional pedestrian feature matrix is obtained. As input to the self-attention layer; S302, using the 8-head decoupled self-attention mechanism to transform the 256-dimensional pedestrian feature matrix Perform implicit feature matching; in the self-attention layer, for the 256-dimensional matrix Projection is performed to obtain the query, key, and value for self-attention calculation, namely Q, K, and V. The self-attention mechanism calculates the similarity between Q and K to obtain the self-attention map , and obtain new features with matching semantics by calculating V , thereby enhancing the similar semantics between the pedestrian features of the previous and next frames, including the similar appearance features of the same pedestrian, the similar position features of adjacent pedestrians, and the similar appearance features of the overlapping and occluded parts of adjacent pedestrians, so that the pedestrian features have potential one-to-many semantic information; S303, from the self-attention map Extract the pedestrian feature similarity part of the two frames , and then find the average value along the spatial dimension and concatenate it with the new features obtained by the self-attention layer , get the pedestrian features with matching semantic enhancement , the formula for this step is as follows: in, is the decoupled self-attention layer, For splicing operation, express Similarity to oneself, express and The similarity of express and The similarity of express Similarity to oneself.
[0021] S4. Separate pedestrian features , get the pedestrian features of the two frames of images , calculate the pedestrian features of the two frames of images The Hadamard product between them is used to obtain a three-dimensional similarity matrix, which is then input into the discriminator to predict the pedestrian matching probability. The number of people flowing in or out is distinguished based on the matching results, and the total flow is obtained by accumulating the number of people flowing in each frame. The specific technical solution is as follows:
[0022] S401, pedestrian features Separate the pedestrian features of the two frames , the pedestrian features Transformed into m and Two-dimensional feature descriptor, c is the number of feature channels; , , and each pedestrian feature and They are expanded into one-dimensional feature vectors, and the Hadamard product of each pedestrian feature is performed to obtain m The three-dimensional similarity matrix is used as the input of the discriminator; S402: The discriminator considers several adjacent pedestrians as small groups and performs a one-to-many matching between the pedestrians in the previous frame and the small groups in the next frame. This matching strategy uses spatial position information to allow a pedestrian to be matched with other pedestrians in the adjacent spatial position. S403, the discriminator contains a multi-layer perceptron with 3 layers of fully connected networks, and outputs a matching result with a channel of 2, i.e., an m The logits vector graph is used to obtain the probability value of each pair of pedestrians matching through the Softmax function. , represents the probability of pedestrian matching. The number of people flowing in or out is distinguished according to the matching results, and the total flow is obtained by accumulating the number of people flowing in each frame.
[0023] in, is the sigmoid function, represents a multilayer perceptron, represents the Hadamard product, is the matching probability of the output, for exist The index number in for exist The index number in ; S404, if the matching probability If the probability is greater than 0.5, the two behaviors are considered to be the same pedestrians in the two frames before and after, and are not counted; if the matching probability is less than or equal to 0.5, the matching is considered to have failed, and the pedestrians not matched in the next frame are regarded as new pedestrians and counted once to get the number of pedestrians flowing in. and the number of people who left .
[0024] in, The rounding symbol.
[0025] S405. Finally, the flow counting result is the accumulation of the number of people entering ,in, hour, That is, the total number of people detected in the first frame of the video stream.
[0026] In some embodiments, the training of the multi-layer perceptron in steps S402 and S403 uses a cross-entropy loss function and introduces a one-to-many positive sample matching strategy, in which pedestrians are allowed to match themselves in adjacent frames and other pedestrians with a normalized distance less than 0.2, so as to implement a one-to-many matching strategy for matching pedestrians with the small group of people they belong to. Compared with the conventional one-to-one matching constraint, the one-to-many matching strategy is more relaxed and robust. The specific difference is shown in the following formula: in, represents the matching matrix, and represent m-dimensional and n-dimensional column vectors respectively, The upper limit for the number of people in a small group of people.
[0027] In some embodiments, in steps S3 and S4, the Kullback-Leibler divergence between the discrete similarity distribution of the actual pedestrian feature matching pair and the discrete similarity distribution of the predicted pedestrian feature matching pair is calculated as one of the loss functions for model training. The mathematical formula is as follows: in, is a smooth term, is the discrete distribution of similarity of predicted pedestrian matching pairs, is the discrete distribution of similarity between actual pedestrian matching pairs.
[0028] The embodiment of the present invention provides a method for estimating pedestrian flow through video based on a one-to-many matching strategy. By processing and analyzing the acquired video samples, the number of pedestrians flowing into and out of a certain place within a period of time, i.e., the flow rate, is inferred. The functions and specific implementation processes of each module of the system are described in detail below. Figure 2 and Figure 3 shown.
[0029] (1) Feature extraction module, such as Figure 2 As shown in (a), the video sequence is first sampled at equal intervals with a sampling interval of 3 seconds. The universal head locator PET in Chengxin Liu, Hao Lu, Zhiguo Cao, and Tongliang Liu. Point-queryquadtree for crowd counting, localization, and more. IEEE International Conference on Computer Visio, 2023, pp. 1676–1685. is used to detect and locate the head in each video frame. The sampled video frame sequence is then paired with its adjacent frames, i.e., each pair of the two frames before and after is composed of the following two frames: Figure 6 Image sequence. According to the coordinates of the head positioning point Cut out the pedestrians Rectangular image blocks, which are and is a rectangular area with the coordinates of the upper left corner and the lower right corner, and then bilinear interpolation is used to interpolate the image block into The cropped pedestrian image is fed into the ConvNeXt-s backbone network for feature extraction, generating pedestrian features at a resolution of 1 / 32 the input resolution. To enable the model to perceive the spatial position of pedestrians, each pedestrian's location coordinates are positionally encoded and mapped through a one-dimensional convolutional layer. This is then concatenated with the pedestrian's appearance features to generate the position-encoded pedestrian features.
[0030] (2) Implicit matching context generation module, such as Figure 2As shown in (b), the pedestrian features after position encoding of the two frames before and after are input, and the pedestrian feature vectors of the two frames are spliced into a matrix. The matrix is then projected and transformed into a matrix that meets the input dimension shape requirements of the multi-head decoupled self-attention layer. The matrix is input into a 6-layer 8-head decoupled self-attention layer. Using the decoupled self-attention mechanism, the similarity between features can be implicitly modeled. As shown in the following formula, the decoupled self-attention response map itself describes the similarity relationship between tokens. In order to further strengthen the similarity matching relationship, the decoupled self-attention map is taken. A subgraph describing the similarity between the pedestrian features of the previous frame and the pedestrian features of the next frame , averaged along its spatial dimension and compressed into one dimension, and spliced together with the new pedestrian features output by the multi-head decoupled self-attention layer, and finally the pedestrian feature expression after matching semantic enhancement is obtained.
[0031] (3) One-to-many pairwise matching module, one-to-many matching strategy such as Figure 4 As shown, the module Figure 2 As shown in (c), the pedestrian feature matrix obtained by (2) is first decomposed into the pedestrian features of the previous frame and the pedestrian features of the next frame. The Hadamard product is calculated for the pedestrian features between the two frames to obtain the similarity matrix between the two frames. The matrix is input into a 3-layer perceptron for judgment. The output result is passed through the Softmax function to obtain the probability of the two matching. If the probability is greater than 0.5, it is considered a match, that is, the two pedestrians are common pedestrians in the previous and next frames and are not counted; otherwise, they are not matched. If a pedestrian in the next frame cannot find a matching object among all the pedestrians in the previous frame, then the pedestrian is a new pedestrian flowing into the next frame and is counted once. Finally, the predicted flow is the sum of the inflow of all detected pedestrians in the first frame of the video sequence and all counts in the subsequent frames.
[0032] (4) One-to-many pairwise matching module, observe that Figure 5 As shown in the figure, pedestrians tend to walk in groups. In order to mine group information, enhance the model's robustness to changes in pedestrian appearance such as occlusion, and obtain more accurate and reliable prediction results, a one-to-many matching strategy is adopted, using the cross-entropy loss function and introducing a one-to-many positive sample matching strategy, which allows pedestrians to match themselves in adjacent frames and other pedestrians with a normalized distance less than 0.2, so as to achieve a one-to-many matching strategy for small groups of pedestrians. When the input image is as follows Figure 6 When shown, the one-to-many matching results are visualized as Figure 7 As shown, the green points represent pedestrians shared by the two frames, the red points represent pedestrians flowing in and out, the blue lines represent pedestrians that match each other, and the red circles represent a small group of people.
[0033] (5) Model training: For the implicit matching context generation module described in (2), the optimal transfer loss function as described in Xinyan Liu, Guorong Li, Yuankai Qi, Ziheng Yan, et al. Weakly supervised video individual counting. IEEE Conference on Computer Vision and Pattern Recognition, 2024, pp. 19228–19237. is used for weak supervision training, so that the model can distinguish the similarity of features; for the one-to-many pairwise matching modules described in (3) and (4), the cross entropy loss function is used, and the training strategy is as described in (4). In addition, in order to better supervise the entire network from a global perspective to improve stability, the Kullback-Leibler divergence between the discrete similarity distribution of the actual pedestrian feature matching pairs and the discrete similarity distribution of the predicted pedestrian feature matching pairs is calculated as one of the loss functions for model training. Its mathematical formula is as follows:
[0034] in, is a smooth term, is the discrete distribution of similarity of predicted pedestrian matching pairs, is the discrete distribution of similarity between actual pedestrian matching pairs.
[0035] The various steps and specific implementation process of the video pedestrian flow estimation method based on one-to-many matching provided by the present invention correspond to the functions of the various modules of the above-mentioned system.
[0036] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A video pedestrian flow estimation method based on a one-to-many matching strategy, characterized in that: The following steps are involved: S1. Periodically sample the video to obtain video frames , two consecutive frames of images in the video frame Perform pedestrian detection and positioning to obtain the pedestrian's head positioning coordinates; S2, according to the head positioning coordinates, from two consecutive frames of image Cropped pedestrian image patches and , using convolutional neural networks to extract image patches The appearance features of the pedestrians are obtained by using the head positioning coordinates for position encoding to obtain the position features, and the appearance features and position features are combined to obtain the pedestrian features. and ; S3, pedestrian features of two consecutive frames of images and Splicing into a matrix , and then use the decoupled multi-head self-attention mechanism to calculate the matrix The similarity between the elements of the self-attention map A is obtained. The pedestrian feature similarity part is extracted according to the self-attention map A, and the mean of the pedestrian feature similarity part is spliced with the semantic matching feature obtained according to the self-attention map A to obtain the pedestrian feature with matching semantic enhancement. ; S4. Separate pedestrian features , get the pedestrian features of the two frames of images , calculate the pedestrian features of the two frames of images The Hadamard product between them is used to obtain a three-dimensional similarity matrix, which is input into the discriminator to predict the pedestrian matching probability. The number of people flowing in or out is distinguished according to the matching probability, and the number of people flowing in each frame is accumulated to obtain the total flow.
2. The video pedestrian flow estimation method based on one-to-many matching strategy according to claim 1 is characterized in that: In the step S1, a universal head locator PET is used to detect and locate pedestrians in the video frame image.
3. The video pedestrian flow estimation method based on one-to-many matching strategy according to claim 1 is characterized in that: The step S2 specifically includes the following steps: S201, using the pedestrian head positioning coordinates as a reference point, from two consecutive frames of image The pixel size of the pedestrians cut out is The rectangular area is obtained by bilinear interpolation Pedestrian image patches of size and ; S202, using convolutional neural network ConvNeXt-s to extract image blocks Pedestrian appearance features at a medium resolution of 1 / 32; S203: Using sinusoidal position coding to obtain a position code with 256 channels based on the coordinates of the pedestrian's head positioning point, inputting a one-dimensional convolution layer with a convolution kernel size of 1 to obtain a position code with 1152 channels, thereby obtaining a position feature; S204: Combine the appearance features and position features of the pedestrian to obtain pedestrian features. and .
4. The video pedestrian flow estimation method based on one-to-many matching strategy according to claim 1 is characterized in that: The step S3 comprises: S301, pedestrian features of two consecutive frames of images and Spliced into a four-dimensional pedestrian feature matrix , the four-dimensional pedestrian feature matrix is transformed into The three-dimensional pedestrian feature matrix is obtained by flattening along the spatial dimension, and the number of channels is reduced to 256 through a two-dimensional convolution layer with a convolution kernel of 1, resulting in a 256-dimensional pedestrian feature matrix. As input to the self-attention layer; S302, using the decoupled 8-head self-attention mechanism to transform the 256-dimensional pedestrian feature matrix Perform implicit feature matching; in the self-attention layer, the 256-dimensional pedestrian feature matrix Projection is performed to obtain query Q, key K, value V. The self-attention graph is obtained by calculating the similarity between query Q and key K using the self-attention mechanism. , and multiplied by the value V to obtain a feature with semantic matching ; S303, from the self-attention map Extract the pedestrian feature similarity part of the two frames , and then the pedestrian feature similarity part Averaging along spatial dimensions , the average value and features Splicing to obtain matching semantically enhanced pedestrian features ,in Where, is the decoupled self-attention layer, For splicing operations, express Similarity to oneself, express and The similarity of express and The similarity of express Similarity to oneself.
5. The video pedestrian flow estimation method based on one-to-many matching strategy according to claim 1 is characterized in that: Step S4 specifically includes the following steps: S401, pedestrian features Separate the pedestrian features of the two frames , the pedestrian features Transformed into m and Two-dimensional feature descriptor, c is the number of feature channels, m and n are the number of pedestrians in the previous and next frames respectively; , , and each pedestrian feature and They are expanded into one-dimensional feature vectors, and the Hadamard product of each pedestrian feature is performed to obtain m The three-dimensional similarity matrix is used as the input of the discriminator; S402: The similarity matrix is passed through the multi-layer perceptron of the discriminator to obtain a logits vector graph, and the probability value of pedestrian matching is obtained through the softmax function. The number of people flowing in or out is distinguished based on the probability value, and the number of people flowing in is accumulated to obtain the total flow.
6. The video pedestrian flow estimation method based on a one-to-many matching strategy according to claim 5 is characterized in that: The step S402 specifically includes: Input the similarity matrix into the multi-layer perceptron of the fully connected network of the discriminator layer, and output m The logits vector graph is used to obtain the probability value of each pair of pedestrians matching through the Softmax function. : in, is the sigmoid function, represents a multilayer perceptron, represents the Hadamard product, is the matching probability of the output, for exist The index number in for exist The index number in ; If the matching probability If the probability is greater than 0.5, the two behaviors are considered to be the same pedestrians in the two frames before and after, and are not counted; if the matching probability is less than or equal to 0.5, the matching is considered to have failed, and the pedestrians not matched in the next frame are regarded as new pedestrians and counted once to get the number of pedestrians flowing in. and the number of people who left ; in, is the rounding symbol; The flow counting result is the accumulation of the number of people flowing in. ,in, hour, That is, the total number of people detected in the first frame of the video stream.
7. The video pedestrian flow estimation method based on one-to-many matching strategy according to claim 5 is characterized in that: In step S403, the multi-layer perceptron is trained using a cross entropy loss function and a one-to-many matching strategy is introduced, in which pedestrians are allowed to match themselves in adjacent frames and other pedestrians with a normalized distance less than 0.
2.
8. The video pedestrian flow estimation method based on one-to-many matching strategy according to claim 7 is characterized in that: The one-to-many matching strategy is as follows: The discriminator regards several adjacent pedestrians as small groups of people and performs a one-to-many match between the pedestrians in the previous frame and the small groups of people in the next frame. This matching strategy uses spatial position information to allow a pedestrian to be matched with other pedestrians adjacent to its spatial position.
Citation Information
Patent Citations
Video crowd counting system and method
CN111860162A
Pedestrian detection method fused with self-supervised semantic learning
CN118692056A
Crowd behavior detection method and apparatus, and electronic device, storage medium and computer program product
WO2022160591A1