A Pedestrian Search Method for Security Monitoring Videos
By generating the temporal and spatial characteristics of pedestrians in security monitoring videos, and organizing the pedestrian spatial characteristics using regional convolutional neural networks and gated loop units, and indexing them in combination with local sensitive hashs, the problem of failure to fully utilize the pedestrian time correlation in the existing technology is solved, and higher pedestrian search accuracy and real-time performance are achieved.
Patent Information
- Application Number
- CN202210682446.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-16
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-06-16
AI Technical Summary
When handling security surveillance videos, existing pedestrian search methods fail to make full use of pedestrian time correlation, resulting in insufficient accuracy and robustness when finding specific pedestrian targets in massive video data.
By generating the spatio-temporal features of pedestrians in the monitoring video, the pre-trained regional convolutional neural network and gated loop unit are used to organize the spatial features of pedestrians, and index the spatio-temporal features with local sensitive hash to improve the accuracy and real-timeness of pedestrian searches.
It enhances the recognition of the same pedestrian and the distinction between different pedestrians, improves the accuracy of pedestrian searches, and ensures the real-time search. It is suitable for the application of actual security surveillance videos.
Smart Images

Figure CN115082854B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and specifically to a pedestrian search method for security monitoring videos. Background Art
[0002] With the continuous advancement of the construction process of smart cities, more and more surveillance cameras are spread throughout the streets and alleys of cities, playing an important role in aspects such as the loss of the elderly and children and the search and positioning of criminal suspects. However, having surveillance videos does not mean that relevant information can be found. Due to the large number of surveillance cameras and long recording times, the amount of video data shows a geometric growth trend. Searching for specific pedestrian targets in a large amount of surveillance videos often requires a large amount of time, manpower, and material resources. At the same time, due to the different installation positions of surveillance cameras, the surveillance scenarios are also different. Especially in large public places such as shopping malls, stations, and convention centers, the flow of people is relatively dense, making the surveillance scenarios more complex and posing greater challenges to pedestrian search. Therefore, how to find relevant pedestrians more quickly and accurately in a large amount of security surveillance videos has become one of the hot topics in the field of computer vision technology.
[0003] The pedestrian search method queries for a specific pedestrian target in an unknown image or video dataset to find the same pedestrian image in the dataset. Currently, most pedestrian search methods usually train on the CHUK-SYSU or PRW image datasets and use the trained neural network model for pedestrian search. Compared with the CHUK-SYSU and PRW image datasets, real-world security surveillance videos not only contain the spatial features of pedestrians but also involve the temporal correlation of pedestrians. Therefore, the pedestrian search methods trained on the CHUK-SYSU and PRW image datasets need to further improve their robustness and reliability when applied to security surveillance videos because they do not consider the temporal correlation of pedestrians.
[0004] Through a literature search of the prior art, it is found that Patent CN 112241682A provides an end-to-end pedestrian search method based on block division and multi-layer information fusion. This method uses a convolutional neural network to extract preliminary features and uses a candidate region extraction network to extract the region where the pedestrian is located, thereby obtaining a high-level feature map. By dividing the high-level features into blocks and fusing them with the middle-level features, the accuracy of pedestrian search is improved. Although this method uses the entire image captured by the surveillance camera as input data, it does not consider the temporal correlation of pedestrians in the front and back frame images and has certain limitations when applied to security surveillance videos.
[0005] Further retrieval found that patent CN 109165540A provides a pedestrian search method based on a prior candidate box selection strategy. This method constructs a pedestrian candidate box vector based on the length and width of all pedestrian bounding boxes in the training set, then obtains prior candidate boxes through the K-means++ clustering algorithm and identifies the identities of pedestrians, and finally determines the positions of pedestrians in the surveillance image through a trained pedestrian search network. This method also only processes the pedestrian features in a single image and does not involve the temporal correlation between pedestrians in multiple frames of surveillance videos. Summary of the Invention
[0006] In view of the above deficiencies in the prior art, the present invention provides a pedestrian search method for security surveillance videos. According to the characteristics of security surveillance videos having spatial invariance and temporal continuity, by generating spatio-temporal features of pedestrians in the surveillance video and making full use of the temporal correlation of pedestrians in the front and back frames of the video, the accuracy of pedestrian search is improved. By organizing all the spatio-temporal features of pedestrians in the surveillance video, the real-time performance of pedestrian search is ensured.
[0007] The present invention is implemented through the following technical solutions, specifically:
[0008] A pedestrian search method for security surveillance videos includes the following steps:
[0009] Step 1, use a pre-trained region convolutional neural network to detect each pedestrian frame by frame in the surveillance video and generate corresponding spatial features.
[0010] Step 2, use the hidden state output by the gated recurrent unit to organize the pedestrian spatial features extracted frame by frame in the surveillance video. An average pooling layer is added at the output end of the gated recurrent unit to reduce the dimension of the hidden state vector and generate corresponding pedestrian spatio-temporal features.
[0011] Step 3, index all the pedestrian spatio-temporal features through locality-sensitive hashing, and determine the final pedestrian search result by calculating the similarity between the spatio-temporal features of the pedestrian to be searched and the spatio-temporal features of pedestrians in the surveillance video.
[0012] Further, Step 1 is executed according to the following steps:
[0013] 1) For the surveillance video V = {v 1 , v 2 , …, v N}, which contains N frames of images, where the i-th frame of image is denoted as v i ;
[0014] 2) Process the surveillance video V frame by frame through a pre-trained region convolutional neural network, and the j-th pedestrian spatial feature s extracted from the i-th frame of image v i ini,j ;
[0015] 3) After processing the N frames of images in the surveillance video V, all pedestrian spatial features are represented as S = {s i,j}, 1 ≤ i ≤ n j , 1 ≤ j ≤ M, where n j represents the number of frame images containing the j-th pedestrian, and M represents the total number of pedestrians appearing in the surveillance video;
[0016] Furthermore, step two is executed according to the following steps:
[0017] 1) Take the spatial feature s i of the j-th pedestrian extracted from the i-th frame image v i,j as the input vector and input it into the gated recurrent unit;
[0018] 2) In the gated recurrent unit, the pedestrian spatial feature s i,j updates the candidate hidden state vector c i,j through the tanh activation function and is expressed as: c i,j = tanh(W n s i,j + U n (r i,j ⊙ h i-1,j ) + b n ), where h i-1,j represents the hidden state vector corresponding to the j-th pedestrian in the (i - 1)-th frame image, r i,j is the weight corresponding to h i-1,j , and W n , U n and b n are the network parameters of the gated recurrent unit;
[0019] 3) In the gated recurrent unit, generate the hidden state vector h i-1,j corresponding to the j-th pedestrian in the i-th frame image according to h i,j and c i,j and express it as: h i,j = z i,j h i-1,j + (1 - z i,j )c i,j , where z i,j is the weight for combining h i-1,j and c i,j ;
[0020] 4) After all the spatial features corresponding to the j-th pedestrian are processed in the gated recurrent unit, obtain the hidden vector sequence h j = {h i,j}, 1 ≤ i ≤ n j, and then all pedestrians in the video are represented as a sequence of hidden vectors H = {h j}, 1 ≤ j ≤ M;
[0021] 5) Add an average pooling layer at the output end of the gated recurrent unit to reduce the dimension of the sequence h j to generate the spatio-temporal feature p j of the j-th pedestrian, and it is represented as All spatio-temporal features of pedestrians are represented as P = {p j}, 1 ≤ j ≤ M.
[0022] Furthermore, step three is executed according to the following steps:
[0023] 1) Map all spatio-temporal features P of pedestrians in the surveillance video to the Hamming vector space. For the spatio-temporal feature p j of the j-th pedestrian, it is mapped to a b-bit hash code;
[0024] 2) After mapping the spatio-temporal feature q of the pedestrian to be searched to the Hamming vector space, calculate the similarity between the spatio-temporal feature of the pedestrian to be searched and the spatio-temporal features of pedestrians in the video;
[0025] 3) After calculating the similarities between the spatio-temporal feature q of the pedestrian to be searched and all spatio-temporal features P of pedestrians, sort the similarities. The pedestrian images corresponding to the top T spatio-temporal features are the final search results.
[0026] The beneficial effects of the present invention are as follows: The present invention uses a region convolutional neural network to extract the spatial features of pedestrians frame by frame from the security surveillance video, avoiding the interference of complex surveillance backgrounds. The gated recurrent unit organizes the temporal correlation of the spatial features of pedestrians in the front and back frames, and generates the spatio-temporal features of pedestrians by adding an average pooling layer at the output end, enhancing the recognition of the same pedestrian and the distinction between different pedestrians. The locality-sensitive hashing organizes a large number of spatio-temporal features of pedestrians, reducing the computational complexity of the search process and ensuring the real-time performance of pedestrian search. Compared with the prior art, the present invention can be applied to actual security surveillance videos, improving the search accuracy while ensuring the real-time performance of the search. Brief Description of the Drawings
[0027] Figure 1 is the flow chart of the present invention.
[0028] Figure 2 is the comparison of the accuracy between the method of the present invention and the search method that only uses the spatial features of pedestrians. Detailed Embodiments
[0029] The present invention will be further described in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by this application.
[0030] Embodiment 1
[0031] The present invention is realized through the following technical solutions, and the specific steps of the present invention are as follows:
[0032] Step 1, the security monitoring video is composed of continuous frame images. The pre-trained region convolutional neural network is used to detect each pedestrian frame by frame in the monitoring video and generate corresponding spatial features. The specific method is as follows:
[0033] 1) For the monitoring video V = {v 1 , v 2 , …, v N}, which contains N frame images, where the i-th frame image is denoted as v i . For the pre-trained region convolutional neural network, it can be regarded as a non-linear equation f RCNN (·);
[0034] 2) The monitoring video V is processed frame by frame through the region convolutional neural network f RCNN (·). The j-th pedestrian spatial feature s i extracted from the i-th frame image v i,j can be expressed as s i,j = f RCNN (v i );
[0035] 3) After processing all N frame images in the monitoring video V, all pedestrian spatial features can be expressed as S = {s i,j}, 1 ≤ i ≤ n j , 1 ≤ j ≤ M, where n j represents the number of frame images containing the j-th pedestrian, and M represents the total number of pedestrians appearing in the video.
[0036] Step 2, the gated recurrent unit is used to organize the pedestrian spatial features extracted from the front and back frames of the video, and an average pooling layer is added at the output end to generate the spatio-temporal features of the pedestrians;
[0037] The organization of the pedestrian spatial features extracted from the front and rear frames of the video by the gated recurrent unit and the addition of an average pooling layer at the output end to generate the spatio-temporal features of the pedestrian means that: Since the gated recurrent unit can effectively process data sequences, the hidden states output by the gated recurrent unit can be used to organize the pedestrian spatial features extracted from the video, thereby transmitting the time correlation information of the pedestrian in the front and rear frames of the video. Since the generated hidden state vector has a high dimension, the computational complexity will increase during the pedestrian search process. Therefore, an average pooling layer is added at the output end of the gated recurrent unit to reduce the dimension of the hidden state vector and generate the corresponding spatio-temporal features of the pedestrian at the same time. This can not only reduce the computational complexity in the subsequent search process but also contain the time correlation information of the pedestrian.
[0038] The steps of organizing the pedestrian spatial features extracted from different frames of the video by the gated recurrent unit and generating the spatio-temporal features of the pedestrian through the average pooling layer include:
[0039] 1) Taking the j-th pedestrian spatial feature s i extracted from the i-th frame image v i,j as the input vector and inputting it into the gated recurrent unit;
[0040] 2) In the gated recurrent unit, the pedestrian spatial feature s i,j updates the candidate hidden state vector c i,j through the tanh activation function and is expressed as: c i,j =tanh(W n s i,j +U n (r i,j ⊙h i-1,j )+b n ), where h i-1,j represents the hidden state vector corresponding to the j-th pedestrian in the (i - 1)-th frame image, r i,j is the weight corresponding to h i-1,j , and W n , U n and b n are the network parameters of the gated recurrent unit;
[0041] 3) In the gated recurrent unit, the hidden state vector h i-1,j corresponding to the j-th pedestrian in the i-th frame image is generated according to h i,j and c i,j and is expressed as: h i,j =z i,j h i-1,j +(1 - z i,j )c i,j , where z i,j is the weight for combining h i-1,j and c i,j . hi,j It not only considers the candidate hidden state of the j-th pedestrian in the i-th frame image, but also includes the hidden state of the j-th pedestrian in the (i - 1)-th frame image, thereby establishing a temporal correlation for the j-th pedestrian between the (i - 1)-th frame image and the i-th frame image;
[0042] 4) After all the spatial features corresponding to the j-th pedestrian are processed in the gated recurrent unit, the hidden vector sequence h corresponding to the j-th pedestrian can be obtained j ={h i,j}, 1 ≤ i ≤ n j , thereby describing the temporal correlation of the pedestrian in different frames of the video. Furthermore, all pedestrians in the video can be represented as the hidden vector sequence H = {h j}, 1 ≤ j ≤ M;
[0043] 5) Since the dimension of the sequence h j is relatively high, an average pooling layer is added at the output end of the gated recurrent unit to reduce the dimension of the sequence h j , and at the same time generate the spatio-temporal feature p j of the j-th pedestrian, and is represented as For the j-th pedestrian, the average pooling layer converts the high-dimensional sequence h j into a single vector p j , and at the same time all pedestrian spatio-temporal features can be represented as P = {p j}, 1 ≤ j ≤ M.
[0044] Step three, index all pedestrian spatio-temporal features through locality-sensitive hashing, and determine the pedestrian search result according to the similarity.
[0045] The indexing of all pedestrian spatio-temporal features through locality-sensitive hashing and determining the pedestrian search result according to the similarity means that: since there are a large number of pedestrians in the surveillance video, and each pedestrian corresponds to a pedestrian spatio-temporal feature, a large number of pedestrian spatio-temporal features will be generated. Directly calculating the similarity of pedestrian spatio-temporal features is difficult to ensure the real-time performance of pedestrian search. Therefore, locality-sensitive hashing is used to index all pedestrian spatio-temporal features, and the final pedestrian search result is given by calculating the similarity between the spatio-temporal feature of the pedestrian to be searched and the spatio-temporal features of pedestrians in the surveillance video.
[0046] The steps of indexing all pedestrian spatio-temporal features through locality-sensitive hashing and determining the pedestrian search result according to the similarity include:
[0047] 1) Map all pedestrian spatio-temporal features P in the surveillance video to the Hamming vector space. For the spatio-temporal feature p j of the j-th pedestrian, it can be mapped to a b-bit hash code, and is represented as H: p j → {0, 1}b ;
[0048] 2) After mapping the pedestrian spatio-temporal feature q to be searched into the Hamming vector space, the similarity between the pedestrian spatio-temporal feature to be searched and the pedestrian spatio-temporal features in the video can be calculated as sim(q, p j ) = P[H(q) = H(p j )] = Jaccard(q, p j );
[0049] 3) After calculating the similarities between the pedestrian spatio-temporal feature q to be searched and all pedestrian spatio-temporal features P, sort the similarities. The pedestrian images corresponding to the top T spatio-temporal features are the final search results.
[0050] Embodiment 2
[0051] This embodiment adopts a pedestrian search method for security surveillance videos. The specific implementation steps are as follows:
[0052] 1. Use a region convolutional neural network to extract the spatial features of pedestrians frame by frame from the surveillance video.
[0053] In the region convolutional neural network, use the conv1 layer to conv4_3 layer in the ResNet-50 model to detect the pedestrian bounding box, and use the conv4_4 layer to conv5_3 layer for pedestrian recognition. After global average pooling and feature mapping, for the j-th pedestrian in the i-th frame image v i , a 256-dimensional spatial feature s i,j can be generated. Furthermore, for all pedestrians in the surveillance video, a set of 256×n j ×M-dimensional spatial feature sequence S can be generated.
[0054] 2. Organize the pedestrian spatial features extracted from the front and back frames of the video through a gated recurrent unit, and add an average pooling layer at the output to generate the pedestrian spatio-temporal features.
[0055] Since the number of frames in the security surveillance video is large, to ensure the efficiency of pedestrian search, only one gated recurrent unit is used here to organize the pedestrian spatial features extracted from the front and back two frame images. Take the spatial feature s i of the j-th pedestrian extracted from the i-th frame image v as the input vector. After being processed by the gated recurrent unit, a 256-dimensional hidden state vector h i,j can be generated. Furthermore, the entire spatial features containing the j-th pedestrian can be represented as a set of 256×n i,j -dimensional hidden vector sequence h j . For all pedestrians in the video, it can be represented as 256×n j . jA sequence of hidden vectors H of dimension ×M.
[0056] Due to the relatively high dimensions of the sequence of hidden vectors h j and H, an average pooling layer is used to reduce the dimension of the sequence h j so as to convert the sequence of hidden vectors h of dimension 256×n j into a pedestrian spatio-temporal feature p of dimension 256. j Furthermore, the M pedestrians appearing in the video can be represented by M independent pedestrian spatio-temporal features P={p j}(1≤j≤M). j
[0057] 3. Index all the pedestrian spatio-temporal features through locality-sensitive hashing, and determine the pedestrian search results according to the similarity.
[0058] To ensure the real-time performance of pedestrian search, locality-sensitive hashing is used to index the M pedestrian spatio-temporal features, and the M pedestrian spatio-temporal features are mapped to the Hamming vector space. For the spatio-temporal feature p j of the j-th pedestrian, it can be mapped to a 128-bit hash code and is denoted as Furthermore, in the Hamming vector space, the similarity between the spatio-temporal feature of the pedestrian to be searched and the spatio-temporal features of the pedestrians in the video is calculated according to the Jaccard coefficient. After sorting all the similarities, the pedestrian images corresponding to the top 5 spatio-temporal features are selected as the final search results.
[0059] The simulation experiment of the method of the present invention is as follows:
[0060] In this experiment, videos taken by 9 surveillance cameras were selected, a total of 15,000 video clips were selected, which included 897 pedestrian targets, thus creating a pedestrian search database. In this database, 11,546 video clips were selected as the training data set, and the remaining 3,454 video clips were used as the test data set. At the same time, the mean average precision (MAP) was selected to test the performance of the pedestrian search method. When changing the number of bits of the hash code, this method was compared with the search method that only uses the pedestrian spatial features. The experimental results are as Figure 2 shown. It can be seen from this that as the hash code increases from 8 bits to 128 bits, the mean average precisions of both this method and the search method that only uses the pedestrian spatial features will increase, but the accuracy of this method is higher than that of the method that only uses the pedestrian spatial features. This is because the pedestrian spatio-temporal features used in this method not only include the pedestrian spatial features of a single frame image, but also organize the temporal correlation of the pedestrian spatial features of the front and back frames, thereby enhancing the discrimination between different pedestrians and being conducive to improving the search accuracy.
Claims
1. A pedestrian search method for security surveillance videos, comprising the following steps: Step 1: Use a pre-trained regional convolutional neural network to detect each pedestrian frame by frame in the surveillance video and generate corresponding spatial features; Step 2: Use the hidden state output by the gated recurrent unit to organize the pedestrian spatial features extracted frame by frame in the surveillance video. An average pooling layer is added at the output end of the gated recurrent unit to reduce the dimension of the hidden state vector and generate corresponding pedestrian spatio-temporal features; Step 3: Index all pedestrian spatio-temporal features through locality-sensitive hashing, and determine the final pedestrian search result by calculating the similarity between the spatio-temporal features of the pedestrian to be searched and the spatio-temporal features of the pedestrians in the surveillance video.
2. The pedestrian search method according to claim 1, wherein, Step 1 is executed according to the following steps: 1) For the surveillance video V = {v 1 , v 2 , …, v N}, which contains N frames of images, where the i-th frame of image is denoted as v i ; 2) Process the surveillance video V frame by frame through a pre-trained region convolutional neural network, and extract the j-th pedestrian spatial feature s in the i-th frame image v i ; i,j ; 3) After processing N frames of images in the surveillance video V, all pedestrian spatial features are represented as S = {s i,j}, where 1 ≤ i ≤ n j , 1 ≤ j ≤ M, and n j represents the number of frame images containing the j-th pedestrian, and M represents the total number of pedestrians appearing in the surveillance video.
3. The pedestrian search method according to claim 1, wherein, Step 2 is executed according to the following steps: 1) Input the j-th pedestrian spatial feature s i extracted from the i-th frame image v i,j as the input vector into the gated recurrent unit; 2) In the gated recurrent unit, the pedestrian spatial feature s i,j updates the candidate hidden state vector c through the tanh activation function i,j , and is expressed as: c i,j = tanh(W n s i,j + U n (r i,j ⊙ h i-1,j ) + b n ), where h i-1,j represents the hidden state vector corresponding to the j-th pedestrian in the (i - 1)-th frame image, r i,j is the weight corresponding to h i-1,j , and W n , U n and b n are the network parameters of the gated recurrent unit; 3) In the gated recurrent unit, according to h i-1,j and c i,j generate the hidden state vector h corresponding to the j-th pedestrian in the i-th frame of the image i,j , and is expressed as: h i,j = z i,j h i-1,j + (1 - z i,j )c i,j , where z i,j is the weight for combining h i-1,j and c i,j ; 4) After all the spatial features corresponding to the j-th pedestrian are processed in the gated recurrent unit, the hidden vector sequence h corresponding to the j-th pedestrian is obtained. j ={h i,j}, 1 ≤ i ≤ n j , and then all the pedestrians in the video are represented as the hidden vector sequence H = {h j}, 1 ≤ j ≤ M; 5) Add an average pooling layer at the output of the gated recurrent unit to reduce the dimension of the sequence h j and generate the spatio-temporal feature p j of the j-th pedestrian, which is denoted as All spatio-temporal features of pedestrians are represented as P = {p j}, where 1 ≤ j ≤ M.
4. The pedestrian search method according to claim 1, wherein, Step 3 is executed according to the following steps: 1) Map all pedestrian spatio-temporal features P in the surveillance video to the Hamming vector space. For the spatio-temporal feature p of the j-th pedestrian j , map it to a b-bit hash code; 2) After mapping the spatio-temporal features q of the pedestrian to be searched into the Hamming vector space, calculate the similarity between the spatio-temporal features of the pedestrian to be searched and the spatio-temporal features of the pedestrians in the video; 3) After calculating the similarities between the spatio-temporal features q of the pedestrian to be searched and all the spatio-temporal features P of the pedestrians, sort the similarities. The pedestrian images corresponding to the top T spatio-temporal features with higher similarity are the final search results.
Citation Information
Patent Citations
A pedestrian search method and device based on a priori candidate box selection strategy
CN109165540A
End-to-end pedestrian searching method based on partitioning and multi-layer information fusion
CN112241682A
Pedestrian hash retrieval based on loss measurement in depth learning networks
CN109241317A
Vision detection system and vision detection method using same
WO2020032506A1