An apparatus and method for reconstructing regional pedestrian trajectories based on deep learning
Through the regional pedestrian trajectory reconstruction device based on deep learning, the problem of slow pedestrian detection speed when intelligent monitoring algorithms process massive video data is solved, real-time high-precision pedestrian detection and trajectory reconstruction are realized, and case handling efficiency is improved.
Patent Information
- Application Number
- CN202210111154.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-27
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-01-27
AI Technical Summary
When existing intelligent monitoring algorithms process massive video data, pedestrian detection speed is slow, which leads to a long time in the case handling process and affects efficiency.
The regional pedestrian trajectory reconstruction device based on deep learning is adopted, including a pedestrian detection module, a pedestrian recognition module, a pedestrian trajectory reconstruction module and a result display module. The features are automatically extracted through deep learning technology to realize pedestrian detection and trajectory reconstruction in high-definition videos.
Real-time high-precision pedestrian detection under high-definition video is realized, and pedestrian trajectory is quickly analyzed and reconstructed, which significantly reduces manpower investment and time consumption and improves the efficiency of intelligent monitoring.
Smart Images

Figure CN114429665B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and particularly relates to a device and method for reconstructing regional pedestrian trajectories based on deep learning. Background Art
[0002] As of 2020, the world's population has exceeded 7.7 billion, and the population of China has exceeded 1.4 billion. How to ensure stable social order has become one of the main tasks of governments around the world. In the era of rapid development of information technology, intelligent monitoring has become one of the most important components of a smart city. As an information carrier, cameras collect a large amount of information every moment. In order to effectively utilize this vast amount of information, more and more intelligent monitoring algorithms have been proposed, such as: pedestrian re-identification, abnormal event monitoring, crowd flow monitoring, etc. And pedestrian detection has become an essential module in most intelligent monitoring algorithms. As a basic tool, pedestrian detection can be embedded and used in various task scenarios. Therefore, the amount of data to be processed is increasing, and the speed requirements for intelligent monitoring algorithms are becoming more and more stringent, which also greatly affects the speed of the entire algorithm.
[0003] The vast amount of video data brings convenience to public security governance while also increasing great difficulties. In the traditional case-handling process, after the case-handling personnel lock in the suspect information, a large amount of manpower is required to search for clues of the suspect in the vast amount of video surveillance data, gradually locate the suspect's activity trajectory, infer the suspect's hiding place, and then carry out the arrest. And the most time-consuming step among them is to review the video surveillance data. The emergence of pedestrian re-identification technology attempts to quickly locate the suspect's position in the vast amount of video surveillance data through artificial intelligence, thereby greatly reducing the investment of manpower and the consumption of time.
[0004] With the success of deep learning, more and more scholars have begun to try to use deep learning to solve problems that were difficult to solve in the past. Deep learning has penetrated into various fields and brought a brand-new solution to each field. The solution based on deep learning is different from the solution of the previous traditional methods. The solution based on deep learning often does not require a large number of features to be extracted manually. Its internal convolutional neural network can automatically learn the way of extracting features through backpropagation, thus having better generalization. In addition, the solution based on deep learning can be transplanted to GPU computing devices and often has a faster computing speed than traditional methods. Summary of the Invention
[0005] To solve the above problems and provide a fast pedestrian detection device that can perform real-time high-precision pedestrian detection in high-definition videos, the present invention adopts the following technical solutions:
[0006] The present invention provides a regional pedestrian trajectory reconstruction device based on deep learning, which is used to process a video stream to reconstruct and display the trajectory information of pedestrians. The device is characterized by including: a camera, a pedestrian detection module, a data storage and update module, a pedestrian re-identification module, a pedestrian trajectory reconstruction module, and a result display module. Among them, the camera collects surveillance videos at corresponding positions according to different installation positions. The pedestrian detection module performs pedestrian detection on the surveillance video RTSP stream data transmitted by the camera to obtain a to-be-reconstructed image containing pedestrians. The data storage and update module stores pedestrian images, face information, and current trajectory data corresponding to the pedestrians in the pedestrian images. The pedestrian re-identification module compares the similarity of pedestrians between the to-be-reconstructed image and the pedestrian images stored in the data storage module. Once the comparison is successful, it generates trajectory record information of the pedestrian. The trajectory record information includes time information, location information, and identity information. The data storage and update module stores the trajectory record information correspondingly into the current trajectory data to generate new trajectory data. The pedestrian trajectory reconstruction module classifies the trajectory record information according to pedestrians based on the identity information, arranges and sorts the new trajectory data according to the time information, and reconstructs the arranged and sorted new trajectory data based on the location information to generate a trajectory map of the pedestrian. The result display module displays the identity information and the trajectory map of the pedestrian.
[0007] A regional pedestrian trajectory reconstruction device based on deep learning provided by the present invention may also have the following technical feature. The pedestrian detection module includes a data decoding unit, a fast downsampling unit, and a multi-scale convolution unit. The data decoding unit decodes the received surveillance video RTSP stream data to generate corresponding video frame images. The fast downsampling unit performs downsampling on the video frame images by a predetermined multiple to obtain the receptive field of the video frame images. The multi-scale convolution unit obtains feature maps under different receptive fields based on the receptive field to generate a to-be-reconstructed image.
[0008] A regional pedestrian trajectory reconstruction device based on deep learning provided by the present invention may also have the following technical feature. The fast downsampling unit uses large-scale convolutions with a convolution kernel size of 7x7 and 5x5 and two pooling layers. When performing downsampling, the stride of the first convolution is 4, and the strides of other convolutions and pooling layers are 2. The predetermined multiple is 32.
[0009] A regional pedestrian trajectory reconstruction device based on deep learning provided by the present invention may also have the following technical feature. The multi-scale convolution unit uses 7 layers of 3x3 convolutions. The multi-scale convolution unit extracts corresponding feature maps after convolving at the third convolution layer, the fifth convolution layer, and the seventh convolution layer respectively.
[0010] A device for reconstructing regional pedestrian trajectories based on deep learning provided by the present invention may further have the following technical feature: it further includes a face recognition module. The face recognition module includes a face detection unit and a face recognition unit. The face detection unit performs face localization detection on the image to be reconstructed, generates a face image based on the detected face and transmits it to the face recognition unit. The face recognition unit extracts the face feature information of the received face image, compares the face feature information with the faces of pedestrians in the pedestrian images stored in the data storage and update module. Once the comparison is successful, the face feature information of the pedestrian is transmitted to the data storage and update module, and the data storage and update module updates and stores the pedestrian image of the pedestrian according to the face feature information.
[0011] A device for reconstructing regional pedestrian trajectories based on deep learning provided by the present invention may further have the following technical feature: it further includes a camera control module. The camera control module realizes the deployment of the camera by regularly reading and monitoring the configuration file of the camera.
[0012] The present invention also provides a method for reconstructing regional pedestrian trajectories based on deep learning, which is characterized in that the above-mentioned device for reconstructing regional pedestrian trajectories based on deep learning is used to process the video stream to obtain and display the trajectory information of pedestrians.
[0013] Functions and effects of the invention
[0014] According to a device for reconstructing regional pedestrian trajectories based on deep learning of the present invention, the device for reconstructing regional pedestrian trajectories adopts a deep learning method to realize the efficient processing and intelligent monitoring of a large amount of surveillance video data. Among them, the pedestrian detection module can detect pedestrians in the collected surveillance video RTSP data stream. The pedestrian re-identification module compares the detected pedestrians with the pedestrian images stored in the data storage and update module, generates trajectory record information according to the successfully compared pedestrians, and then the pedestrian trajectory reconstruction module sorts out the trajectory record information and the stored trajectory data, reconstructs the trajectory map of the pedestrian according to the time sequence, and finally the result display module displays the trajectory map for analysts to view. Therefore, the device for reconstructing regional pedestrian trajectories based on deep learning of the present invention can detect pedestrians in the surveillance video in real time, analyze the identities of pedestrians and obtain the positions of pedestrians, and generate trajectories through the pedestrian trajectory reconstruction module for relevant personnel to analyze. Brief description of the drawings
[0015] Figure 1 It is a schematic structural diagram of a device for reconstructing regional pedestrian trajectories based on deep learning in an embodiment of the present invention;
[0016] Figure 2 It is a schematic structural diagram of the pedestrian detection module in an embodiment of the present invention;
[0017] Figure 3 It is a schematic structural diagram of the face recognition module in the embodiment of the present invention;
[0018] Figure 4 It is a flowchart of the working process of the regional pedestrian trajectory reconstruction device based on deep learning in the embodiment of the present invention;
[0019] Figure 5 It is a flowchart of the method for regional pedestrian trajectory reconstruction based on deep learning in the embodiment of the present invention. Specific Embodiments
[0020] This embodiment is implemented based on a GPU (such as GTX1080ti) computing platform. Utilizing the advantages of deep learning, a set of pedestrian trajectory reconstruction devices including person re-identification technology in a regional scenario is designed to achieve efficient processing and intelligent monitoring of a large amount of surveillance video data.
[0021] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the following specifically elaborates on the regional pedestrian trajectory reconstruction device and method based on deep learning of the present invention in combination with embodiments and drawings.
[0022] <Embodiment>
[0023] The operating platform of the regional pedestrian trajectory reconstruction device based on deep learning in this embodiment is Linux, and this platform is supported by at least one graphics processing unit GPU card.
[0024] Figure 1 It is a schematic structural diagram of the regional pedestrian trajectory reconstruction device based on deep learning in the embodiment of the present invention.
[0025] As Figure 1 shown, the regional pedestrian trajectory reconstruction device 100 based on deep learning has a camera 1, a pedestrian detection module 2, a data storage and update module 3, a person re-identification module 4, a pedestrian trajectory reconstruction module 5, a result display module 6, a face recognition module 7, and a camera control module 8.
[0026] The camera 1 is used to collect surveillance videos corresponding to its installation position according to its different installation positions.
[0027] The pedestrian detection module 2 receives the surveillance video RTSP video stream data transmitted by the camera 1, decodes the RTSP video stream data, and then performs pedestrian detection frame by frame. If no pedestrian is detected, this frame of video is discarded. If a pedestrian is detected, the pedestrian area of this frame of video is retained to obtain a to-be-reconstructed image containing the pedestrian.
[0028] Figure 2 It is a schematic structural diagram of the pedestrian detection module in the embodiment of the present invention.
[0029] In this embodiment, the pedestrian detection module 2 completes the real-time pedestrian detection task based on multiple deep learning units. As Figure 2 shown, the pedestrian detection module 2 includes a data decoding unit 21, a fast downsampling unit 22, and a multi-scale convolution unit 23.
[0030] Among them, the data decoding unit 21 decodes the received monitoring video RTSP stream data to generate corresponding standard video frame images.
[0031] The fast downsampling unit 22 performs downsampling on the video frame images by a predetermined multiple, quickly extracts the visual information in the video frame images to obtain the receptive field of the video frame images.
[0032] In this embodiment, the fast downsampling unit 22 uses large-scale convolutions with a convolution kernel size of 7x7 and 5x5 and two pooling layers. When performing downsampling, the stride of the first convolution is 4, and the strides of other convolutions and pooling layers are 2, quickly downsampling the image by 32 times, so as to improve the inference speed of the network while maintaining a large receptive field.
[0033] The multi-scale convolution unit 23 obtains feature maps under different receptive fields based on the receptive field generated by the fast downsampling unit 22 to generate the image to be reconstructed.
[0034] In this embodiment, the multi-scale convolution unit 23 uses 7 layers of 3x3 convolutions, and extracts corresponding feature maps after convolution in the third convolution layer, the fifth convolution layer, and the seventh convolution layer respectively.
[0035] The data storage and update module 3 stores pedestrian images, face information, and current trajectory data corresponding to the pedestrians in the pedestrian images, and updates and stores the trajectory record information of the pedestrians received from the pedestrian re-identification module 4 and the face feature information received from the face recognition module 7.
[0036] The pedestrian re-identification module 4 uses a feature extractor to perform feature acquisition on the image to be reconstructed received from the pedestrian detection module 2 and compare the similarity of the pedestrians with the pedestrian images stored in the data storage module 3. When the similarity reaches a certain threshold, that is, the comparison is successful, the pedestrian re-identification module 4 generates the trajectory record information of the pedestrian. The trajectory record information includes time information, position information, and identity information. Among them, the position information is obtained based on position prediction of the feature maps generated by the multi-scale convolution unit 23.
[0037] The pedestrian trajectory reconstruction module 5 receives the trajectory record information of the pedestrian from the pedestrian re-identification module 4, classifies the trajectory record information by pedestrian according to the included identity information, arranges and sorts the new trajectory data according to the time information, and based on the position information, plots all the passed position points of the arranged and sorted new trajectory data on a prefabricated trajectory plan view and connects them according to the time sequence information to reconstruct and generate the trajectory map of the pedestrian.
[0038] The result display module 6 displays the trajectory map generated by the pedestrian trajectory reconstruction module 5 and the identity information of the corresponding pedestrian.
[0039] In this embodiment, the result display module 6 can also call the pedestrian trajectory reconstruction module 5 to reconstruct the trajectory maps of different pedestrians in different time periods according to the screening conditions and then display them.
[0040] The face recognition module 7 is used to further identify the face information in the to-be-reconstructed image received from the pedestrian detection module 2 and capture new pedestrian feature data. In the scenario of a high-definition camera, the face recognition module 7 can be called to supplement the insufficient pedestrian re-identification sample data.
[0041] Figure 3 It is a schematic structural diagram of the face recognition module in the embodiment of the present invention.
[0042] As Figure 3 shown, the face recognition module 7 includes a face detection unit 71 and a face recognition unit 72.
[0043] Among them, the face detection unit 71 performs face positioning detection on the to-be-reconstructed image. If a face is detected, a face image is generated based on the detected face and transmitted to the face recognition unit 72 for processing. If no face is detected, the face recognition module 7 no longer performs the face recognition detection task.
[0044] The face recognition unit 72 extracts the face feature information of the received face image, compares the face feature information with the faces of pedestrians in the pedestrian images stored in the data storage and update module. Once the comparison is successful, the face feature information of the pedestrian is transmitted to the data storage and update module 3.
[0045] The camera control module 8 realizes the allocation of adding, deleting, etc. to the camera by regularly reading and monitoring the configuration file of the camera and modifies the corresponding camera position information in the database.
[0046] In this embodiment, the pedestrian detection module 2, the pedestrian re-identification module 4, the pedestrian trajectory reconstruction module 5, and the face recognition module 7 are encapsulated, and the camera control module 8 controls the camera to modify the input data of the device. That is, when the camera control module 8 monitors that the camera configuration file has been modified, it will transmit the new camera information to the pedestrian detection module 2.
[0047] Figure 4 It is the working flowchart of the regional pedestrian trajectory reconstruction device based on deep learning in the embodiment of the present invention.
[0048] As Figure 4 shown, the working process of the regional pedestrian trajectory reconstruction device 100 based on deep learning is as follows:
[0049] Step S1, the camera 1 collects the monitoring videos of the corresponding positions according to different installation positions;
[0050] Step S2, the pedestrian detection module 2 performs pedestrian detection on the monitoring video RTSP data stream transmitted by the camera to obtain the image to be reconstructed containing pedestrians;
[0051] Step S3, the pedestrian re-identification module 4 compares the similarity of pedestrians between the image to be reconstructed and the pedestrian images stored in the data storage and update module 3. Once the comparison is successful, it generates the trajectory record information of the pedestrian. The trajectory record information includes time information, position information, and identity information;
[0052] Step S4, the data storage and update module 3 stores the trajectory record information correspondingly into the current trajectory data to generate new trajectory data;
[0053] Step S5, the pedestrian trajectory reconstruction module 5 sorts out the new trajectory data and reconstructs the trajectory map of the pedestrian according to the chronological information;
[0054] Step S6, the result display module 6 displays the identity information and the trajectory map of the pedestrian.
[0055] In this embodiment, in each of the above working processes, the face recognition module 7 will identify the face information of the image to be reconstructed received from the pedestrian detection module 2, and transmit the successfully generated face feature information to the data storage and update module 3. The camera control module 8 will read the configuration file of the camera to determine whether the camera information is updated. If there is a new camera information update, it will transmit the new camera configuration data to the pedestrian detection module 2, and the pedestrian detection module 2 will detect the new RTSP data stream.
[0056] In practical applications, the deep learning-based regional pedestrian trajectory reconstruction device of this embodiment can be deployed on an embedded device to obtain monitored data through an IP connection with a camera. For example, a police station can implement this system to record the path of a criminal suspect during the interrogation process in the police station.
[0057] Figure 5 It is a flowchart of the deep learning-based regional pedestrian trajectory reconstruction method in an embodiment of the present invention.
[0058] As Figure 5 shown, the deep learning-based regional pedestrian trajectory reconstruction method includes the following steps:
[0059] Step A1: Obtain surveillance videos at different locations;
[0060] Step A2: Perform pedestrian detection on the RTSP stream data of the surveillance video to obtain the images to be reconstructed containing pedestrians;
[0061] Step A3: Compare the similarity of pedestrians between the images to be reconstructed and the pre-stored pedestrian images, and generate trajectory record information for the pedestrians with successful comparison. The trajectory record information includes time information, location information, and identity information;
[0062] Step A4: Correspondingly add the trajectory record information to the pre-stored current trajectory data to generate new trajectory data;
[0063] Step A5: Classify the new trajectory record information by pedestrian according to the identity information, arrange and sort the new trajectory data according to the time information, and sort the new trajectory data in chronological order based on the location information to reconstruct and generate the trajectory map of the pedestrian.
[0064] Functions and effects of the embodiment
[0065] According to the device for regional pedestrian trajectory reconstruction based on deep learning provided by this embodiment, first, the pedestrian detection module performs pedestrian detection on the collected monitoring video RTSP data stream. Then, the pedestrian re-identification module compares the similarity between the detected pedestrians and the pedestrian images stored in the data storage and update module, and generates trajectory record information based on the successfully matched pedestrians. Secondly, the pedestrian trajectory reconstruction module sorts out the trajectory record information and the stored trajectory data, and reconstructs the trajectory map of the pedestrian in chronological order. Finally, the result display module displays the trajectory map for analysts to view. The device for regional pedestrian trajectory reconstruction based on deep learning in this embodiment can use deep learning methods to obtain, detect and analyze pedestrian information in a large amount of monitoring video data, and reconstruct pedestrian trajectories using the feature information of the video data, so as to achieve efficient processing and intelligent monitoring of a large amount of monitoring video data. At the same time, this trajectory reconstruction device can also free criminal investigators from the traditional process of reviewing video monitoring data to find the walking paths of suspects, allowing case handlers to focus more on reasoning and analysis.
[0066] In the embodiment, since the relevant modules are deployed on embedded devices, it can be applied in schools, police stations, and key laboratories, which can help monitor information such as the appearance frequency and appearance paths of pedestrians without awareness, and at the same time ensure relatively low costs and space in implementation.
[0067] In the embodiment, it is also because the pedestrian detection module, pedestrian re-identification module, pedestrian trajectory reconstruction module, and face recognition module are encapsulated, and the input information of the system is modified through the camera control module, thereby reducing the implementation difficulty of the device for regional pedestrian trajectory reconstruction during deployment.
[0068] The above embodiments are only used to illustrate the specific implementation manners of the present invention, and the present invention is not limited to the description scope of the above embodiments.
Claims
1. A device for reconstructing regional pedestrian trajectories based on deep learning, which is used to process video streams to reconstruct pedestrian trajectory information and display it. Characterized in that: It includes: A camera, a pedestrian detection module, a data storage and update module, a pedestrian re-identification module, a pedestrian trajectory reconstruction module, and a result display module. Among them, the camera collects surveillance videos at corresponding positions according to different installation positions. The pedestrian detection module performs pedestrian detection on the surveillance video RTSP stream data transmitted by the camera to obtain a to-be-reconstructed image containing pedestrians. The data storage and update module stores pedestrian images, face information, and current trajectory data corresponding to the pedestrians in the pedestrian images. The pedestrian re-identification module compares the similarity of pedestrians between the to-be-reconstructed image and the pedestrian images stored in the data storage and update module. Once the comparison is successful, it generates trajectory record information for this pedestrian. The trajectory record information includes time information, location information, and identity information. The data storage and update module stores the trajectory record information correspondingly into the current trajectory data to generate new trajectory data. The pedestrian trajectory reconstruction module classifies the trajectory record information by pedestrian according to the identity information, arranges and organizes the new trajectory data according to the time information, and reconstructs the arranged and organized new trajectory data based on the location information to generate a trajectory map of this pedestrian. The result display module displays the identity information of this pedestrian and the trajectory map. The pedestrian detection module includes a data decoding unit, a fast downsampling unit, and a multi-scale convolution unit. The data decoding unit decodes the received surveillance video RTSP stream data to generate corresponding video frame images. The fast downsampling unit performs downsampling on the video frame images by a predetermined multiple to obtain the receptive field of the video frame images. The multi-scale convolution unit obtains feature maps under different receptive fields based on the receptive field to generate the to-be-reconstructed image.
2. The device for reconstructing regional pedestrian trajectories based on deep learning according to claim 1, characterized in that: Among them, The fast downsampling unit uses large-scale convolutions with a convolution kernel size of 7x7 and 5x5 and two pooling layers. When performing the downsampling, the stride of the first convolution is 4, and the strides of other convolutions and pooling layers are 2. The predetermined multiple is 32.
3. The device for reconstructing regional pedestrian trajectories based on deep learning according to claim 1, characterized in that: Among them, The multi-scale convolution unit uses 7 layers of 3x3 convolutions. The multi-scale convolution unit extracts corresponding feature maps after convolution in the third convolution layer, the fifth convolution layer, and the seventh convolution layer respectively.
4. The device for reconstructing regional pedestrian trajectories based on deep learning according to claim 1, Characterized in that: It further includes: A face recognition module. Among them, the face recognition module includes a face detection unit and a face recognition unit. The face detection unit performs face localization detection on the to-be-reconstructed image, generates a face image based on the detected face, and transmits it to the face recognition unit. The face recognition unit extracts the face feature information of the received face image, compares the face feature information with the face of the pedestrian in the pedestrian image stored in the data storage and update module, and once the comparison is successful, transmits the face feature information of the pedestrian to the data storage and update module. The data storage and update module updates and stores the pedestrian image of the pedestrian according to the face feature information.
5. A regional pedestrian trajectory reconstruction device based on deep learning according to claim 1, characterized in that, further comprising: a camera control module, wherein the camera control module realizes the deployment of the camera by periodically reading and monitoring the configuration file of the camera.
6. A regional pedestrian trajectory reconstruction method based on deep learning, characterized in that, using the regional pedestrian trajectory reconstruction device according to any one of claims 1 to 5 to process the video stream to obtain and show the trajectory information of the pedestrian.
Citation Information
Patent Citations
Real-time pedestrian re-identification method and device
CN111291633A