Personnel Trajectory Retrieval Method and Device Based on Multi-Algorithm Fusion Application

By extracting and fusing feature information in video streams or picture streams collected by multiple cameras, and combining multiple recognition algorithms, a complete retrieval of pedestrian trajectories in complex scenarios is achieved, solving the limitations of a single algorithm in complex scenarios.

CN113963399BActive Publication Date: 2025-06-27WUHAN FIBERHOME DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111055598.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-09
Publication Date
2025-06-27
Estimated Expiration
2041-09-09

AI Technical Summary

Technical Problem

The existing technology is difficult to achieve complete retrieval of pedestrian trajectories in complex scenarios. A single algorithm has limitations and cannot automatically integrate multi-algorithm results.

Method used

By extracting the face, human body and gait feature information in video streams or picture streams collected by cameras at multiple different points, combining face recognition, pedestrian re-recognition and gait recognition algorithms, the trajectories of the same target person are automatically connected and fused to form complete trajectory information.

Benefits of technology

The complete search of pedestrian trajectories is realized in complex scenarios, making up for the shortcomings of a single algorithm, expanding the application scenarios, and obtaining more complete pedestrian trajectory information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113963399B_ABST
    Figure CN113963399B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for retrieving personnel trajectories based on the fusion application of multiple algorithms. The method includes the following steps: extracting the feature information of each pedestrian in the video streams or picture streams collected by multiple cameras at different positions, and storing the feature information of each extracted pedestrian together with the corresponding camera position and the capture time as a pedestrian record in a passerby database; extracting the feature information of the target retrieval personnel, and comparing the feature information of the target retrieval personnel with the feature information of the corresponding types of pedestrians in the passerby database using the corresponding recognition algorithm, finding each pedestrian record with feature information consistent with that of the target retrieval personnel and taking the union to obtain the trajectory information of the target retrieval personnel; according to the trajectory information of the target retrieval personnel, connecting the camera positions corresponding to each pedestrian record on a GIS map according to the capture time to obtain the trajectory of the target retrieval personnel. The present invention can solve the problem of personnel retrieval in various complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pedestrian trajectory retrieval, and in particular to a method and device for pedestrian trajectory retrieval based on the fusion application of multiple algorithms. Background Art

[0002] With the widespread construction and use of target recognition algorithms such as face recognition and video structuring in the system, their application deficiencies have also emerged - environmental and human factors such as long distance, improper angle, clothing change, facial occlusion, light influence, low video resolution, far distance, disguise, etc. will all affect the recognition efficiency. Some technologies corresponding to making up for the deficiencies of face recognition technology have emerged, such as gait recognition, pedestrian re-identification technology, etc.

[0003] For pedestrian trajectory retrieval, traditional methods are all to retrieve the trajectory of personnel through a single technology, but single technical means all have their own limitations and can no longer meet the application requirements of complex scenarios and cannot obtain the complete trajectory of pedestrians. For example, in the case of face occlusion or wearing a mask, face retrieval cannot be applied; in the case of across days or personnel changing clothes, pedestrian re-identification cannot be applied for trajectory retrieval, etc. From a single picture, complete gait information cannot be obtained and gait recognition cannot be applied for trajectory retrieval. Therefore, the fusion application of multiple algorithms can make full use of the advantages of each algorithm and make up for the deficiencies of a single algorithm, becoming the future development trend. However, for traditional fusion applications of multiple algorithms, more often, users manually combine multiple algorithms for use, and the system cannot automatically fuse the trajectories of target pedestrians. Therefore, how to fuse multiple algorithms and form a complete trajectory after fusion has become a difficulty in the fusion application of multiple algorithms. Summary of the Invention

[0004] To solve at least some of the above problems existing in the prior art, the present invention provides a method and device for pedestrian trajectory retrieval based on the fusion application of multiple algorithms.

[0005] The present invention is implemented as follows:

[0006] On the one hand, the present invention provides a method for pedestrian trajectory retrieval based on the fusion application of multiple algorithms, including the following steps:

[0007] S1. Extract the feature information of each pedestrian in the video stream or picture stream collected by multiple cameras at different positions. The feature information of each extracted pedestrian includes at least one of face feature information, body feature information, and gait feature information. Store the feature information of each extracted pedestrian together with the corresponding camera position and capture time as a pedestrian record in the passerby database;

[0008] S2. Extract the feature information of the target person to be retrieved from the video clips or pictures of the target person to be retrieved. Compare the face feature information, body feature information, and gait feature information of the target person to be retrieved with the corresponding types of pedestrian feature information in the passerby database using the corresponding face recognition algorithm, person re-identification algorithm, and gait recognition algorithm. Find each pedestrian record with feature information consistent with that of the target person to be retrieved, and take the union of the retrieved pedestrian records to obtain the trajectory information of the target person to be retrieved;

[0009] S3. According to the trajectory information of the target person to be retrieved, connect the camera points corresponding to each pedestrian record on the GIS map according to the capture time to obtain the trajectory of the target person to be retrieved.

[0010] Further, in step S1, the extraction of the face feature information in the video stream or picture stream specifically includes:

[0011] Use the MTCNN network to perform face target detection on each frame of video data or each picture;

[0012] Use the trained face feature extraction model to extract the face features of the detected face targets to obtain face feature information;

[0013] Among them, the training method of the face feature extraction model is as follows:

[0014] Collect face photos in a large number of video surveillance scenarios, perform annotation, and divide them into a training set, a validation set, and a test set;

[0015] Construct an improved faceNet model. Among them, the faceNet backbone network uses the Inception-v4 network and residual connection. In the selection of the loss function, calculate the softmax loss and triplet_loss respectively, and then weight the softmax loss and triplet_loss, and adjust the weights online;

[0016] Use the training set to train the improved faceNet model, use the validation set to verify the convergence of the model, and perform tests through the test set, and finally output the best face feature extraction model.

[0017] Further, in step S1, the extraction of the face feature information in the video stream further includes:

[0018] After performing face target detection, use DeepSort to track the detected face targets and perform moving target detection;

[0019] Use the trained frontal face determination model to determine whether the face target output by tracking is a frontal face and output the determined frontal face data of the face;

[0020] Use the trained face feature extraction model to extract face features from the output frontal face data of the face to obtain face feature information.

[0021] Further, in the step S1, the extraction of human feature information from the video stream or picture stream specifically includes:

[0022] Use the trained yolov5 object detection model to perform pedestrian object detection on each frame of video data or each picture;

[0023] Use the trained human feature extraction model to extract human features from the detected pedestrian objects to obtain human feature information;

[0024] Among them, the training method of the human feature extraction model is as follows:

[0025] Collect and construct multiple pedestrian sequence datasets in the video surveillance scenario, collect the sequences of different pedestrian objects under multiple different camera perspectives, and divide them into a training set, a validation set, and a test set;

[0026] Construct an improved pedestrian FastReID model, the backbone adopts a multi-scale multi-component fusion deep network based on adaptive margin ranking loss, and jointly learns the pedestrian image feature representation and similarity measurement;

[0027] In the data preprocessing stage, downsample the data;

[0028] Use the training set to train the improved pedestrian FastReID model, use the validation set to verify the convergence of the model, and perform tests through the test set, and finally output the best human feature extraction model.

[0029] Further, in the step S1, the extraction of human feature information from the video stream further includes:

[0030] After performing pedestrian object detection, use DeepSort to track the detected pedestrian objects and perform moving object detection;

[0031] Use the trained pedestrian quality assessment model to evaluate the quality of the pedestrian objects output by tracking, and output the pedestrian object with the best quality as a close-up picture;

[0032] Use the trained human feature extraction model to extract human features from the pedestrian object with the best quality output to obtain human feature information.

[0033] Further, in the step S1, the steps of extracting gait feature information from the video stream are as follows:

[0034] Use the trained yolov5 object detection model to perform pedestrian object detection on each frame of video data;

[0035] Use DeepSort to track the detected pedestrian objects and perform moving object detection;

[0036] Use the trained deeplabv3+ segmentation model to perform gait segmentation on the pedestrian sequence output by tracking, and output a black-and-white contour sequence map;

[0037] Use the trained gait feature extraction model to extract gait features from the black-and-white contour sequence map output by gait segmentation;

[0038] Among them, the training method of the gait feature extraction model is as follows:

[0039] Collect and construct multiple pedestrian sequence datasets in the video surveillance scenario, collect sequences of different pedestrian objects from multiple different camera perspectives, and divide them into a training set, a validation set, and a test set;

[0040] Construct an improved GaitSet model; add a CNN module at the beginning of the backbone network, and add MDFG and GMCM modules to the last two layers of the network;

[0041] In the data preprocessing stage, the training data is augmented by random horizontal flipping and random jigsaw puzzles;

[0042] Use the training set to train the improved GaitSet model, use the validation set to verify the convergence of the model, and perform tests through the test set, and finally output the best gait feature extraction model.

[0043] Further, in the step 2, when both face feature information comparison and human body feature information or gait feature information comparison are performed in a pedestrian record, the comparison result of the face feature information shall prevail.

[0044] Further, in the step S2, when a certain type of feature information among face feature information, human body feature information, and gait feature information is not included in the feature information of the retrieved target person extracted from the video segment or picture of the retrieved target person, use the feature information of this type in the retrieved pedestrian record to perform a search in the passerby database to find each pedestrian record corresponding to this type of feature information.

[0045] Further, after the step S3, it further includes:

[0046] Calculate the average speed of the distance between every two adjacent points in the trajectory of the target person to be retrieved. Select the sections with an average speed exceeding 30 km / s as abnormal sections. Calculate the algorithm recognition probabilities of the points on both sides of the abnormal section respectively, select the point with a lower probability as the final abnormal point, and remove it from the trajectory, where point x t The algorithm recognition probability P(x t ) is calculated as follows:

[0047]

[0048] Where M1, M2, and M3 respectively represent the face recognition algorithm, the person re-identification algorithm, and the gait recognition algorithm. P(M i ) represents the recognition accuracy of algorithm M i . P(x t |M i ) represents the probability that the person identified at point x i under algorithm M t is the target person to be retrieved.

[0049] On the other hand, the present invention also provides a person trajectory retrieval device based on multi-algorithm fusion application, including:

[0050] A pedestrian feature information extraction and storage module, which is used to extract the feature information of each pedestrian in the video stream or picture stream collected by cameras at multiple different points. The extracted feature information of each pedestrian includes at least one of face feature information, body feature information, and gait feature information. The extracted feature information of each pedestrian, together with the corresponding camera point and the capture time, is stored as a pedestrian record in the passerby database;

[0051] A trajectory retrieval module, which is used to extract the feature information of the target person to be retrieved according to the video segment or picture of the target person to be retrieved, and compare the face feature information, body feature information, and gait feature information of the target person to be retrieved with the corresponding types of pedestrian feature information in the passerby database by using the corresponding face recognition algorithm, person re-identification algorithm, and gait recognition algorithm, find each pedestrian record with feature information consistent with the target person to be retrieved, and take the union of the retrieved pedestrian records to obtain the trajectory information of the target person to be retrieved;

[0052] A trajectory display module, which is used to connect the camera points corresponding to each pedestrian record in the GIS map according to the capture time according to the trajectory information of the target person to be retrieved to obtain the trajectory of the target person to be retrieved.

[0053] Compared with the prior art, the present invention has the following beneficial effects:

[0054] The method and device for retrieving personnel trajectories based on the fusion application of multiple algorithms provided by the present invention can solve the problem of personnel retrieval in complex scenarios such as long distance, improper angle, clothing change, facial occlusion, light influence, low video resolution, far distance, disguise, etc. The present invention can fuse and apply three algorithms of face recognition, person re-identification, and gait recognition, automatically connect and fuse the trajectories of the same target person, make up for the deficiencies of a single algorithm, make full use of the advantages of each algorithm for fusion application, realize the retrieval of personnel trajectories, and be able to obtain more complete trajectory information of personnel, expanding the application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a flowchart of a method for retrieving personnel trajectories based on the fusion application of multiple algorithms provided by an embodiment of the present invention;

[0056] Figure 2 It is a schematic diagram of the network structure of a gait feature extraction model provided by an embodiment of the present invention;

[0057] Figure 3 It is a block diagram of a device for retrieving personnel trajectories based on the fusion application of multiple algorithms provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0059] As Figure 1 shown, an embodiment of the present invention provides a method for retrieving personnel trajectories based on the fusion application of multiple algorithms, including the following steps:

[0060] S1. Extract the feature information of each pedestrian in the video stream or picture stream collected by multiple cameras at different positions. The extracted feature information of each pedestrian includes at least one of face feature information, body feature information, and gait feature information. When the condition for extracting the corresponding feature information is not met, the feature information is not extracted. For example, if the gait feature information cannot be extracted from the picture stream, the gait feature information of the pedestrians in the picture stream is not extracted. The extracted feature information of each pedestrian, together with the corresponding camera position and capture time, is stored as a pedestrian record in the passerby database, and a unique personnel ID information is generated for each pedestrian record.

[0061] Specifically, in the step S1, various feature information of pedestrians in the video stream or picture stream is extracted respectively. Among them, the steps for extracting the face feature information in the picture stream are as follows:

[0062] Use the MTCNN network to perform face object detection on each picture;

[0063] Use the trained face feature extraction model to extract face features from the detected face objects, and obtain 128-dimensional face feature information.

[0064] Furthermore, in the step S1, the steps for extracting the face feature information in the video stream are as follows:

[0065] Use the MTCNN network to perform face object detection on each frame of video data;

[0066] Use DeepSort to track the detected face objects and perform moving object detection;

[0067] Use the trained face frontal face determination model to perform face frontal face determination on the face objects output by the tracking and output the determined face frontal face data;

[0068] Use the trained face feature extraction model to extract face features from the output face frontal face data, and obtain 128-dimensional face feature information.

[0069] By detecting and tracking the face, ensure that only one optimal face picture of the same target pedestrian is obtained.

[0070] Preferably, the training method of the face feature extraction model adopted in the above two processes of face feature information extraction is as follows:

[0071] Collect face photos in a large number of video surveillance scenarios, perform annotation, and divide them into a training set, a validation set, and a test set;

[0072] Build an improved faceNet model. Among them, the faceNet backbone network adopts the Inception-v4 network and residual connection. In the selection of the loss function, calculate the softmax loss and triplet_loss respectively, and then weight the softmax loss and triplet_loss, and adjust the weights online;

[0073] Use the training set to train the improved faceNet model, use the validation set to verify the convergence of the model, and perform tests through the test set, and finally output the best face feature extraction model.

[0074] In the training method of the above-mentioned face feature extraction model, by constructing an improved faceNet model, the backbone network and loss function are optimized, the gap between target categories is widened, the training speed is accelerated, and the robustness of face features is enhanced.

[0075] Further, in the step S1, the steps of extracting human feature information from the picture stream are as follows:

[0076] Use the trained yolov5 object detection model to detect pedestrian targets in each picture;

[0077] Use the trained human feature extraction model to extract human features from the detected pedestrian targets, and obtain 2048-dimensional human feature information.

[0078] Further, in the step S1, the steps of extracting human feature information from the video stream are as follows:

[0079] Use the trained yolov5 object detection model to detect pedestrian targets in each frame of video data;

[0080] Use DeepSort to track the detected pedestrian targets and perform moving object detection;

[0081] Use the trained pedestrian quality assessment model to evaluate the quality of the pedestrian targets output by the tracking, and output the pedestrian target with the best quality as a close-up picture;

[0082] Use the trained human feature extraction model to extract human features from the pedestrian target with the best quality output, and obtain 2048-dimensional human feature information.

[0083] By detecting and tracking pedestrian targets, it is ensured that only one optimal pedestrian human picture of the same target pedestrian is obtained.

[0084] Preferably, the training method of the human feature extraction model adopted in the above two human feature extraction processes is as follows:

[0085] Collect and construct a 1000 pedestrian sequence dataset in the video surveillance scenario, collect sequences of different pedestrian targets from 5 different camera perspectives, and divide them into a training set, a validation set, and a test set;

[0086] Construct an improved pedestrian FastReID model. The backbone adopts a multi-scale multi-component fusion deep network based on adaptive margin ranking loss, and jointly learns the pedestrian image feature expression and similarity measurement;

[0087] In the data preprocessing stage, based on the existing preprocessing, downsample the data to enrich the multi-dimensionality of the data and reduce the sensitivity of the model to the data;

[0088] Use the training set to train the improved pedestrian FastReID model, use the validation set to verify the convergence of the model, and conduct tests through the test set, and finally output the best human feature extraction model.

[0089] In the above training method of the human feature extraction model, the data set is preprocessed diversely, enriching the diversity and multi-dimensionality of the data, and reducing the sensitivity of the model to the data; an improved pedestrian FastReID model is constructed, and the backbone adopts a multi-scale multi-component fusion deep network based on adaptive margin ranking loss, jointly learning the pedestrian image feature expression and similarity measurement, enhancing the robustness of pedestrian features, and greatly improving the accuracy of pedestrian comparison and recognition.

[0090] Furthermore, in the step S1, the steps of extracting gait feature information from the video stream are as follows:

[0091] Use the trained yolov5 object detection model to detect pedestrian targets for each frame of video data;

[0092] Use DeepSort to track the detected pedestrian targets and conduct moving target detection;

[0093] Use the trained deeplabv3+ segmentation model to perform gait segmentation on the pedestrian sequence output by tracking, output a black and white contour sequence map, and crop the contour map size to 64*64;

[0094] Use the trained gait feature extraction model to extract gait features from the black and white contour sequence map output by gait segmentation.

[0095] Preferably, the training method of the gait feature extraction model adopted in the above process of gait feature extraction is as follows:

[0096] Collect and construct a data set of 1000 pedestrian sequences in a video surveillance scenario, collect sequences of different pedestrian targets from 5 different camera perspectives, and divide them into a training set, a validation set, and a test set;

[0097] Construct an improved GaitSet model, see Figure 2As shown in the figure, a CNN module is added at the beginning of the backbone network, and the number of channels is increased to 256, which can obtain and extract more information. In the last two layers of the network, an MDFG module and a GMCM module are newly added. The MDFG module is used to generate visual cues in different regions for fine-grained feature learning, and the GMCM module updates the existing aggregation strategy to map the frame-level global information feature vector into a feature vector.

[0098] In the data preprocessing stage, the training data is augmented by random horizontal flipping and random jigsaw methods to enhance the robustness of the model.

[0099] The improved GaitSet model is trained using the training set, the convergence of the model is verified using the validation set, and it is tested using the test set. Finally, the best gait feature extraction model is output.

[0100] In the above training method of the gait feature extraction model, an improved GaitSet model is constructed, adding a convolutional neural network module (CNN), a target fine-grained learning module (MDFG), and an aggregation strategy update module (GMCM), which greatly improves the robustness of the gait features, widens the gap between different target gait features, and improves the accuracy of gait recognition.

[0101] S2. Extract the feature information of the retrieval target person according to the video segment or picture of the retrieval target person. The extracted feature information of the retrieval target person also includes at least one of face feature information, human body feature information, and gait feature information. According to the face feature information, human body feature information, and gait feature information of the retrieval target person, use the corresponding face recognition algorithm, person re-identification algorithm, and gait recognition algorithm to compare with the corresponding type of pedestrian feature information in the passerby library, find each pedestrian record with feature information consistent with the retrieval target person, and take the union of the retrieved pedestrian records to obtain the trajectory information of the retrieval target person.

[0102] The above comparison process specifically includes:

[0103] Using the face recognition algorithm, compare the face feature information of the retrieval target person with the face feature information in the passerby library, find the pedestrian records with the face feature information similarity greater than the set threshold, record their corresponding person ID information and similarity, and set it as set L1.

[0104] Using the person re-identification algorithm, compare the human body feature information of the retrieval target person with the human body feature information in the passerby library, find the pedestrian records with the human body feature information similarity greater than the set threshold, record their corresponding person ID information and similarity, and set it as set L2.

[0105] Using a gait recognition algorithm, compare the gait feature information of the target person to be retrieved with the gait feature information in the passerby database, find the pedestrian records with a gait feature information similarity greater than the set threshold, record their corresponding person ID information and similarity, and set it as set L3.

[0106] Preferably, in step S2, when both face feature information comparison and human body feature information or gait feature information comparison are performed in a pedestrian record, considering that face recognition is relatively mature and has high accuracy, the comparison result of the face feature information shall prevail.

[0107] Preferably, in step S2, when a certain type of feature information among face feature information, human body feature information, and gait feature information is not included in the feature information of the target person to be retrieved extracted from the video segment or picture of the target person to be retrieved, use this type of feature information in the retrieved pedestrian records to perform a search in the passerby database to find each pedestrian record corresponding to this type of feature information.

[0108] After the above process is completed, take the union of the above three groups of person ID information, L = L1 ∪ L2 ∪ L3, then L is the trajectory information of the preliminary screening of the target person to be retrieved. This trajectory information includes the person's ID information, camera location information, and capture time, and may also include information such as corresponding person pictures, gait sequence diagrams, similarity, etc.

[0109] S3. According to the trajectory information of the target person to be retrieved, connect the camera locations corresponding to each pedestrian record on the GIS map according to the capture time to obtain the trajectory of the target person to be retrieved.

[0110] Specifically, according to the obtained trajectory information of the target person to be retrieved, on the GIS map, connect the camera locations of the corresponding records in chronological order and display the picture information of the target person at the corresponding locations, which is the complete trajectory information of the person, and information such as gait sequence diagrams and similarity can also be displayed.

[0111] Considering that the maturity of the current pedestrian re-identification algorithm and gait recognition algorithm is not as good as that of the face recognition algorithm, and the recognition accuracy is not particularly high, through comprehensive analysis by combining the spatio-temporal information in the trajectory and the recognition accuracy of the algorithm, some abnormal trajectory points in the trajectory Trac can be removed, and the trajectory Trac can be further screened for effectiveness.

[0112] In actual situations, the legitimacy of a node in a person's trajectory can be judged by calculating the average speed of the people at adjacent nodes in the trajectory. For any two adjacent nodes x i , x i+1 The distance length between them in the network G is denoted as S(x i , x i+1), the average speed of the person between x i , x i+1 can be expressed as:

[0113]

[0114] Preferably, after the step S3, it further includes:

[0115] Calculate the average speed v of the distance between every two adjacent points in the trajectory of the retrieved target person i,i+1 , select the sections where the average speed exceeds 30 km / s as abnormal sections, calculate the algorithm recognition probabilities of the points on both sides of the abnormal section respectively, select the points with lower probabilities as the final abnormal points, and remove them from the trajectory,

[0116] At the point x where the retrieved target person is located at time t in the trajectory Trac t , there is a point x t The probability is jointly confirmed by M1, M2, and M3. The algorithm recognition probability P(x t ) is calculated according to the total probability formula as: t ) is:

[0117]

[0118] Where M1, M2, and M3 respectively represent the face recognition algorithm, the person re-identification algorithm, and the gait recognition algorithm, and P(M i ) represents the recognition accuracy of the algorithm M i , and P(x t |M i ) represents the probability that the person identified at the point x i under the algorithm M t is the retrieved target person (i.e., the similarity value).

[0119] Based on the trajectory fusion of multiple algorithms, the present invention can effectively filter out the pedestrians misdetected due to the low recognition accuracy of some algorithms by comprehensively analyzing the spatio-temporal information of each algorithm, ensure the relative accuracy of the pedestrian trajectory, and reduce the time investment of users in manual screening.

[0120] The following uses a specific embodiment to illustrate the method for retrieving a person's trajectory based on the application of multi-algorithm fusion in the embodiments of the present invention.

[0121] Set N = 90 days, and the camera positions in key sections are X1, X2, X3... XN. Start the analysis task to detect and track people in the video stream or picture stream, and extract pedestrian feature information, including: face feature information, body feature information, and gait feature information. The feature information of each pedestrian includes at least one of these three types of feature information. Form a passerby database and generate unique person ID information. Among them, N can be adjusted larger or smaller according to the actual situation, and the camera positions to be analyzed are configured according to the local actual situation.

[0122] Suppose a person passes by camera X1 at 2021-6-28 11:01:01, and clear face, body, and gait information can be obtained under this camera.

[0123] Suppose a person passes by camera X2 at 2021-6-28 11:11:01, and body and gait information can be obtained under this camera, but clear face information cannot be obtained.

[0124] Suppose a person passes by camera X3 at 2021-6-28 11:21:01, and face and body information can be obtained under this camera, but complete gait information cannot be obtained due to the short passing time.

[0125] Suppose a person passes by camera X4 at 2021-6-29 9:21:01, and clear face, body, and gait information can be obtained under this camera.

[0126] Store all the obtained feature information corresponding to the above pedestrians. The primary key records the person number, and the specific person information list is as follows:

[0127]

[0128] Retrieve the trajectory of the target person to be retrieved during the time range from 【2021-6-28 00:00:00】 to 【2021-6-29 12:59:59】. Suppose the picture information of the target person to be retrieved corresponds to the record ID1, and retrieve its trajectory information.

[0129] Through the face recognition algorithm, compare the face feature RL1 of person ID1 with the face features of other people in the person list to obtain the person numbers that meet the set threshold. Suppose ID3 meets the threshold requirement, that is, L1 = {ID1, ID3}.

[0130] Through the person re-identification algorithm, the human body features of person ID1 are compared with the human body features of other persons in the person list to obtain the person numbers that meet the set threshold. Suppose ID2 and ID3 meet the requirements, that is, L2 = {ID1, ID2, ID3}. Through the gait recognition algorithm, the gait features of person ID1 are compared with the gait features of other persons in the person list to obtain the person numbers that meet the set threshold. Suppose ID4 meets the threshold requirement, that is, L3 = {ID1, ID4}.

[0131] Take the union of the above three sets to obtain the trajectory L of person ID1, L = L1 ∪ L2 ∪ L3 = {ID1, ID2, ID3, ID4}.

[0132] Furthermore, screen out the abnormal points in the trajectory L, and calculate the average speed of each section of the journey in the trajectory L, that is, v 1,2 、v 2,3 、v 3,4 Suppose the section with an average speed exceeding 30 km / h is selected as the abnormal section. For example, the section {ID2, ID3} where v 2,3 is located.

[0133] Use the probability calculation formula to calculate the recognition probabilities P(ID2) and P(ID3) of ID2 and ID3 respectively, and eliminate the points with lower probabilities, such as ID2. In this way, the final trajectory in L is obtained as L = {ID1, ID3, ID4}.

[0134] Perform GIS map display on the trajectory information L of person ID1: Obtain the detailed information of the persons in the trajectory L = {ID1, ID3, ID4}, including: person ID information, capture time, camera number, camera longitude and latitude coordinates, corresponding person picture information, etc. And sort them in chronological order, and connect each trajectory point in turn. This connecting line is the specific trajectory information of the person.

[0135] Based on the same inventive concept, the embodiment of the present invention also provides a person trajectory retrieval device based on the fusion application of multiple algorithms. Since the principle of the problem solved by this device is similar to the method of the foregoing embodiment, the implementation of this device can refer to the implementation of the foregoing method, and the repeated parts will not be described again.

[0136] As Figure 3 shown, a person trajectory retrieval device based on the fusion application of multiple algorithms provided by the embodiment of the present invention can be used to execute the above method embodiment. The device includes:

[0137] The pedestrian feature information extraction and storage module 11 is used to extract the feature information of each pedestrian in the video stream or picture stream collected by cameras at multiple different positions. The feature information of each extracted pedestrian includes at least one of face feature information, body feature information, and gait feature information. The feature information of each extracted pedestrian, together with the corresponding camera position and capture time, is stored as a pedestrian record in the passerby database.

[0138] The trajectory retrieval module 12 is used to extract the feature information of the retrieved target person according to the video segment or picture of the retrieved target person, and compare the face feature information, body feature information, and gait feature information of the retrieved target person with the corresponding types of pedestrian feature information in the passerby database by using the corresponding face recognition algorithm, person re-identification algorithm, and gait recognition algorithm. Each pedestrian record with feature information consistent with that of the retrieved target person is found, and the union of the retrieved pedestrian records is taken to obtain the trajectory information of the retrieved target person.

[0139] The trajectory display module 13 is used to connect the camera positions corresponding to each pedestrian record on the GIS map according to the capture time based on the trajectory information of the retrieved target person to obtain the trajectory of the retrieved target person.

[0140] In summary, the method and device for retrieving personnel trajectories based on the fusion application of multiple algorithms provided by the present invention can solve the problem of personnel retrieval in complex scenarios such as long distance, improper angle, clothing change, facial occlusion, light influence, low video resolution, far distance, and disguise. The present invention can fuse and apply three algorithms of face recognition, person re-identification, and gait recognition, automatically connect and fuse the trajectories of the same target person, make up for the deficiencies of a single algorithm, make full use of the advantages of each algorithm for fusion application, realize the retrieval of personnel trajectories, and can obtain more complete trajectory information of personnel, expanding the application scenarios.

[0141] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.

[0142] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for retrieving personnel trajectories based on the integration and application of multiple algorithms, characterized in that, Including the following steps: S1. Extract the feature information of each pedestrian in the video stream or picture stream collected by multiple cameras at different positions. The extracted feature information of each pedestrian includes at least one of face feature information, body feature information, and gait feature information. Store the extracted feature information of each pedestrian together with the corresponding camera position and capture time as a pedestrian record in the passerby database; S2. Extract the feature information of the retrieved target person according to the video segment or picture of the retrieved target person. Compare the face feature information, body feature information, and gait feature information of the retrieved target person with the corresponding types of pedestrian feature information in the passerby database using the corresponding face recognition algorithm, person re-identification algorithm, and gait recognition algorithm. Find each pedestrian record with feature information consistent with the retrieved target person, and take the union of the retrieved pedestrian records to obtain the trajectory information of the retrieved target person; S3. According to the trajectory information of the retrieved target person, connect the camera points corresponding to each pedestrian record on the GIS map according to the capture time to obtain the trajectory of the retrieved target person, calculate the average speed of the distance between every two adjacent points in the trajectory of the retrieved target person, select the sections with an average speed exceeding 30 km / h as abnormal sections, calculate the algorithm recognition probabilities of the points on both sides of the abnormal section respectively, select the points with a lower probability as the final abnormal points, and remove them from the trajectory, where the points The algorithm recognition probability The calculation method is as follows: ; Among them respectively represent the face recognition algorithm, the person re-identification algorithm, and the gait recognition algorithm represents the algorithm recognition accuracy represents at the point under the algorithm the probability that the person identified at the point is the target person to be retrieved 2. The method for retrieving personnel trajectories based on the fusion application of multiple algorithms according to claim 1, characterized in that, In step S1, the extraction of face feature information from the video stream or picture stream specifically includes: Use the MTCNN network to perform face target detection on each frame of video data or each picture; Use the trained face feature extraction model to extract face features from the detected face targets to obtain face feature information; Among them, the training method of the face feature extraction model is as follows: Collect face photos in a large number of video surveillance scenarios, perform annotation, and divide them into a training set, a validation set, and a test set; Construct an improved faceNet model. The faceNet backbone network uses the Inception-v4 network and residual connection. In the selection of the loss function, calculate the softmax loss and triplet_loss respectively, and then weight the softmax loss and triplet_loss, and adjust the weights online; Use the training set to train the improved faceNet model, use the validation set to verify the convergence of the model, and perform tests through the test set, and finally output the best face feature extraction model.

3. The method for retrieving personnel trajectories based on the fusion application of multiple algorithms according to claim 2, wherein In step S1, the extraction of face feature information from the video stream also includes: After performing face target detection, use DeepSort to track the detected face targets and perform moving target detection; Use the trained face frontal face determination model to perform face frontal face determination on the face targets output by the tracking and output the determined face frontal face data; Use the trained face feature extraction model to extract face features from the output face frontal face data to obtain face feature information.

4. The method for retrieving personnel trajectories based on the fusion application of multiple algorithms according to claim 1, wherein In step S1, the extraction of body feature information from the video stream or picture stream specifically includes: Use the trained yolov5 target detection model to perform pedestrian target detection on each frame of video data or each picture; Use the trained body feature extraction model to extract body features from the detected pedestrian targets to obtain body feature information; Among them, the training method of the body feature extraction model is as follows: Collect and construct multiple pedestrian sequence datasets in the video surveillance scenario, collect sequences of different pedestrian targets from multiple different camera perspectives, and divide them into training set, validation set and test set; Construct an improved pedestrian FastReID model. The backbone adopts a multi-scale multi-component fusion deep network based on adaptive margin ranking loss, and jointly learns pedestrian image feature representation and similarity measurement; In the data preprocessing stage, downsample the data; Use the training set to train the improved pedestrian FastReID model, use the validation set to verify the convergence of the model, and conduct tests through the test set, and finally output the best human feature extraction model.

5. The method for retrieving personnel trajectories based on the fusion application of multiple algorithms according to claim 4, wherein In step S1, the extraction of human feature information from the video stream further includes: After pedestrian target detection, use DeepSort to track the detected pedestrian targets and conduct moving target detection; Use the trained pedestrian quality assessment model to evaluate the quality of the pedestrian targets output by the tracking, and output the pedestrian target with the best quality as a close-up image; Use the trained human feature extraction model to extract human feature information from the pedestrian target with the best quality output.

6. The method for retrieving personnel trajectories based on the fusion application of multiple algorithms according to claim 1, characterized in that In step S1, the steps for extracting gait feature information from the video stream are as follows: Use the trained yolov5 target detection model to detect pedestrian targets in each frame of video data; Use DeepSort to track the detected pedestrian targets and conduct moving target detection; Use the trained deeplabv3+ segmentation model to segment the gait of the pedestrian sequence output by the tracking and output a black and white contour sequence map; Use the trained gait feature extraction model to extract gait features from the black and white contour sequence map output by the gait segmentation; Among them, the training method of the gait feature extraction model is as follows: Collect and construct multiple pedestrian sequence datasets in the video surveillance scenario, collect sequences of different pedestrian targets from multiple different camera perspectives, and divide them into training set, validation set and test set; Construct an improved GaitSet model; add a CNN module at the beginning of the backbone network, and add MDFG and GMCM modules to the last two layers of the network; In the data preprocessing stage, the training data is augmented by random horizontal flipping and random jigsaw puzzles; Use the training set to train the improved GaitSet model, use the validation set to verify the convergence of the model, and conduct tests through the test set, and finally output the best gait feature extraction model.

7. The method for retrieving personnel trajectories based on the fusion application of multiple algorithms according to claim 1, wherein In step S2, when face feature information comparison and human feature information or gait feature information comparison are both performed in a pedestrian record, the comparison result of the face feature information shall prevail.

8. A personnel trajectory retrieval device based on the fusion application of multiple algorithms, characterized in that Include: The pedestrian feature information extraction and storage module is used to extract the feature information of each pedestrian in the video stream or picture stream collected by cameras at multiple different locations. The feature information of each extracted pedestrian includes at least one of face feature information, body feature information, and gait feature information. The feature information of each pedestrian extracted is stored in the passerby database as a pedestrian record together with the corresponding camera location and the capture time; The trajectory retrieval module is used to extract the feature information of the retrieval target person according to the video segment or picture of the retrieval target person, and compare the face feature information, body feature information, and gait feature information of the retrieval target person with the corresponding types of pedestrian feature information in the passerby database using the corresponding face recognition algorithm, pedestrian re-identification algorithm, and gait recognition algorithm to find each pedestrian record with feature information consistent with that of the retrieval target person, and take the union of the retrieved pedestrian records to obtain the trajectory information of the retrieval target person; The trajectory display module is used to connect the camera locations corresponding to each pedestrian record according to the capture time on the GIS map based on the trajectory information of the retrieval target person to obtain the trajectory of the retrieval target person; It is also used to calculate the average speed of the distance between every two adjacent points in the trajectory of the target person to be retrieved, select the sections with an average speed exceeding 30 km / h as abnormal sections, calculate the algorithm recognition probabilities of the points on both sides of the abnormal section respectively, select the points with lower probabilities as the final abnormal points, and remove them from the trajectory, where the points algorithm recognition probability is calculated as follows: ; Among them respectively represent the face recognition algorithm, the person re-identification algorithm, and the gait recognition algorithm represents the algorithm recognition accuracy represents at the algorithm the probability that the person identified at the point is the target person to be retrieved

Citation Information

Patent Citations

  • Personnel trajectory searching method and system

    CN110942003A

  • Smart park-oriented personnel trajectory analysis method and device, and medium

    CN113269091A