A monitoring method, identification method, related device and system
By performing preliminary gait recognition and screening at the monitoring device end, reducing the amount of data, and then performing accurate recognition at the recognition device end, the problem of monitoring video transmission and recognition latency is solved, and fast and accurate target identification and tracking monitoring is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2018-06-19
- Publication Date
- 2026-04-03
AI Technical Summary
Existing surveillance methods suffer from long latency in cloud servers due to the large amount of video data and high complexity of recognition when performing gait identification, making it impossible to identify targets in sensitive scenarios such as criminal suspects in a timely manner.
By performing preliminary identification on the monitoring video at the monitoring device, video frames containing human figures are selected and gait energy maps are extracted. After reducing the amount of data, the human figure image sequence of the target area is uploaded to the recognition device for accurate identification. Pre-stored features are used for secondary identification to improve accuracy.
It reduces data transmission volume and the burden on recognition devices, improves the speed and accuracy of target person recognition, and supports the tracking and coordinated monitoring of target objects.
Smart Images

Figure CN114049681B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of surveillance, and more particularly to a surveillance method, identification method, related device and system. Background Technology
[0002] Gait is the way people walk and can be used for identification. Compared to biometrics such as fingerprints, faces, or irises, which require contact or close proximity for identification, gait has the characteristic of being able to be identified from a distance. Therefore, gait recognition is an important research direction in the field of surveillance.
[0003] In gait recognition, the monitoring device captures video and sends it to the recognition device, such as a cloud server. The cloud server then performs gait recognition on the received video. Because gait is a dynamic process, the algorithm for gait recognition is highly complex, so a high-performance cloud server is required for its implementation.
[0004] Because the amount of video data transmitted from the monitoring device to the cloud server is large, the transmission process takes a long time. The cloud server also takes a long time to identify the received video. Therefore, there is a large delay when the cloud server performs gait recognition on the video. For time-sensitive monitoring scenarios such as the need to promptly identify criminal suspects based on the monitoring video, the existing monitoring methods are not timely enough. Summary of the Invention
[0005] This application provides a monitoring method, an identification method, related devices, and a system that can reduce data transmission volume and accelerate the identification of target individuals.
[0006] In a first aspect, embodiments of this application provide a monitoring method, the method comprising:
[0007] The surveillance video is subjected to a first identification screening process to select video frames containing human figures, thereby obtaining a first video frame sequence.
[0008] Human silhouette images are extracted from each video frame in the first video frame sequence to obtain a gait energy map;
[0009] The gait energy map is first identified to obtain the similarity between the gait energy map and the preset gait energy map of the target person;
[0010] When the similarity is greater than the second threshold, the first video frame sequence is identified as the target video frame sequence.
[0011] The target video frame sequence includes multiple target video frames;
[0012] A target region human image sequence is determined based on the target video frame sequence, wherein the target region human image sequence includes a region in at least one of the plurality of target video frames that may contain a target person;
[0013] The target area human figure image sequence is uploaded to the recognition device. The target area human figure image sequence is used by the recognition device to determine whether the target person is included in the target area human figure image sequence through a second recognition. The recognition accuracy of the second recognition is greater than the recognition accuracy of the first recognition.
[0014] The technical solution provided in this application embodiment uploads a sequence of human-shaped images of the target area obtained after the first identification to the identification device. This is much smaller than the surveillance video, which reduces the amount of data transmitted to the identification device and alleviates the identification burden on the identification device, thus helping to speed up the identification of the target person by the identification device.
[0015] Based on the first aspect, in some possible embodiments of this application, the step of performing a first identification on the surveillance video to obtain a target video frame sequence includes:
[0016] Perform the first identification on each video frame in the surveillance video to obtain the probability that each video frame in the surveillance video includes the target person;
[0017] The target video frame sequence is formed by combining video frames with a probability greater than a first preset threshold.
[0018] The technical solution provided in this application embodiment limits the target video frame sequence. In this embodiment, each video frame in the surveillance video is identified to obtain the probability that each video frame includes the target person, and the video frames with a corresponding probability greater than a first preset threshold are combined to form the target video frame sequence.
[0019] The technical solution provided in this application embodiment limits the target video frame sequence. In this embodiment, the target video frame sequence is determined by the similarity between the gait energy map obtained from the surveillance video and the gait energy map of the preset target person.
[0020] Based on the first aspect, in some possible embodiments of this application, the method further includes:
[0021] The appearance features of the target person identified by the recognition device are obtained;
[0022] The first identification of the surveillance video to obtain the target video frame sequence includes:
[0023] The target video frame sequence is obtained by performing the first identification on the surveillance video based on the appearance features.
[0024] The technical solution provided in this application embodiment uses the appearance features of the target person identified by the recognition device to identify surveillance video. These appearance features can be clothing features, accessory features, age information, gender information, or hairstyle features, etc. Using the appearance features of the target person identified by the recognition device for initial identification can improve the accuracy of the identification, reduce the amount of data uploaded to the recognition device, alleviate the recognition burden on the recognition device, and improve the recognition efficiency of the recognition device.
[0025] Based on the first aspect, in some possible embodiments of this application, the method further includes:
[0026] Obtain the tracking command sent by the identification device;
[0027] Adjust the monitoring angle according to the tracking instructions.
[0028] The technical solutions provided in this application are beneficial for tracking and monitoring target objects.
[0029] Secondly, embodiments of this application provide an identification method, the method comprising:
[0030] The system acquires a sequence of human-shaped images of a target area uploaded by a monitoring device. The sequence of human-shaped images of the target area is determined by a sequence of target video frames, which includes multiple target video frames. The sequence of target video frames is obtained by the monitoring device through a first identification of the monitoring video. The sequence of human-shaped images of the target area includes an area in at least one of the multiple target video frames that may contain a target person.
[0031] Based on the pre-stored features of the target person, a second identification is performed on the human image sequence of the target area to determine whether the human image sequence of the target area includes the target person. The identification accuracy of the second identification is greater than that of the first identification.
[0032] The technical solution provided in this application embodiment identifies a sequence of human-shaped images of the target area uploaded by the monitoring device. This is much smaller than the monitoring video, which reduces the amount of data transmitted and alleviates the recognition burden on the identification device, thus accelerating the identification of the target person.
[0033] Based on the second aspect, in some possible embodiments of this application, the method further includes:
[0034] A tracking command is sent to the monitoring device, which instructs the monitoring device to adjust the monitoring angle.
[0035] The technical solutions provided in this application are beneficial for tracking and monitoring target objects.
[0036] Based on the second aspect, in some possible embodiments of this application, the pre-stored features of the target person include one or more of the following features: the gait energy map of the target person, the facial features of the target person, the humanoid features of the target person, the skeletal features of the target person, and the optical flow features of the target person.
[0037] Based on the second aspect, in some possible embodiments of this application, when it is determined that the target person is included in the target region human image sequence, the method further includes:
[0038] Motion features of the target person are extracted from the human image sequence of the target region, and the motion features include the direction of movement;
[0039] Predict the activity area of the target person based on the motion characteristics;
[0040] Send monitoring instructions, including the characteristics of the target person, to the monitoring devices within the activity area.
[0041] The technical solution provided in this application embodiment is beneficial for the coordinated monitoring of target objects.
[0042] Based on the second aspect, in some possible embodiments of this application, when it is determined that the target person is included in the target region human image sequence, the method further includes:
[0043] Extract the appearance features of the target person from the human image sequence of the target region;
[0044] The appearance features of the target person are sent to the monitoring device, and the appearance features of the target person are used by the monitoring device to perform the first identification on the monitoring video.
[0045] The technical solution provided in this application, when determining that a target person is included in a human image sequence of the target area, extracts the appearance features of the target person from the human image sequence of the target area, and sends the appearance features of the target person to the monitoring device for first identification. The appearance features can be: clothing features, accessory features, age information, gender information, or hairstyle features, etc. Using the appearance features of the target person for first identification can improve the accuracy of identification, help reduce the amount of data uploaded to the identification device, reduce the identification burden on the identification device, and improve the identification efficiency of the identification device.
[0046] Based on the second aspect, in some possible embodiments of this application, when the pre-stored features of the target person include two or more, the second identification of the human image sequence of the target region based on the pre-stored features of the target person includes:
[0047] The second identification is performed on the human image sequence of the target region by extracting features of two or more pre-stored target persons using the extracted features of the two or more pre-stored target persons.
[0048] The technical solution provided in this application embodiment limits the implementation of the second identification.
[0049] Thirdly, embodiments of this application provide a monitoring device, the monitoring device comprising:
[0050] The first identification unit is used to perform a first identification process on the surveillance video, filter out video frames containing human figures in the surveillance video to obtain a first video frame sequence; extract human figure contour images from each video frame in the first video frame sequence to obtain a gait energy map; perform a first identification on the gait energy map to obtain the similarity between the gait energy map and the gait energy map of the preset target person; when the similarity is greater than a second threshold, confirm the first video frame sequence as the target video frame sequence.
[0051] The target video frame sequence includes multiple target video frames;
[0052] The first processing unit is configured to determine a target region human image sequence based on the target video frame sequence, wherein the target region human image sequence includes a region in at least one of the plurality of target video frames that may contain a target person;
[0053] The first uploading unit is used to upload the target area human figure image sequence to the recognition device. The target area human figure image sequence is used by the recognition device to determine whether the target person is included in the target area human figure image sequence through a second recognition. The recognition accuracy of the second recognition is greater than the recognition accuracy of the first recognition.
[0054] The technical solution provided in this application embodiment uploads a sequence of human-shaped images of the target area obtained after the first identification to the identification device. This is much smaller than the surveillance video, which reduces the amount of data transmitted to the identification device and alleviates the identification burden on the identification device, thus helping to speed up the identification of the target person by the identification device.
[0055] Based on this third aspect, in some possible implementations of this application...
[0056] The first identification unit is specifically used to perform the first identification on each video frame in the surveillance video to obtain the probability that each video frame in the surveillance video includes the target person; and to form the target video frame sequence by combining the video frames with the corresponding probability greater than a first preset threshold.
[0057] The technical solution provided in this application embodiment limits the target video frame sequence. In this embodiment, each video frame in the surveillance video is identified to obtain the probability that each video frame includes the target person, and the video frames with a corresponding probability greater than a first preset threshold are combined to form the target video frame sequence.
[0058] Based on the third aspect, in some possible embodiments of this application, the first identification unit is specifically used to: filter out video frames containing human figures in the surveillance video to obtain a first video frame sequence; extract human figure contour images from each video frame in the first video frame sequence to obtain a gait energy map; perform a first identification on the gait energy map to obtain the similarity between the gait energy map and the gait energy map of the preset target person; when the similarity is greater than a second threshold, confirm the first video frame sequence as the target video frame sequence.
[0059] The technical solution provided in this application embodiment limits the target video frame sequence. In this embodiment, the target video frame sequence is determined by the similarity between the gait energy map obtained from the surveillance video and the gait energy map of the preset target person.
[0060] Based on a third aspect, in some possible embodiments of this application, the monitoring device further includes:
[0061] The first acquisition unit is used to acquire the appearance features of the target person identified by the recognition device;
[0062] When the first identification unit performs a first identification on the surveillance video to obtain a target video frame sequence, it is specifically used to perform a first identification on the surveillance video using the appearance features of the target person.
[0063] The technical solution provided in this application embodiment uses the appearance features of the target person identified by the recognition device to identify surveillance video. These appearance features can be clothing features, accessory features, age information, gender information, or hairstyle features, etc. Using the appearance features of the target person identified by the recognition device for initial identification can improve the accuracy of the identification, reduce the amount of data uploaded to the recognition device, alleviate the recognition burden on the recognition device, and improve the recognition efficiency of the recognition device.
[0064] Based on the third aspect, in some possible embodiments of this application, the first acquisition unit is further configured to acquire the tracking instructions sent by the identification device;
[0065] The monitoring device also includes:
[0066] The first adjustment unit is used to adjust the monitored angle according to the tracking command.
[0067] The technical solutions provided in this application are beneficial for tracking and monitoring target objects.
[0068] The technical solutions provided in this application are beneficial for tracking and monitoring target objects.
[0069] Fourthly, embodiments of this application provide an identification device, the identification device comprising:
[0070] The second acquisition unit is used to acquire a sequence of human figures in a target area uploaded by the monitoring device; the sequence of human figures in the target area is determined by a sequence of target video frames, the sequence of target video frames includes multiple target video frames, the sequence of target video frames is obtained by the monitoring device through a first identification of the monitoring video, and the sequence of human figures in the target area includes an area in at least one of the multiple target video frames that may contain a target person;
[0071] The second recognition unit is used to perform a second recognition on the target area human image sequence based on the pre-stored features of the target person, and to determine whether the target area human image sequence includes the target person. The recognition accuracy of the second recognition is greater than that of the first recognition.
[0072] The technical solution provided in this application embodiment identifies a sequence of human-shaped images of the target area uploaded by the monitoring device. This is much smaller than the monitoring video, which reduces the amount of data transmitted and alleviates the recognition burden on the identification device, thus accelerating the identification of the target person.
[0073] Based on the fourth aspect, in some possible embodiments of this application, the identification device further includes:
[0074] The first sending unit is used to send a tracking command to the monitoring device, the tracking command being used to instruct the monitoring device to adjust the monitoring angle.
[0075] The technical solutions provided in this application are beneficial for tracking and monitoring target objects.
[0076] Based on the fourth aspect, in some possible embodiments of this application, the pre-stored features of the target person include one or more of the following features: the gait energy map of the target person, the facial features of the target person, the humanoid features of the target person, the skeletal features of the target person, and the optical flow features of the target person.
[0077] Based on the fourth aspect, in some possible embodiments of this application, the identification device may further include:
[0078] The second processing unit is configured to, when the second recognition unit determines that the target person is included in the target region human image sequence, extract motion features of the target person from the target region human image sequence, the motion features including a direction of movement; and predict the activity area of the target person based on the motion features.
[0079] The first sending unit is used to send a monitoring instruction, including the characteristics of the target person, to the monitoring device in the activity area.
[0080] The technical solution provided in this application embodiment is beneficial for the coordinated monitoring of target objects.
[0081] Based on the fourth aspect, in some possible embodiments of this application, the identification device may further include:
[0082] The second processing unit is used to extract the appearance features of the target person from the target region human image sequence when the second recognition unit determines that the target region human image sequence includes the target person;
[0083] The first sending unit is used to send the appearance features of the target person to the monitoring device, and the appearance features of the target person are used by the monitoring device to perform a first identification on the monitoring video.
[0084] The technical solution provided in this application, when determining that a target person is included in a human image sequence of the target area, extracts the appearance features of the target person from the human image sequence of the target area, and sends the appearance features of the target person to the monitoring device for first identification. The appearance features can be: clothing features, accessory features, age information, gender information, or hairstyle features, etc. Using the appearance features of the target person for first identification can improve the accuracy of identification, help reduce the amount of data uploaded to the identification device, reduce the identification burden on the identification device, and improve the identification efficiency of the identification device.
[0085] Based on the fourth aspect, in some possible embodiments of this application, when the second identification unit has two or more pre-stored features of the target person, and is used to perform a second identification on the target region human figure image sequence based on the pre-stored features of the target person, it is specifically used to extract two or more pre-stored features of the target person from the target region human figure image sequence, and to identify the target region human figure image sequence based on the extracted two or more pre-stored features of the target person.
[0086] The technical solution provided in this application embodiment limits the implementation method of the second identification.
[0087] Fifthly, embodiments of this application provide a computer-readable storage medium storing a program that, when executed, implements the monitoring method described in the first aspect or any possible implementation thereof.
[0088] Sixthly, embodiments of this application provide a computer-readable storage medium storing a program that, when executed, implements the identification method described in the second aspect or any possible implementation thereof.
[0089] In a seventh aspect, embodiments of this application provide a monitoring device, including: a camera, a communication unit, a processor, a memory, and a bus; wherein,
[0090] The camera is used to acquire surveillance video;
[0091] The communication unit is used to communicate with the identification device;
[0092] The camera, the communication unit, the processor, and the memory are connected via the bus and communicate with each other.
[0093] The memory stores executable program code;
[0094] The processor runs a program corresponding to the executable program code stored in the memory to execute the monitoring method described in the first aspect of this application or any possible implementation thereof.
[0095] Eighthly, embodiments of this application provide an identification device, including: a communication unit, a memory, a processor, and a bus; wherein,
[0096] The communication unit is used to communicate with the monitoring device;
[0097] The memory stores executable program code;
[0098] The processor runs a program corresponding to the executable program code stored in the memory to execute the identification method described in the second aspect of this application or any possible implementation thereof.
[0099] In a ninth aspect, embodiments of this application provide a monitoring system, including a monitoring device and an identification device, wherein the monitoring device is the monitoring device described in the seventh aspect of this application, and the identification device is the identification device described in the eighth aspect of this application.
[0100] The technical solution provided in this application embodiment is that the monitoring device uploads a sequence of human-shaped images of the target area obtained after the first identification to the identification device. This is much smaller than the monitoring video. This reduces the amount of data transmitted from the monitoring device to the identification device and also reduces the identification burden on the identification device, which helps to speed up the identification of the target person by the identification device. Attached Figure Description
[0101] To more clearly illustrate the technical solutions in the embodiments of this application or the background art, the accompanying drawings used in the embodiments of this application or the background art will be described below.
[0102] Figure 1 This is a schematic diagram of the application scenario architecture of a monitoring system provided in an embodiment of this application.
[0103] Figure 2A This is a schematic diagram of the interaction process of a monitoring method provided in an embodiment of this application.
[0104] Figure 2B This is a schematic diagram of a process for obtaining a target video frame sequence provided in an embodiment of this application.
[0105] Figure 2C This is a schematic diagram of a process for obtaining a target video frame sequence provided in an embodiment of this application.
[0106] Figure 3A This is a functional structure diagram of a monitoring device provided in an embodiment of this application.
[0107] Figure 3B This is a functional structure diagram of a monitoring device provided in another embodiment of this application.
[0108] Figure 3C This is a functional structure diagram of a monitoring device provided in another embodiment of this application.
[0109] Figure 4A This is a functional structure diagram of an identification device provided in an embodiment of this application.
[0110] Figure 4B This is a functional structure diagram of an identification device provided in another embodiment of this application.
[0111] Figure 4C This is a functional structure diagram of an identification device provided in another embodiment of this application.
[0112] Figure 5 This is a schematic diagram of the structure of a monitoring device provided in an embodiment of this application.
[0113] Figure 6 This is a schematic diagram of the structure of an identification device provided in an embodiment of this application. Detailed Implementation
[0114] The terminology used in the embodiments of this application is only for explaining specific embodiments of this application and is not intended to limit this application.
[0115] Figure 1 This is a schematic diagram of the application scenario architecture of a monitoring system provided in one embodiment of this application, as shown below. Figure 1 As shown, the monitoring system 100 includes a monitoring device 101 and an identification device 103. The monitoring device 101 transmits data with the identification device 103 via a network 102. The monitoring device 101 is used to acquire monitoring video and perform a first identification on the acquired monitoring video to obtain a target video frame sequence, which includes multiple target video frames. The monitoring device 101 is also used to determine a target area human image sequence based on the target video frame sequence, which includes an area in at least one of the multiple target video frames that may contain a target person. The monitoring device 101 uploads the target area human image sequence to the identification device 103 via the network 102. The identification device 103 performs a second identification on the target area human image sequence based on pre-stored characteristics of the target person to determine whether the target area human image sequence includes the target person. The accuracy of the second identification is greater than the accuracy of the first identification.
[0116] By adopting the technical solution provided in the embodiments of this application, since the target area human image sequence obtained after the first recognition is uploaded to the recognition device, it is much smaller than the surveillance video. This reduces the amount of data transmitted to the recognition device and alleviates the recognition burden of the recognition device, which helps to speed up the recognition of the target person by the recognition device.
[0117] Please see Figure 2A , Figure 2A An interactive flowchart of a monitoring method provided for one embodiment of this application is shown below. Figure 2A As shown, a monitoring method provided in one embodiment of this application may include the following:
[0118] 201. The monitoring device acquires the monitoring video.
[0119] In some possible embodiments of this application, the monitoring device can obtain monitoring instructions from the identification device and acquire monitoring video according to the monitoring instructions. The monitoring instructions may include the characteristics of the target person. This is used by the monitoring device to identify the acquired monitoring video.
[0120] 202. Perform a first identification on the acquired surveillance video to obtain a target video frame sequence, wherein the target video frame sequence includes multiple target video frames.
[0121] In one possible implementation of this application, when step 202 performs a first identification on the acquired surveillance video to obtain the target video frame sequence, it can be done by performing a first identification on each video frame in the surveillance video, such as... Figure 2B As shown, the process of identifying the acquired surveillance video may include the following steps (2021-2022): Wherein,
[0122] 2021. Perform a first identification on each video frame in the surveillance video to obtain the probability that each video frame in the surveillance video includes the target person.
[0123] 2022. The target video frame sequence is formed by combining video frames with a probability greater than a first preset threshold.
[0124] In another possible implementation of this application, when performing a first identification on the acquired surveillance video to obtain the target video sequence in step 202, gait recognition can be used to identify the acquired surveillance video, such as... Figure 2C As shown, the process of identifying the acquired surveillance video may include the following steps 2023-2026, wherein,
[0125] 2023. Filter out video frames containing human figures from the surveillance video to obtain the first video frame sequence.
[0126] In some possible implementations, a human detection model can be used to detect whether each video frame in the surveillance video contains a human figure. The first video frame is obtained from the video frames containing human figures. Specifically, the human detection model can be flexibly selected according to the hardware configuration of the surveillance device. To ensure real-time processing, the human detection model can employ simple image processing methods, such as a deformable parts model (DPM), which performs target matching and classification based on the trained model; or a miniaturized or compressed convolutional neural network (CNN) model, which automatically extracts features from the image through a series of convolutions and nonlinear activation transformations, performs regression calculations on the human figure's position, and achieves end-to-end human detection. Considering the limited computing resources of the surveillance device, the convolutional neural network can be appropriately pruned and compressed to reduce the number of model parameters and computational complexity.
[0127] 2024. Extract the human silhouette image from each video frame in the first video frame sequence to obtain the gait energy map.
[0128] In some possible implementations of this application, a binary contour map of the human figure can be extracted using traditional image processing methods or a compressed CNN human figure segmentation model, and the pixel values of multiple contour maps can be added together to obtain a gait energy image (GEI).
[0129] 2025. Perform a first identification on the gait energy map to obtain the similarity between the gait energy map and the gait energy map of the preset target person.
[0130] In some possible implementations of this application, the canonical correlation analysis (CCA) algorithm can be used to perform projection-based matching between the GEI obtained in step 2024 and the preset GEI feature template of the target person (finding the projection deformation of each of the two sets of data and maximizing the similarity value of the two sets of data), calculate the similarity value, and compare the similarity value with a preset threshold.
[0131] The CCA algorithm can be solved by solving the generalized eigenvalue problem, making it simple and efficient to implement. By setting a threshold, it can filter out most of the video data that does not contain the target person from the monitoring device side.
[0132] 2026. When the similarity is greater than the second threshold, the first video frame sequence is confirmed as the target video frame sequence.
[0133] 203. Determine the target region human figure image sequence based on the target video frame sequence.
[0134] The target region human image sequence includes at least one region in the plurality of target video frames that may contain the target person;
[0135] In some possible implementations of this application, if step 202 identifies that the surveillance video does not include the target video frame sequence, the surveillance video is directly discarded and no further processing is performed. If the identification result in step 202 is that the target video frame sequence is included, the position of the rectangular box is obtained according to the human detection algorithm, the original video frame is cropped, and a rectangular region including the human figure is obtained, which is the target region (ROI) human figure image, thereby obtaining the target region human figure image sequence.
[0136] 204. Upload the human-shaped image sequence of the target area to the recognition device.
[0137] In some possible embodiments of this application, the identification device may be a device that communicates with the monitoring device via a network, and has strong processing capabilities, which can further identify the target area human image sequence sent by the monitoring device and determine whether the target area human image sequence includes the target person.
[0138] 205. The recognition device acquires a sequence of human-shaped images of the target area.
[0139] 206. The recognition device performs a second recognition on the human image sequence of the target area based on the pre-stored features of the target person, to determine whether the human image sequence of the target area includes the target person. The accuracy of the second recognition is greater than the accuracy of the first recognition.
[0140] In a possible implementation of this application, the recognition device can accurately identify human image sequences in a target area. Specifically, the recognition device can execute a gait recognition algorithm based on any of the following dimensions:
[0141] GEI feature extraction and recognition involves using a CNN model to perform pixel-level segmentation on human image sequences. The binary human contour maps segmented by the CNN model are then superimposed and averaged to obtain gait energy maps (GEI features). These GEI features are further fed into another CNN for recognition.
[0142] In some possible implementations of this application, key point feature extraction and recognition can be adopted. Through CNN, the positions of the skeletal key points of the human figure in each frame (such as head, neck, left shoulder, left elbow, left wrist, right shoulder, right elbow, right wrist, left and right hips, left and right knees, or left and right ankles, etc.) are extracted. The positional information of all frames constitutes a time series feature, and the sequence is fed into a recurrent neural network (RNN) to further calculate the key point identity feature vector. Then, the similarity between the feature vector and the saved preset key point identity feature vector is calculated for matching and recognition.
[0143] In some possible implementations of this application, optical flow feature extraction and recognition can be employed. By leveraging the correspondence between pixels in the previous and current frames, optical flow information of objects between adjacent frames can be calculated. Optical flow is the instantaneous velocity of pixel movement, representing the motion information of pixels changing in the time domain within an image sequence. In some possible implementations of this application, a CNN can be used to directly estimate the optical flow information map between frames from the image sequence. This optical flow information map is then fed into another CNN to calculate the optical flow identity feature vector. Finally, matching and recognition are performed by calculating the similarity between the feature vector and a pre-stored preset optical flow identity feature vector.
[0144] To improve the recognition accuracy of human image sequences in the target area, in some possible embodiments of this application, a gait recognition method that fuses multiple recognition methods can be used for recognition. Specifically, features from three dimensions—human features (GEI features), key point features (such as skeletal key points), and motion features (optical flow features)—can be fused together for gait recognition.
[0145] In practical implementation, the fusion method can include two approaches. One approach is to fuse only at the recognition result level, i.e., executing gait recognition algorithms for multiple different dimensional features separately and then weighting the similarity of each feature. The other approach is to use the same convolutional neural network to extract common low-level features across the three dimensions during the feature extraction stage. These common low-level features are then further input into three independent feature extraction modules to obtain human contour maps, key point locations, and optical flow features, respectively. Subsequently, gait recognition algorithms for each different dimensional feature are executed separately, and the similarity of each feature is weighted and averaged. After passing through the common low-level feature extraction module, the gait features of each dimension are no longer completely independent. The correlation and constraints between features (e.g., known human contours can help more accurately locate key points, and vice versa) can be utilized to improve the effectiveness of each feature extraction, leading to more accurate recognition. After the convolutional neural network shares the output of the above features, it further performs identification in the following ways: The GEI image is input into the second CNN to calculate its similarity s1 with a preset GEI feature image; a time series composed of keypoint location information from multiple frames is input into the RNN to calculate the keypoint identity feature vector, and the similarity s2 between this feature vector and a preset keypoint identity feature vector; the optical flow image is input into the third CNN to calculate the optical flow identity feature vector, and the similarity s3 between this feature vector and a preset optical flow identity feature vector. The final similarity S is obtained by weighted averaging of s1, s2, and s3, where the weights of s1, s2, and s3 are w1, w2, and w3, respectively. Therefore, S = w1*s1 + w2*s2 + w3*s3, where (w1 + w2 + w3 = 1).
[0146] In some possible embodiments of this application, when the recognition device determines that the target person is included in the human image sequence of the target region, the recognition device can make a decision, such as... Figure 2A As shown, the method may further include step 207: sending a tracking command to the monitoring device, instructing the monitoring device to adjust the monitoring angle, such as instructing the monitoring device to rotate clockwise by 10 degrees per second. Step 208: the monitoring device receives the tracking command sent by the identification device. Step 209: the monitoring device adjusts the monitoring angle according to the tracking command, such as rotating clockwise by 10 degrees per second.
[0147] In some possible embodiments of this application, the identification device may also issue monitoring instructions to other related monitoring devices, such as extracting motion features including the direction of movement of the target person from the human image sequence of the target area, predicting the activity area of the target person based on the motion features, and sending monitoring instructions including the features of the target person to the monitoring devices in the activity area to jointly monitor the target person.
[0148] In some possible implementations of this application, if the similarity of the identification result is greater than the second threshold, the identification device can send the target area human figure map containing real-time information such as the target person's clothing attributes and the target person's facial features to the monitoring devices in the adjacent area, and work together to focus on monitoring, tracking and identifying the target person. (For example, using a camera with human attribute recognition capabilities for attribute-based identification and tracking, or using a camera with facial recognition capabilities for facial capture, identification and tracking of the target person, etc.).
[0149] In some possible embodiments of this application, when the recognition device determines that the target area human image sequence includes a target person, the recognition device can also extract the appearance features of the target person from the target area human image sequence. The appearance features of the target person may include: clothing features, accessory features, age information, gender information, or hairstyle features, etc. The recognition device can send the appearance features to a monitoring device for the monitoring device to recognize the surveillance video.
[0150] Please see Figure 3A , Figure 3A This is a schematic diagram of a monitoring device provided in one embodiment of this application. The monitoring device 300 includes:
[0151] The first identification unit 301 is used to perform a first identification on the surveillance video to obtain a target video frame sequence, the target video frame sequence including multiple target video frames;
[0152] The first processing unit 302 is configured to determine a target region human image sequence based on the target video frame sequence, wherein the target region human image sequence includes at least one region in the plurality of target video frames that may contain a target person.
[0153] The first uploading unit 303 is used to upload a sequence of human figures in a target area to a recognition device. The sequence of human figures in the target area is used by the recognition device to determine whether the target person is included in the sequence of human figures in the target area through a second recognition. The recognition accuracy of the second recognition is greater than that of the first recognition.
[0154] In some possible implementations of this application, the first identification unit 301 is specifically used to perform the first identification on each video frame in the surveillance video to obtain the probability that each video frame in the surveillance video includes the target person; and to form the target video frame sequence by combining the video frames with corresponding probabilities greater than a first preset threshold.
[0155] In some possible implementations of this application, the first identification unit 301 is specifically used to: filter out video frames containing human figures in the surveillance video to obtain a first video frame sequence; extract human contour images from each video frame in the first video frame sequence to obtain a gait energy map; perform a first identification on the gait energy map to obtain the similarity between the gait energy map and the gait energy map of the preset target person; and when the similarity is greater than a second threshold, confirm the first video frame sequence as the target video frame sequence.
[0156] In some possible implementations of this application, such as Figure 3B As shown, the monitoring device 300 may further include:
[0157] The first acquisition unit 304 is used to acquire the appearance features of the target person identified by the recognition device; the appearance features of the target person are the features of the target person extracted from the human image sequence of the target region when the recognition device identifies that the target person is included in the human image sequence of the target region.
[0158] When the first identification unit 301 performs first identification on the surveillance video to obtain the target video frame sequence, it is specifically used to perform first identification on the surveillance video using the appearance features of the target person.
[0159] In some possible implementations of this application, such as Figure 3C As shown, the first acquisition unit 304 is also used to acquire tracking instructions sent by the identification device.
[0160] The monitoring device 300 may further include: a first adjustment unit 305, used to adjust the monitoring angle of the monitoring device according to the tracking command.
[0161] Please see Figure 4A , Figure 4A This is a schematic diagram of the structure of an identification device provided in one embodiment of this application. The identification device 400 may include:
[0162] The second acquisition unit 401 is used to acquire a human image sequence of a target area uploaded by the monitoring terminal; the human image sequence of the target area is determined by a target video frame sequence, the target video frame sequence includes multiple target video frames, the target video frame sequence is obtained by the monitoring terminal performing a first identification on the monitoring video, and the human image sequence of the target area includes an area in at least one of the multiple target video frames that may contain a target person.
[0163] The second recognition unit 402 is used to perform a second recognition on the target area human image sequence based on the pre-stored features of the target person, and to determine whether the target area human image sequence includes the target person. The recognition accuracy of the second recognition is greater than that of the first recognition.
[0164] In some possible implementations of this application, such as Figure 4B As shown, the identification device 400 may further include:
[0165] The first sending unit 403 is used to send a tracking instruction to the monitoring terminal, the tracking instruction being used to instruct the monitoring terminal to adjust the monitoring angle.
[0166] In some possible implementations of this application, the pre-stored features of the target person may include one or more of the following features: the target person's gait energy map, the target person's facial features, the target person's humanoid features, the target person's skeletal features, and the target person's optical flow features, etc.
[0167] In some possible implementations of this application, such as Figure 4C As shown, the identification device 400 may further include:
[0168] The second processing unit 404 is configured to, when the second recognition unit determines that the target person is included in the target region human image sequence, extract motion features of the target person from the target region human image sequence, the motion features including a direction of movement; and predict the activity area of the target person based on the motion features.
[0169] The first sending unit 403 is used to send a monitoring instruction including the characteristics of the target person to the monitoring device in the activity area.
[0170] In some possible implementations of this application, such as Figure 4C As shown, the identification device 400 may further include:
[0171] The second processing unit 404 is used to extract the appearance features of the target person from the target region human image sequence when the second recognition unit 402 determines that the target region human image sequence includes the target person.
[0172] The first sending unit 403 is used to send the appearance features of the target person to the monitoring terminal, and the appearance features of the target person are used by the monitoring terminal to perform a first identification on the monitoring video.
[0173] In some possible embodiments of this application, when the second identification unit 402 has two or more pre-stored features of the target person and is used to perform a second identification on the target region human figure image sequence based on the pre-stored features of the target person, it is specifically used to extract two or more pre-stored features of the target person from the target region human figure image sequence, and to identify the target region human figure image sequence based on the extracted two or more pre-stored features of the target person.
[0174] This application also provides a computer-readable storage medium storing a program thereon, which, when the program is run, implements the monitoring method executed by any of the preceding embodiments of the monitoring device.
[0175] This application also provides a computer-readable storage medium storing a program thereon, which, when run, implements the identification method executed by any of the preceding embodiments of the identification device.
[0176] This application also provides a monitoring device, such as... Figure 5 As shown, the monitoring device 500 includes: a camera 501, a communication unit 502, a processor 503, a memory 504, and a bus 505; wherein, the camera 501 is used to acquire monitoring video; the communication unit 502 is used to communicate with the identification device; the camera 501, the communication unit 502, the processor 503, and the memory 504 are connected through the bus 505 and complete communication between them; the memory 504 stores executable program code; the processor 503 reads the executable program code stored in the memory 504 to run the program corresponding to the executable program code, so as to execute the monitoring method executed by the monitoring device in any of the preceding embodiments.
[0177] This application also provides an identification device, such as... Figure 6 As shown, the identification device 600 includes: a communication unit 601, a memory 602, a processor 603, and a bus 604; wherein, the communication unit 601 is used to communicate with a monitoring device; the memory 602 stores executable program code; the processor 603 reads the executable program code stored in the memory 602 to run a program corresponding to the executable program code, so as to execute the identification method performed by the identification device in any of the preceding embodiments.
[0178] This application also provides a monitoring system, including a monitoring device and an identification device. The monitoring device can be any of the monitoring devices described in the preceding embodiments, and the identification device can be any of the identification devices described in the preceding embodiments.
[0179] Those skilled in the art will recognize that the functions described in the embodiments of the present invention in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transmission of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0180] The above detailed embodiments further illustrate the purpose, technical solution, and beneficial effects of the embodiments of the present invention. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the embodiments of the present invention. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solutions of the embodiments of the present invention should be included within the scope of protection of the embodiments of the present invention.
Claims
1. A monitoring method, characterized in that, Applied to a monitoring device, the method includes: A human detection model is used to perform a first identification process on the surveillance video to filter out video frames containing human figures, thus obtaining a first video frame sequence; wherein, the human detection model is determined based on the hardware configuration of the surveillance device; Human silhouette images are extracted from each video frame in the first video frame sequence to obtain a gait energy map; The gait energy map is first identified to obtain the similarity between the gait energy map and the gait energy map of a preset target person; When the similarity is greater than the second threshold, the first video frame sequence is identified as the target video frame sequence, and the target video frame sequence includes multiple target video frames. A target region human image sequence is determined based on the target video frame sequence, wherein the target region human image sequence includes a region in at least one of the plurality of target video frames that may contain a target person; The target area human figure image sequence is uploaded to the recognition device. The target area human figure image sequence is used by the recognition device to determine whether the target person is included in the target area human figure image sequence through a second recognition. The recognition accuracy of the second recognition is greater than that of the first recognition. The second recognition integrates at least one dimensional feature gait recognition method. The second gait recognition method that integrates at least one dimension feature includes: using the same convolutional neural network to extract common low-level features during the feature extraction stage; inputting the extracted common low-level features into multiple independent feature extraction modules to obtain multiple features of different dimensions; and executing gait recognition algorithms for multiple features of different dimensions to perform a weighted average of the similarity corresponding to each feature.
2. The method according to claim 1, characterized in that, The first identification of the surveillance video to obtain the target video frame sequence includes: Perform the first identification on each video frame in the surveillance video to obtain the probability that each video frame in the surveillance video includes the target person; The target video frame sequence is formed by combining video frames with a probability greater than a first preset threshold.
3. The method according to claim 1, characterized in that, The method further includes: The appearance features of the target person identified by the recognition device are obtained; The step of performing a first identification on the surveillance video to obtain the target video frame sequence includes: performing the first identification on the surveillance video based on the appearance features to obtain the target video frame sequence.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Obtain the tracking command sent by the identification device; Adjust the monitoring angle according to the tracking instructions.
5. A monitoring method, characterized in that, Applied to a monitoring system, the monitoring system including a monitoring device and an identification device, the method includes: The monitoring device uses a human detection model to perform a first recognition process on the monitoring video, filtering out video frames containing human figures to obtain a first video frame sequence; wherein, the human detection model is determined based on the hardware configuration of the monitoring device; The monitoring device extracts a human silhouette image from each video frame in the first video frame sequence to obtain a gait energy map; The monitoring device performs a first identification on the gait energy map to obtain the similarity between the gait energy map and the gait energy map of a preset target person; When the similarity is greater than the second threshold, the monitoring device confirms the first video frame sequence as the target video frame sequence, and the target video frame sequence includes multiple target video frames; The monitoring device determines a human image sequence in the target area based on the target video frame sequence, wherein the human image sequence in the target area includes an area in at least one of the plurality of target video frames that may contain a target person; The monitoring device uploads the sequence of human-shaped images of the target area to the recognition device; The identification device acquires a sequence of human-shaped images of the target area uploaded by the monitoring device; the target video frame sequence includes multiple target video frames. The recognition device performs a second recognition on the human image sequence of the target area based on the pre-stored features of the target person, and determines whether the human image sequence of the target area includes the target person. The recognition accuracy of the second recognition is greater than that of the first recognition. The second recognition integrates gait recognition methods with at least one dimension of features. The second gait recognition method that integrates at least one dimension feature includes: using the same convolutional neural network to extract common low-level features during the feature extraction stage; inputting the extracted common low-level features into multiple independent feature extraction modules to obtain multiple features of different dimensions; and executing gait recognition algorithms for multiple features of different dimensions to perform a weighted average of the similarity corresponding to each feature.
6. The method according to claim 5, characterized in that, The method further includes: The identification device sends a tracking command to the monitoring device, which instructs the monitoring device to adjust the monitoring angle.
7. The method according to claim 5, characterized in that, The pre-stored features of the target person include one or more of the following features: the target person's gait energy map, the target person's facial features, the target person's humanoid features, the target person's skeletal features, and the target person's optical flow features.
8. The method according to claim 5, characterized in that, When it is determined that the target person is included in the human image sequence of the target region, the method further includes: Motion features of the target person are extracted from the human image sequence of the target region, and the motion features include the direction of movement; Predict the activity area of the target person based on the motion characteristics; Send monitoring instructions, including the characteristics of the target person, to the monitoring devices within the activity area.
9. The method according to claim 5, characterized in that, The method further includes: when it is determined that the target person is included in the human image sequence of the target region, the method further includes: Extract the appearance features of the target person from the human image sequence of the target region; The physical characteristics of the target person are sent to the monitoring device, and the physical characteristics of the target person are used by the monitoring device to perform the first identification on the monitoring video.
10. The method according to any one of claims 5-9, characterized in that, When the pre-stored features of the target person include two or more, the second identification of the human-shaped image sequence of the target region based on the pre-stored features of the target person includes: The second identification is performed on the human image sequence of the target region by extracting features of two or more pre-stored target persons using the extracted features of the two or more pre-stored target persons.
11. A monitoring device, characterized in that, The monitoring device includes: A first identification unit is configured to perform a first identification process on the surveillance video using a human detection model, filtering out video frames containing human figures to obtain a first video frame sequence; wherein the human detection model is determined based on the hardware configuration of the surveillance device; extracting human contour images from each video frame in the first video frame sequence to obtain a gait energy map; performing a first identification on the gait energy map to obtain the similarity between the gait energy map and the gait energy map of a preset target person; when the similarity is greater than a second threshold, confirming the first video frame sequence as a target video frame sequence, wherein the target video frame sequence includes multiple target video frames; The first processing unit is configured to determine a target region human image sequence based on the target video frame sequence, wherein the target region human image sequence includes a region in at least one of the plurality of target video frames that may contain a target person; The first uploading unit is used to upload the target area human image sequence to the recognition device. The target area human image sequence is used by the recognition device to determine whether the target person is included in the target area human image sequence through a second recognition. The recognition accuracy of the second recognition is greater than that of the first recognition. The second recognition integrates at least one dimensional feature gait recognition method. The second gait recognition method that integrates at least one dimension feature includes: using the same convolutional neural network to extract common low-level features during the feature extraction stage; inputting the extracted common low-level features into multiple independent feature extraction modules to obtain multiple features of different dimensions; and executing gait recognition algorithms for multiple features of different dimensions to perform a weighted average of the similarity corresponding to each feature.
12. The monitoring device according to claim 11, characterized in that, The first identification unit is specifically used to perform the first identification on each video frame in the surveillance video to obtain the probability that each video frame in the surveillance video includes the target person; and to form the target video frame sequence by combining the video frames with the corresponding probability greater than a first preset threshold.
13. The monitoring device according to claim 11, characterized in that, The monitoring device also includes: The first acquisition unit is used to acquire the appearance features of the target person identified by the recognition device; When the first identification unit performs a first identification on the surveillance video to obtain a target video frame sequence, it is specifically used to perform a first identification on the surveillance video using the appearance features of the target person.
14. The monitoring device according to claim 13, characterized in that, The first acquisition unit is further configured to acquire tracking instructions sent by the identification device; The monitoring device also includes: The first adjustment unit is used to adjust the monitored angle according to the tracking command.
15. An identification system, the identification system comprising a monitoring device and an identification device, characterized in that, The monitoring device includes: A first identification unit is configured to perform a first identification process on the surveillance video using a human detection model, filtering out video frames containing human figures to obtain a first video frame sequence; wherein the human detection model is determined based on the hardware configuration of the surveillance device; extracting human contour images from each video frame in the first video frame sequence to obtain a gait energy map; performing a first identification on the gait energy map to obtain the similarity between the gait energy map and the gait energy map of a preset target person; when the similarity is greater than a second threshold, confirming the first video frame sequence as a target video frame sequence, wherein the target video frame sequence includes multiple target video frames; The first processing unit is configured to determine a target region human image sequence based on the target video frame sequence, wherein the target region human image sequence includes a region in at least one of the plurality of target video frames that may contain a target person; The first uploading unit is used to upload the human image sequence of the target area to the recognition device; The identification device includes: The second acquisition unit is used to acquire a sequence of human-shaped images of the target area uploaded by the monitoring device; the target video frame sequence includes multiple target video frames. The second recognition unit is used to perform a second recognition on the human image sequence of the target area based on the pre-stored features of the target person, and to determine whether the human image sequence of the target area includes the target person. The recognition accuracy of the second recognition is greater than that of the first recognition. The second recognition integrates at least one dimensional feature gait recognition method. The second gait recognition method that integrates at least one dimension feature includes: using the same convolutional neural network to extract common low-level features during the feature extraction stage; inputting the extracted common low-level features into multiple independent feature extraction modules to obtain multiple features of different dimensions; and executing gait recognition algorithms for multiple features of different dimensions to perform a weighted average of the similarity corresponding to each feature.
16. The identification system according to claim 15, characterized in that, The identification device also includes: The first sending unit is used to send a tracking command to the monitoring device, the tracking command being used to instruct the monitoring device to adjust the monitoring angle.
17. The identification system according to claim 15, characterized in that, The pre-stored features of the target person include one or more of the following features: the target person's gait energy map, the target person's facial features, the target person's humanoid features, the target person's skeletal features, and the target person's optical flow features.
18. The identification system according to claim 15, characterized in that, The identification device also includes: The second processing unit is configured to, when the second recognition unit determines that the target person is included in the target region human image sequence, extract motion features of the target person from the target region human image sequence, the motion features including a direction of movement; and predict the activity area of the target person based on the motion features. The first sending unit is used to send a monitoring instruction, including the characteristics of the target person, to the monitoring device in the activity area.
19. The identification system according to claim 15, characterized in that, The identification device further includes: The second processing unit is used to extract the appearance features of the target person from the target region human image sequence when the second recognition unit determines that the target region human image sequence includes the target person; The first sending unit is used to send the appearance features of the target person to the monitoring device, and the appearance features of the target person are used by the monitoring device to perform a first identification on the monitoring video.
20. The identification system according to any one of claims 15-19, characterized in that, When the second recognition unit has two or more pre-stored features of the target person, and is used to perform a second recognition on the target region human figure image sequence based on the pre-stored features of the target person, it is specifically used to extract two or more pre-stored features of the target person from the target region human figure image sequence, and to recognize the target region human figure image sequence based on the extracted two or more pre-stored features of the target person.
21. A computer-readable storage medium, characterized in that, The storage medium stores a program that, when executed, implements the monitoring method as described in any one of claims 1-4.
22. A monitoring device, characterized in that, include: Camera, communication unit, processor, memory, and bus; among which, The camera is used to acquire surveillance video; The communication unit is used to communicate with the identification device; The camera, the communication unit, the processor, and the memory are connected via the bus and communicate with each other. The memory stores executable program code; The processor runs a program corresponding to the executable program code stored in the memory to perform the monitoring method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Method for extracting and identifying small sample character contour feature
CN101635031A
Large-scale distributed monitoring video data processing method and device
CN107071344A