Target filtering and merging processing method in short-time domain based on face video structured data

By combining YOLO and SORT algorithms, faces in video streams are quickly detected, identified and tracked, which solves the problems of invalid faces, small faces and repeated faces, and reduces the system memory footprint, achieving fast and efficient face target processing.

CN119992616APending Publication Date: 2025-05-13WUHAN DAQIAN INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411836447.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the process of face detection, recognition and target tracking in video streaming/video files, there are problems such as invalid face target detection, small face target detection, face target repetition, and slow post-processing of face video structured and excessive memory usage.

Method used

Using the method based on YOLO object detection algorithm and SORT face target tracking algorithm, video stream/video files are decoded and image output, and feature vector extraction and trajectory merging are used to quickly filter invalid faces, face classification and trajectory merging, reducing computing resource occupation.

Benefits of technology

In real-time or near-real-time scenarios, quickly detect, identify and track faces in video streams, effectively filter invalid faces and small faces, reduce face target repetition, and reduce system memory usage to avoid system lag.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992616A_ABST
    Figure CN119992616A_ABST
Patent Text Reader

Abstract

The invention relates to a target filtering and merging processing method in a short time domain based on face video structured data, and the method comprises the steps: firstly detecting a face target in an output image of a video file, and extracting a feature vector of the face target; tracking the face target according to the frame position and size of the face target, extracting the trajectory of the face target, and outputting a tracking result; reprocessing the face target according to the tracking result; and finally, carrying out similarity comparison on the reprocessed face target and the feature vector of the face target, attributing similar faces to the same face group, carrying out merging sorting on tracks, and attributing dissimilar faces to a new face group. According to the method, in a real-time or near-real-time scene, the face in the video stream / video file is rapidly detected, recognized and tracked, the method focuses on instant response to the target in the dynamic video, and the method does not need to depend on accumulation and analysis of multi-point data in a long time period.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the processing of face information in a video, and specifically refers to a method for filtering and merging targets in a short time domain based on structured data of face videos, and belongs to the field of public security and intelligent monitoring. Background Art

[0002] Currently, there are the following technical problems in the rapid detection, recognition and target tracking of faces in video streams / video files:

[0003] 1. Some invalid face targets are detected. Specifically, some faces with aspect ratios greater than 1.6 and less than 0.6 are detected, but these faces are actually invalid faces; faces are detected near the image boundary with only one track, but these faces are actually invalid faces; faces with yaw angles greater than or equal to 60 and pitch angles greater than or equal to 20 are detected, but these faces are not frontal faces.

[0004] 2. Small faces are detected. For example, some faces are detected when their width is less than 30 and their height is less than 30. In real-time or near-real-time scenarios, such small faces generally do not exist.

[0005] 3. Repeated face targets. There are many face targets with only one track, and some of these targets are repeated. Similar faces appear repeatedly and are not merged into the same face.

[0006] 4. After the face video is structured, the processing speed of the face target is slower than the speed of calling back the face target after face structuring. Over time, it will cause the computer system to occupy too much memory.

[0007] Each face target has an original image data. The original image corresponding to the face in the face target attribute is relatively large, and only uses memory storage to facilitate the subsequent fast image clipping. When the face target reprocessing speed is slower than the face structure callback face target time, the amount of data is too large, which will cause the computer system memory to occupy too much, causing the system to freeze. Summary of the invention

[0008] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a short-time domain target filtering and merging processing method based on structured data of face video. The method first decodes the video stream / video file and outputs an image at a certain interval, then uses the YOLO target detection algorithm to locate all faces in the image, aligns the detected faces, standardizes the facial posture and size, and extracts feature vectors from the aligned face image; then SORT is used to track the face target, extract the face target trajectory, output the face target, and finally reprocess the output face target to achieve target merging. The method of the present invention quickly detects, identifies and tracks faces in video streams / video files in real-time or near real-time scenarios, focusing on immediate response to targets in dynamic videos without relying on the accumulation and analysis of multi-point data over a long period of time.

[0009] The technical solution used to achieve the purpose of the present invention is: a method for filtering and merging targets in a short-time domain based on structured face video data, the method comprising:

[0010] S1, detecting the face target in the output image of the video file and extracting the feature vector of the face target;

[0011] S2, tracking the face target according to the frame position and size of the face target, extracting the trajectory of the face target, and outputting the tracking result;

[0012] S3, reprocessing the face target according to the tracking result, including filtering invalid face targets and filtering the smallest face target;

[0013] S4. Compare the similarity between the reprocessed face target and the feature vector of the face target, assign similar faces to the same face group, merge and sort the trajectories, and assign dissimilar faces to a new face group.

[0014] In the current technology, the face video structuring and target tracking methods rely on the accumulation and analysis of multi-point data in a long period of time. Although faces and tracking targets can be detected, the amount of calculation is large, the resources are high, and the time is long. The present invention makes full use of the behavior and changes of the face target frame position and size and feature data on a relatively short time scale, adopts the SORT real-time multi-target tracking algorithm, which combines YOLOv5 target detection and Kalman filter data prediction, and solves the identity matching problem in target tracking through the Hungarian algorithm. At the same time, frame skipping decoding video is adopted to solve the problem of large computing resource occupation. Some invalid targets are filtered out by conditions such as aspect ratio, yaw angle, and effective range of pitch angle. According to the conditions of feature similarity, frame difference, and the number of times the same face group is not found, the situation where the same target is classified into multiple targets is classified and merged. The present invention makes full use of the behavior and changes of face feature data on a relatively short time scale to achieve rapid filtering of invalid faces, face classification, and trajectory merging. It can provide great help for more extensive and rapid capture of face targets and whereabouts retrieval, and achieve the effect of quickly locating the whereabouts of people. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 The present invention is a flow chart of a method for filtering and merging targets in a short time domain based on structured face video data.

[0016] Figure 2 Flowchart of face detection and feature extraction.

[0017] Figure 3 Flowchart of the target tracking SORT algorithm.

[0018] Figure 4 Flowchart for filtering invalid face targets.

[0019] Figure 5 Flowchart for face object merging. DETAILED DESCRIPTION

[0020] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] like Figure 1 As shown, the present invention is based on the short-time domain target filtering and merging processing method of face video structured data, comprising the following steps:

[0022] S1, face target detection, feature extraction, such as Figure 2 As shown, the following operations are included:

[0023] S1.1. Decode the video file and output an image at a certain interval (for example, at an interval of 5 frames). The interval can be a number of frames selected between 1-15 frames, and in this embodiment, an interval of 5 frames is selected.

[0024] S1.2. Use the YOLOv5 detection algorithm to locate all faces in the image.

[0025] S1.3. Align the detected faces and standardize facial pose and size to facilitate feature extraction.

[0026] S1.4. Extract feature vectors (attributes, position, size, confidence, etc.) from the aligned face images.

[0027] S2, face target trajectory tracking

[0028] SORT is an efficient and popular target tracking algorithm, especially suitable for tracking multiple targets in real-time surveillance video. SORT combines the Kalman filter to predict the target state and the Hungarian algorithm to associate the detection results with the predicted target, thereby achieving continuous tracking and identity preservation of the target.

[0029] SORT uses the position and size of the target box to estimate the motion of the target and associate data. It uses the IOU (intersection-over-union) between the predicted position of the target in the current frame and the target detection box in the current frame. When the IOU is lower than a certain threshold, the existing track is marked as unmatched and a new track is started when a new detection arrives.

[0030] The present invention uses SORT to track human face targets and extracts the trajectory of human face targets. KalmanBoxTracker is a core component used in the SORT algorithm to predict and track targets. It uses a Kalman filter to predict the next state of the target and update the target's location information. The present invention uses the workflow of KalmanBoxTracker in SORT. Figure 3 As shown, the following steps are included:

[0031] S2.1. Initialization

[0032] When a new object is detected, a KalmanBoxTracker instance is created and its state is initialized to the position and velocity of the detected object.

[0033] S2.2. State prediction

[0034] KalmanBoxTracker estimates the position of the target in the next frame using the prediction step of the Kalman filter, which includes state transfer and noise addition.

[0035] S2.3. Measurement Update

[0036] If the object is detected in the next frame, the KalmanBoxTracker uses the Kalman filter update step to adjust the predicted state of the object to match the actual detected position.

[0037] S2.4. Status Update

[0038] After the update, the state of the KalmanBoxTracker is corrected to reflect the latest position and velocity information of the target.

[0039] S2.5 Prediction and Matching

[0040] In the next frame, the predicted state of KalmanBoxTracker is used to match the new detection result. If the match is successful, the state is updated; if it is not matched, it may mean that the target is occluded or disappeared. The predicted target position is matched with the detection result of the current frame using the Hungarian algorithm. The Hungarian algorithm can associate the detection result with the predicted target in an optimal way, even if the detection result is not completely accurate.

[0041] S2.6. Time Update

[0042] If the KalmanBoxTracker does not match a detection result within a certain period of time, it will be removed from the tracking list, and it is considered that the target has left the field of view or is blocked.

[0043] S2.7. Identity maintenance

[0044] Assigning a unique ID to each target allows its identity to remain unchanged even if the target disappears temporarily and then reappears, which is crucial for continuous tracking of the target.

[0045] S2.8, Output

[0046] When the target is considered to have left the field of view or is blocked, the target tracking result is output. The tracking result includes: bounding box position, target unique ID, frame number, original image corresponding to the face, face position (left, top, right, bottom), confidence, yaw angle, pitch angle, features, and trajectory data.

[0047] S3. Reprocess the face target according to the tracking result output by S2.8, including filtering invalid face targets and filtering the smallest face target. The specific operations are as follows:

[0048] S3.1、Reprocessing of face targets to filter out invalid face targets, such as Figure 4 As shown, the following operations are included:

[0049] S3.1.1. Filter invalid faces by aspect ratio. If the aspect ratio is in the range of [0.6, 1.6], it is considered a valid face; otherwise, it is an invalid face.

[0050] S3.1.2. The validity of a face is determined by key points (five key points on the face: the center of the left eye, the center of the right eye, the tip of the nose, the center of the mouth, and the center of the chin). All five key points must meet the following requirements: the x coordinate of the point must be smaller than the width, and the y coordinate of the point must be smaller than the height; otherwise, the face is invalid.

[0051] S3.1.3 Determination of yaw angle and pitch angle.

[0052] If the yaw angle or pitch angle is equal to nan or -nan, it is considered an invalid face. When the result of an operation is too large or too small to be represented by a floating point number, nan or -nan is used to represent it.

[0053] If the yaw angle is less than 60 and the pitch angle is less than 20, it is considered a valid face, otherwise it is an invalid face.

[0054] S3.1.4 When there is only one face target track, and the target frame distance is less than 10 from the frame, if the aspect ratio is greater than 1.2, it is considered an invalid face.

[0055] S3.2. Face target reprocessing: filtering the smallest face target

[0056] In real-time or near-real-time scenarios, the minimum face size ranges from 30x30 pixels to 60x60 pixels. However, this is not a fixed standard and the specific value needs to be adjusted according to the following factors:

[0057] Camera parameters: The camera's resolution, focal length, and distance to the subject will affect the size of the face in the image.

[0058] Application scenarios: Different application scenarios have different requirements for the minimum face range. For example, in a surveillance camera, it may be necessary to detect a small-sized face at a long distance; while in a mobile phone selfie scenario, the face may occupy most of the screen, and the minimum size can be set larger.

[0059] When creating a target filtering and merging task, enter the minimum face width and height parameters. Usually, the minimum face width is set to 30 and the minimum face height is set to 30. During the face target processing, face target filtering is performed based on this parameter. If the face width and height are both less than the minimum value, it is considered an invalid face; otherwise, it is a valid face.

[0060] S4. Merging of face targets.

[0061] For the reprocessed face targets, similar faces are assigned to the same face group (primary face target, secondary face target list, maximum feature similarity score, target unique ID value corresponding to the maximum score, number of times no similar face group is found), and the trajectories are merged and sorted. Dissimilar faces are assigned to a new face group. When the number of times no similar face group is found is greater than the threshold, the face group is removed. Figure 5 As shown, the following steps are included:

[0062] Compare the new face target with all the faces in all the existing face groups, specifically: compare the similarity of the new face features with all the face features in all the existing face groups (the feature vector extracted from the face image in S1.4), and traverse all the faces in all the existing face groups.

[0063] If the frame number of the current face in all existing face groups is the same as the frame number of the new face, the comparison with the current face is skipped (a face can only be in one frame, and faces with the same frame number do not need to be compared).

[0064] The current face features are compared with the new face features for similarity and the similarity score is calculated. The similarity score is greater than or equal to the tracking target similarity threshold (default: 0.55). At the same time, the absolute value of the difference between the main face frame number of the face group and the new face frame number is less than or equal to the minimum interval between the target frame number and the previous frame (default: 500).

[0065] If the conditions are met, it is considered that the new face has found a similar face group, and the maximum face similarity and the face target unique ID value are calculated. If the conditions are not met, it is considered that no similar face group has been found, and the number of times no similar face group has been found is accumulated.

[0066] If a similar face group is found and the maximum face similarity is greater than the maximum similarity of the found face group, update the main data of the face group (maximum similarity, target unique ID value, and the number of times the same face group is not found is set to 0), and insert the face target into the found face group.

[0067] If no similar face group is found, a new face group is generated, the main face data is used as the new face target, the number of times the same face group is not found is set to 1, and the face group is inserted into the face group queue.

[0068] Traverse the face group queue and determine the number of times the same face group is not found. If the number is greater than the maximum number of face targets to be tracked (default: 100), remove the face group. Process the removed face group data, merge the face targets of the face group and sort the tracks, and save the face cutout / face corresponding original image / confidence / track, etc.

[0069] As a preferred embodiment of the present invention, the present invention stores data of face targets in the following manner:

[0070] The attributes of the face target include: target unique ID value, frame number, original image corresponding to the face, face position (left, top, right, bottom), confidence, yaw angle, pitch angle, features, and trajectory data. The original image corresponding to the face is relatively large, so it is stored in memory + disk. Other data is stored in memory.

[0071] If the original images corresponding to the face are stored in the memory, when the face target trajectory tracking and merging processing speed is slower than the face structure callback of the face target, over time, it will cause the computer system memory to be too high, and eventually cause the system to slow down. In order to solve the problem of high memory usage, optimize the data storage strategy:

[0072] For the same frame of face corresponding to multiple faces, only the original image data corresponding to one face is retained.

[0073] The memory image data of the original face image is stored as a disk file. When the face target image queue exceeds the maximum number of target images (for example, 300 images), the memory data of the face target image is stored in the disk file to release the memory.

[0074] When saving the original image corresponding to the face and the face cutout, reload them into the memory.

Claims

1. A short-time domain target filtering and merging processing method based on face video structured data, characterized in that: include: S1, detecting the face target in the output image of the video file and extracting the feature vector of the face target; S2, tracking the face target according to the frame position and size of the face target, extracting the trajectory of the face target, and outputting the tracking result; S3, reprocessing the face target according to the tracking result, including filtering invalid face targets and filtering the smallest face target; S4. Compare the similarity between the reprocessed face target and the feature vector of the face target, assign similar faces to the same face group, merge and sort the trajectories, and assign dissimilar faces to a new face group.

2. According to claim 1, the method for filtering and merging targets in the short-term domain based on structured face video data is characterized in that Step S1 includes: S1.1, decode the video file and output an image at a certain interval; S1.

2. Use the YOLO detection algorithm to locate all faces in the image; S1.3, align the detected faces and standardize the facial pose and size; S1.

4. Extract feature vectors from the aligned face images.

3. According to claim 1, the method for filtering and merging targets in the short-term domain based on structured face video data is characterized in that Step S2 includes: S2.

1. Initialization When a new target is detected, a KalmanBoxTracker instance is created and its state is initialized to the position and velocity of the detected target; S2.

2. State prediction KalmanBoxTracker uses the prediction step of the Kalman filter to estimate the position of the target in the next frame, which includes state transfer and noise addition; S2.

3. Measurement Update If the object is detected in the next frame, KalmanBoxTracker uses the Kalman filter update step to adjust the predicted state of the object to match the actual detected position; S2.

4. Status Update After the update, the state of the KalmanBoxTracker is corrected to reflect the latest position and velocity information of the target; S2.5 Prediction and Matching In the next frame, the predicted state of KalmanBoxTracker is used to match the new detection result. If the match is successful, the state is updated. If it is not matched, it may mean that the target is occluded or disappeared. The predicted target position is matched with the detection result of the current frame using the Hungarian algorithm. The Hungarian algorithm can associate the detection result with the predicted target in the most optimized way. S2.

6. Time Update If the KalmanBoxTracker does not match a detection result within a certain period of time, it will be removed from the tracking list, and it is considered that the target has left the field of view or is blocked; S2.

7. Identity maintenance Assign a unique ID to each target; S2.8, Output When the target is considered to have left the field of view or is blocked, the target tracking result is output. The tracking result includes: bounding box position, target unique ID, frame number, original image corresponding to the face, face position, confidence, yaw angle, pitch angle, features, and trajectory data.

4. The method for filtering and merging targets in the short-time domain based on structured face video data according to claim 1, characterized in that: The filtering of invalid face targets includes: Filter invalid faces by aspect ratio. If the aspect ratio is in the range of [0.6, 1.6], it is considered a valid face, otherwise it is an invalid face. The validity of a face is determined by five key points: the center of the left eye, the center of the right eye, the tip of the nose, the center of the mouth, and the center of the chin. All five key points must meet the following requirements: the x coordinate of the point must be smaller than the width, and the y coordinate of the point must be smaller than the height, otherwise it is an invalid face. If the yaw angle and pitch angle are equal to nan or -nan, it is considered an invalid face. When the result of an operation is too large or too small to be represented by a floating point number, nan or -nan is used to represent it. If the yaw angle is less than 60 and the pitch angle is less than 20, it is considered a valid face, otherwise it is an invalid face. When there is only one face target track, and the target frame distance is less than 10 from the frame, if the aspect ratio is greater than 1.2, it is considered an invalid face.

5. The method for filtering and merging targets in the short-time domain based on structured face video data according to claim 1, characterized in that: The filtering minimum face target includes: When creating a target filtering and merging task, enter the minimum face width and height parameters. During the face target processing, face target filtering is performed based on this parameter. If the face width and height are both less than the minimum value, it is considered an invalid face, otherwise it is a valid face.

6. The method for filtering and merging targets in the short-time domain based on face video structured data according to claim 1, characterized in that: The step S4 comprises: Traverse the face group, compare the new face features with all the face features in the face group for similarity, calculate the similarity score, the similarity score is greater than or equal to the tracking threshold of the same target similarity, and the absolute value of the difference between the main face frame number of the face group and the new face frame number is less than or equal to the target frame number and the minimum interval between the previous frame; if the conditions are met, it is considered that the face has found the same face group, and the maximum face similarity and the face target unique ID value are calculated. If the conditions are not met, it is considered that the same face group has not been found, and the number of times the same face group has not been found is accumulated; if the same face group is found, and the maximum face similarity is greater than the maximum similarity of the found face group, update the main data of the face group, and insert the face target into the found same face group; if the same face group is not found, generate a new face group, the main face data is the new face target, the number of times the same face group has not been found is set to 1, and insert the face group into the face group queue; Traverse the face group queue and determine the number of times the same face group is not found. If the number is greater than the maximum number of face targets to be tracked, remove the face group; process the removed face group data, merge the face targets of the face group and sort the trajectories, and save the face cutout / face corresponding original image / confidence / trajectory.

7. The method for filtering and merging targets in a short time domain based on structured face video data according to claims 1 to 6, characterized in that The data storage for face targets is as follows: For the original image of the same frame corresponding to multiple faces, only the original image data corresponding to one face is retained; The memory image data corresponding to the original image of the face is stored in a disk file. When the face target image queue exceeds the maximum number of target images, the memory data of the face target image is stored in a disk file to release the memory; When saving the original image corresponding to the face and the face cutout, reload them into the memory.