A face recognition and tracking method based on video understanding

By pre-acquisitioning the calculation consumption-time curve of images with different resolutions and dynamically adjusting the resource allocation strategy, the problem of real-time and accuracy of face recognition in resource-limited environments is solved, and efficient and accurate face recognition tracking is achieved.

CN119919985BActive Publication Date: 2025-06-20BEIJING ZHONGSHITONG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510398138.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-06-20
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

In the prior art, in facial recognition tracking, especially in environments with limited resources, it is difficult to meet real-time and high-precision requirements, resulting in inaccurate recognition results or system crashes.

Method used

By acquiring the operation consumption-time curves corresponding to images with different resolutions in advance, the resource requirements of images with different resolutions in the face recognition process are determined, and the resource allocation strategy is dynamically adjusted according to the real-time computing resource status, and the target face image group is divided to optimize the face recognition process.

Benefits of technology

It improves the real-time and accuracy of facial recognition tasks, ensures that different resolutions and computing consumption needs can be flexibly respond to different resolutions and computing consumption needs in different scenarios, and achieves the optimal balance of accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919985B_ABST
    Figure CN119919985B_ABST
Patent Text Reader

Abstract

The present invention discloses a face recognition and tracking method based on video understanding, belonging to the technical field of face recognition. Specifically, it includes: pre-collecting the same face images at different resolutions in advance, and; performing face recognition on each pre-collected image in turn and recording the recognition time each time. At the same time, during the recognition time of the pre-collected image, the operation consumption C of the system is obtained in real time to obtain the change curve of the operation consumption with the recognition time at each resolution; obtaining the monitoring area image and generating a monitoring area image sequence, and performing face area detection on the monitoring area image in turn to obtain several face images to be recognized. Extract the resolution of any face image to be recognized and obtain the corresponding change curve of this resolution to obtain the change curves corresponding to all face images to be recognized; divide all face images to be recognized into several target face image groups according to the change curves. The present invention improves the efficiency and accuracy of face recognition and tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face recognition, and in particular to a face recognition and tracking method based on video understanding. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, especially in the field of computer vision, the application scenarios of target detection and tracking algorithms are constantly expanding. At present, many deep learning-based target detection and tracking models, such as YOLO, Faster R-CNN, DeepSORT, etc., although they have high accuracy, have high computational complexity and often require a lot of computing resources, especially in real-time applications where the demand for computing power is more stringent, especially in edge devices or resource-limited environments, traditional deep learning models are difficult to meet real-time and high-precision requirements.

[0003] In the prior art, some methods reduce the amount of computation and resource consumption by designing lightweight models, such as YOLO-tiny and MobileNet. Although these methods improve the recognition speed, they usually sacrifice a certain degree of accuracy. Other methods improve the processing speed by adjusting the image resolution, reducing the tracking frequency, etc., and by adjusting the resolution or the number of captures of each captured face image to reduce the load of the recognition system, further improving the computing speed. However, in real life, the flow of people in the monitoring area changes dynamically. For example, when a shopping mall is promoting a promotion, the surveillance camera will capture a large number of faces with a large flow of people. Even if the resolution is adjusted or the capture frequency is reduced, it may still face insufficient resources. Especially in an environment with limited resources, the large amount of image data that needs to be processed may still exceed the system processing capacity, resulting in a slow processing speed, the inability of the recognition algorithm to fully operate, and deviations in the recognition calculation process, which in turn leads to inaccurate recognition results, and even system crashes or freezes. Summary of the invention

[0004] The purpose of the present invention is to provide a face recognition and tracking method based on video understanding to solve the following technical problems:

[0005] In the prior art, some methods reduce the computational amount and resource consumption by designing lightweight models, such as YOLO-tiny and MobileNet. Although the recognition speed is improved, these models usually sacrifice a certain degree of accuracy. Other methods improve the processing speed by adjusting the image resolution, reducing the tracking frequency, etc., and reduce the load of the recognition system by adjusting the resolution of each captured face image or the number of captures, further improving the operation speed. However, in real life, the number of people in the monitoring area is dynamically changing. For example, during a promotion in a shopping mall, there is a large number of people, and the surveillance cameras will capture a large number of faces. Even if the resolution is adjusted or the capture frequency is reduced, the situation of insufficient resources may still be faced. Especially in an environment with limited resources, a large amount of image data to be processed may still exceed the system's processing capacity, resulting in a slow processing speed, the recognition algorithm cannot fully operate, and there are deviations in the recognition calculation process, which further leads to inaccurate recognition results and even system crashes or freezes.

[0006] The object of the present invention can be achieved by the following technical solutions:

[0007] A face recognition and tracking method based on video understanding, characterized by comprising the following steps:

[0008] S1. Pre-collect the same face image at different resolutions in advance, and mark each pre-collected image as F1, F2,..., Fn respectively; perform face recognition on each pre-collected image in turn and record the recognition time each time. At the same time, during the recognition time of the pre-collected image, the operation consumption C of the system is obtained in real time. Taking the recognition time as the abscissa and the operation consumption as the ordinate, generate a change curve FC(t) of the operation consumption with respect to the recognition time, and obtain the change curves of the operation consumption with respect to the recognition time at each resolution; where F represents the image resolution.

[0009] S2. Monitor and obtain the images in the monitoring area in real time and generate a sequence of monitoring area images, and perform face area detection on the monitoring area images in turn to obtain a number of face images to be recognized and mark them as Y1, Y2,..., Yh. Extract the resolution of any face image to be recognized and obtain the corresponding change curve of this resolution, and obtain the change curves corresponding to all face images to be recognized and mark them as F1C(t), F2C(t),..., FhC(t);

[0010] S3. Divide all face images to be recognized into several target face image groups according to the change curves, and the target face image groups are used for face images to be recognized for face recognition simultaneously.

[0011] As a further solution of the present invention: in the S1, the specific determination process of the recognition time further includes:

[0012] For any pre-collected image of any resolution, perform face recognition N times, obtain the recognition time Ti for each face recognition, and calculate it according to the formula The average recognition time Tv is calculated, and the preferred score of each face recognition is calculated according to the calculation formula ΔT=|Ti-Tv|. The face recognition time with the smallest preferred score is selected as the recognition time of the pre-collected image at this resolution, where N is the preset recognition number threshold.

[0013] As a further solution of the present invention: if there are two or more preferred scores that are the same and are all minimum values, the average recognition time Tv is used as the recognition time of the pre-collected image at the resolution.

[0014] As a further solution of the present invention: the specific calculation process of the computing consumption is:

[0015] S11, periodically obtain the system computing consumption in each recognition process, with the recognition time as the horizontal axis and the computing consumption as the vertical axis, and generate a curve of computing consumption versus recognition time. According to the calculation formula The correction coefficient Qm of the change curve m is calculated, and any data point (tz, Cz) in the change curve m is corrected to (tz*Qm, Cz*Qm) to obtain the corrected change curve m, where z is the zth data point in the change curve m;

[0016] S12, repeat S11, obtain all the modified change curves and mark them as y1, y2, ..., yN, according to the calculation formula The average consumption at any time point is calculated to obtain the average consumption at all time points and generate a corresponding change curve, which is used as the corresponding change curve at the resolution.

[0017] As a further solution of the present invention: S2 further includes screening the face images to be identified, and the specific screening process is:

[0018] The monitoring area image is input into a preset OpenCV recognition model, all face areas in the monitoring area image are extracted, the horizontal length is taken as the length, and the vertical length is taken as the width, the minimum circumscribed rectangular frame of any face area is obtained, the monitoring area image is cropped according to the minimum circumscribed rectangular frame to obtain a number of face images to be identified, the aspect ratio of any face image to be identified is calculated and compared with a preset face aspect ratio threshold, if the face aspect ratio is greater than or equal to the face aspect ratio threshold, the area of ​​the face image to be identified is obtained and compared with the preset face area threshold, if the area of ​​the face image to be identified is greater than or equal to the preset face area threshold, the face image to be identified is retained; otherwise, the face image to be identified is discarded.

[0019] As a further solution of the present invention: in S3, the specific process of dividing several target face image groups is as follows:

[0020] S21. Obtain the change curve lengths corresponding to all the face images to be recognized, select the face image to be recognized with the largest change curve length as the reference face image, obtain the change curve corresponding to the reference face image, and uniformly determine X time nodes at a preset time interval Δt based on the starting point of the change curve;

[0021] S22. Starting from the first time node, calculate the total consumption of computing resources at this time node , where C j (t) represents the operation consumption at the time node corresponding to the jth face image to be recognized, t represents the time node, and h represents the total number of face images to be recognized;

[0022] If the operation consumption C 总 is greater than or equal to C - ΔC, then eliminate the face image to be recognized corresponding to the largest operation consumption at the current time node, and recalculate the operation consumption C 总 , until the operation consumption at the first time node satisfies C 总 is less than C - ΔC; where C 总 represents the total operation consumption at any time point, and ΔC represents the preset operation consumption threshold;

[0023] Repeat the above process until the operation consumption C of all the uneliminated face images to be recognized at each reference time node is less than C - ΔC, and classify all the uneliminated face images to be recognized into a target face image group;

[0024] S23. Obtain all the eliminated face images to be recognized, and repeat S21 and S22 to obtain several target face image groups.

[0025] As a further solution of the present invention: it further includes obtaining the number of face images to be recognized in any target face image group, sorting them in descending order according to the number of face images to be recognized, and performing face recognition on all the target face image groups according to the sorting order.

[0026] As a further solution of the present invention: in S2, it further includes that if the number h of the face images to be recognized is less than the preset threshold, then perform face recognition on all the face images to be recognized simultaneously.

[0027] The beneficial effects of the present invention:

[0028] By pre-acquiring the operation consumption-time curves corresponding to images of different resolutions, it is possible to determine the resource requirements of images of different resolutions during the face recognition process, thereby providing data support for subsequent image processing. After obtaining the resource requirements of each face image to be recognized, the resource allocation strategy is dynamically adjusted according to the real-time computing resource status. By precisely controlling the operation consumption of each group of tasks, it is possible to efficiently schedule computing resources, minimize overall recognition latency while reducing resource waste and overcrowding. This not only improves the real-time performance and accuracy of the face recognition task, but also ensures that in different scenarios, it is possible to flexibly handle face image recognition tasks with different resolution and operation consumption requirements, achieving the optimal balance between accuracy and efficiency. The present invention improves the efficiency and accuracy of face recognition tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The present invention will be further described below with reference to the accompanying drawings.

[0030] Figure 1 It is a schematic flowchart of a face recognition and tracking method based on video understanding according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0032] Please refer to Figure 1 As shown, the present invention is a face recognition and tracking method based on video understanding, including the following steps:

[0033] S1. Pre-acquire the same face image at different resolutions in advance, and label each pre-acquired image as F1, F2,..., Fn; perform face recognition on each pre-acquired image in turn and record the recognition time each time. At the same time, obtain the operation consumption C of the system in real time during the recognition time of the pre-acquired image. Taking the recognition time as the abscissa and the operation consumption as the ordinate, generate a change curve FC(t) of the operation consumption with the recognition time, and obtain the change curves of the operation consumption with the recognition time at each resolution; where F represents the image resolution.

[0034] S2. Obtain the images of the monitoring area in real time through monitoring and generate a sequence of images of the monitoring area. Then, perform face area detection on the images of the monitoring area in sequence to obtain a number of face images to be recognized, marked as Y1, Y2, ..., Yh. Extract the resolution of any face image to be recognized and obtain the corresponding change curve of the resolution, and obtain the change curves corresponding to all face images to be recognized, marked as F1C(t), F2C(t), ..., FhC(t);

[0035] S3. Divide all face images to be recognized into several target face image groups according to the change curves. The target face image groups are used for face images to be recognized for face recognition simultaneously.

[0036] By pre-obtaining the operation consumption-time curves corresponding to images of different resolutions, it is possible to determine the resource requirements of images of different resolutions during the face recognition process, thereby providing data support for subsequent image processing. After obtaining the resource requirements of each face image to be recognized, the resource allocation strategy is dynamically adjusted according to the real-time computing resource status. By precisely controlling the operation consumption of each group of tasks, it is possible to efficiently schedule computing resources, reduce resource waste and overcrowding while minimizing the overall recognition delay to the greatest extent. This not only improves the real-time performance and accuracy of the face recognition task, but also ensures that in different scenarios, it is possible to flexibly handle face image recognition tasks with different resolution and operation consumption requirements, achieving the optimal balance between accuracy and efficiency. The present invention improves the efficiency and accuracy of face recognition tracking.

[0037] In a preferred embodiment of the present invention, in S1, the specific determination process of the recognition time further includes:

[0038] Perform face recognition on the pre-acquired images of any resolution N times respectively, obtain the recognition time Ti of each face recognition, and calculate according to the calculation formula to obtain the average recognition time Tv. Calculate the preferred score of each face recognition according to the calculation formula ΔT = |Ti - Tv|, and select the face recognition time with the smallest preferred score as the recognition time of the pre-acquired image at this resolution, where N is the preset recognition times threshold.

[0039] In another preferred embodiment of the present invention, if there are two or more preferred scores that are the same and are all the minimum values, then use the average recognition time Tv as the recognition time of the pre-acquired image at this resolution.

[0040] In another preferred embodiment of the present invention, the specific calculation process of the operation consumption is:

[0041] S11, periodically obtain the system computing consumption in each recognition process, with the recognition time as the horizontal axis and the computing consumption as the vertical axis, and generate a curve of computing consumption versus recognition time. According to the calculation formula The correction coefficient Qm of the change curve m is calculated, and any data point (tz, Cz) in the change curve m is corrected to (tz*Qm, Cz*Qm) to obtain the corrected change curve m, where z is the zth data point in the change curve m;

[0042] S12, repeat S11, obtain all the modified change curves and mark them as y1, y2, ..., yN, according to the calculation formula The average consumption at any time point is calculated to obtain the average consumption at all time points and generate a corresponding change curve, which is used as the corresponding change curve at the resolution.

[0043] In another preferred embodiment of the present invention, the step S2 further includes screening the face image to be identified, and the specific screening process is as follows:

[0044] The monitoring area image is input into a preset OpenCV recognition model, all face areas in the monitoring area image are extracted, the horizontal length is taken as the length, and the vertical length is taken as the width, the minimum circumscribed rectangular frame of any face area is obtained, the monitoring area image is cropped according to the minimum circumscribed rectangular frame to obtain a number of face images to be identified, the aspect ratio of any face image to be identified is calculated and compared with a preset face aspect ratio threshold, if the face aspect ratio is greater than or equal to the face aspect ratio threshold, the area of ​​the face image to be identified is obtained and compared with the preset face area threshold, if the area of ​​the face image to be identified is greater than or equal to the preset face area threshold, the face image to be identified is retained; otherwise, the face image to be identified is discarded.

[0045] In another preferred embodiment of the present invention, the specific process of dividing the target face image groups in S3 is as follows:

[0046] S21, obtaining the lengths of the change curves corresponding to all the face images to be identified, selecting the face image to be identified with the largest change curve length as the reference face image, obtaining the change curve corresponding to the reference face image, and using the starting point of the change curve as a reference, evenly determining X time nodes at a preset time interval Δt;

[0047] S22, starting from the first time node, calculate the total computing resource consumption of the time node , where C j (t) represents the computational consumption of the jth face image to be recognized corresponding to the time node, t represents the time node, and h represents the total number of face images to be recognized;

[0048] If the operation consumption C 总 is greater than or equal to C - ΔC, then the face image to be recognized corresponding to the largest operation consumption at the current time node is removed, and the operation consumption C is recalculated 总 , until the operation consumption at the first time node satisfies that C 总 is less than C - ΔC; where C 总 represents the total operation consumption at any time point, and ΔC represents a preset operation consumption threshold;

[0049] Repeat the above process until the operation consumption C of all the face images to be recognized that are not removed is less than C - ΔC at each reference time node, and all the face images to be recognized that are not removed are classified into a target face image group;

[0050] S23. Obtain all the face images to be recognized that are removed, and repeat S21 and S22 to obtain several target face image groups.

[0051] In another preferred embodiment of the present invention, it further includes obtaining the number of face images to be recognized in any target face image group, sorting them in descending order according to the number of face images to be recognized, and performing face recognition on all target face image groups according to the sorting order.

[0052] In another preferred embodiment of the present invention, in S2, it further includes that if the number h of face images to be recognized is less than a preset threshold, then all face images to be recognized are simultaneously subjected to face recognition.

[0053] The above has described in detail an embodiment of the present invention, but the above content is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.

Claims

1. A face recognition and tracking method based on video understanding, characterized in that: The following steps are involved: S1, pre-collect different resolutions of the same face image in advance, and mark each pre-collected image as F1, F2, ..., Fn; perform face recognition on each pre-collected image in turn and record the recognition time of each time, and at the same time obtain the system's computing consumption C in real time during the recognition time of the pre-collected image, with the recognition time as the horizontal axis and the computing consumption as the vertical axis, generate a curve FC(t) of the change of computing consumption versus recognition time, and obtain the curve of the change of computing consumption versus recognition time at each resolution; wherein F represents the image resolution; S2, acquiring images of the monitoring area in real time through monitoring and generating a sequence of images of the monitoring area, and sequentially performing face region detection on the images of the monitoring area to obtain a number of face images to be identified and marked as Y1, Y2, ..., Yh, extracting the resolution of any face image to be identified and obtaining a change curve corresponding to the resolution, obtaining change curves corresponding to all face images to be identified and marking them as F1C(t), F2C(t), ..., FhC(t); S3, dividing all the facial images to be identified into a plurality of target facial image groups according to the change curve, wherein the target facial image groups are used for simultaneously performing facial recognition on the facial images to be identified; In S3, the specific process of dividing the target face image groups is as follows: S21, obtaining the lengths of the change curves corresponding to all the face images to be identified, selecting the face image to be identified with the largest change curve length as the reference face image, obtaining the change curve corresponding to the reference face image, and using the starting point of the change curve as a reference, evenly determining X time nodes at a preset time interval Δt; S22, starting from the first time node, calculate the total computing resource consumption of the time node , where C' j (t) represents the computational consumption of the jth face image to be recognized corresponding to the time node, t represents the time node, and h represents the total number of face images to be recognized; If the computational consumption C 总 When it is greater than or equal to C-ΔC, the face image to be recognized corresponding to the maximum computational consumption at the current time node is eliminated, and the consumption C is recalculated. 总 , until the computational consumption at the first time node meets C 总 Less than C-ΔC; where C 总 represents the total computing consumption at any time point, and ΔC represents the preset computing consumption threshold; Repeat the above process until the computational consumption C of all unrejected face images to be recognized at each reference time node is less than C-ΔC, and all unrejected face images to be recognized are classified into a target face image group; S23, obtaining all the eliminated face images to be recognized, repeating S21 and S22 to obtain several target face image groups.

2. The face recognition and tracking method based on video understanding according to claim 1, characterized in that: In S1, the specific determination process of the identification time is also included: For any pre-collected image of any resolution, perform face recognition N times, obtain the recognition time Ti for each face recognition, and calculate it according to the formula The average recognition time Tv is calculated, and the preferred score of each face recognition is calculated according to the calculation formula ΔT=|Ti-Tv|. The face recognition time with the smallest preferred score is selected as the recognition time of the pre-collected image at this resolution, where N is the preset recognition number threshold.

3. The face recognition and tracking method based on video understanding according to claim 2, characterized in that: If there are two or more preferred scores that are the same and are all minimum values, the average recognition time Tv is used as the recognition time of the pre-collected image at the resolution.

4. The face recognition and tracking method based on video understanding according to claim 3, characterized in that: The specific calculation process of computing consumption is as follows: S11, periodically obtain the system computing consumption in each recognition process, with the recognition time as the horizontal axis and the computing consumption as the vertical axis, and generate a curve of computing consumption versus recognition time. According to the calculation formula The correction coefficient Qm of the change curve m is calculated, and any data point (tz, Cz) in the change curve m is corrected to (tz*Qm, Cz*Qm) to obtain the corrected change curve m, where z is the zth data point in the change curve m; S12, repeat S11, obtain all the modified change curves and mark them as y1, y2, ..., yN, according to the calculation formula The average consumption at any time point is calculated to obtain the average consumption at all time points and generate a corresponding change curve, which is used as the corresponding change curve at the resolution.

5. The face recognition and tracking method based on video understanding according to claim 1, characterized in that: In the step S2, the face image to be identified is screened, and the specific process of the screening is as follows: The monitoring area image is input into a preset OpenCV recognition model, all face areas in the monitoring area image are extracted, the horizontal length is taken as the length, and the vertical length is taken as the width, the minimum circumscribed rectangular frame of any face area is obtained, the monitoring area image is cropped according to the minimum circumscribed rectangular frame to obtain a number of face images to be identified, the aspect ratio of any face image to be identified is calculated and compared with a preset face aspect ratio threshold, if the face aspect ratio is greater than or equal to the face aspect ratio threshold, the area of ​​the face image to be identified is obtained and compared with the preset face area threshold, if the area of ​​the face image to be identified is greater than or equal to the preset face area threshold, the face image to be identified is retained; otherwise, the face image to be identified is discarded.

6. The face recognition and tracking method based on video understanding according to claim 1, characterized in that: It also includes obtaining the number of face images to be identified in any target face image group, sorting the face images to be identified in descending order according to the number of face images to be identified, and performing face recognition on all target face image groups in sequence according to the arrangement order.

7. The face recognition and tracking method based on video understanding according to claim 1, characterized in that: In the S2, if the number h of the face images to be recognized is less than a preset threshold, face recognition is performed on all the face images to be recognized at the same time.

Citation Information

Patent Citations

  • Face sequence cooperative recognition method based on monitoring video

    CN111652070A

  • Intelligent flight navigation method based on face recognition technology

    CN117974092A