Method and device for extracting optimal face image based on multi-face video
By detecting and feature extraction of multi-face videos, combined with the similarity calculation of the face database, real-time detection of newly emerging faces and extraction of optimal face images are achieved, solving the problem of insufficient real-time performance in the prior art and enhancing adaptability.
Patent Information
- Application Number
- CN202510324065.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-19
AI Technical Summary
In the prior art, hierarchical clustering algorithms require a large number of iterative calculations in multiface video processing, resulting in the inability to meet the real-time requirements and are not suitable for real-time video analysis.
By acquiring multi-face videos, detecting video frames, using the pre-trained face detection model to extract the face feature vector, and performing similarity calculations with the preset face database to determine whether a new face has been detected. If a new face is detected, extract the best face image based on confidence and key points and update the face database.
Real-time detection of newly emerging faces in multi-face videos and extraction of optimal face images, solving the problem of real-time requirements and enhancing adaptability.
Smart Images

Figure CN119851329B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio - video processing, and in particular to a method and device for extracting the best face image based on a multi - face video. Background Art
[0002] With the development of face recognition technology, face recognition technology has been widely applied in fields such as security, transportation, and education. In the application of face recognition technology, a hierarchical clustering algorithm is used to extract the best face image from a multi - face video; however, the hierarchical clustering algorithm is a method of clustering similar face features together. Although it can effectively process multi - angle features and select the best face image, it cannot meet the real - time requirement. This is because the hierarchical clustering algorithm needs to perform iterative calculations on a large amount of data, resulting in too long processing time and being unsuitable for application in real - time video analysis.
[0003] Regarding the problem in the related technology that a large amount of iterative calculations are required, resulting in not meeting the real - time requirement and being inapplicable to real - time video analysis, no effective solution has been proposed yet. Summary of the Invention
[0004] In this embodiment, a method and device for extracting the best face image based on a multi - face video are provided to solve the problem in the related technology that a large amount of iterative calculations are required, resulting in not meeting the real - time requirement and being inapplicable to real - time video analysis.
[0005] In the first aspect, in this embodiment, a method for extracting the best face image based on a multi - face video is provided, including:
[0006] Obtain a multi - face video, detect video frames in the multi - face video, and obtain a detection result;
[0007] Detect the first face feature vector in the detection result against a preset face database to determine whether a new face is detected in the video frame;
[0008] When a new face is detected in the video frame, extract the best face image corresponding to the new face according to the confidence level and key points in the detection result; and update the best face image to the face database; the key points are used to indicate the occlusion of key parts of the face.
[0009] In some embodiments, obtaining a multi - face video, detecting video frames in the multi - face video, and obtaining a detection result includes:
[0010] Obtain a multi - face video, pre - process the multi - face video to obtain corresponding video frames;
[0011] Input the video frame into a pre-trained face detection model for recommendation to obtain a detection result.
[0012] In some embodiments, the method further includes:
[0013] After obtaining the detection result, perform L2 normalization processing on the first face feature vector in the detection result.
[0014] In some embodiments, detecting the first face feature vector in the detection result against a preset face database to determine whether a new face is detected in the video frame includes:
[0015] Calculate the similarity between the first face feature vector and the second face feature vector in the face database to obtain a similarity result;
[0016] According to the similarity result and a preset similarity threshold, assign an identity ID to the face in the video frame and determine whether a new face is detected in the video frame.
[0017] In some embodiments, according to the similarity result and a preset similarity threshold, assigning an identity ID to the face in the video frame and determining whether the video frame detects a new face includes:
[0018] When the similarity result is greater than the first threshold among the similarity thresholds, the face to be detected in the video frame is the old face, and assign the existing identity ID to the face to be detected;
[0019] When the similarity result is less than or equal to the first threshold and greater than the second threshold among the similarity thresholds, the face to be detected in the video frame is the old face, assign the existing identity ID to the face to be detected, and store the first face feature vector in the face database;
[0020] When the similarity result is less than or equal to the second threshold, determine that the video frame detects a new face, and assign a new identity ID to the new face.
[0021] In some embodiments, calculating the similarity between the first face feature vector and the second face feature vector in the face database to obtain a similarity result includes:
[0022] Calculate the similarity between the first face feature vector and the second face feature vector in the face database through the cosine similarity formula to obtain a similarity result.
[0023] In some embodiments, the method further includes:
[0024] When storing the first face feature vector in the face database, the face database is dynamically updated based on a preset update strategy.
[0025] In a second aspect, an apparatus for extracting the best face image based on a multi-face video is provided in this embodiment, including: a processing module, a detection module, and an extraction module;
[0026] The processing module is configured to obtain a multi-face video, detect video frames in the multi-face video, and obtain a detection result;
[0027] The detection module is configured to detect the first face feature vector in the detection result with a preset face database to determine whether a new face is detected in the video frame;
[0028] The extraction module is configured to, when a new face is detected in the video frame, extract the best face image corresponding to the new face according to the confidence level and key points in the detection result; and update the best face image to the face database; the key points are used to indicate the occlusion of key parts of the face.
[0029] In a third aspect, a computer device is provided in this embodiment, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for extracting the best face image based on a multi-face video described in the first aspect above is implemented.
[0030] In a fourth aspect, a storage medium is provided in this embodiment, on which a computer program is stored. When the program is executed by a processor, the method for extracting the best face image based on a multi-face video described in the first aspect above is implemented.
[0031] Compared with the related art, in the method and apparatus for extracting the best face image based on a multi-face video provided in this embodiment, by obtaining a multi-face video, detecting video frames in the multi-face video to obtain a detection result; detecting the first face feature vector in the detection result with a preset face database to determine whether a new face is detected in the video frame; when a new face is detected in the video frame, extracting the best face image corresponding to the new face according to the confidence level and key points in the detection result; and updating the best face image to the face database; the key points are used to indicate the occlusion of key parts of the face, the problems in the related art that a large amount of iterative calculations are required, resulting in non-meeting the real-time requirement and being inapplicable to real-time video analysis are solved. It can automatically detect newly emerging faces in a multi-face video and extract the best face image of the face, realizing the recognition requirement of real-time processing of new people and enhancing the adaptability ability.
[0032] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0034] Figure 1 is a hardware structure block diagram of a terminal device for an optimal face image extraction method based on a multi-face video provided by an embodiment of the present application;
[0035] Figure 2 is a flowchart of an optimal face image extraction method based on a multi-face video provided by an embodiment of the present application;
[0036] Figure 3 is a flowchart of step S210;
[0037] Figure 4 is a flowchart of step S220;
[0038] Figure 5 is a structure block diagram of an optimal face image extraction device based on a multi-face video provided by an embodiment of the present application.
[0039] In the figure: 102, a processor; 104, a memory; 106, a transmission device; 108, an input / output device; 210, a processing module; 220, a detection module; 230, an extraction module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] To more clearly understand the purpose, technical solution, and advantages of the present application, the present application is described and illustrated below in conjunction with the drawings and embodiments.
[0041] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meanings understood by those with ordinary skills in the technical field to which this application belongs. In this application, words such as "a", "an", "one kind", "the", "these" and the like do not indicate a limitation in quantity, and they can be singular or plural. The terms "including", "containing", "having" and any variants thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device containing a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connected", "linked", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The term "plurality" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third" and the like involved in this application only distinguish similar objects and do not represent a specific order for the objects.
[0042] The method embodiment provided in this embodiment can be executed on a terminal, a computer or a similar computing device. For example, when running on a terminal, Figure 1 is the hardware structure block diagram of the terminal of the best face image extraction method based on multi-face video in this embodiment. As Figure 1 shown, the terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 and a memory 104 for storing data. Among them, the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA. The above terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include more or fewer components than those shown in Figure 1 the figure, or have a different configuration from that shown in Figure 1 the figure.
[0043] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the method for extracting the best face image based on multi-face video in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0044] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by a communication provider of the terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0045] In this embodiment, a method for extracting the best face image based on multi-face video is provided. Figure 2 is a flowchart of the method for extracting the best face image based on multi-face video in this embodiment, as Figure 2 shown, this process includes the following steps:
[0046] Step S210, obtain a multi-face video, detect video frames in the multi-face video, and obtain a detection result;
[0047] Step S220, detect the first face feature vector in the detection result with a preset face database to determine whether a new face is detected in the video frame;
[0048] Step S230, when a new face is detected in the video frame, extract the best face image corresponding to the new face according to the confidence level and key points in the detection result; and update the best face image to the face database; the key points are used to indicate the occlusion of key parts of the face.
[0049] Specifically, a multi-face video refers to a video that contains multiple human faces; such as videos at the face punching points of a company, videos at the attendance checkpoints, videos at the security checkpoints, etc. In actual applications, the method for obtaining a multi-face video in the embodiments of the present application includes, but is not limited to, obtaining from a database pre-stored, obtaining a multi-face video that meets the above requirements; it can also be downloading a multi-face video that meets the requirements from a network platform; it can also be generating a corresponding multi-face video according to requirements, etc. The embodiments of the present application do not limit the method for obtaining a multi-face video.
[0050] After obtaining the multi-face video, the multi-face video can be decoded to obtain video frames in the multi-face video; it can be considered that each video frame also has multiple human faces, and one video frame is a face image. Then, using a trained face detection model or face detection algorithm, the video frames in the multi-face video are detected to obtain a detection result. Among them, the detection result at least includes a first face feature vector (embedding1), a confidence level, and key points (kps); the first face feature vector is a vector composed of a set of representative numerical features obtained after the video frame is detected. It can be regarded as a mathematical representation of a face image in a high-dimensional feature space, which can capture key feature information of the face, such as the angles, shapes, positions, textures of facial organs, etc., and can identify multiple angles or states of the face. The key points are used to indicate the occlusion of key parts of the face; for example, if obvious occlusions occur in key parts such as eyes, nose, and mouth, the priority of this video frame will be reduced. The confidence level is the score of the face in the video frame, with the highest confidence level for a frontal photo; the higher the score, the higher the credibility of this face photo. In other embodiments, the detection result may further include a bounding box; each bounding box represents a face image.
[0051] Next, using the first face feature vector in the detection result as the detection basis, a preset face database is traversed. If no second face feature vector similar to the first face feature vector is detected, then it can be considered that a new face is detected in the video frame. Among them, the face database is preset, and it stores the identity ID, the second face feature vector, and the confidence level corresponding to each face, etc.
[0052] Finally, when a new face is detected in a video frame, since the confidence level is the score of the face in the video frame and the key points are used to indicate the occlusion of key parts of the face, the best face image corresponding to the new face is extracted from each video frame according to the confidence level and key points in the detection result. That is, in this embodiment, the clear and unoccluded best face image can be automatically selected from the video frames (face images) at different times, and the best face image is updated to the face database to ensure that the stored images have high quality and are easy to recognize. And in this embodiment, the relevant data in the detection result is used to automatically detect the newly appeared face in the multi-face video and extract the best face image of this face, realizing the recognition requirement of real-time processing of new people and enhancing the adaptability.
[0053] In the related art, a hierarchical clustering algorithm is used to extract the best face image from a multi-face video. However, the hierarchical clustering algorithm is a method of clustering similar face features together. Although it can effectively process multi-angle features and select the best face image, it cannot meet the real-time requirement. This is because the hierarchical clustering algorithm needs to perform iterative calculations on a large amount of data, resulting in too long processing time and being not suitable for application in real-time video analysis. In this embodiment, by obtaining a multi-face video, detecting the video frames in the multi-face video to obtain a detection result, detecting the first face feature vector in the detection result with a preset face database to determine whether a new face is detected in the video frame. When a new face is detected in the video frame, according to the confidence level and key points in the detection result, the best face image corresponding to the new face is extracted and the best face image is updated to the face database. The key points are used to indicate the occlusion of key parts of the face, solving the problem in the related art that a large amount of iterative calculations are required, resulting in not meeting the real-time requirement and being not applicable to real-time video analysis, being able to automatically detect the newly appeared face in the multi-face video and extract the best face image of this face, realizing the recognition requirement of real-time processing of new people and enhancing the adaptability.
[0054] The above steps will be described in detail below:
[0055] In some of the embodiments, as Figure 3 shown, obtaining a multi-face video, detecting the video frames in the multi-face video to obtain a detection result in step S210 includes the following steps:
[0056] Step S211, obtaining a multi-face video and preprocessing the multi-face video to obtain corresponding video frames;
[0057] Step S212, inputting the video frames into a pre-trained face detection model for recommendation to obtain a detection result.
[0058] Specifically, since multi-face videos are usually stored after encoding and compression, a decoding operation is required to convert them into an image frame format that can be directly processed by a computer (such as a pixel matrix in RGB or BGR format). This process can be automatically completed by calling a video codec plugin. Then, the video frames in the multi-face video are read frame by frame in sequence. A specific frame rate extraction rule can be set, for example, extracting one frame every certain number of frames, to reduce the amount of data processed and improve processing efficiency. In practical applications, the appropriate extraction strategy is determined according to specific requirements and the content characteristics of the multi-face video, and there is no limitation on this.
[0059] Among them, the face detection model is pre-trained, and it can adopt YOLO series models, SSD (Single Shot MultiBox Detector) models, Mask R–CNN, etc. In this embodiment, the face detection model can be MTCNN (Multi-task Cascaded Convolutional Networks); it adopts a cascaded structure and consists of three convolutional neural networks (P-Net, R-Net, and O-Net). P-Net is used to generate the bounding boxes of candidate face regions, R-Net screens and refines the candidate regions, and O-Net further improves the detection accuracy and outputs the key points of the face, face feature vectors, and confidence levels. Then, by inputting the video frame into the pre-trained face detection model for recommendation, the detection result can be obtained. The detection result includes bounding boxes, key points, face feature vectors, and confidence levels.
[0060] Through this embodiment, the processing of multi-face videos can be completed efficiently and accurately, which can be applied to real-time video analysis and prepare for subsequent processing tasks at the same time.
[0061] In some of these embodiments, based on the best face image extraction method for multi-face videos, the method further includes:
[0062] After obtaining the detection result, perform L2 normalization processing on the first face feature vector in the detection result.
[0063] Specifically, perform L2 normalization processing on the first face feature vector in the detection result; L2 normalization processing will scale the L2 norm (i.e., the Euclidean length of the vector) of the first face feature vector to 1, so that the modulus lengths of each vector are equal. L2 normalization ensures that the direction information of the feature vector remains unchanged; so that in the subsequent calculation of cosine similarity, it will not be affected by the size of the vector, only focusing on direction consistency, thereby improving the accuracy of cosine similarity calculation. In other embodiments, other normalization methods can also be adopted, and there is no limitation on this.
[0064] It should be noted that the first face feature vector and the second face feature vector in the face database are obtained by the same processing method. For example, if the first face feature vector is processed by L2 normalization, then the second face feature vector is also processed by L2 normalization.
[0065] In some of these embodiments, as Figure 4 shown, the step of detecting the first face feature vector in the detection result against a preset face database to determine whether a new face is detected in the video frame includes the following steps:
[0066] Step S221, calculate the similarity between the first face feature vector and the second face feature vector in the face database to obtain a similarity result;
[0067] Step S222, according to the similarity result and a preset similarity threshold, assign an identity ID to the face in the video frame and determine whether a new face is detected in the video frame.
[0068] In practical applications, the first face feature vector and the second face feature vector are data of the same type, but only in different positions. The first face feature vector is obtained by processing the current video frame, and the second face feature vector is stored in the face database. Similarity calculation formulas such as cosine similarity, Pearson correlation coefficient, Euclidean distance, or Manhattan distance can be used to calculate the similarity between the first face feature vector and the second face feature vector. Specifically, the corresponding similarity calculation formula can be selected according to different application scenarios. For example, in the fields of data mining, machine learning, etc., the Euclidean distance similarity formula can be used. In the field of face recognition, the cosine similarity formula can be used, and its expression can be:
[0069] Cosine Similarity(A, B)=A×B;
[0070] In the formula, A represents the first face feature vector; B represents the second face feature vector.
[0071] Among them, the cosine similarity formula can be considered that the cosine similarity is equal to the dot product of two feature vectors (the sum of the products of corresponding elements). The greater the cosine similarity between two faces, the more similar they are. Then, after obtaining the similarity result, compare the similarity result with the preset similarity threshold to obtain the comparison result of whether a new face is detected in the video frame; for the comparison result of detecting a new face, create a new identity ID and assign it to the new face; for the comparison result of detecting an old face, assign the old identity ID to the old face.
[0072] In this embodiment, the similarity between the first face feature vector and the second face feature vector is first calculated to obtain a similarity result; then, based on the similarity result and a preset similarity threshold, an identity ID is assigned to the face in the video frame, and whether a new face is detected in the video frame is determined; thus, automatic identity matching and identity ID assignment are achieved.
[0073] In some of these embodiments, according to the similarity result and the preset similarity threshold in step S222, an identity ID is assigned to the face in the video frame, and whether a new face is detected in the video frame includes the following steps:
[0074] When the similarity result is greater than the first threshold of the similarity threshold, the face to be detected in the video frame is an old face, and the existing identity ID is assigned to the face to be detected;
[0075] When the similarity result is less than or equal to the first threshold and greater than the second threshold of the similarity threshold, the face to be detected in the video frame is an old face, the existing identity ID is assigned to the face to be detected, and the first face feature vector is stored in the face database;
[0076] When the similarity result is less than or equal to the second threshold, it is determined that a new face is detected in the video frame, and a new identity ID is assigned to the new face.
[0077] Specifically, the similarity threshold is preset. In this embodiment, there are two similarity thresholds; including the first threshold and the second threshold; the first threshold is greater than the second threshold. The specific parameters of the first threshold and the second threshold can be set according to the usage scenario, and no limitation is imposed thereon.
[0078] Exemplarily: The first threshold is 0.8; the second threshold is 0.5; when the similarity result of the face to be detected in the video frame is greater than 0.8, it is considered that the face to be detected in the video frame is an old face and is the same face as the face corresponding to the second face feature vector in the face data volume. Then, the identity ID of the old face is assigned to the face to be detected. When the similarity result is less than or equal to 0.8 and greater than 0.5, it is considered that the face to be detected in the video frame is an old face, and it is considered that the face to be detected in the video frame is an old face and is the same face as the face corresponding to the second face feature vector in the face data volume. Then, the identity ID of the old face is assigned to the face to be detected; however, it is considered that there are differences between the face image of the face to be detected and the face images in the face database (which may be different angles or expression changes of known faces). Therefore, the first face feature vector is stored in the face database to further improve the face images in the face database and continuously optimize the face database to improve the recognition effect. When the similarity result is less than or equal to 0.5, it is considered that the face to be detected in the video frame is a new face, that is, it is determined that a new face is detected in the video frame. At this time, a new identity ID will be created for the new face, and the created identity ID will be assigned to the new face.
[0079] Through this embodiment, it is possible to fuse face feature vectors from multiple angles, improve the recognition accuracy across scenarios, and store the first face feature vector in the face database, enabling multi-angle and multi-expression features for each identity, effectively reducing the occurrence of misidentifications.
[0080] In some of these embodiments, the method for extracting the best face image based on a multi-face video further includes the following steps:
[0081] When storing the first face feature vector in the face database, the face database is dynamically updated based on a preset update strategy.
[0082] Specifically, the list of second face feature vectors for each identity ID in the face database will be restricted to no more than N; N represents N second face feature vectors with different angles or states. In this embodiment, N can be 10.
[0083] When the similarity result is less than or equal to the first threshold and greater than the second threshold among the similarity thresholds, it is considered that the first face feature vector corresponding to the new face has high representativeness. At this time, the first face feature vector will be stored in the feature list of the face database (each face corresponds to a feature list and an identity ID; N second face feature vectors are stored in this feature list).
[0084] If the feature list has reached the upper limit, the corresponding feature vector is removed according to a predetermined update strategy. The update strategy is that in the feature list of each identity ID, when a new first face feature vector is stored, the length of the feature list of this identity ID will be checked. If the feature list has reached the maximum number, the second face feature vector will be removed according to the principle of first in first out or according to the confidence level, etc., and then the new first face feature vector will be stored and become the second face feature vector.
[0085] Through this embodiment, it is possible to dynamically update the multi-angle and multi-expression face feature vectors of each identity, effectively reducing the occurrence of misidentification; at the same time, it can improve the comparison speed and efficiency.
[0086] It should be noted that the steps shown in the above process or the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0087] In this embodiment, there is also provided an apparatus for extracting the best face image based on a multi-face video. This apparatus is used to implement the above embodiment and the preferred implementation manner, and those that have been described will not be repeated here. The following terms "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0088] Figure 5 is the structural block diagram of the apparatus for extracting the best face image based on a multi-face video in this embodiment, as Figure 5 shown, the apparatus includes: a processing module 210, a detection module 220, and an extraction module 230;
[0089] The processing module 210 is configured to obtain a multi-face video, detect video frames in the multi-face video, and obtain a detection result;
[0090] The detection module 220 is configured to detect the first face feature vector in the detection result with a preset face database to determine whether a new face is detected in the video frame;
[0091] The extraction module 230 is configured to, when a new face is detected in the video frame, extract the best face image corresponding to the new face according to the confidence level and key points in the detection result; and update the best face image to the face database; the key points are used to indicate the occlusion of key parts of the face.
[0092] Through the above device, the problems in the related art that a large number of iterative calculations are required, resulting in non - meeting the real - time requirements and being inapplicable to real - time video analysis are solved. It can automatically detect newly - emerged faces in a multi - face video and extract the best face image of the face, realizing the recognition requirements for real - time processing of new people and enhancing the adaptability ability.
[0093] In some of the embodiments, the processing module 210 is further configured to obtain a multi - face video, pre - process the multi - face video to obtain corresponding video frames.
[0094] Input the video frames into a pre - trained face detection model for recommendation to obtain detection results.
[0095] In some of the embodiments, the device further includes: a normalization module.
[0096] The normalization module is configured to perform L2 normalization processing on the first face feature vector in the detection results after obtaining the detection results.
[0097] In some of the embodiments, the detection module 220 is further configured to calculate the similarity between the first face feature vector and the second face feature vector in the face database to obtain a similarity result.
[0098] According to the similarity result and a preset similarity threshold, assign an identity ID to the face in the video frame and determine whether a new face is detected in the video frame.
[0099] In some of the embodiments, when the similarity result is greater than the first threshold among the similarity thresholds, the face to be detected in the video frame is an old face, and an existing identity ID is assigned to the face to be detected.
[0100] When the similarity result is less than or equal to the first threshold and greater than the second threshold among the similarity thresholds, the face to be detected in the video frame is an old face, an existing identity ID is assigned to the face to be detected, and the first face feature vector is stored in the face database.
[0101] When the similarity result is less than or equal to the second threshold, it is determined that a new face is detected in the video frame, and a new identity ID is assigned to the new face.
[0102] In some of the embodiments, the detection module 220 is further configured to calculate the similarity between the first face feature vector and the second face feature vector in the face database through the cosine similarity formula to obtain a similarity result.
[0103] In some of the embodiments, the device further includes: an update module.
[0104] An update module, configured to dynamically update the face database based on a preset update policy when storing the first face feature vector into the face database.
[0105] It should be noted that the above-mentioned various modules can be functional modules or program modules, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned various modules can be located in the same processor; or the above-mentioned various modules can also be located in different processors in any combined form.
[0106] In this embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0107] Optionally, the above computer device may further include a transmission device and an input / output device. Among them, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0108] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:
[0109] S1. Obtain a multi-face video, detect video frames in the multi-face video, and obtain a detection result;
[0110] S2. Detect the first face feature vector in the detection result against a preset face database to determine whether a new face is detected in the video frame;
[0111] S3. When a new face is detected in the video frame, extract the best face image corresponding to the new face according to the confidence level and key points in the detection result; and update the best face image to the face database; the key points are used to indicate the occlusion of key parts of the face.
[0112] It should be noted that specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated in this embodiment.
[0113] In addition, in combination with the method for extracting the best face image based on a multi-face video provided in the above embodiments, a storage medium can also be provided to implement it in this embodiment. A computer program is stored on the storage medium; when the computer program is executed by a processor, it implements any one of the above methods for extracting the best face image based on a multi-face video.
[0114] It should be noted that the information and data involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and will be used legally.
[0115] It should be understood that the specific embodiments described herein are for explaining this application rather than limiting it. All other embodiments obtained by those of ordinary skill in the art without creative efforts according to the embodiments provided in this application fall within the protection scope of this application.
[0116] Obviously, the accompanying drawings are only some examples or embodiments of this application. For those of ordinary skill in the art, this application can also be applied to other similar situations based on these drawings without creative efforts. Additionally, it can be understood that although the work done during the development process here may be complex and time-consuming, for those of ordinary skill in the art, certain design, manufacturing, or production changes based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient disclosure of this application.
[0117] The term "embodiment" in this application means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of this application. The phrase appears in various positions in the specification and does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in this application can be combined with other embodiments without conflict.
[0118] The above-described embodiments only represent several implementation manners of this application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.
Claims
1. An optimal face image extraction method based on multi-face video, characterized in that: include: Acquire a multi-face video, and detect video frames in the multi-face video to obtain a detection result; Detecting the first face feature vector in the detection result with a preset face database to determine whether a new face is detected in the video frame; When a new face is detected in the video frame, extracting the best face image corresponding to the new face according to the confidence and key points in the detection result; and updating the best face image into the face database; The key points are used to indicate the occlusion of key parts of the face; When the first facial feature vector is stored in the facial database, the facial database is dynamically updated based on a preset update strategy; the update strategy is: in the feature list of each identity ID in the facial database, when a new first facial feature vector is stored, if the feature list of the identity ID has reached the maximum number, the second facial feature vector in the feature list will be eliminated according to the first-in-first-out principle or according to the level of confidence, and the new first facial feature vector will be stored to become the second facial feature vector.
2. The optimal face image extraction method based on multi-face video according to claim 1 is characterized in that: Acquire a multi-face video, detect video frames in the multi-face video, and obtain a detection result, including: Acquire a multi-face video, and pre-process the multi-face video to obtain corresponding video frames; The video frame is input into a pre-trained face detection model for recommendation to obtain a detection result.
3. The optimal face image extraction method based on multi-face video according to claim 1 is characterized in that: The method further comprises: After obtaining the detection result, L2 normalization processing is performed on the first facial feature vector in the detection result.
4. The optimal face image extraction method based on multi-face video according to any one of claims 1 to 3, characterized in that: Detecting the first face feature vector in the detection result with a preset face database to determine whether a new face is detected in the video frame includes: Calculate the similarity between the first face feature vector and the second face feature vector in the face database to obtain a similarity result; According to the similarity result and a preset similarity threshold, an identity ID is assigned to the face in the video frame, and it is determined whether a new face is detected in the video frame.
5. The optimal face image extraction method based on multi-face video according to claim 4 is characterized in that: According to the similarity result and a preset similarity threshold, an identity ID is assigned to the face in the video frame, and it is determined whether a new face is detected in the video frame, including: When the similarity result is greater than a first threshold in the similarity thresholds, the face to be detected in the video frame is an old face, and an existing identity ID is assigned to the face to be detected; When the similarity result is less than or equal to the first threshold, and the similarity result is greater than a second threshold in the similarity thresholds, the face to be detected in the video frame is an old face, an existing identity ID is assigned to the face to be detected, and the first face feature vector is stored in the face database; When the similarity result is less than or equal to the second threshold, it is determined that a new face is detected in the video frame, and a new identity ID is assigned to the new face.
6. The optimal face image extraction method based on multi-face video according to claim 4 is characterized in that: Calculating the similarity between the first face feature vector and the second face feature vector in the face database to obtain a similarity result includes: The similarity between the first facial feature vector and the second facial feature vector in the face database is calculated by using a cosine similarity formula to obtain a similarity result.
7. An optimal face image extraction device based on multi-face video, characterized in that: include: A processing module, a detection module, and an extraction module; The processing module is used to obtain a multi-face video, detect video frames in the multi-face video, and obtain a detection result; The detection module is used to detect the first face feature vector in the detection result with a preset face database to determine whether a new face is detected in the video frame; The extraction module is used to extract the best face image corresponding to the new face according to the confidence and key points in the detection result when the new face is detected in the video frame; and update the best face image into the face database; The key points are used to indicate the occlusion of key parts of the face; When the first facial feature vector is stored in the facial database, the facial database is dynamically updated based on a preset update strategy; the update strategy is: in the feature list of each identity ID in the facial database, when a new first facial feature vector is stored, if the feature list of the identity ID has reached the maximum number, the second facial feature vector in the feature list will be eliminated according to the first-in-first-out principle or according to the level of confidence, and the new first facial feature vector will be stored to become the second facial feature vector.
8. A computer device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps of the optimal face image extraction method based on multi-face video according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the optimal face image extraction method based on multi-face video according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Safety monitoring method based on face identification, apparatus thereof and storage medium
CN108399665A
Face recognition method and device, electronic equipment and readable storage medium
CN114898416A
Human body detection and face recognition system and method
CN119541026A