Video person identification method, device and equipment and storage medium

By applying quality scoring and adjusting voting weights to facial images in videos, the problem of decreased recognition accuracy and trajectory breakage caused by differences in facial image quality in existing technologies is solved, achieving higher recognition accuracy and recall.

CN117523625BActive Publication Date: 2026-07-31TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-07-29
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies for video person recognition have low accuracy, especially when face images of varying quality are present. They ignore the decrease in recognition accuracy caused by differences in face image quality, and the trajectory of people in the video is easily broken, reducing sample diversity.

Method used

By scoring the quality of face images to distinguish between high-quality and low-quality faces, adjusting the voting weights to increase the voting weight of high-quality faces and decrease the weight of low-quality faces, and simultaneously performing trajectory association and cascade voting, the sample diversity of the overall results is enhanced.

Benefits of technology

It improves the accuracy and recall of video person recognition, ensures the accuracy of recognition results for high-quality face images, mitigates the adverse effects of low-quality faces, and enhances the integrity of trajectory association and recognition recall.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117523625B_ABST
    Figure CN117523625B_ABST
Patent Text Reader

Abstract

This application discloses a video person recognition method, apparatus, device, and storage medium, relating to the field of artificial intelligence technology. The method includes: extracting a sequence of face images from a video, the sequence comprising multiple face images; obtaining preliminary recognition results and quality scores corresponding to each of the multiple face images in the sequence, the preliminary recognition results indicating candidate persons corresponding to the face images, and the quality scores indicating the image quality of the face images; and voting on the candidate persons corresponding to the multiple face images based on their respective quality scores to determine the target person corresponding to the face image sequence. This method uses the quality scores of the face images to influence the voting results of the candidate persons corresponding to the face images, thus improving the accuracy of the identified target person and simultaneously increasing the recall rate of persons in the video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a video person recognition method, apparatus, device, and storage medium. Background Technology

[0002] Video person recognition is an important task in computer vision and artificial intelligence, encompassing several fundamental image processing technologies.

[0003] In related technologies, multiple facial images obtained from a video are typically used to identify people in the video. Facial trajectories are formed based on the extracted facial images, and the recognition result of the facial trajectory is determined based on the recognition result of each facial image.

[0004] Furthermore, the accuracy of the target person corresponding to the trajectory determined by the recognition results of each face image in the related technologies is relatively low. Summary of the Invention

[0005] This application provides a video person recognition method, apparatus, device, and storage medium, which can assign different quality scores to different face images. The quality score of the face image affects the voting weight of the preliminary recognition result corresponding to the face image, thereby determining the target person corresponding to the face image sequence, resulting in higher accuracy in identifying the target person. The technical solution is as follows:

[0006] According to one aspect of the embodiments of this application, a video person recognition method is provided, the method comprising:

[0007] Extract a sequence of facial images from a video, the sequence of facial images including multiple facial images;

[0008] The preliminary recognition results and quality scores corresponding to multiple face images included in the face image sequence are obtained respectively. The preliminary recognition results are used to indicate the candidate objects corresponding to the face images, and the quality scores are used to indicate the image quality of the face images.

[0009] Based on the quality scores corresponding to the multiple face images, a vote is taken on the candidate individuals corresponding to the multiple face images to determine the target person corresponding to the face image sequence.

[0010] According to one aspect of the embodiments of this application, a video person recognition device is provided, the device comprising:

[0011] An image extraction module is used to extract a sequence of face images from a video, wherein the sequence of face images includes multiple face images;

[0012] The result acquisition module is used to acquire the preliminary recognition results and quality scores corresponding to multiple face images included in the face image sequence. The preliminary recognition results are used to indicate the candidate objects corresponding to the face images, and the quality scores are used to indicate the image quality of the face images.

[0013] The person identification module is used to vote on the candidate persons corresponding to the multiple face images based on the quality scores corresponding to the multiple face images, and to determine the target person corresponding to the face image sequence.

[0014] According to one aspect of the embodiments of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the above-described method.

[0015] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein a computer program is stored in the storage medium, the computer program being loaded and executed by a processor to implement the above-described method.

[0016] According to one aspect of the embodiments of this application, a computer program product is provided, the computer program product including a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the computer device to perform the method described above.

[0017] The technical solution provided in this application embodiment can include the following beneficial effects: extracting face image sequences from a video, each face image sequence including multiple face images; voting on candidate individuals corresponding to multiple face images based on the quality scores of multiple face images belonging to the same face image sequence; and determining the target person corresponding to the face image sequence based on the voting results. This application uses the image quality of face images to vote on the recognition results of face images, and determines the target person corresponding to the face image sequence composed of multiple face images based on the voting results. In other words, the image quality of face images affects the voting results of the candidate individuals corresponding to the face images. Therefore, the technical solution provided in this application embodiment, based on the preliminary recognition results of face images, uses the image quality of face images as an important factor affecting the voting results of the candidate individuals corresponding to the face images, making the finally determined target person more accurate, and thus the recall rate of people in the video is also higher. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the implementation environment of a solution provided in one embodiment of this application;

[0019] Figure 2This is a schematic diagram of video person recognition results provided in one embodiment of this application;

[0020] Figure 3 This is a flowchart of a video person recognition method provided in one embodiment of this application;

[0021] Figure 4 This is a schematic diagram illustrating the quality scoring of a face image provided in one embodiment of this application;

[0022] Figure 5 This is a flowchart of a video person recognition method provided in another embodiment of this application;

[0023] Figure 6 This is a schematic diagram of in-trajectory voting provided in one embodiment of this application;

[0024] Figure 7 This is a flowchart of a video person recognition method provided in another embodiment of this application;

[0025] Figure 8 This is a schematic diagram of trajectory merging provided in one embodiment of this application;

[0026] Figure 9 This is a schematic diagram of inter-trajectory voting provided in one embodiment of this application;

[0027] Figure 10 This is a block diagram of a video person recognition method provided in one embodiment of this application;

[0028] Figure 11 This is a block diagram of a video person recognition device provided in one embodiment of this application;

[0029] Figure 12 This is a block diagram of a video person recognition device provided in another embodiment of this application;

[0030] Figure 13 This is a structural block diagram of a computer device provided in one embodiment of this application. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0032] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0033] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0034] Computer vision (CV) is the science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D (three-dimensional) technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0035] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning.

[0036] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, and smart customer service. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.

[0037] The technical solutions provided in this application relate to technologies such as computer vision in artificial intelligence, and are specifically illustrated through the following embodiments.

[0038] Before introducing the embodiments of this application, the following explanations are provided for the terms appearing in this solution to facilitate understanding of the solution.

[0039] MOT (Multiple Object Tracking) technology: It acquires a single video and splits it into discrete frames at a specific frame rate (fps) for output. It detects which objects exist in each frame, marks the position of the objects in each frame, and associates object images in different frames to determine whether they belong to the same target object or different target objects. In the field of face recognition, the typical workflow of the MOT algorithm is as follows: (1) Given the original frames of the video; (2) Run the object detector to obtain the bounding boxes of the face images; (3) For each detected face image, calculate different features, usually visual and motion features; (4) Then, the similarity calculation step calculates the probability that two face images belong to the same target person; (5) Finally, the association step assigns a digital identifier to each target person.

[0040] Clustering is the process of dividing a dataset into different classes or clusters based on a specific criterion (such as distance), maximizing the similarity of data objects within the same cluster and maximizing the differences between data objects in different clusters. After clustering, data of the same class should be grouped together as much as possible, while data of different classes should be separated as much as possible. Common clustering methods include K-Means clustering, mean-shift clustering, and clustering using Gaussian mixture models, etc. This application does not limit the clustering method.

[0041] Please refer to Figure 1 The diagram illustrates an implementation environment for a solution provided in one embodiment of this application. This implementation environment may include: a terminal device 10 and a server 20.

[0042] Terminal device 10 includes, but is not limited to, mobile phones, tablets, smart voice interaction devices, game consoles, wearable devices, multimedia playback devices, PCs (Personal Computers), in-vehicle terminals, smart home appliances, and other electronic devices. The client for the target application can be installed on terminal device 10.

[0043] In this embodiment, the target application can be any application capable of providing video feed content services. Typically, the application is a video application. Of course, other types of applications besides video applications can also provide feed content services. For example, news applications, social applications, interactive entertainment applications, browser applications, shopping applications, content sharing applications, virtual reality (VR) applications, augmented reality (AR) applications, etc., are not limited in this embodiment. In addition, the videos pushed by different applications will be different, and the corresponding functions will also be different. These can be pre-configured according to actual needs, and are not limited in this embodiment. Optionally, the terminal device 10 runs a client of the above-mentioned application. In some embodiments, the above-mentioned feed content service covers many vertical content such as variety shows, movies, news, finance, sports, entertainment, and games, and users can enjoy a variety of content services such as articles, pictures, short videos, live broadcasts, special topics, and columns through the above-mentioned feed content service.

[0044] Server 20 is used to provide backend services for the client of the target application in terminal device 10. For example, server 20 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, but it is not limited to these.

[0045] Terminal device 10 and server 20 can communicate with each other via a network. This network can be a wired network or a wireless network.

[0046] The method provided in this application embodiment can be executed by a computer device in each step. The computer device can be any electronic device capable of data storage and processing. For example, the computer device can be... Figure 1 Server 20 in the middle can be Figure 1 The terminal device 10 can also be another device other than the terminal device 10 and the server 20.

[0047] Please refer to Figure 2 The figure shows a schematic diagram of a video person recognition result provided in an embodiment of this application. In the figure, 200 is one of the image frames of the video, 201 is a face image obtained from image frame 200, and 202 is the recognition result "Celebrity X" obtained based on face image 201.

[0048] In related technologies, the process generally includes steps such as face detection, facial feature extraction, face indexing, data association, and result synthesis. First, face detection identifies all faces appearing in the video frames. Then, a fixed-length facial feature is extracted from each detected face, and the feature is indexed in a constructed facial feature database to obtain preliminary recognition results. Subsequently, facial appearance or motion information is used to associate the faces to form their motion trajectories in the video. Finally, all face recognition results are combined in each motion trajectory to obtain the final recognition result for that trajectory. The data association step typically employs MOT (Modal of the Time) technology or clustering methods. While these video face recognition technologies use MOT and clustering methods to associate multiple faces in a video sequence and ultimately combine the recognition results through voting to improve accuracy and recall, these methods still have several problems. One issue is that these technologies do not differentiate between the quality of face images. Videos contain faces of varying quality due to factors such as sharpness, resolution, motion blur, and lighting conditions. Different quality faces present different challenges to the accuracy and robustness of the recognition methods. High-quality faces are relatively easy for face recognition methods to use, resulting in higher accuracy and precision in the model output. Conversely, low-quality faces are less accurate. Related technologies maintain a consistent voting weight for each face when aggregating trajectory results, neglecting the impact of face quality differences on accuracy. This leads to the indexing results of low-quality faces negatively affecting the overall voting result. Furthermore, video footage often features scene transitions and target occlusion, particularly in movies, TV series, and variety shows. These situations cause the motion trajectory of the same person to be broken into multiple fragmented trajectories during the data association step. This reduces sample diversity during the voting aggregation of individual trajectories, hindering the voting aggregation process.

[0049] Unlike other methods, the technical solution provided in this application can distinguish between high-quality and low-quality faces, thereby improving the accuracy of comprehensive voting. First, the face quality of all faces in the trajectory is estimated, and a certain threshold is set to distinguish between high-quality and low-quality faces. Subsequently, the voting weight of high-quality faces is increased while the voting weight of low-quality faces is decreased, thereby mitigating the adverse effects of low-quality faces and ultimately improving the accuracy of comprehensive voting.

[0050] In addition, the technical solution provided in this application can associate multiple trajectories and combine them through voting again, increasing the sample diversity of the combined results and improving the recognition recall rate. Multiple separated trajectories are re-clustered to form a more complete motion trajectory. Simultaneously, the technical solution provided in this application proposes a cascaded voting mechanism, dividing the voting and combining process into two steps: intra-trajectory voting and inter-trajectory voting. This increases the number of samples referenced in the voting and combining process, resulting in the recognition and recall of more faces.

[0051] The technical solution provided in this application will be described in detail below through several embodiments. This invention provides a video face recognition method for building celebrity recognition capabilities in video types such as movies, TV series, variety shows, and animations. It outputs the face position and celebrity information for each frame of video, and finally stores the face recognition data in structured video data. This data is applied to editing individual videos of target celebrities, editing videos of similar celebrities, etc., enabling the separate editing of videos of target celebrities based on comprehensive videos, or timely annotation of facial images appearing in video frames to provide prompts to viewers.

[0052] Please refer to Figure 3 The diagram illustrates a flowchart of a video person recognition method according to an embodiment of this application. The execution entity for each step of the method can be a computer device. The method may include at least one of the following steps (310-330):

[0053] Step 310: Extract a sequence of face images from the video. The sequence of face images includes multiple face images.

[0054] This application does not limit the type of video. It can be a video that has undergone post-production processing, such as a movie, TV series, variety show, or animation. It can also be a video of the road surface captured by a road camera, a video of a person's face captured by a home camera, etc. Any video that constitutes a continuous frame can be included in the protection scope of this application.

[0055] In some embodiments, video decoding can obtain video frames with a temporal relationship. Optionally, a subset of video frames can be extracted from these frames to extract face images. Optionally, video frames can be extracted at fixed intervals, such as the first frame, the second frame, the third frame, and so on. Of course, to reduce processing volume and save processing costs, video frames can also be extracted at intervals of n frames (n is an integer greater than 1). However, reducing the number of extracted video frames may lead to a decrease in the accuracy of the recognition results to some extent. Therefore, the frame extraction interval can be determined by comprehensively considering recognition accuracy and processing cost.

[0056] In some embodiments, the extracted video frames are detected to obtain the coordinates of the face bounding box and the key point recognition results in each video frame. Optionally, in some embodiments, the face bounding box is a rectangle, and the face bounding box coordinates can be the coordinates of the four vertices of the rectangle, or the coordinates of one vertex of the rectangle and the length and width values. In some embodiments, the key point recognition results can be key point coordinates, where key points can be understood as points that can represent a face, such as facial features. Optionally, the key point recognition results are the coordinates of at least 10 points extracted from the facial features on the face. This application does not limit the definition of key points; any point that can represent a face can be called a key point.

[0057] In some embodiments, the face image in each video frame can be determined based on the face bounding box coordinates in each video frame. Optionally, the number of face images in a video frame can be one or more. In some embodiments, each identified face image can be mapped to one face image sequence among multiple face image sequences corresponding to previous video frames. Optionally, based on the first i (i is a positive integer) video frames, m (m is a positive integer) face image sequences can be obtained. Then, based on the face image of the (i+1)th video frame, the face image can be mapped to at least one of the m face image sequences, or it can stand alone as the first face image in the (m+1)th face image sequence. In some embodiments, the face image sequence corresponding to the face image is determined based on the appearance information of the face image. In some embodiments, the face image sequence to which the face image of the current video frame belongs can also be determined based on the position information of the face bounding box of the current video frame and the face bounding box of the previous frame. This application does not limit the method for determining the face image sequence to which the face image belongs; it can be based on appearance information or on motion information (face bounding box information).

[0058] In some embodiments, a face detection model can be used to obtain the coordinates of the face bounding box and the coordinates of the face key points corresponding to each video frame. In some embodiments, the face detection model includes, but is not limited to, at least one of RetinaFace and MTCNN; the specific detection principles are not elaborated here.

[0059] In this application, the sequence of face images may also be referred to as face trajectory or trajectory, and this application does not limit this. In some embodiments, different trajectories correspond to different numbers. Optionally, the numbering starts from 1, and each trajectory corresponds to a number that can be uniquely represented by that number.

[0060] Step 320: Obtain the preliminary recognition results and quality scores corresponding to the multiple face images included in the face image sequence. The preliminary recognition results are used to indicate the candidate objects corresponding to the face images, and the quality scores are used to indicate the image quality of the face images.

[0061] Preliminary Recognition Result: The preliminary recognition result inferred from the face image. This result includes candidate objects and their corresponding confidence scores, which can also be understood as candidate objects and their corresponding probability values ​​or similarities. Optionally, the preliminary recognition result includes multiple candidate objects and their respective confidence scores. Optionally, the preliminary recognition result of a face image is "Celebrity a, 0.99; Celebrity b, 0.88; Celebrity c, 0.76;...". Here, celebrity a, celebrity b, and celebrity c represent candidate objects, 0.99 indicates that the face image is celebrity a with a confidence / similarity score of 0.99, 0.88 indicates that the face image is celebrity b with a confidence / similarity score of 0.88, and 0.76 indicates that the face image is celebrity c with a confidence / similarity score of 0.76. In some embodiments, the preliminary recognition result of the face image is obtained through a face recognition model.

[0062] Quality Score: A score characterizing the image quality of a face image. In some embodiments, the quality score of a face image is determined based on information such as brightness and sharpness, and different levels of image quality are assigned according to the quality score. In other embodiments, the quality score of a face image is determined based on whether the face image can be correctly classified, and a level of image quality is assigned according to the quality score. In some embodiments, image quality and quality score are directly proportional; the higher the quality score, the higher the image quality, and the higher the corresponding image quality level. In some embodiments, image quality is a piecewise function; for example, face images with a quality score of 80 or higher are determined to have high quality, and face images with a quality score below 80 are determined to have low quality. In some embodiments, the quality score of a face image is determined using a quality assessment model, and the quality level of the face image is determined based on the quality score. In some embodiments, the quality assessment model is trained on multiple face images with correct classification labels. The quality score is adjusted based on whether the face image is correctly classified. Gradient descent is then used to adjust the parameters in the model. When a face image is correctly classified, the quality score of the face image is increased; when a face image is misclassified, the quality score of the face image is decreased. The quality assessment model trained on multiple face images with correct classification labels can be used to evaluate unlabeled face images, and is used for the quality score of face images in the embodiments of this application.

[0063] refer to Figure 4This diagram illustrates a quality scoring method for face images according to an embodiment of this application. Images 41, 42, 43, and 44 are multiple face images. Through a quality assessment model, quality scores corresponding to these multiple face images and face quality levels determined based on these scores can be obtained. The face quality level reflects the image quality of the face image. Optionally, based on the image quality, the face quality level is divided into multiple levels, including high quality, medium quality, and low quality. Figure 4 As can be seen from the data, the quality scores of face images 41 and 42 are 80 and 82 respectively, which are high quality scores. Therefore, the face quality level of face images 41 and 42 is high quality. The quality scores of face images 43 and 44 are 18 and 0 respectively, which are low quality scores. Therefore, the face quality level of face images 43 and 44 is low quality.

[0064] In some embodiments, such as Figure 4 The quality assessment model shown can evaluate face quality. First, it obtains a quality score for the face image. The input to the quality assessment model is the face image after face alignment. After evaluating the quality of the face image, the model outputs the corresponding quality score, which is an integer ranging from 0 to 100. In some embodiments, all detected faces are evaluated using the quality assessment model to obtain a quality score for each face, which is used for subsequent trajectory clustering and cascade voting processes, as detailed in the following embodiments. Second, it distinguishes between high-quality and low-quality faces. A certain threshold is set; faces greater than the threshold are defined as high-quality faces, and faces less than the threshold are defined as low-quality faces.

[0065] In some embodiments, step 320 includes at least one of the following steps (320-2 to 320-6, not shown in the figure).

[0066] Step 320-2: Based on the key point recognition results of the face image, perform key point alignment to obtain the aligned face image.

[0067] In some embodiments, the key point recognition results of a face image are obtained based on a face detection model, wherein the key point recognition results include the coordinate information of multiple key points.

[0068] Keypoint alignment refers to aligning key points obtained from a face image with key points of a standard frontal face shape, that is, transforming non-frontal face images, such as those in profile, into frontal face images. In some embodiments, keypoint alignment can transform non-frontal face images into frontal face images, thereby improving the accuracy of face recognition and increasing the recognition rate of the target person.

[0069] Step 320-4: Obtain the feature information corresponding to the face image from the aligned face image using the face recognition model. Based on the feature information corresponding to the face image and the feature information of each object contained in the feature library, determine the preliminary recognition result corresponding to the face image.

[0070] In some embodiments, feature information of the aligned face image can be extracted based on a face recognition model. In some embodiments, the feature information is a feature vector. Optionally, a 512-dimensional feature vector corresponding to the face image can be obtained from the aligned face image using a face recognition model. This 512-dimensional feature vector can characterize the face image.

[0071] In some embodiments, the feature information of the aforementioned face image is compared with the feature information of each object contained in the feature library, and a preliminary recognition result of the face image is determined based on the comparison result. In some embodiments, the feature library includes multiple objects and feature information corresponding to each of the multiple objects, wherein the multiple objects correspond to correct recognition results or target persons. Optionally, the feature information of the aforementioned face image is compared with the feature information of each object contained in the feature library, and the object closest to the aforementioned face image is determined based on the similarity between the feature information. Optionally, the similarity between the feature information includes, but is not limited to, cosine similarity and Euclidean distance. In some embodiments, the preliminary recognition result of the face image is also referred to as the face index result.

[0072] In some embodiments, the 512-dimensional feature vector corresponding to the face image is normalized into a one-dimensional vector, and the similarity is calculated with the vector in the feature library, for example, the cosine similarity is calculated. The calculated similarity results are sorted in descending order, and the objects corresponding to the top N (the first N, where N is an integer greater than 1) similarity results are used to determine the preliminary recognition result of the face image.

[0073] Step 320-6: Process the aligned face image using a quality assessment model to obtain the corresponding quality score for the face image.

[0074] In some embodiments, the aligned face image is input into a quality assessment model, which can obtain a quality score corresponding to the face image. Optionally, the quality score can be in the form of 0-100 or 0-10. In some embodiments, the quality level corresponding to the face image can also be obtained through the quality assessment model. For example, inputting an aligned face image can result in the face image being classified as high quality.

[0075] In the technical solution provided in this application embodiment, by extracting key point recognition results from face images, obtaining feature information of face images based on key point recognition results, and determining preliminary recognition results of face images based on feature information, the preliminary recognition results can be made more accurate, and further make the target person corresponding to the face image sequence determined based on the preliminary recognition results closer to the prototype person in the image.

[0076] Step 330: Based on the quality scores corresponding to the multiple face images, vote on the candidate figures corresponding to the multiple face images to determine the target person corresponding to the face image sequence.

[0077] This application does not limit the number of target individuals; there may be only one or at least two.

[0078] In some embodiments, a face image sequence contains multiple face images, each corresponding to a preliminary recognition result, meaning each image corresponds to multiple candidate individuals. If a face image has a higher quality score, its image quality is considered higher, and consequently, the voting weight of the preliminary recognition result corresponding to that face image is also higher. In other words, the voting weight of the preliminary recognition result of a face image is directly proportional to its image quality. In some embodiments, the face image sequence includes four face images, with voting weights q1, q2, q3, and q4, respectively. Taking the first face image as an example, the preliminary recognition result of the first face image includes two candidate individuals with confidence levels z1 and z2, respectively. Therefore, for the first face image, its preliminary recognition result is voted on by multiplying q1 by z1 and z2 respectively, yielding the voted result for the first face image. By combining the voting results of the recognition results of multiple face images in the face image sequence, the target person corresponding to the face image sequence is determined.

[0079] The technical solution provided in this application embodiment can ultimately obtain the face location information corresponding to each video frame and the face recognition information identified based on the face image. Here, the face recognition information can be understood as trajectory information and target person information.

[0080] The recall rate mentioned in this application refers to the ratio of the number of video frames recalled from a video for the same target person to the number of video frames in the video in which the target person appears. Of course, the recall rate can also be understood as the ratio of the types of target persons recalled from a video to the types of target persons actually appearing in the video for the same video.

[0081] The technical solution provided in this application extracts facial image sequences from videos. Each facial image sequence includes multiple facial images. Based on the quality scores of the multiple facial images belonging to the same sequence, a vote is cast for the candidate individuals corresponding to the multiple facial images. The target person corresponding to the facial image sequence is determined based on the voting results. This application uses the image quality of facial images to vote on the recognition results, and determines the target person corresponding to the facial image sequence composed of multiple facial images based on the voting results. In other words, the image quality of the facial images affects the voting results for the candidate individuals corresponding to the facial images. Therefore, the technical solution provided in this application, based on the preliminary recognition results of facial images, uses the image quality of facial images as an important factor affecting the voting results for the candidate individuals corresponding to the facial images, making the final determined target person more accurate, and thus resulting in a higher recall rate for people in the video.

[0082] Please refer to Figure 5 This illustrates a flowchart of a video person recognition method according to another embodiment of this application. The execution entity for each step of this method can be a computer device. The method may include at least one of the following steps (310-336):

[0083] Step 310: Extract a sequence of face images from the video. The sequence of face images includes multiple face images.

[0084] Step 322: Obtain the preliminary recognition results and quality scores corresponding to the multiple face images included in the face image sequence. The preliminary recognition results are used to indicate the candidate objects corresponding to the face images, and the quality scores are used to indicate the image quality of the face images. The preliminary recognition results include: at least one candidate object corresponding to the face image, and the confidence level corresponding to the at least one candidate object.

[0085] Step 332: Determine the voting weights of the multiple face images based on their respective quality scores.

[0086] Optionally, the voting weight corresponding to face images with quality scores less than a threshold is set as a first value; the voting weight corresponding to face images with quality scores greater than a threshold is set as a second value; wherein the second value is greater than the first value.

[0087] In some embodiments, the voting weights are divided into two levels: the voting weight corresponding to face images with quality scores below a threshold is determined to be a lower first value, and the voting weight corresponding to face images with quality scores above the threshold is determined to be a higher second value. Optionally, the threshold is 80, the first value is 0, and the second value is the reciprocal of the number of face images above the threshold. The voting weight corresponding to face images with quality scores below 80 is determined to be the first value 0, and the voting weight corresponding to face images with quality scores above 80 is determined to be the higher second value. Optionally, if there are four face images in the face image sequence, three of which have quality scores above 80 and only one has a quality score below 80, the voting weight of the face image with a quality score below 80 is determined to be 0, and the voting weights of the other three are each determined to be 1 / 3, where 3 represents the number of face images above the threshold.

[0088] In some embodiments, when the quality score equals a threshold, the voting weight of the face image can be considered as a first value or a second value. This application does not limit the value of the voting weight of the face image when the quality score equals the threshold; it can be at least one of the first and second values, or other values. In some embodiments, considering the accuracy of the voting results, the voting weight of the face image can be considered as the first value when the quality score equals the threshold. In some embodiments, considering the diversity of the samples participating in the voting, the voting weight of the face image can be considered as the second value when the quality score equals the threshold.

[0089] In some embodiments, the quality score can be directly used as the voting weight, and the voting weight is proportional to the quality score. In some embodiments, the voting weight is not limited to the first and second values, but can be determined based on different quality scores, with at least three voting weights.

[0090] The technical solution provided in this application reduces the voting weight of face images with scores below a threshold and increases the voting weight of face images with scores above a threshold. This ensures that face images with low quality scores are excluded from voting as much as possible, meaning that voting is primarily based on high-quality face images for initial recognition results. Voting based on high-quality face images can improve the accuracy of the final face recognition result. Furthermore, differentiating voting weights based on face quality makes the voting method novel and diverse.

[0091] Step 334: Determine the target confidence level for each candidate based on the voting weights corresponding to the multiple face images, the at least one candidate corresponding to each of the multiple face images, and the confidence level corresponding to each candidate.

[0092] In some embodiments, step 334 includes at least one of the following steps (334-2-334-4, not shown in the figure).

[0093] Step 334-2: For each face image in the multiple face images, multiply the confidence level of at least one candidate for each face image by the voting weight of the face image to obtain the intermediate confidence level of at least one candidate.

[0094] In some embodiments, a face image has three candidates r1, r2, and r3, with confidence levels d1, d2, and d3 (d1, d2, and d3 are all positive numbers), and the voting weight of the face image is q5 (q5 is a positive number). Then, the confidence level d1 of candidate r1 is multiplied by q5, the confidence level d2 of candidate r2 is multiplied by q5, and the confidence level d3 of candidate r3 is multiplied by q5 to obtain the median confidence level of candidate r1 as d1*q5, the median confidence level of candidate r2 as d2*q5, and the median confidence level of candidate r3 as d3*q5.

[0095] Step 334-4: Sum the intermediate confidence scores of each candidate object corresponding to multiple face images to obtain the target confidence score of each candidate object.

[0096] In some embodiments, in a face image sequence, there are two face images, both of which contain a candidate person P. The median confidence of P in the first face image is p1, and the median confidence of P in the second face image is p2. Then the target confidence of candidate P is p1 + p2 (p1 and p2 are both positive numbers).

[0097] Step 336: The candidate objects whose target confidence level meets the first condition are identified as the target objects corresponding to the face image sequence.

[0098] In some embodiments, based on the preliminary recognition results of the face images and the voting weights, the target confidence level of each candidate is obtained, and the candidate whose target confidence level meets the first condition is identified as the target person corresponding to the face image sequence.

[0099] In some embodiments, the first condition is the maximum value of the target confidence score. Optionally, the candidate corresponding to the maximum value of the target confidence score is used to determine the target person corresponding to the face image sequence.

[0100] In some embodiments, the first condition is to sort the target confidence scores in descending order, with the target confidence score ranking among the first X (X being an integer greater than 1) in the sorted queue. Optionally, the candidate individuals corresponding to the confidence scores at the first 5 positions are determined as the target individuals corresponding to the face image sequence. The technical solution provided by the embodiments of this application enriches the methods for determining target individuals and improves the accuracy of the determined target individuals.

[0101] refer to Figure 6 This illustrates a schematic diagram of in-trajectory voting provided in one embodiment of this application. Figure 6 As shown, the method includes at least one of the following steps (61-63).

[0102] Step 61: Obtain preliminary recognition results and face quality scores, and distinguish between high-quality and low-quality faces based on thresholds.

[0103] First, preliminary recognition results (face index results) are obtained for each of the four face images. Each recognition result includes multiple candidate objects and their corresponding confidence scores. For example, the first five face index results for the first face image in the figure are "Celebrity a1, 0.92; Celebrity a2, 0.87; Celebrity a3, 0.84; Celebrity a4, 0.83; Celebrity a5, 0.80", where Celebrity a1, Celebrity a2, Celebrity a3, Celebrity a4, and Celebrity a5 represent the first five candidate objects. It can be observed that their corresponding confidence scores decrease in that order. In this embodiment, the top N (the first N, where N is a positive integer) recognition results can be selected, or all recognition results can be selected as preliminary recognition results. This application does not limit this choice.

[0104] Step 62: Assign different weights to faces of different qualities according to a certain strategy.

[0105] Figure 6 Of the four face images, the first three have a quality score of 60 or higher, therefore they are classified as high quality. The fourth face image has a quality score of only 10, therefore it is classified as low quality. The voting weight for the three high-quality face images is set at 33.3% (1 / 3), while the voting weight for the low-quality face images is set to 0. In other words, low-quality face images are not considered, and equal weighting is applied to the high-quality face images (all voting weights are the same).

[0106] Step 63: Calculate a weighted average of the top five recognition results for all faces according to their weights, and select the celebrity with the highest confidence score as the celebrity for trajectory recognition. This confidence score is then used as the trajectory recognition confidence score.

[0107] from Figure 6As can be seen, the results of voting on the three face images are as follows: Celebrity a1, 0.93; Celebrity a2, 0.88; Celebrity a3, 0.85; Celebrity a4, 0.82; Celebrity a6, 0.84; Celebrity a7, 0.80; Celebrity a5, 0.80; Celebrity a8, 0.80. Among them, Celebrity a1 has the highest confidence score, so Celebrity a1 is selected as the celebrity (or target person) for this trajectory, and its confidence score of 0.93 is taken as the recognition confidence score (or target confidence score) for this trajectory.

[0108] The technical solution provided in this application sets a threshold value to distinguish the voting weights corresponding to face images. Based on the candidates and their corresponding confidence levels in the preliminary recognition results, the target person in the face image sequence is determined according to the product of the confidence level and the voting weight. This can reduce the voting weight for low-quality faces, so that low-quality faces do not participate in voting as much as possible, thereby improving the accuracy of the recognition results.

[0109] Please refer to Figure 7 The diagram illustrates a flowchart of a video person recognition method according to another embodiment of this application. The execution entity for each step of the method can be a computer device. The method may include at least one of the following steps (310-370):

[0110] Step 310: Extract a sequence of face images from the video. The sequence of face images includes multiple face images.

[0111] Step 320: Obtain the preliminary recognition results and quality scores corresponding to the multiple face images included in the face image sequence. The preliminary recognition results are used to indicate the candidate objects corresponding to the face images, and the quality scores are used to indicate the image quality of the face images.

[0112] Step 330: Based on the quality scores corresponding to the multiple face images, vote on the candidate figures corresponding to the multiple face images to determine the target person corresponding to the face image sequence.

[0113] Step 340: The number of face image sequences extracted from the video is multiple. Based on the quality scores corresponding to the multiple face images included in each face image sequence, the representative face image corresponding to each face image sequence is determined.

[0114] In some embodiments, for each face image sequence, the face image with the highest quality score in the sequence is identified as the representative face image corresponding to the sequence. Using the face image with the highest quality score as the representative face image for clustering ensures accurate representation of the face image sequence, resulting in more accurate final clustering results.

[0115] In some embodiments, the top X face images with the highest quality scores in each face image sequence may be determined as representative face images corresponding to the face image sequence. This application does not limit the number of representative face images.

[0116] Step 350: Cluster the representative face images corresponding to each face image sequence to obtain at least one cluster, and each cluster includes at least one representative face image.

[0117] Based on the feature information corresponding to each representative face image, clustering is performed based on the similarity between the feature information to obtain at least one cluster; among them, representative face images with a quality score greater than a threshold participate in clustering, while representative face images with a quality score less than a threshold do not participate in clustering.

[0118] In some embodiments, representative face images of certain trajectories may have relatively low image quality. For accuracy considerations, a threshold is set: representative face images below the threshold are not included in clustering, while those above the threshold are. This is to prevent inaccurate clustering results caused by clustering representative face images with very low quality scores. In some embodiments, the threshold is set to 80, meaning representative face images with a quality score greater than 80 are included in clustering, while those with a quality score less than 80 are not. In some embodiments, representative face images with a quality score equal to the threshold may or may not participate in clustering; this application does not limit this. In some embodiments, considering clustering accuracy, representative face images with a quality score equal to the threshold are not included in clustering. In other embodiments, considering the diversity of samples participating in clustering, representative face images with a quality score equal to the threshold are included in clustering.

[0119] The technical solution provided in this application only clusters representative face images. Compared with clustering all face images, this can effectively reduce the number of face images participating in clustering, reduce processing costs, and reduce overhead.

[0120] In some embodiments, the feature information is feature vectors, and clustering is performed based on the similarity between feature vectors representing face images. Optionally, the similarity includes, but is not limited to, cosine similarity and Euclidean distance.

[0121] Step 360: Merge the face image sequences to which at least one representative face image belonging to the same cluster belongs to obtain a set of face image sequences.

[0122] refer to Figure 8 This illustrates a schematic diagram of trajectory merging provided in one embodiment of this application. Figure 8The method shown includes at least one of the following steps (81-83).

[0123] Step 81: Select representative face images from each trajectory based on the quality score.

[0124] The face quality scores for each face trajectory are sorted, and one or more faces with the highest quality scores are selected as the representative face images for that trajectory. If the highest quality score is lower than a pre-set quality score, the trajectory is not further aggregated. For example... Figure 8 In the process, one representative face image is selected for each of trajectories 1, 2, and 3.

[0125] Step 82: Use a clustering algorithm to perform face clustering on the features representing the face images.

[0126] First, obtain the facial feature information of each trajectory representing a face image. Then, apply a clustering algorithm (such as DBSCAN) to cluster these facial feature information, divide them into different face clusters based on facial similarity, and aggregate faces with high similarity.

[0127] like Figure 8 As shown, three representative face images are clustered to obtain two clusters. The first cluster includes the representative face images of trajectory 1 and trajectory 2, while the second cluster only includes the representative face image of trajectory 3.

[0128] Step 83: Merge the trajectories corresponding to the representative face images of the same cluster.

[0129] Within each cluster, the trajectories corresponding to each representative face image are merged to obtain a more complete trajectory for the same target person. This ultimately achieves further aggregation of trajectories.

[0130] like Figure 8 As shown, the trajectories 1 and 2 corresponding to the representative face images of the first cluster are merged to obtain the merged trajectory.

[0131] The technical solution provided in this application can obtain a more complete trajectory based on the same target person by merging the scattered trajectories together. At the same time, by extracting representative face images, some face images can participate in the recognition, which improves the sample diversity for target person recognition to a certain extent.

[0132] Step 370: Determine the target person corresponding to the face image sequence set based on the target person corresponding to each face image sequence in the face image sequence set.

[0133] In some embodiments, step 370 includes at least one of the following steps (370-2-370-6, not shown in the figure).

[0134] Step 370-2: Determine the voting weight of each face image sequence in the face image sequence set.

[0135] In some embodiments, each face image sequence in the face image sequence set is voted on equally, that is, the voting weight of each face image sequence is the same.

[0136] In some embodiments, the voting weight of a face image sequence is determined based on the quality score of a representative face image in the face image sequence. Optionally, the quality score of the representative face image is used as the voting weight of the face image sequence, or the voting weight of the face image sequence is proportional to the quality score of the face image.

[0137] Step 370-4: Based on the voting weights corresponding to each face image sequence and the confidence levels of the target individuals corresponding to each face image sequence, determine the confidence levels of at least one candidate target individual.

[0138] The voting weights corresponding to each face image sequence are multiplied by the confidence scores of the target individuals corresponding to each face image sequence to obtain the weighted confidence scores of each target individual; the weighted confidence scores of the same target individual are summed to obtain the confidence scores of at least one candidate target individual.

[0139] Step 370-6: The target person whose confidence level meets the second condition is identified as the target person corresponding to the set of face image sequences.

[0140] In this application embodiment, the second condition is similar to the first condition. The second condition has been discussed in detail in the above embodiments, so the first condition will not be repeated here. The number of target individuals corresponding to the merged trajectory or face image sequence set is not limited in this application; it can be only one or multiple.

[0141] refer to Figure 9 This illustrates a schematic diagram of inter-trajectory voting provided in one embodiment of this application. Figure 9 The method shown includes at least one of the following steps (91-94).

[0142] The cascaded voting mentioned in this application refers to voting within the trajectory plus voting between trajectories. In related technologies, voting is only performed within the trajectory, and voting is never performed between trajectories, let alone the final recognition result of the trajectory is readjusted based on the result of voting between trajectories.

[0143] Step 91: Obtain the voting results within the trajectory.

[0144] Step 92: Obtain the results of trajectory clustering.

[0145] Step 93: In each cluster of trajectory clustering, vote on the identification results between trajectories to complete the trajectory clustering.

[0146] like Figure 9 As shown, equal-weighted voting is performed on each trajectory. For example, the voting result for merging trajectory 1 is calculated as follows: 0.90*1 / 2+0.92*1 / 2=0.91. The specific calculation method is similar to that of voting within a trajectory, so it will not be elaborated here.

[0147] Step 94: Complete the voting between trajectories to obtain the final recognition result of the merged trajectory.

[0148] like Figure 9 As shown, the star with the highest voting confidence is finally selected as the star identification result of the merged trajectory, and this confidence level is used as the identification confidence of the merged trajectory. The voting results between trajectories are used as the final output of the cascaded voting.

[0149] In this embodiment, because faces may be obscured in the video, causing discontinuities in the facial trajectory, related technologies do not re-merge these trajectories when they are discontinuous. This results in multiple trajectories corresponding to the same person being output. This embodiment uses clustering to group similar trajectories together, forming more complete trajectories. Related technologies only use intra-trajectory voting, not inter-trajectory voting. This means that if the intra-trajectory recognition result is incorrect, the output will be incorrect and cannot be corrected. This embodiment combines intra-trajectory and inter-trajectory voting. Inter-trajectory voting can, to some extent, correct errors in intra-trajectory voting, utilizing more facial trajectory information for comprehensive voting, thus improving the accuracy of the voting mechanism in trajectory recognition. Furthermore, faces not recalled in some trajectories can be recalled through inter-trajectory voting, thereby improving the recall rate.

[0150] Please refer to Figure 10 The diagram illustrates a block diagram of a video person recognition method according to an embodiment of this application. The entity executing each step of this method can be... Figure 1 In the implementation environment of the scheme shown, the terminal device 10 can be either the client of the target application or the execution subject of each step. Figure 1Server 20 is shown in the implementation environment of the scheme. In the following method embodiments, for ease of description, only the execution subject of each step is described as a "computer device". The method may include at least one of the following steps (S1 to S8).

[0151] Step S1: Video decoding. The video is decoded to obtain video frames with temporal relationships. In particular, in order to balance recognition accuracy and processing speed, fixed-interval frame extraction can be performed during video decoding to reduce the number of video frames processed. Here, fixed intervals of 1, 2, and 3 frames are usually used.

[0152] Step S2: Face detection. All decoded video frame images are input into the face detection model. The model detects faces in the frame and outputs the bounding box coordinates and facial landmark coordinates for each frame. The face detection model can be, but is not limited to, RetinaFace, MTCNN, etc.

[0153] Step S3: Face Feature Extraction and Face Indexing. First, based on the detected face bounding box coordinates and keypoint coordinates from Step 2, the face image is cropped and deformed to achieve face alignment. Then, the aligned face image is input into a face recognition model to obtain fixed-length face features, which serve as the appearance representation of the face image. Finally, face indexing is performed in a pre-built database of celebrity face features. The extracted face features are compared with features in the database, and the celebrity information with the highest similarity in the database is used as the face indexing result for this face image. This similarity is used as the confidence score of the index. A fixed threshold is set; faces with an index confidence score greater than the threshold are considered celebrity faces, and faces with an index confidence score less than the threshold are considered non-celebrity faces. The face recognition model used here can be, but is not limited to, models such as CosFace and ArcFace, and the face indexing tool can be, but is not limited to, tools such as Faiss.

[0154] Step S4: Face Quality Assessment. Input the face images (after face alignment in Step 3) into the face quality model. The model assesses the image quality and outputs a quality score for each face. Face quality is positively correlated with the quality score. Models such as EQFace can be used here, but are not limited to these.

[0155] Step S5: Face Trajectory. Detected faces with temporal relationships are correlated to obtain a series of face trajectories. Each trajectory records the state of one person in the video, while preserving the bounding box, facial features, and face index of each face within the trajectory. Methods used here include, but are not limited to, DeepSort.

[0156] Step S6: Trajectory Clustering. A certain number of faces are selected from each trajectory according to a specific strategy. Subsequently, a clustering operation is performed on all selected faces, merging the trajectories corresponding to faces in the same cluster to obtain a more complete motion trajectory. The clustering methods used here include, but are not limited to, DBSCAN.

[0157] Step S7: Cascaded voting, which includes intra-trajectory voting and inter-trajectory voting. Intra-trajectory voting combines the recognition results of each trajectory in step five to obtain the trajectory's star recognition result and recognition confidence. Inter-trajectory voting combines the recognition results of the trajectories after trajectory clustering in step six, further aggregating the combined results of intra-trajectory voting to obtain the final recognition result.

[0158] Step S8: Output the final recognition result and store it in the video structured information.

[0159] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0160] Please refer to Figure 11 The diagram illustrates a block diagram of a video person recognition device according to an embodiment of this application. The device 900 may include: an image extraction module 910, a result acquisition module 920, and a person determination module 930.

[0161] The image extraction module 910 is used to extract a sequence of face images from a video, the sequence of face images including multiple face images.

[0162] The result acquisition module 920 is used to acquire the preliminary recognition results and quality scores corresponding to the multiple face images included in the face image sequence. The preliminary recognition results are used to indicate the candidate objects corresponding to the face images, and the quality scores are used to indicate the image quality of the face images.

[0163] The person determination module 930 is used to vote on the candidate persons corresponding to the multiple face images based on the quality scores corresponding to the multiple face images respectively, and determine the target person corresponding to the face image sequence.

[0164] In some embodiments, the preliminary identification result includes: at least one candidate object corresponding to the face image, and the confidence level corresponding to each of the at least one candidate object.

[0165] In some embodiments, such as Figure 12 As shown, the character determination module 930 includes a weight determination unit 932, a confidence determination unit 934, and a character determination unit 936.

[0166] The weight determination unit 932 is used to determine the voting weights corresponding to the multiple face images based on the quality scores corresponding to the multiple face images respectively.

[0167] The confidence determination unit 934 is used to determine the target confidence level of each candidate based on the voting weights corresponding to the plurality of face images, the at least one candidate corresponding to each of the plurality of face images, and the confidence level of each candidate.

[0168] The person determination unit 936 is used to determine the candidate person whose target confidence level meets the first condition as the target person corresponding to the face image sequence.

[0169] In some embodiments, the weight determination unit 932 is used to set the voting weight corresponding to the face image with a quality score less than a threshold value as a first value; and to set the voting weight corresponding to the face image with a quality score greater than a threshold value as a second value; wherein the second value is greater than the first value.

[0170] In some embodiments, the confidence determination unit 934 is configured to, for each of the plurality of face images, multiply the confidence level corresponding to at least one candidate for each face image by the voting weight corresponding to the face image to obtain the intermediate confidence level corresponding to each of the at least one candidate.

[0171] The confidence determination unit 934 is used to add up the intermediate confidence scores corresponding to each of the candidate objects corresponding to the plurality of face images to obtain the target confidence score corresponding to each of the candidate objects.

[0172] In some embodiments, such as Figure 10 As shown, the result acquisition module 920 includes a key point alignment unit 922, a result determination unit 924, and a score determination unit 926.

[0173] The key point alignment unit 922 is used to perform key point alignment based on the key point recognition results of the face image to obtain an aligned face image.

[0174] The result determination unit 924 is used to obtain the feature information corresponding to the face image from the aligned face image through a face recognition model, and determine the preliminary recognition result corresponding to the face image based on the feature information corresponding to the face image and the feature information of each object contained in the feature library.

[0175] The scoring determination unit 926 is used to process the aligned face image through a quality evaluation model to obtain a quality score corresponding to the face image.

[0176] In some embodiments, the number of facial image sequences extracted from the video is multiple.

[0177] In some embodiments, such as Figure 12 As shown, the device further includes: a representative image determination module 940, an image clustering module 950, and a merging module 960.

[0178] The representative image determination module 940 is used to determine the representative face image corresponding to each face image sequence based on the quality scores corresponding to the multiple face images included in each face image sequence.

[0179] The image clustering module 950 is used to cluster the representative face images corresponding to each of the face image sequences to obtain at least one cluster, and each cluster includes at least one representative face image.

[0180] The merging module 960 is used to merge the face image sequences to which at least one representative face image belonging to the same cluster belongs, to obtain a set of face image sequences.

[0181] The person determination module 930 is further configured to determine the target person corresponding to the face image sequence set based on the target person corresponding to each face image sequence in the face image sequence set.

[0182] In some embodiments, the representative image determination module 940 is used to determine, for each face image sequence, the face image with the highest quality score in the face image sequence as the representative face image corresponding to the face image sequence.

[0183] In some embodiments, the image clustering module 950 is used to cluster based on the similarity between the feature information corresponding to each of the representative face images to obtain at least one cluster; wherein, representative face images with a quality score greater than a threshold participate in clustering, and representative face images with a quality score less than the threshold do not participate in clustering.

[0184] In some embodiments, the weight determination unit 932 is used to determine the voting weight corresponding to each face image sequence in the face image sequence set.

[0185] The confidence determination unit 934 is used to determine the confidence level of at least one candidate target person based on the voting weight corresponding to each of the face image sequences and the confidence level of the target person corresponding to each of the face image sequences.

[0186] The person determination unit 936 is used to determine the target person whose confidence level meets the second condition as the target person corresponding to the set of face image sequences.

[0187] In some embodiments, the confidence determination unit 934 is used to multiply the voting weight corresponding to each of the face image sequences by the confidence of the target person corresponding to each of the face image sequences to obtain the weighted confidence of each target person; and to sum the weighted confidence of the same target person to obtain the confidence of at least one candidate target person.

[0188] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0189] Please refer to Figure 13 The diagram shows a structural block diagram of a computer device 2100 provided in one embodiment of this application.

[0190] Typically, computer device 2100 includes a processor 2101 and a memory 2102.

[0191] Processor 2101 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 2101 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). Processor 2101 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 2101 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 2101 may also include an AI processor for handling computational operations related to machine learning.

[0192] The memory 2102 may include one or more computer-readable storage media, which may be non-transitory. The memory 2102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 2102 is used to store a computer program configured to be executed by one or more processors to implement the video person recognition method described above.

[0193] Those skilled in the art will understand that Figure 13 The structure shown does not constitute a limitation on computer device 2100 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0194] In an exemplary embodiment, a computer-readable storage medium is also provided, wherein a computer program is stored therein, which, when executed by a processor, implements the above video person recognition method.

[0195] Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or optical disc, etc. The random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).

[0196] In an exemplary embodiment, a computer program product is also provided, the computer program product including a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the computer device to perform the video person recognition method described above.

[0197] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.

[0198] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A video person recognition method, characterized in that, The method includes: Extract a sequence of facial images from a video, the sequence of facial images including multiple facial images; The preliminary recognition results and quality scores corresponding to multiple face images included in the face image sequence are obtained respectively. The preliminary recognition results are used to indicate the candidate objects corresponding to the face images, and the quality scores are used to indicate the image quality of the face images. The preliminary recognition results include: at least one candidate object corresponding to the face image, and the confidence level corresponding to the at least one candidate object. Based on the quality scores corresponding to the multiple face images, the voting weights corresponding to the multiple face images are determined. For each face image among the plurality of face images, the confidence level corresponding to at least one candidate for each face image is multiplied by the voting weight corresponding to the face image to obtain the intermediate confidence level corresponding to each at least one candidate. The intermediate confidence scores corresponding to each candidate object corresponding to the plurality of face images are summed to obtain the target confidence score corresponding to each candidate object. Candidates whose target confidence level meets the first condition are identified as the target individuals corresponding to the face image sequence.

2. The method of claim 1, wherein, The step of determining the voting weights for each of the multiple face images based on their respective quality scores includes: The voting weight corresponding to the face image whose quality score is less than the threshold value is set to the first value; The voting weight corresponding to the face image whose quality score is greater than the threshold value is set as the second value; The second value is greater than the first value.

3. The method of claim 1, wherein, The step of obtaining the preliminary recognition results and quality scores corresponding to the multiple face images included in the face image sequence includes: Based on the key point recognition results of the face image, key point alignment is performed to obtain the aligned face image; The face recognition model obtains the feature information corresponding to the face image from the aligned face image, and determines the preliminary recognition result corresponding to the face image based on the feature information corresponding to the face image and the feature information of each object contained in the feature library. The aligned face image is processed using a quality assessment model to obtain a quality score corresponding to the face image.

4. The method of claim 1, wherein, The method further includes extracting multiple facial image sequences from the video, and the method also includes: Based on the quality scores corresponding to the multiple face images included in each face image sequence, the representative face image corresponding to each face image sequence is determined. Cluster the representative face images corresponding to each of the face image sequences to obtain at least one cluster, and each cluster includes at least one representative face image; The face image sequences to which at least one representative face image belonging to the same cluster belong are merged to obtain a set of face image sequences; The target person corresponding to the set of face image sequences is determined based on the target person corresponding to each face image sequence in the set of face image sequences.

5. The method of claim 4, wherein, The step of determining the representative face image corresponding to each face image sequence based on the quality scores corresponding to the multiple face images included in each face image sequence includes: For each face image sequence, the face image with the highest quality score in the face image sequence is determined as the representative face image corresponding to the face image sequence.

6. The method of claim 4, wherein, The step of clustering representative face images corresponding to each of the face image sequences to obtain at least one cluster includes: Based on the feature information corresponding to each of the representative face images, clustering is performed based on the similarity between the feature information to obtain at least one cluster; Among them, representative face images with a quality score greater than the threshold participate in clustering, while representative face images with a quality score less than the threshold do not participate in clustering.

7. The method of claim 4, wherein, The step of determining the target person corresponding to the face image sequence set based on the target person corresponding to each face image sequence in the face image sequence set includes: Determine the voting weight corresponding to each face image sequence in the face image sequence set; Based on the voting weights corresponding to each of the face image sequences and the confidence levels of the target individuals corresponding to each of the face image sequences, at least one confidence level corresponding to a candidate target individual is determined. The target person whose confidence level meets the second condition is identified as the target person corresponding to the set of face image sequences.

8. The method of claim 7, wherein, The step of determining the confidence level of at least one candidate target person based on the voting weights corresponding to each of the face image sequences and the confidence level of the target person corresponding to each of the face image sequences includes: The voting weight corresponding to each of the face image sequences is multiplied by the confidence level of the target person corresponding to each of the face image sequences to obtain the weighted confidence level of each target person. The weighted confidence scores corresponding to the same target person are summed to obtain the confidence scores corresponding to at least one candidate target person.

9. A video person recognition apparatus, characterized by comprising: The device includes: An image extraction module is used to extract a sequence of face images from a video, wherein the sequence of face images includes multiple face images; The result acquisition module is used to acquire preliminary recognition results and quality scores corresponding to multiple face images included in the face image sequence. The preliminary recognition results are used to indicate the candidate objects corresponding to the face images, and the quality scores are used to indicate the image quality of the face images. The preliminary recognition results include: at least one candidate object corresponding to the face image, and the confidence level corresponding to the at least one candidate object. The person identification module is used to determine the voting weight corresponding to each of the multiple face images based on the quality scores corresponding to each of the multiple face images; for each face image among the multiple face images, the confidence level corresponding to at least one candidate for each face image is multiplied by the voting weight corresponding to the face image to obtain the intermediate confidence level corresponding to the at least one candidate for each face image; the intermediate confidence levels corresponding to each candidate for each of the multiple face images are added together to obtain the target confidence level corresponding to each candidate for each face image; and the candidate whose target confidence level satisfies a first condition is identified as the target person corresponding to the face image sequence.

10. The apparatus of claim 9, wherein, The result acquisition module includes: The key point alignment unit is used to align key points based on the key point recognition results of the face image to obtain an aligned face image. The result determination unit is used to obtain the feature information corresponding to the face image from the aligned face image through a face recognition model, and determine the preliminary recognition result corresponding to the face image based on the feature information corresponding to the face image and the feature information of each object contained in the feature library. The scoring determination unit is used to process the aligned face image through a quality evaluation model to obtain the quality score corresponding to the face image.

11. The apparatus of claim 9, wherein, The number of facial image sequences extracted from the video is multiple, and the device further includes: The representative image determination module is used to determine the representative face image corresponding to each face image sequence based on the quality scores corresponding to the multiple face images included in each face image sequence. The image clustering module is used to cluster the representative face images corresponding to each of the face image sequences to obtain at least one cluster, and each cluster includes at least one representative face image; The merging module is used to merge the face image sequences to which at least one representative face image belonging to the same cluster belongs, to obtain a set of face image sequences. The person determination module is further configured to determine the target person corresponding to the face image sequence set based on the target person corresponding to each face image sequence in the face image sequence set.

12. A computer device, comprising: The computer device includes a processor and a memory, the memory storing a computer program that is loaded and executed by the processor to implement the method as claimed in any one of claims 1 to 8.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that is loaded and executed by a processor to implement the method as described in any one of claims 1 to 8.

14. A computer program product, characterised in that, The computer program product includes a computer program stored in a computer-readable storage medium, which a processor reads from and executes to implement the method as described in any one of claims 1 to 8.