A face recognition method and device, electronic equipment and storage medium

CN116469146BActive Publication Date: 2026-08-07BEIJING IQIYI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING IQIYI TECH CO LTD
Filing Date
2023-04-18
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

在实际应用中,客户端所截取的视频帧中包含的人脸可能是模糊人脸(比如因切换镜头,或者视频拍摄时拍摄的设备、光线以及角度等引起的人脸模糊),或者侧脸、背影等,使得发送给云端进行人脸识别的图像的质量无法保证,进而导致云端无法准确识别图像中包含的人物或者无法识别图像中包含的人物

Benefits of technology

[0067]This invention provides a face recognition method, apparatus, electronic device, and storage medium. Face recognition is divided into two parts: face detection and face feature comparison. A client device acquires multiple video frames within a preset time range of the target timestamp in a target video playback. Then, using target detection and face classification technologies, it determines the target video frame from the multiple video frames to be sent to a cloud device for face feature comparison. Compared to existing methods that send a captured video frame from a movie or TV show playback to the cloud device for face recognition, this method improves the image quality and robustness of the video frame sent to the cloud device for face recognition. It also fully utilizes the computing resources of the client device. Furthermore, the cloud device can more accurately identify faces in the target video frame based on the high-quality image and feed back the corresponding face recognition result to the client device, thus improving the accuracy and recall rate of face recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116469146B_ABST
    Figure CN116469146B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a face recognition method and device, electronic equipment and storage medium, the method comprises: a client acquires the video frame in the preset time range of the target timestamp in the target playing video, obtains the video frame to be processed, carries out target detection on each video frame in the video frame to be processed to obtain the target detection result of each video frame, classifies the face belonging to different objects in each video frame of the video frame to be processed based on each target detection result, obtains the face classification result, determines the target video frame from the video frame to be processed based on the face classification result and sends it to the cloud; the cloud extracts the face feature in the target video frame to obtain the target face feature, matches the face feature in the face feature library with the target face feature, and when matching, the corresponding character information is sent to the client as the face recognition result of the target video frame; the client receives the face recognition result, improves the accuracy of face recognition and the recall rate of recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a face recognition method, device, electronic device, and storage medium. Background Technology

[0002] Facial recognition is a biometric technology that identifies individuals based on their facial features. With the development of artificial intelligence, facial recognition is becoming increasingly widely used. For example, it's used to identify people in movies and TV shows, or to recognize people in images captured by electronic devices.

[0003] In scenarios where related technologies are used to identify characters in movies and TV dramas, the user triggers the client to capture a specific video frame during playback. The client then sends the captured frame to the cloud. Upon receiving the frame, the cloud extracts facial features and compares them with features in a pre-set facial feature database. The comparison result is then fed back to the client as the facial recognition result, which is then displayed to the user. However, in practical applications, the faces captured in the video frame may be blurry (e.g., due to camera transitions, or blurry faces caused by the shooting equipment, lighting, or angle), or they may be profiles or back views. This compromises the quality of the image sent to the cloud for facial recognition, leading to the cloud's inability to accurately identify individuals within the image or to recognize individuals at all. Summary of the Invention

[0004] The purpose of this invention is to provide a face recognition method, device, electronic device, and storage medium to improve the accuracy and recall rate of face recognition results. The specific technical solution is as follows:

[0005] In a first aspect of this invention, a face recognition method is provided, applied to a client device, the method comprising:

[0006] Obtain video frames within a preset time range of the target timestamp in the target playback video to obtain the video frames to be processed;

[0007] Target detection is performed on each video frame in the video frame to be processed to obtain the target detection results of each video frame.

[0008] Based on the target detection results, the faces belonging to different objects in each video frame of the video frame to be processed are classified to obtain the face classification results belonging to different objects;

[0009] Based on the face classification results, the target video frame is determined from the video frames to be processed;

[0010] The target video frame is sent to a cloud device so that the cloud device can extract facial features from the target video frame to obtain target facial features. The target facial features are then matched with facial features contained in a preset facial feature library. The person information corresponding to the facial features that match the target facial features in the preset facial feature library is sent to the client device as the facial recognition result of the target video frame. The preset facial feature library includes facial features and the person information corresponding to the facial features.

[0011] Receive the face recognition results sent by the cloud device.

[0012] In one possible implementation, the step of classifying faces belonging to different objects in each video frame of the video frame to be processed based on the target detection results, to obtain face classification results belonging to different objects, includes:

[0013] According to the playback order of each video frame in the target playback video, face tracking is performed on each of the target detection results in each video frame of the video frame to be processed to obtain the face tracking results of each target detection result; wherein, the face tracking results include: front face, side face, back view, global or local;

[0014] For each of the aforementioned face tracking results, the video frames corresponding to frontal faces are retained to obtain candidate video frames;

[0015] The faces belonging to different objects in each of the candidate video frames are classified to obtain the face classification results belonging to different objects.

[0016] In one possible implementation, the method further includes:

[0017] Cache the video frames played by the target video;

[0018] The step of obtaining video frames within a preset time range of the target timestamp in the target playback video to obtain the video frames to be processed includes:

[0019] Upon receiving a face recognition instruction, the system acquires cached video frames within a first preset time range before the current timestamp of the target video, and video frames played and cached within a second preset time range after the current timestamp of the target video, to obtain the video frames to be processed.

[0020] In one possible implementation, determining the target video frame from the video frames to be processed based on the face classification result includes:

[0021] For each face classification result, select one frame image that meets the preset quality requirements from the video frames corresponding to that face classification result, and use it as the target video frame corresponding to that face classification result.

[0022] In one possible implementation, the target detection result includes the position of the target detection box corresponding to the detected target in the video frame, and the method further includes:

[0023] The location of the target detection box corresponding to the target in the target video frame, the timestamp of the target video frame, and the identification information of the target video being played are sent to the cloud device.

[0024] In a second aspect of this invention, a face recognition method is also provided, applied to a cloud device, the method comprising:

[0025] The client device receives a target video frame sent by the client device. The target video frame is a video frame obtained by the client device within a preset time range of the target timestamp in the target playback video, which is then processed. The client device performs target detection on each video frame in the video frame to be processed, and classifies the faces of different objects in each video frame of the video frame to be processed based on the target detection results. The target video frame is then determined from the video frame to be processed based on the face classification results.

[0026] Extract facial features from the target video frame to obtain the target facial features;

[0027] The target facial features are matched with facial features contained in a preset facial feature database;

[0028] If a face feature matching the target face feature exists in the preset face feature library, the person information corresponding to the face feature matching the target face feature is used as the face recognition result of the target video frame, and the face recognition result is sent to the client device. The preset face feature library includes face features and the person information corresponding to the face feature.

[0029] In one possible implementation, the method further includes:

[0030] The system receives the position of the target detection box corresponding to the target in the target video frame sent by the client device, the timestamp corresponding to the target video frame, and the identification information of the target playing video.

[0031] Determine whether the identification information of the target video corresponds to a data relationship table;

[0032] If so, query the data relationship table to see if there is a video frame timestamp that is the same as the timestamp corresponding to the target video frame, and a detection box position that matches the position of the target detection box corresponding to the detected target in the target video frame; wherein, the data relationship table includes: the correspondence between the video frame timestamp and the detection box position and the person recognition result; the detection box position matching the position of the target detection box corresponding to the detected target in the target video frame means that: the detection box position is the same as the position of the target detection box corresponding to the detected target in the target video frame, or the distance between the detection box position and the position of the target detection box corresponding to the detected target in the target video frame is less than a preset distance threshold;

[0033] If it exists, the person recognition result corresponding to the timestamp of the target video frame and the position of the target detection box in the target video frame in the data relationship table is determined as the face recognition result of the target video frame and sent to the client device; otherwise, the face features in the target video frame are extracted to obtain the target face features.

[0034] In one possible implementation, the preset facial feature library includes: a first facial feature library and a second facial feature library, wherein the facial features contained in the first facial feature library are used to characterize the facial features of the target object in the role it plays in the video, and the facial features contained in the second facial feature library are used to characterize the facial features of the target object's real identity.

[0035] The step of matching the target facial feature with facial features contained in a preset facial feature database, and then, if a facial feature matching the target facial feature exists in the preset facial feature database, using the person information corresponding to the facial feature matching the target facial feature as the facial recognition result of the target video frame, and sending the facial recognition result to the client device, includes:

[0036] Based on the identification information of the target video and the correspondence table between the video identification information and the first face feature database, it is determined whether there is a first face feature database corresponding to the target video.

[0037] If a first facial feature library exists that corresponds to the target video, the target facial features are matched with the facial features contained in the first facial feature library.

[0038] If the target facial feature matches the facial features contained in the first facial feature library, the person information corresponding to the facial feature that matches the target facial feature in the first facial feature library is used as the facial recognition result of the target video frame and sent to the client device.

[0039] If the target facial feature does not match the facial features contained in the first facial feature library, or if there is no first facial feature library corresponding to the target video, the target facial feature will be matched with the facial features contained in the second facial feature library.

[0040] If the target facial feature matches a facial feature contained in the second facial feature library, the person information corresponding to the facial feature in the second facial feature library that matches the target facial feature is used as the facial recognition result of the target video frame and sent to the client device.

[0041] In one possible implementation, the preset facial feature library includes: a first facial feature library and a second facial feature library, wherein the facial features contained in the first facial feature library are used to characterize the facial features of the target object in the role it plays in the video, and the facial features contained in the second facial feature library are used to characterize the facial features of the target object's real identity.

[0042] The step of matching the target facial feature with facial features contained in a preset facial feature database, and then, if a facial feature matching the target facial feature exists in the preset facial feature database, using the person information corresponding to the facial feature matching the target facial feature as the facial recognition result of the target video frame, and sending the facial recognition result to the client device, includes:

[0043] Based on the identification information of the target video and the correspondence table between the video identification information and the first face feature database, it is determined whether there is a first face feature database corresponding to the target video.

[0044] If a first facial feature library exists that corresponds to the target video, the target facial features are matched with the facial features contained in the first facial feature library to obtain a first matching result.

[0045] The target facial features are matched with the facial features contained in the second facial feature database to obtain a second matching result;

[0046] Based on the first matching result and the second matching result, the face recognition result of the target video frame is obtained, and the face recognition result is sent to the client device.

[0047] In one possible implementation, the method further includes:

[0048] If the target facial features do not match the facial features contained in the second facial feature library, the target video frame and the position of the target detection box corresponding to the detected target in the target video frame are sent to the operation terminal so that the operation terminal updates the facial features of the real identity of the detected target in the target video frame and the person information corresponding to the facial features to the second facial feature library.

[0049] If the target facial features do not match the facial features contained in the first facial feature library corresponding to the target video, or if there is no first facial feature library corresponding to the target video, the target video frame, the position of the target detection box corresponding to the detected target in the target video frame, and the identification information of the target video are sent to the operator, so that the operator updates the facial features of the detected target in the target video frame under the role played in the target video, and the person information corresponding to the facial features to the first facial feature library.

[0050] In a third aspect of the present invention, a face recognition device is also provided, applied to a client device, the device comprising:

[0051] The video frame acquisition module is used to acquire video frames within a preset time range of the target timestamp in the target playback video, and obtain the video frames to be processed.

[0052] The target detection module is used to perform target detection on each video frame in the video frame to be processed, and obtain the target detection results of each video frame.

[0053] The face classification module is used to classify faces belonging to different objects in each video frame of the video frame to be processed based on the detection results of each target, and obtain face classification results belonging to different objects;

[0054] The video frame determination module is used to determine the target video frame from the video frames to be processed based on the face classification result;

[0055] A video frame sending module is used to send the target video frame to a cloud device, so that the cloud device can extract facial features from the target video frame, obtain target facial features, match the target facial features with facial features contained in a preset facial feature library, and send the person information corresponding to the facial features that match the target facial features in the preset facial feature library as the facial recognition result of the target video frame to a client device. The preset facial feature library includes facial features and the person information corresponding to the facial features.

[0056] The recognition result receiving module is used to receive the face recognition results sent by the cloud device.

[0057] In a fourth aspect of this invention, a face recognition device is also provided, applied to a cloud device, the device comprising:

[0058] The video frame receiving module is used to receive target video frames sent by the client device. The target video frame is: the video frame to be processed obtained by the client device from the video frame with the target timestamp within a preset time range in the target playback video, and the target detection is performed on each video frame in the video frame to be processed. After classifying the faces of different objects in each video frame of the video frame to be processed based on the target detection results, the target video frame is determined from the video frame to be processed based on the face classification results.

[0059] The face feature extraction module is used to extract face features from the target video frame to obtain the target face features;

[0060] A face feature matching module is used to match the target face features with face features contained in a preset face feature database;

[0061] A face recognition module is used to, when a face feature matching the target face feature exists in the preset face feature library, use the person information corresponding to the face feature matching the target face feature as the face recognition result of the target video frame, and send the face recognition result to the client device. The preset face feature library includes face features and the person information corresponding to the face feature.

[0062] In another aspect of the present invention, an electronic device is also provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus.

[0063] Memory, used to store computer programs;

[0064] When a processor executes a program stored in memory, it implements the steps of any of the above-described face recognition methods.

[0065] In another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements the steps of any of the above-described face recognition methods.

[0066] In another aspect of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the steps of any of the face recognition methods described above.

[0067] This invention provides a face recognition method, apparatus, electronic device, and storage medium. Face recognition is divided into two parts: face detection and face feature comparison. A client device acquires multiple video frames within a preset time range of the target timestamp in a target video playback. Then, using target detection and face classification technologies, it determines the target video frame from the multiple video frames to be sent to a cloud device for face feature comparison. Compared to existing methods that send a captured video frame from a movie or TV show playback to the cloud device for face recognition, this method improves the image quality and robustness of the video frame sent to the cloud device for face recognition. It also fully utilizes the computing resources of the client device. Furthermore, the cloud device can more accurately identify faces in the target video frame based on the high-quality image and feed back the corresponding face recognition result to the client device, thus improving the accuracy and recall rate of face recognition. Attached Figure Description

[0068] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0069] Figure 1 This is a flowchart illustrating a face recognition method according to an embodiment of the present invention;

[0070] Figure 2 This is a schematic diagram illustrating a face recognition result in an embodiment of the present invention;

[0071] Figure 3 This is a flowchart illustrating another face recognition method in an embodiment of the present invention;

[0072] Figure 4 This is a schematic diagram illustrating the process of the client acquiring the target video frame in an embodiment of the present invention;

[0073] Figure 5 This is a flowchart illustrating another face recognition method in an embodiment of the present invention;

[0074] Figure 6 This is a flowchart illustrating another face recognition method in an embodiment of the present invention;

[0075] Figure 7 This is a schematic diagram of a facial feature matching implementation method in an embodiment of the present invention;

[0076] Figure 8 This is a schematic diagram of the cloud-based face recognition process in an embodiment of the present invention;

[0077] Figure 9 This is a schematic diagram illustrating the update of the facial feature database on the operational side in an embodiment of the present invention;

[0078] Figure 10 This is an interactive schematic diagram of a face recognition method according to an embodiment of the present invention;

[0079] Figure 11 This is an interactive schematic diagram of a face recognition method according to an embodiment of the present invention;

[0080] Figure 12 This is a schematic diagram of the structure of a face recognition device according to an embodiment of the present invention;

[0081] Figure 13 This is a schematic diagram of another face recognition device in an embodiment of the present invention;

[0082] Figure 14 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0083] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.

[0084] With the development of artificial intelligence technology, the application of artificial intelligence algorithms (such as visual recognition, OCR (optical character recognition), NLP (natural language processing), and speech and voiceprint recognition algorithms) to identify various entity information in film and television content is increasing, such as the identification of characters, clothing, decorations, and scenic spots in film and television content.

[0085] On the one hand, the rapid pace of film and television production updates means that characters, clothing, decorations, and scenic spots within the content are constantly being updated, leading to a growing demand from users to identify these elements. On the other hand, while most shots in film and television productions feature faces, these faces are not always shown directly; there may be side profiles, back views, or even just voices. Furthermore, characters may appear with makeup and costumes, making character identification in film and television productions quite challenging.

[0086] In related technologies, character recognition in film and television content is mainly achieved through facial feature comparison, with a large facial feature database stored in the cloud. The client acquires the current frame image to be identified and sends it to the cloud. The cloud extracts facial features from the current frame image and compares them with those in the facial feature database, sending the comparison result back to the client, which then displays the result to the user. However, in reality, the faces in the current frame image acquired by the client may be blurry, or only in profile or from behind, making it difficult to guarantee the quality of the image sent to the cloud for facial feature comparison. Consequently, the cloud may be unable to accurately identify the characters in the current frame image, or may be unable to identify the characters at all.

[0087] To improve the accuracy and recall rate of face recognition, embodiments of the present invention provide a face recognition method, apparatus, electronic device, and storage medium. The face recognition method provided by this embodiment of the present invention, applied to a client device, includes:

[0088] Obtain video frames within a preset time range of the target timestamp in the target playback video to obtain the video frames to be processed;

[0089] Target detection is performed on each video frame in the video frame to be processed to obtain the target detection results of each video frame.

[0090] Based on the target detection results, the faces belonging to different objects in each video frame of the video frame to be processed are classified to obtain the face classification results belonging to different objects;

[0091] Based on the face classification results, the target video frame is determined from the video frames to be processed;

[0092] The target video frame is sent to a cloud device so that the cloud device can extract facial features from the target video frame to obtain target facial features. The target facial features are then matched with facial features contained in a preset facial feature library. The person information corresponding to the facial features that match the target facial features in the preset facial feature library is sent to the client device as the facial recognition result of the target video frame. The preset facial feature library includes facial features and the person information corresponding to the facial features.

[0093] Receive the face recognition results sent by the cloud device.

[0094] This invention provides a face recognition method in which a client device acquires multiple video frames within a preset time range of a target timestamp in a target video playback. Then, using target detection and face classification technologies, it determines the target video frame from the multiple video frames to be sent to a cloud device for face feature comparison. Compared to existing methods that send a captured video frame from a movie or TV show playback to the cloud device for face recognition, this method improves the image quality and robustness of the video frame sent to the cloud device for face recognition. It also fully utilizes the computing resources of the client device. Furthermore, it enables the cloud device to more accurately identify faces in the target video frame based on the high-quality image and to feed back the corresponding face recognition result to the client device, thereby improving the accuracy and recall rate of face recognition.

[0095] The following is a detailed description of a face recognition method provided by an embodiment of the present invention:

[0096] The face recognition method provided in this invention can be applied to face recognition scenarios in videos such as movies and TV dramas.

[0097] In scenarios involving facial recognition in videos, the varying performance levels of client devices can lead to delays in receiving recognition results from the client device, negatively impacting the user experience. Conversely, deploying the entire facial recognition process on a cloud-based device, where the client device only acquires the image frames to be recognized, makes the process reliant on the cloud, resulting in high service costs. Therefore, this invention divides facial recognition into two parts: face detection and face feature comparison. The face detection process is deployed on the client device to fully utilize its computing resources and caching capabilities, while the face feature comparison process is deployed on the cloud, reducing the cloud's service costs for face detection. This allows the client device and cloud device to collaborate in completing the facial recognition process.

[0098] like Figure 1 As shown in the figure, the face recognition method provided in this embodiment of the invention, applied to a client device, can be implemented through the following steps:

[0099] S101, Obtain video frames within a preset time range of the target timestamp in the target playback video to obtain the video frames to be processed.

[0100] In this embodiment of the invention, the video frame caching module in the client device is used to cache the video frames required for face recognition, thereby fully utilizing the caching capabilities of the client device. The target playback video can be the video currently playing in the client device's player, or any playable video in the client device's player, etc. The preset time range can be set according to the client device's computing power, or based on empirical values, etc. For example, the preset time range can be two seconds before, two seconds after, or one or two seconds before and after the target timestamp, etc. The target timestamp can be the current timestamp.

[0101] In one example, upon receiving a face recognition command, the client device retrieves video frames within a time range of two seconds before, two seconds after, or one or two seconds before and after the current timestamp of the target playback video, thus obtaining the video frames to be processed. For instance, if one second of the target playback video corresponds to 10 video frames, retrieving video frames within a time range of two seconds before or after the current timestamp yields 20 video frames to be processed.

[0102] S102, Perform target detection on each video frame in the video frame to be processed, and obtain the target detection results for each video frame.

[0103] For each video frame in the video to be processed, a pre-trained object detection model is used to detect the objects contained in that video frame, and the object detection result for that video frame is obtained. The pre-trained object detection model is trained based on sample video frames and the object detection results of those sample video frames.

[0104] The detected target can be a human image or face. The target detection result may include: the location information of the anchor box (or target detection box) corresponding to the human image or face, and information such as the number of human images or faces contained in the video frame. In one example, a single target detection result contains the location information of the target detection box corresponding to one human image or face; a single video frame may correspond to multiple target detection results.

[0105] S103, based on the detection results of each target, classify the faces belonging to different objects in each video frame of the video frame to be processed, and obtain the face classification results belonging to different objects.

[0106] A video frame may contain multiple faces, and the faces appearing in each video frame may belong to different objects (i.e., detection targets). Therefore, after obtaining the target detection results for each video frame, a pre-trained classification model or an object clustering algorithm can be used to classify the faces belonging to different objects in each video frame to be processed, and obtain the face classification results belonging to different objects.

[0107] The pre-trained classification model can be trained based on the target detection results of the samples and the corresponding classification results. Target clustering algorithms can include, for example, K-means (k-means clustering algorithm), DBSCAN (Density-Based Spatial Clustering of Applications with Noise), ISODATA (Iterative Self-organizing Data Analysis Techniques Algorithm), and BIRCH (Balanced Iterative Reducing and Clustering using Hierarchies), etc.

[0108] For example, the video frame to be processed contains 20 video frames, each containing 5 faces corresponding to 5 objects. The 20 video frames correspond to 100 object detection results. These 100 object detection results are input into a pre-trained classification model, or the 100 object detection results are clustered using an object clustering algorithm, so as to classify the faces belonging to different objects in the 20 video frames and obtain five categories of face classification results.

[0109] S104, Based on the face classification results, determine the target video frame from the video frames to be processed.

[0110] For each class in the face classification results, select one video frame from the video frames corresponding to that class as the target video frame. For example, one frame can be randomly selected from the video frames corresponding to each class as the target video frame, or the frame with the highest clarity and / or the face being frontal and / or the face not being obscured can be selected as the target video frame.

[0111] S105, the target video frame is sent to the cloud device so that the cloud device can extract the facial features in the target video frame, obtain the target facial features, match the target facial features with the facial features contained in the preset facial feature library, and send the person information corresponding to the facial features that match the target facial features in the preset facial feature library as the facial recognition result of the target video frame to the client device.

[0112] The preset facial feature database includes facial features and the corresponding personal information for each facial feature. For example, the personal information corresponding to a facial feature may include, but is not limited to: personal name, photo, age, place of birth, zodiac sign, stage name, resume, and acting experience.

[0113] After the client device determines the target video frame, it sends the target video frame to the cloud device, so that the cloud device can extract and match the features of the face contained in the target video frame, realize face recognition, and feed back the face recognition result to the client device.

[0114] In one possible implementation, the target detection result includes the position of the target detection box corresponding to the detected target in the video frame. While sending the target video frame to the cloud device, the position of the target detection box corresponding to the detected target in the target video frame, the timestamp of the target video frame, and the identification information of the target playing video can also be sent to the cloud device.

[0115] The target detection result includes the position of the target detection box corresponding to the detected target in the video frame. That is, after performing target detection on the video frame, the position information of the anchor box (or target detection box) corresponding to the portrait or face in the video frame is obtained. The detected target is the object being detected. Each video frame in the target playback video acquired by the client device has a corresponding timestamp within a preset time range of the target timestamp. The identification information of the target playback video is used to identify the target playback video. This identification information can be the name, encoding, or tag information of the target playback video, etc.

[0116] In this embodiment of the invention, the position of the target detection box corresponding to the target in the target video frame, the timestamp of the target video frame, and the identification information of the target video are sent to the cloud device so that the cloud device can more accurately identify the face contained in the target video frame, thereby further improving the accuracy of face recognition and the recall rate of the recognition results.

[0117] S106 receives the face recognition results sent by the cloud device.

[0118] In one example, after the client device receives the facial recognition result sent by the cloud device, it can also display the facial recognition result at any location on the playback interface of the target video. For example, such as... Figure 2 As shown, the face recognition result is displayed on the playback interface of the target video. The recognition result includes person A, and person A's information is displayed on one side of the playback interface of the target video.

[0119] This invention provides a face recognition method in which a client device acquires multiple video frames within a preset time range of a target timestamp in a target video playback. Then, using target detection and face classification technologies, it determines the target video frame from the multiple video frames to be sent to a cloud device for face feature comparison. Compared to existing methods that send a captured video frame from a movie or TV show playback to the cloud device for face recognition, this method improves the image quality and robustness of the video frame sent to the cloud device for face recognition. It also fully utilizes the computing resources of the client device. Furthermore, it enables the cloud device to more accurately identify faces in the target video frame based on the high-quality image and to feed back the corresponding face recognition result to the client device, thereby improving the accuracy and recall rate of face recognition.

[0120] In one possible implementation, step S103, based on the detection results of each target, classifies faces belonging to different objects in each video frame of the video frame to be processed, and obtains the face classification results belonging to different objects. This implementation may include:

[0121] According to the playback order of each video frame in the target playback video, face tracking is performed on each target detection result in each video frame of the video frame to be processed to obtain the face tracking result of each target detection result;

[0122] For each face tracking result, the video frames corresponding to the frontal faces are retained to obtain candidate video frames;

[0123] The faces belonging to different objects in each candidate video frame are classified to obtain the face classification results belonging to different objects.

[0124] In this embodiment of the invention, the above-mentioned target detection is performed on each video frame in the video frame to be processed. The target detection results of each video frame include: the position of the target detection box corresponding to the portrait or face in the video frame, and one target detection result corresponds to the position of the target detection box of a portrait or face in the video frame. For each target detection result, according to the playback order of each video frame in the target playback video (i.e., the order of the timestamps corresponding to each video frame), face tracking is performed on the position of the target detection box in the target detection result in each video frame of the video frame to be processed to obtain the face tracking result of the target detection result. The face tracking result may include: frontal face, side face, back view, global or local, background, blurred, etc.

[0125] In one example, methods such as motion segmentation, optical flow, or stereo vision can be used to perform face tracking on the target detection results in each video frame of the video frame to be processed. The face tracking results can include: the identifier of the video frame where the face is located, and information such as whether the face is a frontal face, a side face, a back view, a global or local view, a background, or a blurred view in each video frame.

[0126] After obtaining the face tracking results, video frames corresponding to each target detection result that contain faces in profile, back view, partial view, background, or blurred, or that are redundant (or repetitive), are discarded. Only video frames corresponding to frontal faces are retained, resulting in candidate video frames. Then, for each candidate video frame, faces belonging to different objects are classified according to the target detection results, yielding classification results for faces belonging to different objects. In one example, a pre-trained classification model or a target clustering algorithm can be used to classify faces belonging to different objects in each candidate video frame; this will not be elaborated further in this embodiment.

[0127] In this embodiment of the invention, according to the playback order of each video frame in the target playback video, face tracking is performed on the target detection results in each video frame of the video frame to be processed to obtain the face tracking results of each target detection result. The video frames corresponding to the face tracking results of the frontal face are retained, while video frames with low image quality, face tracking results of side face, back view, partial, background or blurry, and redundant video frames are discarded. The faces of different objects in each candidate video frame are classified so that the client device can determine the target video frame with higher image quality from the video frames corresponding to the face classification results of different objects after face classification and send it to the cloud device for face recognition. This improves the image quality and robustness of the video frames sent to the cloud device for face recognition.

[0128] In one possible implementation, the above method may further include: caching the video frames played by the target playback video;

[0129] Accordingly, step S101 above, which obtains video frames within a preset time range of the target timestamp in the target playback video to obtain the video frame to be processed, may include: upon receiving a face recognition instruction, obtaining video frames that have been cached within a first preset time range before the current timestamp of the target playback video, and video frames that have been played and cached within a second preset time range after the current timestamp of the target playback video to obtain the video frame to be processed.

[0130] In this embodiment of the invention, a video frame caching module in the client device is used to cache the video frames played by the target video in real time, so as to make full use of the caching capability of the client device. The first preset time range and the second preset time range may be the same or different. The first preset time range and the second preset time range may be 1 second, 2 seconds, or 3 seconds, etc.

[0131] For example, both the first and second preset time ranges are 1 second. Upon receiving a face recognition command, the client device acquires the cached video frames within 1 second before the current timestamp of the target playback video, and the video frames played and cached within 1 second after the current timestamp of the target playback video, to obtain the video frames to be processed. If 1 second corresponds to 10 frames, then the video frames within 1 second before and 1 second after the current timestamp are acquired, i.e., 20 video frames within 1 second before and after the current timestamp, to obtain the video frames to be processed.

[0132] In this embodiment of the invention, video frames played by the target video are cached, making full use of the caching capability of the client device. When a face recognition instruction is received, video frames within a preset time range (a first preset time range and / or a second preset time range) before and after the current timestamp in the target video are obtained to improve the quality and robustness of the face recognition images sent by the client device to the cloud device.

[0133] In one possible implementation, the above step S104, which determines the target video frame from the video frames to be processed based on the face classification result, may include:

[0134] For each face classification result, select one frame image that meets the preset quality requirements from the video frames corresponding to that face classification result, and use it as the target video frame corresponding to that face classification result.

[0135] The preset quality requirements may include meeting at least one of the following: highest clarity, face is frontal, face is not obscured, etc.

[0136] For example, the face classification result contains faces in 5 categories (corresponding to 5 people). For each face classification result, if there is only one video frame corresponding to the face classification result, that video frame is used as the target video frame corresponding to the face classification result; if there are multiple video frames corresponding to the face classification result, the frame with the highest clarity, the face being frontal, or the face not being occluded is selected from the video frames corresponding to the face classification result as the target video frame corresponding to the face classification result, and the other video frames corresponding to the face classification result are discarded.

[0137] In this embodiment of the invention, the image with the highest clarity and the face being frontal or unobstructed is selected from the video frames corresponding to the face classification result as the target video frame corresponding to the face classification result, so as to improve the image quality sent to the cloud device for face recognition.

[0138] like Figure 3 As shown, another face recognition method provided in this embodiment of the invention, applied to a client device, can be implemented through the following steps:

[0139] S301, buffer the video frames played by the target video;

[0140] S302, upon receiving a face recognition instruction, obtain the video frames that have been cached within a first preset time range before the current timestamp of the target video, and the video frames that have been played and cached within a second preset time range after the current timestamp of the target video, to obtain the video frames to be processed;

[0141] S303, Perform target detection on each video frame in the video frame to be processed, and obtain the target detection results of each video frame;

[0142] S304, according to the playback order of each video frame in the target playback video, perform face tracking on each target detection result in each video frame of the video frame to be processed, and obtain the face tracking result of each target detection result;

[0143] The face tracking results include: frontal face, side face, back view, global or local;

[0144] S305, For each face tracking result, retain the video frame corresponding to the frontal face to obtain candidate video frames;

[0145] S306, Classify the faces belonging to different objects in each candidate video frame to obtain the face classification results belonging to different objects;

[0146] S307, Based on the face classification results, determine the target video frame from the video frames to be processed;

[0147] S308, send the target video frame to the cloud device so that the cloud device can extract the facial features in the target video frame, obtain the target facial features, match the target facial features with the facial features contained in the preset facial feature library, and send the person information corresponding to the facial features that match the target facial features in the preset facial feature library as the facial recognition result of the target video frame to the client device;

[0148] The preset facial feature database includes facial features and the corresponding person information.

[0149] S309 receives the face recognition results sent by the cloud device.

[0150] This invention provides a face recognition method in which a client device acquires multiple video frames within a preset time range of the current timestamp of a target video being played. Then, using object detection and face classification technologies, it determines the target video frame from the multiple video frames to be sent to a cloud device for face feature comparison. Compared to existing methods that send a captured video frame from the playback of a movie or TV show to the cloud device for face recognition, this method improves the image quality and robustness of the video frame sent to the cloud device for face recognition. Furthermore, it fully utilizes the caching capabilities and computing resources of the client device. This enables the cloud device to more accurately identify faces in the target video frame based on the high-quality image and to feed back the corresponding face recognition result to the client device, thereby improving the accuracy and recall rate of face recognition.

[0151] For example, such as Figure 4 As shown, the client device includes a video frame caching module, a face tracking and detection module, and a face classification and merging module. The video frame caching module caches the video frames played by the target video. The client device has a user-triggerable face recognition function entry. When the user triggers face recognition, the video frame caching module receives the face recognition command, obtains the video frames cached within 1 second before the current timestamp of the target video, and the video frames played and cached within 1 second after the current timestamp of the target video, to obtain the video frame to be processed. Figure 4 Extracting video frames. For example, if 1 second corresponds to 10 frames, then 20 video frames (i.e., the video frames to be processed) within 1 second before and 1 second after the current timestamp are obtained. The video frames to be processed and their timestamps are stored. Figure 4 (Stores video frames and timestamps).

[0152] The face tracking and detection module performs face detection on each video frame in the video to be processed. Figure 4 (Face detection in the middle) is performed to obtain the face detection results for each video frame. Each video frame's face detection result can be represented as a list of [frames with faces, timestamps]. Then, following the playback order of the video frames in the target video, face tracking is performed on each face detection result in each video frame of the video to be processed, yielding the face tracking result for each face detection result. Figure 4 (Based on face tracking between consecutive frames), for each face tracking result, the video frame corresponding to the frontal face is retained to obtain candidate video frames.

[0153] The face classification and merging module classifies faces belonging to different objects in each candidate video frame, obtaining face classification results for each object. This filters out similar frontal face frames, assuming they belong to the same face entity. Further, based on the face classification results, the client device determines the target video frame from the video frames to be processed and sends it to the cloud device, enabling the cloud device to recognize faces in the target video frame.

[0154] like Figure 5 As shown, another face recognition method provided in this embodiment of the invention, applied to cloud devices, can be implemented through the following steps:

[0155] S501 receives the target video frame sent by the client device.

[0156] The target video frame is defined as follows: the client device acquires video frames within a preset time range of the target timestamp in the target playback video to obtain the video frame to be processed. Target detection is then performed on each video frame in the video frame to be processed. Based on the target detection results, faces belonging to different objects in each video frame of the video frame to be processed are classified, and the target video frame is determined from the video frame to be processed based on the face classification results. The specific implementation process of the client device determining and sending the target video frame is described above, and will not be repeated here in this embodiment of the invention.

[0157] S502, extract facial features from the target video frame to obtain the target facial features.

[0158] In the process of face recognition, facial features are compared. Since discrete features are more suitable for vector operations, they can achieve fast feature comparison. In this embodiment of the invention, the facial features are discrete facial features, which are stored through discrete sequences, thus saving storage space.

[0159] In one example, the object detection result includes the location of the object detection box corresponding to the detected object in the video frame. When the client device sends the target video frame to the cloud device, it also sends the location of the object detection box corresponding to the detected object in the target video frame. The cloud device receives the target video frame and the location of the object detection box corresponding to the detected object in the target video frame from the client device. Based on the location of the object detection box, it can accurately extract the facial features contained within the object detection box of the target video frame to obtain the target facial features.

[0160] S503, match the target facial features with the facial features contained in the preset facial feature database.

[0161] The cloud device stores a preset facial feature database, which contains a large number of facial features. For example, the facial features included in the preset facial feature database can be the facial features of the target object in the role played in the video, or facial features that can represent the real identity of the target object, such as the facial features of the target object in different age groups, etc. The target object can be, for example, a person.

[0162] In one example, the similarity between the target facial feature and each facial feature in a pre-defined facial feature database is calculated. The facial feature with the highest similarity to the target facial feature, or the facial feature with a similarity greater than a target threshold, is identified as the facial feature that matches the target facial feature. The target threshold could be, for example, 0.8, 0.9, or 0.95, etc.

[0163] S504: If a face feature matching the target face feature exists in the preset face feature database, the person information corresponding to the face feature matching the target face feature is used as the face recognition result of the target video frame, and the face recognition result is sent to the client device.

[0164] The preset facial feature database includes facial features and the corresponding personal information for each facial feature. For example, the personal information corresponding to a facial feature may include, but is not limited to: personal name, photo, age, place of birth, zodiac sign, stage name, resume, and acting experience.

[0165] In one example, if no matching facial feature exists in the preset facial feature database, the system sends an unrecognized result to the client device.

[0166] This invention provides a face recognition method in which a client device acquires multiple video frames within a preset time range of a target timestamp from a target video being played. Then, using target detection and face classification technologies, it identifies the target video frame from the multiple video frames to be sent to a cloud device for face feature comparison. Compared to existing methods that send a captured video frame from a movie or TV show to the cloud device for face recognition, this method improves the image quality and robustness of the video frame sent to the cloud device for face recognition, and fully utilizes the computing resources of the client device. Furthermore, the cloud device can compare the target face features in the target video frame with those in a preset face feature library based on the high-quality image of the target video frame, thereby more accurately identifying faces in the target video frame and feeding back the face recognition result to the client device. This improves the accuracy and recall rate of face recognition. Moreover, since face detection and classification are performed on the client device, the computational load on the cloud device is reduced, thus lowering the cost of face detection and classification on the cloud device.

[0167] like Figure 6 As shown, another face recognition method provided in this embodiment of the invention, applied to cloud devices, can be implemented through the following steps:

[0168] S601 receives the target video frame, the position of the target detection box corresponding to the detected target in the target video frame, the timestamp corresponding to the target video frame, and the identification information of the target video being played, sent by the client device.

[0169] S602, determine whether the identification information of the target playback video corresponds to a data relationship table.

[0170] Historical facial recognition results are stored in the cloud device, and these results are stored in the form of a data relationship table. In one example, a film or television program (i.e., a target video) corresponds to a data relationship table.

[0171] When the cloud device receives the target video frame, the position of the target detection box corresponding to the detected target in the target video frame, the timestamp of the target video frame, and the identification information of the target video being played, sent by the client device, it queries the historical face recognition results to see if there is a data relationship table corresponding to the identification information of the target video being played, based on the correspondence between the video identification information and the data relationship table. If it exists, it means that the cloud device has performed face recognition on the detected target contained in the target video being played; if it does not exist, it means that the cloud device has not performed face recognition on the detected target contained in the target video being played.

[0172] S603, if the identification information of the target video corresponds to a data relationship table, query whether the data relationship table contains: a video frame timestamp that is the same as the timestamp corresponding to the target video frame, and a detection box position that matches the position of the target detection box corresponding to the target in the target video frame.

[0173] The data relationship table includes: the correspondence between video frame timestamps and detection box positions and the person recognition results; the matching of the detection box position with the position of the target detection box corresponding to the detected target in the target video frame indicates that: the detection box position is the same as the position of the target detection box corresponding to the detected target in the target video frame, or the distance between the detection box position and the position of the target detection box corresponding to the detected target in the target video frame is less than a preset distance threshold. The preset distance threshold can be set according to needs or experience, such as 0.4, 0.5, or 0.6 milliseconds or centimeters, etc.

[0174] S604, if the data relationship table contains: a video frame timestamp that is the same as the timestamp corresponding to the target video frame, and a detection box position that matches the position of the target detection box corresponding to the detected target in the target video frame, then the person recognition result corresponding to the timestamp corresponding to the target video frame and the position of the target detection box corresponding to the detected target in the target video frame in the data relationship table is determined as the face recognition result of the target video frame and sent to the client device.

[0175] The data relationship table contains timestamps that are the same as the timestamps corresponding to the target video frame, as well as detection box positions that match the positions of the detection boxes corresponding to the detected targets in the target video frame. This indicates that the person in the target video frame to be identified belongs to the same person as the person already identified in the data relationship table. That is, the cloud device has previously performed face recognition on the same person corresponding to the timestamps of video frames with the same timestamps as the target video frame in the target video playback video. At this time, the corresponding person recognition result is determined as the face recognition result of the target video frame.

[0176] S605, if the identification information of the target video does not have a corresponding data relationship table, or if the data relationship table does not contain: a video frame timestamp with the same timestamp as the target video frame, and a detection box position that matches the position of the target detection box in the target video frame, then extract the face features in the target video frame to obtain the target face features.

[0177] The process of extracting facial features from the target video frame to obtain the target facial features can be referred to the implementation process of step S502 above, and will not be repeated here in this embodiment of the invention.

[0178] S606, Match the target facial features with the facial features contained in the preset facial feature database.

[0179] S607: If a face feature matching the target face feature exists in the preset face feature library, the person information corresponding to the face feature matching the target face feature is used as the face recognition result of the target video frame, and the face recognition result is sent to the client device.

[0180] The preset facial feature database includes facial features and the corresponding person information. The implementation process of steps S606-S607 can be referred to the implementation process of steps S503-S504 above, and will not be repeated here.

[0181] This invention provides a face recognition method in which a client device acquires multiple video frames within a preset time range of a target timestamp from a target video playback. Then, using target detection and face classification technologies, it identifies the target video frame from the multiple video frames to be sent to a cloud device for face feature comparison. Compared to existing methods that send a captured video frame from a movie or TV show playback to the cloud device for face recognition, this method improves the image quality and robustness of the video frame sent to the cloud device for face recognition and fully utilizes the computing resources of the client device. Furthermore, the cloud device first checks whether face recognition has already been performed on the same person at the same time point in the same target video playback. If so, it directly returns the stored face recognition result; otherwise, based on the high-quality target video frame, it compares the face features contained in a preset face feature library with the target face features in the target video frame to more accurately identify the face in the target video frame and feeds back the corresponding face recognition result to the client device, thus improving the accuracy and recall rate of face recognition.

[0182] In one possible implementation, the aforementioned preset facial feature database includes: a first facial feature database and a second facial feature database. The facial features in the first facial feature database are used to characterize the facial features of the target object in the role it plays in the video, and the facial features in the second facial feature database are used to characterize the facial features representing the true identity of the target object. The first facial feature database includes first facial features and the corresponding person information, and the second facial feature database also includes second facial features and the corresponding person information.

[0183] In one example, the first facial feature library contains facial features of the target object in different videos under different attire (such as hair accessories, headwear / hats, sunglasses, makeup, etc.). These facial features under different attire can be obtained from images of the target object in the roles they play in the videos. The second facial feature library contains facial features of the target object at different ages (i.e., facial features throughout the entire life cycle). These facial features at different ages can be obtained from images used to represent the target object's true identity, such as casual photos or ID photos, taken at different ages. The second facial feature library is a global feature library and is not affected by different videos, while the first facial feature library is different for different videos (i.e., one first facial feature library corresponds to one video).

[0184] Accordingly, see Figure 7 The above involves matching the target facial features with facial features contained in a preset facial feature database; if a facial feature matching the target facial feature exists in the preset facial feature database, the person information corresponding to the facial feature matching the target facial feature is used as the facial recognition result of the target video frame, and the facial recognition result is sent to the client device, including:

[0185] S701, based on the identification information of the target video and the correspondence table between the video identification information and the first face feature database, determine whether there is a first face feature database corresponding to the target video.

[0186] The cloud device stores a table showing the correspondence between video identification information and the first facial feature database. By querying the table showing the correspondence between video identification information and the first facial feature database, it can be determined whether there is a first facial feature database corresponding to the target video being played.

[0187] S702, if a first face feature library corresponding to the target video exists, match the target face feature with the face features contained in the first face feature library.

[0188] In one example, given a first facial feature library corresponding to the target video, the similarity between the target facial feature and each facial feature in the first facial feature library is calculated. The facial feature with the highest similarity to the target facial feature, or the facial feature with a similarity greater than a target threshold, is identified as the facial feature that matches the target facial feature. The target threshold could be, for example, 0.8, 0.9, or 0.95, etc.

[0189] S703, if the target face feature matches the face features contained in the first face feature database, the person information corresponding to the face feature that matches the target face feature in the first face feature database is used as the face recognition result of the target video frame and sent to the client device.

[0190] S704, if the target face feature does not match the face features contained in the first face feature database, or if there is no first face feature database corresponding to the target video, the target face feature is matched with the face features contained in the second face feature database.

[0191] S705, if the target face feature matches the face features contained in the second face feature library, the person information corresponding to the face feature that matches the target face feature in the second face feature library is used as the face recognition result of the target video frame and sent to the client device.

[0192] In this embodiment of the invention, the cloud device uses facial features from a first facial feature library containing facial features representing the role played by the target object in the video, and a second facial feature library containing facial features representing the true identity of the target object, to compare with the target facial features in the target video frame to be recognized. Compared with using only facial features representing the true identity of the target object for facial recognition, this improves the accuracy and recall rate of facial recognition. Furthermore, the first facial feature library is different for different videos, so there is no information interference between different first facial feature libraries.

[0193] In one possible implementation, the aforementioned preset facial feature database includes: a first facial feature database and a second facial feature database. The facial features contained in the first facial feature database are used to characterize the facial features of the target object in the role it plays in the video, and the facial features contained in the second facial feature database are used to characterize the facial features of the target object's true identity. The first facial feature database includes a first facial feature and the person information corresponding to the first facial feature, and the second facial feature database also includes a second facial feature and the person information corresponding to the second facial feature.

[0194] Accordingly, the target facial features are matched with facial features contained in a preset facial feature database; if a facial feature matching the target facial features exists in the preset facial feature database, the person information corresponding to the facial feature matching the target facial features is used as the facial recognition result of the target video frame, and the facial recognition result is sent to the client device, including:

[0195] Based on the identification information of the target video and the correspondence table between the video identification information and the first face feature database, it is determined whether there is a first face feature database corresponding to the target video.

[0196] If a first face feature library exists that corresponds to the target video, the target face feature is matched with the face features contained in the first face feature library to obtain the first matching result.

[0197] The target facial features are matched with the facial features contained in the second facial feature database to obtain the second matching result;

[0198] Based on the first and second matching results, the face recognition result of the target video frame is obtained and sent to the client device.

[0199] Based on the first matching result and the second matching result, the face recognition result of the target video frame is obtained and sent to the client device.

[0200] In one example, given the existence of a first facial feature library corresponding to the target video, the system identifies the first facial feature in the first facial feature library that has the highest similarity to the target facial feature. The person information corresponding to this first facial feature is used as the first matching result, and the similarity between this first facial feature and the target facial feature is determined as the first matching similarity. Similarly, the system identifies the second facial feature in the second facial feature library that has the highest similarity to the target facial feature. The person information corresponding to this second facial feature is used as the second matching result, and the similarity between this second facial feature and the target facial feature is determined as the second matching similarity. Further, the first matching similarity and the second matching similarity are compared, and the matching result with the higher similarity is determined as the facial recognition result of the target video frame. Alternatively, both the first matching result and the second matching result can be determined as the facial recognition result of the target video frame.

[0201] In this embodiment of the invention, the cloud device uses facial features from a first facial feature library containing facial features representing the role a target object plays in a video, and a second facial feature library containing facial features representing the true identity of the target object, to compare with the target facial features in the target video frame to be recognized. Compared to using only facial features representing the true identity of the target object for facial recognition, this improves the accuracy and recall rate of facial recognition. Furthermore, the first facial feature library is different for different videos, so there is no information interference between different first facial feature libraries. Combining the first matching result and the second matching result to obtain the facial recognition result of the target video frame further improves the recall rate of the recognition result.

[0202] For example, such as Figure 8 As shown, the cloud device includes a timestamp comparison module, a face feature extraction module, and a face feature comparison module. The cloud device receives the target video frame sent by the client device, the position of the target detection box corresponding to the detected target in the target video frame, the timestamp corresponding to the target video frame, and the identification information of the target video being played.

[0203] The timestamp comparison module determines whether the identifier information of the target video corresponds to a data relationship table. If the identifier information of the target video corresponds to a data relationship table, it checks whether the data relationship table contains: a video frame timestamp with the same timestamp as the target video frame, and a detection box position that matches the position of the detection box corresponding to the detected target in the target video frame. If these exist, the person recognition result corresponding to the timestamp and the position of the detection box corresponding to the detected target in the target video frame is determined as the face recognition result of the target video frame. The cloud device then sends the face recognition result to the client device.

[0204] If the identification information of the target video does not have a corresponding data relationship table, or if the data relationship table does not contain: a video frame timestamp with the same timestamp as the target video frame, and a detection box position that matches the position of the target detection box in the target video frame, the face feature extraction module extracts the face features in the target video frame to obtain the target face features.

[0205] The facial feature comparison module, based on the identifier information of the target video and the correspondence table between the video identifier information and the first facial feature database, determines whether a first facial feature database corresponding to the target video exists. If it does, it matches the target facial features with the facial features contained in the first facial feature database to obtain a first matching result; and it matches the target facial features with the facial features contained in a second facial feature database to obtain a second matching result. Based on the first and second matching results, the facial recognition result of the target video frame is obtained. The cloud device then sends the facial recognition result to the client device.

[0206] In one possible implementation, when the face recognition result of the target video frame is obtained based on the first face feature library and / or the second face feature library, the position of the target detection box corresponding to the target in the target video frame, the timestamp corresponding to the target video frame, the identification information of the target video being played, and the face recognition result of the target video frame can be further updated to the data relationship table corresponding to the identification information of the target video being played.

[0207] In one possible implementation, the above method may further include:

[0208] If the target facial features do not match the facial features contained in the second facial feature database, the target video frame and the position of the target detection box corresponding to the detected target in the target video frame are sent to the operation terminal so that the operation terminal can update the facial features of the real identity of the detected target in the target video frame and the person information corresponding to the facial features to the second facial feature database.

[0209] If the target facial features do not match the facial features contained in the first facial feature library corresponding to the target video, or if there is no first facial feature library corresponding to the target video, the target video frame, the position of the target detection box corresponding to the detected target in the target video frame, and the identification information of the target video are sent to the operation terminal so that the operation terminal updates the facial features of the detected target in the target video frame under the role played in the target video, and the person information corresponding to the facial features to the first facial feature library.

[0210] If the target facial features do not match the facial features contained in the second facial feature database, the cloud device cannot identify the person information in the target video frame. In this case, the target video frame and the position of the target detection box corresponding to the detected target in the target video frame are sent to the operation terminal so that the operation terminal can update the facial features of the real identity of the detected target in the target video frame and the person information corresponding to the facial features to the second facial feature database, thereby realizing the real-time update of the second facial feature database.

[0211] If the target facial features do not match the facial features contained in the first facial feature library corresponding to the target video, or if there is no first facial feature library corresponding to the target video, the cloud device cannot accurately identify the person in the video. In this case, the target video frame, the position of the target detection box corresponding to the detected target in the target video frame, and the identification information of the target video are sent to the operation terminal. This allows the operation terminal to update the facial features of the detected target in the target video frame under the role played in the target video and the person information corresponding to the facial features to the first facial feature library, thereby realizing the real-time update and creation of the first facial feature library.

[0212] For example, such as Figure 9As shown, the operator can pre-extract facial features of characters in film and television images using a facial feature extraction module, and then create a first facial feature library based on these facial features and the corresponding character information. Similarly, based on everyday photos of the target object, the operator can extract facial features from these photos and create a second facial feature library based on these features and the corresponding character information. Furthermore, the first and second facial feature libraries are stored on a cloud device. The operator maintains both libraries and updates the second facial feature library upon receiving a target video frame from the cloud device, along with the location of the target detection box corresponding to the detected target within the video frame. The operator also updates the first facial feature library upon receiving a target video frame from the cloud device, the location of the target detection box corresponding to the detected target within the video frame, and the identifier information of the target video being played.

[0213] In this embodiment of the invention, the operator maintains a first facial feature database and a second facial feature database. When the cloud device cannot identify the person information in the target video frame, the relevant information of the target video frame is sent to the operator. This allows the operator to update the second facial feature database with the facial features of the detected target's real identity and the corresponding person information, thus achieving real-time updates to the second facial feature database. Similarly, when the cloud device cannot accurately identify the person in the video, the relevant information of the target video frame is sent to the operator. This allows the operator to update the first facial feature database with the facial features of the detected target's role in the target video and the corresponding person information, thus achieving real-time updates and creation of the first facial feature database.

[0214] For example, such as Figure 10 As shown in the figure, a face recognition method provided by an embodiment of the present invention may include:

[0215] The client device caches the video frames played by the target video. Upon receiving a face recognition command, it acquires the cached video frames within a first preset time range before the current timestamp of the target video, and the video frames played and cached within a second preset time range after the current timestamp of the target video, to obtain the video frames to be processed. It performs target detection on each video frame to obtain the target detection results for each video frame. Based on the target detection results, it classifies the faces belonging to different objects in each video frame to obtain the face classification results for different objects. Based on the face classification results, it determines the target video frame from the video frames to be processed and sends the target video frame, the position of the target detection box corresponding to the detected target in the target video frame, the timestamp corresponding to the target video frame, and the identification information of the target video to the cloud device.

[0216] The cloud device receives the target video frame, the position of the target detection box corresponding to the detected target in the target video frame, the timestamp of the target video frame, and the identification information of the target video being played, sent by the client device. It determines whether the identification information of the target video corresponds to a data relationship table. If so, it checks whether the data relationship table contains: a video frame timestamp with the same timestamp as the target video frame, and a detection box position that matches the position of the target detection box corresponding to the detected target in the target video frame. If these exist, the person recognition result corresponding to the timestamp and the position of the target detection box in the target video frame is identified as the face recognition result of the target video frame and sent to the client device. Otherwise, it extracts the face features from the target video frame to obtain the target face features. It matches the target face features with face features contained in a preset face feature library. If a face feature matching the target face features exists in the preset face feature library, the person information corresponding to the face feature matching the target face features is taken as the face recognition result of the target video frame, and the face recognition result is sent to the client device.

[0217] The client device receives the facial recognition results sent by the cloud device.

[0218] This invention provides a face recognition method in which a client device acquires multiple video frames within a preset time range of the target timestamp in a target video, and then uses target detection and face classification technologies to determine the target video frame to be sent to a cloud device for face feature comparison. Compared with existing methods that send a captured video frame from the playback of a movie or TV series to the cloud device for face recognition, this method improves the image quality and robustness of the video frame sent to the cloud device for face recognition, and fully utilizes the caching capabilities and computing resources of the client device. Furthermore, the cloud device prioritizes checking whether face recognition has already been performed on the same person at the same time in the same target video. If so, it directly returns the stored face recognition result. Otherwise, based on the high-quality target video frame, it compares the face features contained in the preset face feature library with the target face features in the target video frame to more accurately identify the face in the target video frame and feed back the face recognition result corresponding to the target video frame to the client device. This improves the accuracy and recall rate of face recognition. Moreover, face detection and classification are completed on the client device, reducing the computational load on the cloud device and thus lowering the cost of face detection and classification on the cloud device.

[0219] For example, such as Figure 11 As shown in the figure, a face recognition method provided by an embodiment of the present invention may include:

[0220] The client's video frame caching module caches the video frames played by the target video. Upon receiving a face recognition command, it retrieves the cached video frames within a first preset time range before the current timestamp of the target video, and the cached video frames within a second preset time range after the current timestamp of the target video, thus obtaining the video frame to be processed. The client's face detection and tracking module performs target detection on each video frame in the video frame to be processed, obtaining the target detection results for each video frame. Following the playback order of the video frames in the target video, it performs face tracking on each target detection result in each video frame of the video frame to be processed, obtaining the face tracking results for each target detection result. For each face tracking result, it retains the video frame corresponding to a frontal face, obtaining candidate video frames. The client's face classification module classifies faces belonging to different objects in each candidate video frame, obtaining face classification results for different objects. Based on the face classification results, it determines the target video frame from the video frames to be processed. The client sends the target video frame, the position of the target detection box corresponding to the detected target in the target video frame, the timestamp of the target video frame, and the identification information of the target video to the cloud. When there are multiple target video frames, the list of video frames corresponding to the target video frames and the list of timestamps consisting of the timestamps corresponding to each target video frame in the list of video frames are sent to the cloud.

[0221] The cloud-based system receives the target video frame, the location of the target detection box corresponding to the detected target within the target video frame, the timestamp of the target video frame, and the identification information of the target video being played, all sent by the client. The cloud-based timestamp comparison module determines whether the identification information of the target video corresponds to a data relationship table. If so, it checks if the data relationship table contains a video frame timestamp identical to the timestamp corresponding to the target video frame, and a detection box location matching the location of the target detection box in the target video frame. If these exist, the system identifies the person recognition result corresponding to the timestamp and the location of the target detection box in the target video frame as the face recognition result of the target video frame and sends it to the client device. The cloud-based face feature extraction module extracts the face features from the target video frame if the identification information of the target video does not correspond to a data relationship table, or if the data relationship table does not contain a video frame timestamp identical to the timestamp corresponding to the target video frame, or a detection box location matching the location of the target detection box in the target video frame.

[0222] The cloud-based facial feature comparison module determines whether a first facial feature database exists that corresponds to the target video, based on the identification information of the target video and the correspondence table between the video's identification information and the first facial feature database. If it exists, the module matches the target facial features with the facial features contained in the first facial feature database. If the target facial features match the facial features contained in the first facial feature database, the person information corresponding to the facial features in the first facial feature database that match the target facial features is used as the facial recognition result of the target video frame and sent to the client. If the target facial features do not match the facial features contained in the first facial feature database, or if a first facial feature database that matches the target video does not exist, the module matches the target facial features with the facial features contained in a second facial feature database. If the target facial features match the facial features contained in the second facial feature database, the person information corresponding to the facial features in the second facial feature database that match the target facial features is used as the facial recognition result of the target video frame and sent to the client.

[0223] If the target facial features do not match those in the second facial feature database, the cloud sends the target video frame and the location of the target detection box within the target video frame to the operations team. The operations team obtains a snapshot image of the face corresponding to the target detection box in the target video frame, extracts the facial features from the snapshot image using a facial feature extraction module, and updates the second facial feature database with the facial features and the corresponding person information. If the target facial features do not match those in the first facial feature database corresponding to the target video, or if a first facial feature database does not exist, the cloud sends the target video frame, the location of the target detection box within the target video frame, and the identification information of the target video to the operations team. The operations team obtains a film / TV image of the character played by the target in the target video frame, extracts the facial features from the film / TV image using a facial feature extraction module, and updates the first facial feature database with the facial features and the corresponding person information.

[0224] This invention provides a face recognition method in which a client device acquires multiple video frames within a preset time range of the target timestamp in a target video, and then uses target detection and face classification technologies to determine the target video frame to be sent to a cloud device for face feature comparison. Compared with existing methods that send a captured video frame from the playback of a movie or TV series to the cloud device for face recognition, this method improves the image quality and robustness of the video frame sent to the cloud device for face recognition, and fully utilizes the caching capabilities and computing resources of the client device. Furthermore, the cloud device first checks whether facial recognition has already been performed on the same person at the same time in the same video. If so, it directly returns the stored facial recognition result. Otherwise, it uses facial features from a first facial feature library containing facial features representing the role the target person plays in the video, and a second facial feature library containing facial features representing the target person's true identity, to compare with the target facial features in the target video frame to be facially recognized. Compared to using only the facial features representing the target person's true identity for facial recognition, this improves the accuracy and recall rate of facial recognition. Moreover, the first facial feature library is different for different videos, so there is no information interference between different first facial feature libraries.

[0225] Corresponding to the above method embodiments, the present invention also provides corresponding device embodiments.

[0226] like Figure 12 As shown, this embodiment of the invention provides a face recognition device applied to a client device, the device comprising:

[0227] The video frame acquisition module 1201 is used to acquire video frames within a preset time range of the target timestamp in the target playback video to obtain the video frames to be processed.

[0228] The target detection module 1202 is used to perform target detection on each video frame in the video frame to be processed, and obtain the target detection results of each video frame.

[0229] The face classification module 1203 is used to classify faces belonging to different objects in each video frame of the video frame to be processed based on the detection results of each object, and obtain the face classification results belonging to different objects.

[0230] The video frame determination module 1204 is used to determine the target video frame from the video frames to be processed based on the face classification results.

[0231] The video frame sending module 1205 is used to send the target video frame to the cloud device so that the cloud device can extract the facial features in the target video frame, obtain the target facial features, match the target facial features with the facial features contained in the preset facial feature library, and send the person information corresponding to the facial features that match the target facial features in the preset facial feature library as the facial recognition result of the target video frame to the client device. The preset facial feature library includes facial features and the person information corresponding to the facial features.

[0232] The recognition result receiving module 1206 is used to receive the face recognition results sent by the cloud device.

[0233] In one possible implementation, the face classification module 1203 is specifically used for:

[0234] According to the playback order of each video frame in the target playback video, face tracking is performed on each target detection result in each video frame of the video frame to be processed to obtain the face tracking result of each target detection result; among which, the face tracking result includes: front face, side face, back view, global or local;

[0235] For each face tracking result, the video frames corresponding to the frontal faces are retained to obtain candidate video frames;

[0236] The faces belonging to different objects in each candidate video frame are classified to obtain the face classification results belonging to different objects.

[0237] In one possible implementation, the above-described apparatus further includes:

[0238] The video frame caching module is used to cache the video frames played by the target video.

[0239] The aforementioned video frame acquisition module 1201 is specifically used to acquire, upon receiving a face recognition instruction, video frames that have been cached within a first preset time range before the current timestamp of the target playback video, and video frames that have been played and cached within a second preset time range after the current timestamp of the target playback video, to obtain the video frame to be processed.

[0240] In one possible implementation, the video frame determination module 1204 is specifically used to select a frame image that meets the preset quality requirements from the video frames corresponding to each face classification result, and use it as the target video frame corresponding to the face classification result.

[0241] In one possible implementation, the target detection result includes the position of the target detection box corresponding to the detected target in the video frame, and the apparatus further includes:

[0242] The video information sending module is used to send the position of the target detection box corresponding to the target in the target video frame, the timestamp of the target video frame, and the identification information of the target playing video to the cloud device.

[0243] like Figure 13 As shown, this embodiment of the invention provides a face recognition device applied to a cloud device, the device comprising:

[0244] The video frame receiving module 1301 is used to receive target video frames sent by the client device. The target video frame is: the video frame obtained by the client device within a preset time range of the target timestamp in the target playback video, to obtain the video frame to be processed, and to perform target detection on each video frame in the video frame to be processed, and to classify the faces of different objects in each video frame of the video frame to be processed based on the target detection results, and to determine the face from the video frame to be processed based on the face classification results.

[0245] The face feature extraction module 1302 is used to extract face features from the target video frame to obtain the target face features;

[0246] The face feature matching module 1303 is used to match the target face features with the face features contained in the preset face feature library;

[0247] The face recognition module 1304 is used to take the person information corresponding to the face feature that matches the target face feature as the face recognition result of the target video frame when there is a face feature in the preset face feature library that matches the target face feature, and send the face recognition result to the client device. The preset face feature library includes face features and the person information corresponding to the face feature.

[0248] In one possible implementation, the above-described apparatus further includes:

[0249] The video information receiving module is used to receive the position of the target detection box corresponding to the target in the target video frame sent by the client device, the timestamp corresponding to the target video frame, and the identification information of the target playing video;

[0250] The first determining module is used to determine whether the identification information of the target playback video corresponds to a data relationship table;

[0251] The query module is used to query the data relationship table to determine whether, when the first determining module determines that the identification information of the target video corresponds to a data relationship table, the following exist in the data relationship table: a video frame timestamp that is the same as the timestamp corresponding to the target video frame, and a detection box position that matches the position of the target detection box corresponding to the detected target in the target video frame; wherein, the data relationship table includes: the correspondence between the video frame timestamp and the detection box position and the person recognition result; the matching of the detection box position with the position of the target detection box corresponding to the detected target in the target video frame means that: the detection box position is the same as the position of the target detection box corresponding to the detected target in the target video frame, or the distance between the detection box position and the position of the target detection box corresponding to the detected target in the target video frame is less than a preset distance threshold;

[0252] The second determining module is used to determine the face recognition result of the target video frame as the face recognition result of the target video frame when the data relationship table contains a video frame timestamp that is the same as the timestamp corresponding to the target video frame and a detection box position that matches the position of the detection box corresponding to the target in the target video frame, and send it to the client device; otherwise, it triggers the face feature extraction module 1302 to extract the face features in the target video frame to obtain the target face features.

[0253] In one possible implementation, the aforementioned preset facial feature database includes: a first facial feature database and a second facial feature database. The facial features contained in the first facial feature database are used to characterize the facial features of the target object in the role it plays in the video, and the facial features contained in the second facial feature database are used to characterize the facial features of the target object's real identity.

[0254] The aforementioned face feature matching module 1303 and face recognition module 1304 are specifically used for:

[0255] Based on the identification information of the target video and the correspondence table between the video identification information and the first face feature database, it is determined whether there is a first face feature database corresponding to the target video.

[0256] If a first facial feature library exists that corresponds to the target video, the target facial features are matched with the facial features contained in the first facial feature library.

[0257] If the target facial features match the facial features contained in the first facial feature database, the person information corresponding to the facial features that match the target facial features in the first facial feature database is used as the facial recognition result of the target video frame and sent to the client device.

[0258] If the target facial features do not match the facial features contained in the first facial feature database, or if there is no first facial feature database corresponding to the target video, the target facial features will be matched with the facial features contained in the second facial feature database.

[0259] If the target facial features match the facial features contained in the second facial feature database, the person information corresponding to the facial features that match the target facial features in the second facial feature database is used as the facial recognition result of the target video frame and sent to the client device.

[0260] In one possible implementation, the aforementioned preset facial feature database includes: a first facial feature database and a second facial feature database. The facial features contained in the first facial feature database are used to characterize the facial features of the target object in the role it plays in the video, and the facial features contained in the second facial feature database are used to characterize the facial features of the target object's real identity.

[0261] The aforementioned face feature matching module 1303 and face recognition module 1304 are specifically used for:

[0262] Based on the identification information of the target video and the correspondence table between the video identification information and the first face feature database, it is determined whether there is a first face feature database corresponding to the target video.

[0263] If a first face feature library exists that corresponds to the target video, the target face feature is matched with the face features contained in the first face feature library to obtain the first matching result.

[0264] The target facial features are matched with the facial features contained in the second facial feature database to obtain the second matching result;

[0265] Based on the first and second matching results, the face recognition result of the target video frame is obtained and sent to the client device.

[0266] In one possible implementation, the above-described apparatus further includes:

[0267] The first update module is used to send the target video frame and the position of the target detection box corresponding to the detected target in the target video frame to the operation terminal when the target facial features do not match the facial features contained in the second facial feature library, so that the operation terminal updates the facial features of the real identity of the detected target in the target video frame and the person information corresponding to the facial features to the second facial feature library.

[0268] The second update module is used to send the target video frame, the position of the target detection box corresponding to the detected target in the target video frame, and the identification information of the target video to the operation terminal when the target facial features do not match the facial features contained in the first facial feature library corresponding to the target video, or when there is no first facial feature library corresponding to the target video. This allows the operation terminal to update the facial features of the detected target in the target video frame under the role played in the target video, as well as the person information corresponding to the facial features, to the first facial feature library.

[0269] This invention also provides an electronic device, such as... Figure 14 As shown, it includes a processor 1401, a communication interface 1402, a memory 1403, and a communication bus 1404, wherein the processor 1401, the communication interface 1402, and the memory 1403 communicate with each other through the communication bus 1404.

[0270] Memory 1403 is used to store computer programs;

[0271] When the processor 1401 executes the program stored in the memory 1403, it implements the steps of any of the above method embodiments to achieve the same technical effect.

[0272] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0273] The communication interface is used for communication between the aforementioned terminal and other devices.

[0274] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0275] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0276] In another embodiment of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements the steps of any of the methods described in the above embodiments to achieve the same technical effect.

[0277] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the steps of any of the methods described in the above embodiments to achieve the same technical effect.

[0278] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0279] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0280] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device / electronic device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0281] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A face recognition method, characterized in that, Applied to a client device, the method includes: Upon receiving a face recognition instruction, video frames within a preset time range of the target timestamp in the target playback video are obtained to obtain the video frames to be processed; the target timestamp is the timestamp at the time the face recognition instruction is received. Target detection is performed on each video frame in the video frame to be processed to obtain the target detection results of each video frame. Based on the target detection results, the faces belonging to different objects in each video frame of the video frame to be processed are classified to obtain the face classification results belonging to different objects; Based on the face classification results, the target video frame is determined from the video frames to be processed; The target video frame is sent to a cloud device, which extracts facial features from the target video frame to obtain target facial features. These target facial features are then matched with facial features in a preset facial feature library. The person information corresponding to the facial features matching the target facial features in the preset facial feature library is sent to the client device as the facial recognition result of the target video frame. The preset facial feature library includes facial features and the person information corresponding to those facial features. The preset facial feature library includes a first facial feature library and a second facial feature library. The facial features in the first facial feature library represent the facial features of the target object in the role it plays in the video, while the facial features in the second facial feature library represent the facial features of the target object's true identity. Receive the face recognition results sent by the cloud device.

2. The method according to claim 1, characterized in that, Based on the target detection results, the faces belonging to different objects in each video frame of the video frame to be processed are classified to obtain face classification results belonging to different objects, including: According to the playback order of each video frame in the target playback video, face tracking is performed on each of the target detection results in each video frame of the video frame to be processed to obtain the face tracking results of each target detection result; wherein, the face tracking results include: front face, side face, back view, global or local; For each of the aforementioned face tracking results, the video frames corresponding to frontal faces are retained to obtain candidate video frames; The faces belonging to different objects in each of the candidate video frames are classified to obtain the face classification results belonging to different objects.

3. The method according to claim 1, characterized in that, The method further includes: Cache the video frames played by the target video; The step of obtaining video frames within a preset time range of the target timestamp in the target playback video to obtain the video frames to be processed includes: Upon receiving a face recognition instruction, the system acquires cached video frames within a first preset time range before the current timestamp of the target video, and video frames played and cached within a second preset time range after the current timestamp of the target video, to obtain the video frames to be processed.

4. The method according to claim 1, characterized in that, The step of determining the target video frame from the video frames to be processed based on the face classification result includes: For each face classification result, select one frame image that meets the preset quality requirements from the video frames corresponding to that face classification result, and use it as the target video frame corresponding to that face classification result.

5. The method according to any one of claims 1-4, characterized in that, The target detection result includes the position of the target detection box corresponding to the detected target in the video frame, and the method further includes: The location of the target detection box corresponding to the target in the target video frame, the timestamp of the target video frame, and the identification information of the target video being played are sent to the cloud device.

6. A face recognition method, characterized in that, Applied to cloud devices, the method includes: The client device receives a target video frame sent by a client device. The target video frame is defined as follows: upon receiving a face recognition instruction, the client device acquires video frames within a preset time range of the target timestamp in the target playback video to obtain a video frame to be processed. The client device then performs target detection on each video frame in the video frame to be processed, and classifies faces belonging to different objects in each video frame based on the target detection results. The target timestamp is the timestamp at which the face recognition instruction is received. Extract facial features from the target video frame to obtain the target facial features; The target facial features are matched with facial features contained in a preset facial feature library; the preset facial feature library includes: a first facial feature library and a second facial feature library, the facial features contained in the first facial feature library are used to characterize the facial features of the target object in the role played in the video, and the facial features contained in the second facial feature library are used to characterize the facial features of the target object's real identity; If a face feature matching the target face feature exists in the preset face feature library, the person information corresponding to the face feature matching the target face feature is used as the face recognition result of the target video frame, and the face recognition result is sent to the client device. The preset face feature library includes face features and the person information corresponding to the face feature.

7. The method according to claim 6, characterized in that, The method further includes: The system receives the position of the target detection box corresponding to the target in the target video frame sent by the client device, the timestamp corresponding to the target video frame, and the identification information of the target playing video. Determine whether the identification information of the target video corresponds to a data relationship table; If so, query the data relationship table to see if there is a video frame timestamp that is the same as the timestamp corresponding to the target video frame, and a detection box position that matches the position of the target detection box corresponding to the detected target in the target video frame; wherein, the data relationship table includes: the correspondence between the video frame timestamp and the detection box position and the person recognition result; the detection box position matching the position of the target detection box corresponding to the detected target in the target video frame means that: the detection box position is the same as the position of the target detection box corresponding to the detected target in the target video frame, or the distance between the detection box position and the position of the target detection box corresponding to the detected target in the target video frame is less than a preset distance threshold; If it exists, the person recognition result corresponding to the timestamp of the target video frame and the position of the target detection box in the target video frame in the data relationship table is determined as the face recognition result of the target video frame and sent to the client device; otherwise, the face features in the target video frame are extracted to obtain the target face features.

8. The method according to claim 7, characterized in that, The step of matching the target facial feature with facial features contained in a preset facial feature database, and then, if a facial feature matching the target facial feature exists in the preset facial feature database, using the person information corresponding to the facial feature matching the target facial feature as the facial recognition result of the target video frame, and sending the facial recognition result to the client device, includes: Based on the identification information of the target video and the correspondence table between the video identification information and the first face feature database, it is determined whether there is a first face feature database corresponding to the target video. If a first facial feature library exists that corresponds to the target video, the target facial features are matched with the facial features contained in the first facial feature library. If the target facial feature matches the facial features contained in the first facial feature library, the person information corresponding to the facial feature that matches the target facial feature in the first facial feature library is used as the facial recognition result of the target video frame and sent to the client device. If the target facial feature does not match the facial features contained in the first facial feature library, or if there is no first facial feature library corresponding to the target video, the target facial feature will be matched with the facial features contained in the second facial feature library. If the target facial feature matches a facial feature contained in the second facial feature library, the person information corresponding to the facial feature in the second facial feature library that matches the target facial feature is used as the facial recognition result of the target video frame and sent to the client device.

9. The method according to claim 7, characterized in that, The step of matching the target facial feature with facial features contained in a preset facial feature database, and then, if a facial feature matching the target facial feature exists in the preset facial feature database, using the person information corresponding to the facial feature matching the target facial feature as the facial recognition result of the target video frame, and sending the facial recognition result to the client device, includes: Based on the identification information of the target video and the correspondence table between the video identification information and the first face feature database, it is determined whether there is a first face feature database corresponding to the target video. If a first facial feature library exists that corresponds to the target video, the target facial features are matched with the facial features contained in the first facial feature library to obtain a first matching result. The target facial features are matched with the facial features contained in the second facial feature database to obtain a second matching result; Based on the first matching result and the second matching result, the face recognition result of the target video frame is obtained and sent to the client device.

10. The method according to claim 8, characterized in that, The method further includes: If the target facial features do not match the facial features contained in the second facial feature library, the target video frame and the position of the target detection box corresponding to the detected target in the target video frame are sent to the operation terminal so that the operation terminal updates the facial features of the real identity of the detected target in the target video frame and the person information corresponding to the facial features to the second facial feature library. If the target facial features do not match the facial features contained in the first facial feature library corresponding to the target video, or if there is no first facial feature library corresponding to the target video, the target video frame, the position of the target detection box corresponding to the detected target in the target video frame, and the identification information of the target video are sent to the operator, so that the operator updates the facial features of the detected target in the target video frame under the role played in the target video, and the person information corresponding to the facial features to the first facial feature library.

11. A face recognition device, characterized in that, Applied to a client device, the device includes: The video frame acquisition module is used to acquire video frames within a preset time range of a target timestamp in the target playback video when a face recognition instruction is received, and obtain the video frames to be processed; the target timestamp is the timestamp when the face recognition instruction is received; The target detection module is used to perform target detection on each video frame in the video frame to be processed, and obtain the target detection results of each video frame. The face classification module is used to classify faces belonging to different objects in each video frame of the video frame to be processed based on the detection results of each target, and obtain face classification results belonging to different objects; The video frame determination module is used to determine the target video frame from the video frames to be processed based on the face classification result; A video frame sending module is used to send the target video frame to a cloud device, so that the cloud device can extract facial features from the target video frame, obtain target facial features, match the target facial features with facial features contained in a preset facial feature library, and send the person information corresponding to the facial features that match the target facial features in the preset facial feature library as the facial recognition result of the target video frame to a client device. The preset facial feature library includes facial features and the person information corresponding to those facial features; the preset facial feature library includes a first facial feature library and a second facial feature library, wherein the facial features contained in the first facial feature library are used to characterize the facial features of the target object in the role it plays in the video, and the facial features contained in the second facial feature library are used to characterize the facial features of the target object's true identity. The recognition result receiving module is used to receive the face recognition results sent by the cloud device.

12. A face recognition device, characterized in that, The device is applied to cloud devices and includes: The video frame receiving module is used to receive target video frames sent by the client device. The target video frame is determined by the client device, upon receiving a face recognition instruction, acquiring video frames within a preset time range of the target timestamp in the target playback video to obtain a video frame to be processed. The client device then performs target detection on each video frame in the video frame to be processed, and classifies faces belonging to different objects in each video frame based on the target detection results, and determines the target timestamp from the video frame to be processed based on the face classification results. The target timestamp is the timestamp at the time the face recognition instruction is received. The face feature extraction module is used to extract face features from the target video frame to obtain the target face features; A facial feature matching module is used to match the target facial features with facial features contained in a preset facial feature library; the preset facial feature library includes: a first facial feature library and a second facial feature library, wherein the facial features contained in the first facial feature library are used to characterize the facial features of the target object in the role it plays in the video, and the facial features contained in the second facial feature library are used to characterize the facial features of the target object's real identity. A face recognition module is used to, when a face feature matching the target face feature exists in the preset face feature library, use the person information corresponding to the face feature matching the target face feature as the face recognition result of the target video frame, and send the face recognition result to the client device. The preset face feature library includes face features and the person information corresponding to the face feature.

13. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-10.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1-10.

Citation Information

Patent Citations

  • Face recognition method, device and system based on cloud edge fusion

    CN109190532A

  • Face recognition method and device, equipment and storage medium

    CN112101216A

  • Character recognition method, electronic equipment, storage medium and device

    CN112949427A