Method, device and electronic whiteboard for face registration based on video data
Through the face registration method based on video data, the problem of poor confidentiality of electronic whiteboards during use is solved, and automated face registration is realized, simplifying the registration process and improving confidentiality.
Patent Information
- Application Number
- CN202080003669.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-25
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-12-25
AI Technical Summary
Electronic whiteboards have poor confidentiality during use, because anyone can modify their content at any distance, resulting in unsafe content.
Through the face registration method based on video data, the face features are detected using the image frame sequence in the video data, and whether the face is the same object is determined based on the relative position and feature similarity of the face detection frame, thereby realizing automated face registration.
This method does not require users to have complex interactions during the registration process, simplifies registration operations, shortens registration time, improves user experience, and enhances the confidentiality of electronic whiteboards.
Smart Images

Figure CN115053268B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of face recognition, and in particular to a method, device and electronic whiteboard for face registration based on video data. Background Art
[0002] With the gradual popularization of paperless meetings and paperless offices, the application of electronic whiteboards is becoming more and more widespread. An electronic whiteboard can receive the content written on the whiteboard and transmit the received content to a computer, so as to conveniently record and store the content on the whiteboard. When using an electronic whiteboard, in order to be able to conveniently operate the electronic whiteboard at any distance, the function of locking the electronic whiteboard is not set. Therefore, anyone can modify the content on the electronic whiteboard, which leads to the problem of poor confidentiality of the electronic whiteboard during use. Summary of the invention
[0003] The embodiments of the present disclosure provide a method and device for performing face registration based on video data, and an electronic whiteboard.
[0004] According to a first aspect of an embodiment of the present disclosure, a method for face registration based on video data is provided, comprising: receiving video data; acquiring a first image frame sequence from the video data, each image frame in the first image frame sequence each including a face detection frame including complete facial features; determining whether the image frame reaches a preset clarity based on the relative positions between the face detection frames in each image frame; when it is determined that the image frame reaches the preset clarity, extracting multiple groups of facial features based on image information of the multiple face detection frames, and determining whether the faces represent the same object based on the multiple groups of facial features; and when it is determined that the faces represent the same object, registering the object according to the first image frame sequence.
[0005] In some embodiments, obtaining a first image frame sequence from the video data includes: obtaining multiple image frames from the video data in the order in which the video is captured; determining whether the image frames contain faces based on a face detection model; and when it is determined that the image frames contain faces, determining a face detection frame containing the face in each of the multiple image frames.
[0006] In some embodiments, acquiring a first image frame sequence from the video data further includes: determining whether the acquired image frame contains complete facial features; if the image frame contains complete facial features, storing the image frame as a frame in the first image frame sequence; and ending acquisition of the image frame if the stored first image frame sequence includes a predetermined number of frames.
[0007] In some embodiments, determining whether the acquired image frame contains complete facial features includes: determining whether the face is a frontal face based on a face posture detection model; if it is determined that the face contained in the image frame is a frontal face, determining whether the face is occluded based on a face occlusion detection model; if it is determined that the face contained in the image frame is not occluded, determining that the image frame contains complete facial features; and otherwise, determining that the image frame does not contain complete facial features.
[0008] In some embodiments, determining whether the image frame has reached a preset clarity based on the relative position between the face detection frames in each image frame includes: determining a first ratio of the area of the intersection area of the face detection frames in two image frames in the first image frame sequence relative to the area of the union area of the face detection frames in the two image frames; and when the determined first ratios are all greater than a first threshold, determining that the image frame has reached the preset clarity.
[0009] In some embodiments, determining whether the image frame has reached a preset clarity based on the relative positions between the face detection frames in each image frame includes: determining a first ratio of the area of the intersection area of the face detection frames in two image frames in the first image frame sequence to the area of the union area of the face detection frames in the two image frames; determining a second ratio of the number of the first ratios greater than a first threshold to the total number of the first ratios; and determining that the image frame has reached the preset clarity when the second ratio is greater than or equal to a second threshold.
[0010] In some embodiments, determining whether the faces represent the same object based on the multiple groups of facial features includes: determining the similarity between the facial features in any two adjacent image frames in the first image frame sequence; and when the determined similarities are all greater than a third threshold, determining that the faces represent the same object.
[0011] In some embodiments, the facial features include facial feature vectors, and wherein determining the similarity between facial features in any two adjacent image frames in the first image frame sequence includes: determining the distance between the facial feature vectors in any two adjacent image frames in the first image frame sequence.
[0012] In some embodiments, registering the object according to the first image frame sequence includes: registering the object using a designated image frame in the first image frame sequence as registration data.
[0013] In some embodiments, the method further includes: storing registered data obtained by registering the object according to the first image frame sequence in a face library; and recognizing faces in the received video data based on the face library.
[0014] In some embodiments, recognizing faces in the received video data based on the face library includes: acquiring a second image frame sequence from the received video data, each image frame in the second image frame sequence each including a face detection frame including complete facial features; determining whether the image frame includes a living face based on the relative position between the face detection frames in each image frame; when it is determined that the image frame includes a living face, extracting facial features based on the face detection frame; and determining whether the facial features match the registration data in the face library to identify the face.
[0015] In some embodiments, determining whether an image frame includes a living face based on the relative positions between face detection frames in each image frame includes: determining a face detection frame that meets an overlap condition among the face detection frames in each image frame; determining a third ratio of the number of face detection frames that meet the overlap condition relative to the number of all face detection frames in the face detection frames; and determining that the face is a non-living face when the third ratio is greater than or equal to a fourth threshold; and determining that the face is a living face when the third ratio is less than the fourth threshold.
[0016] In some embodiments, determining a face detection frame that meets the overlap condition in the face detection frame in each image frame includes: determining a fourth ratio of the area of an intersection area between any two face detection frames in the face detection frame to the area of each face detection frame in the any two face detection frames; when the determined fourth ratio is greater than a fifth threshold, determining that the any two face detection frames are face detection frames that meet the overlap condition; and when the determined fourth ratio is less than the fifth threshold, determining that the any two face detection frames are face detection frames that do not meet the overlap condition.
[0017] In some embodiments, determining whether an image frame includes a living face based on the relative position between face detection frames in each image frame also includes: determining that the face is a non-living face when one of the determined fourth ratios is greater than the fifth threshold and the other fourth ratio is less than or equal to the fifth threshold.
[0018] According to a second aspect of an embodiment of the present disclosure, a device for performing face registration based on video data is provided, comprising: a memory configured to store instructions; and a processor configured to execute the instructions to execute the method provided according to the first aspect of an embodiment of the present disclosure.
[0019] According to a third aspect of the embodiments of the present disclosure, an electronic whiteboard is provided, comprising the device provided according to the second aspect of the embodiments of the present disclosure.
[0020] According to the method for face registration based on video data in the embodiment of the present disclosure, face registration can be achieved without the user having to perform complex interactions during the registration process, thereby simplifying the steps of the registration operation, shortening the registration time, and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the description of the embodiments of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0022] Figure 1 A flow chart of a method for performing face registration based on video data according to an embodiment of the present disclosure is shown;
[0023] Figure 2 The process of acquiring a first image frame sequence from video data according to an embodiment of the present disclosure is shown;
[0024] Figure 3A and Figure 3B Examples of determining whether an image frame reaches a preset definition based on the relative positions between face detection frames according to an embodiment of the present disclosure are respectively shown;
[0025] Figure 4 An example of calculating the intersection of face detection frames based on the coordinates and sizes of the face detection frames according to an embodiment of the present disclosure is shown;
[0026] Figure 5 A flowchart of a method for identifying and unlocking a face in received video data based on a face library according to another embodiment of the present disclosure is shown;
[0027] Figure 6 The process of determining a face detection frame that meets the overlap condition among multiple face detection frames according to an embodiment of the present disclosure is shown;
[0028] Figure 7 A block diagram of an apparatus for performing face registration based on video data according to another embodiment of the present disclosure is shown; and
[0029] Figure 8 A block diagram of an electronic whiteboard according to another embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of them. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present disclosure. It should be noted that throughout the drawings, the same elements are represented by the same or similar figure marks. In the following description, some specific embodiments are only used for descriptive purposes and should not be understood as any limitation to the present disclosure, but are only examples of the embodiments of the present disclosure. Conventional structures or constructions will be omitted when they may cause confusion in the understanding of the present disclosure. It should be noted that the shapes and sizes of the components in the figures do not reflect the actual size and proportion, but only illustrate the contents of the embodiments of the present disclosure.
[0031] Unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the usual meanings understood by those skilled in the art. The words "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components.
[0032] Figure 1 FIG. 1 is a flow chart of a method 100 for performing face registration based on video data according to an embodiment of the present disclosure. Figure 1 As shown, the method 100 for performing face registration based on video data may include the following steps.
[0033] In step S110 , video data is received.
[0034] In step S120, a first image frame sequence is acquired from the video data, and each image frame in the first image frame sequence includes a face detection frame including complete facial features.
[0035] In step S130, it is determined whether the image frame reaches a preset definition according to the relative positions between the face detection frames in each image frame.
[0036] In step S140, when it is determined that the image frame reaches a preset definition, multiple groups of facial features are extracted based on the image information of multiple face detection frames, and it is determined whether the faces represent the same object based on the multiple groups of facial features.
[0037] In step S150, when it is determined that the faces represent the same object, the object is registered according to the first image frame sequence.
[0038] According to an embodiment, in step S110, the video data of the object may be captured by a video acquisition device such as a camera. In other embodiments, the video data of the object may also be captured by a camera with a timed photo function. Any video acquisition device or image acquisition device capable of obtaining continuous image frames may be used. In addition, in the embodiments of the present disclosure, the format of the video data is not limited.
[0039] According to an embodiment, in step S120, after receiving the video data, a first image frame sequence is obtained from the video data. Each image frame in the first image frame sequence includes a face detection frame including complete facial features. Image frames that do not include a face detection frame including complete facial features cannot be used in the face registration process.
[0040] According to an embodiment, if a video acquisition device captures multiple objects in one image frame, the objects can be screened according to preset rules. By screening the registered objects, it is ensured that only one object is registered. According to an embodiment, multiple image frames are obtained from the video data in the order in which the video is captured, and whether the image frame contains a face is determined based on a face detection model. When it is determined that the image frame contains a face, a face detection frame is determined in each of the multiple image frames. The disclosed embodiment does not limit the face detection model used, and any face detection model can be used, or a special detection model can be established through model training. The parameters of the face detection frame can be in the form of a quaternion array, which records the coordinates of the reference point and the two side lengths of the face detection frame respectively, so as to determine the position and size of the face detection frame (or face). According to an embodiment, the process of screening the registered objects can include: determining a face detection frame including each object face in the image frame, and comparing the area of the area surrounded by each face detection frame, selecting the face detection frame with the largest area of the surrounded area, and taking the face contained in the face detection frame as the registered object. In other embodiments of the present disclosure, when the video acquisition device captures the video, a video capture window may be provided through a graphical user interface (GUI) to prompt the subject to place his or her face in the video capture window to complete the video acquisition.
[0041] According to an embodiment, in step S130, the action behaviors of the object in the multiple image frames sequentially arranged in the first image frame sequence are analyzed by analyzing the relative positions between the face detection frames. For example, it is possible to analyze whether the object is moving, the direction of the movement, and the amplitude of the movement. If the object's movement amplitude is too large, it is possible to cause the image frames captured by the video acquisition device to be blurred. Blurred image frames cannot be used for authentication during the registration process, nor can they be stored as final registration data for the object. Therefore, in an embodiment of the present disclosure, by analyzing the relative positions between the face detection frames to determine whether the movement amplitude of the face is within a predetermined range, it is possible to determine whether the captured image frames have reached a preset clarity.
[0042] According to an embodiment, in step S140, if it can be determined that the movement amplitude of the face is within a predetermined range, that is, the image frame reaches a preset clarity, it can be further determined based on the image frame whether the faces in each image frame belong to the same object. According to an embodiment, a face feature extraction model can be used to extract multiple groups of face features, and the extracted face features are feature vectors with a certain dimension.
[0043] According to an embodiment, in step S150, after ensuring that clear image frames containing complete facial features have been used for registration and authentication, and that the faces in each image frame belong to the same object, a specified image frame in the first image frame sequence can be stored as registration data for the object.
[0044] According to the embodiments of the present disclosure, the registration and authentication process can be completed only by analyzing the received video data without the need for the registered object to cooperate with interactions such as blinking and opening the mouth, which greatly simplifies the registration and authentication process.
[0045] Figure 2 FIG. 2 shows a process of acquiring a first image frame sequence from video data according to an embodiment of the present disclosure. Figure 2 As shown, in step S201, image frames are sequentially acquired from a plurality of image frames, where the plurality of image frames are continuous image frames acquired from video data in a video capture order, and the extracted image frame sequence can be temporarily stored in a cache.
[0046] Next, in step S202, parameters for extracting the first image frame sequence may be set, including setting an initial value i=1 of a loop variable i.
[0047] Next, in step S203, starting from the first frame of the multiple image frames, the i-th image frame is acquired in sequence. Next, it is determined whether the acquired image frame contains complete facial features. This is because the model for processing faces has certain requirements on the quality of input data. If the face in the image frame is blocked, or the face deviates greatly from the frontal posture, it is not conducive to the model's processing of the data.
[0048] Next, in step S204, it is determined whether the face is a frontal face based on the face posture detection model. For example, the face key points can be obtained by training the face key points using a deep alignment network (DAN), a fine-tuned convolutional neural network (TCNN), etc. And the obtained face key points are input into the face posture detection model to estimate the posture of the face in the image frame according to the face key points. The face posture detection model can calculate the pitch angle, yaw angle and roll angle of the face respectively, and determine whether the face is a frontal face based on the pitch angle, yaw angle and roll angle, or whether the deflection range of the face is within the allowed range.
[0049] Next, in step S205, if it is determined that the face is a frontal face, whether the face is occluded is determined based on the face occlusion detection model. For example, the face occlusion model of seetaface can be used to determine whether the face is occluded. Alternatively, a lightweight network such as shuffleNet, mobileNet, etc. can be used to classify and train the frontal face and the occluded face to obtain a face occlusion model to determine whether the face is occluded.
[0050] Next, in step S206, when it has been determined that the extracted image frame contains a frontal face and is not obstructed, it is determined that the extracted image frame contains complete facial features, and the extracted image frame, i.e., the i-th image frame, is stored as a frame in the first image frame sequence S1.
[0051] Next, in step S207, it is determined whether the stored first image frame sequence S1 includes a predetermined number of image frames. Here, the predetermined number of image frames can be determined according to the computing power of the computing device that performs the registration. For example, if the computing power of the computing device is strong, the predetermined number of frames can be appropriately increased, for example, the predetermined number of frames can be determined to be 30 or 50 frames or more. If the computing power of the computing device is weak, the predetermined number of frames can be determined to be 20 frames or less. The predetermined number of frames can be determined by weighing the authentication accuracy requirement in the registration process, the computing power of the device, and the registration authentication time requirement. If it is determined that the predetermined number of image frames has been stored in the first image frame sequence S1, the process of continuing to extract image frames is exited to obtain a first image frame sequence S1 including a plurality of image frames of the predetermined number of frames. If it is determined that the predetermined number of image frames has not been stored in the first image frame sequence S1, then in step S208, the loop variable i is increased by 1, that is, i=i+1 is set, and then the process returns to step S203, and the i-th image frame is continuously obtained from the plurality of image frames until the predetermined number of image frames is stored in the first image frame sequence S1.
[0052] The first image frame sequence obtained by the method according to the embodiment of the present disclosure includes multiple image frames each including complete facial features, which can be used for analyzing facial action behaviors and identifying facial features during the registration process.
[0053] According to an embodiment of the present disclosure, determining whether an image frame has reached a preset clarity based on the relative positions between face detection frames in each image frame includes: determining a first ratio of an area of an intersection region of face detection frames in two image frames in a first image frame sequence relative to an area of a union region of face detection frames in the two image frames, and determining that the image frame has reached a preset clarity when the determined first ratios are all greater than a first threshold.
[0054] According to another embodiment of the present disclosure, determining whether an image frame has reached a preset clarity based on the relative positions between face detection frames in each image frame includes: determining a first ratio of an area of an intersection area of face detection frames in two image frames in a first image frame sequence to an area of a union area of the face detection frames in the two image frames, determining a second ratio of the number of first ratios greater than a first threshold to the total number of first ratios, and determining that the image frame has reached a preset clarity when the second ratio is greater than or equal to a second threshold.
[0055] According to an embodiment, the two image frames in the first image frame sequence for performing the calculation may be adjacent image frames or spaced image frames. For example, the first image frame sequence S1 includes image frames F 1 、F 2 、F 3 、F 4、F 5 、F 6 ...and other image frames. In the embodiment of calculating the first ratio for adjacent image frames, the first ratio can be calculated in F 1 and F 2 The first ratio is calculated between 2 and F 3 The first ratio is calculated between 3 and F 4 In another embodiment of calculating the first ratio for the image frames at intervals, the calculation may be performed at intervals of one image frame, for example, at F 1 and F 3 The first ratio is calculated between 3 and F 5 In another embodiment of calculating the first ratio for an image frame at intervals, the calculation may be performed at intervals of two or more image frames, for example, between F 1 and F 4 The first ratio is calculated between, … and so on.
[0056] Figure 3A and Figure 3B Examples of determining whether an image frame reaches a preset definition based on the relative positions between face detection frames according to an embodiment of the present disclosure are respectively shown. Figure 3A and Figure 3B In the description, only the case of calculating the first ratio for adjacent image frames is taken as an example for description.
[0057] like Figure 3A As shown, the first image frame sequence includes multiple image frames, and the ratio of the area of the intersection area of the face detection frames in two image frames to the area of the union area is calculated to analyze the action behavior of the object. Figure 3A As shown in the figure, the ratio of the area of the intersection of two adjacent face detection frames to the area of the union can be calculated as F 12 / (F 1 +F 2 -F 12 ), where F 1 represents the face detection frame in the first image frame, F 2 Indicates the face detection frame in the second image frame, and at the same time F 1 and F 2 Represents the face detection frame F 1 and F 2 The area of F 12 Represents the face detection frame F 1 and F 2 The area of the intersection region.
[0058] According to an embodiment, the first threshold value may be set according to the reliability requirement of the registration and the image clarity requirement. If the first threshold value is set to be larger, the image quality may be improved, that is, the image may be clearer, but it may cause multiple registration authentications to fail to proceed. Conversely, if the first threshold value is set to be smaller, the registration authentication may be smoother, but more unclear images may be introduced, thereby affecting the reliability of the registration authentication. According to an embodiment, the image quality may be ensured by adjusting the first threshold value.
[0059] like Figure 3B As shown in FIG. 1 , the process of calculating the ratio of the area of the intersection region of the face detection frames in adjacent image frames to the area of the union region is similar to Figure 3A The process is the same as shown in Figure 3A Calculate F 12 / (F 1 +F 2 -F 12 ).exist Figure 3B , the number N of first ratios greater than the first threshold 1 Perform statistics and then calculate the number N of first ratios greater than the first threshold 1 The second ratio N relative to the total number N of the first ratio 1 / N. If N 1 / N is greater than or equal to the second threshold, it is determined that the image frame reaches the preset definition.
[0060] In this embodiment, even if the clarity of some image frames does not reach the preset first threshold, for example, F 23 / (F 2 +F 3 -F 23 ) is less than the first threshold, and the image frame is not considered to have failed to reach the preset definition. According to an embodiment, when the image frames reaching the preset definition reach a certain scale, that is, the number N of the first ratios greater than the first threshold is 1 The ratio of the total number N of the first ratios meets a certain requirement, that is, the number N of the first ratios greater than the first threshold 1 The second ratio N relative to the total number N of the first ratio 1 When / N is greater than or equal to the second threshold, it is considered that the image frame reaches the preset definition. According to the embodiment, the quality of the image can be guaranteed by coordinating and adjusting the first threshold and the second threshold. By introducing two adjustment parameters, it is more flexible and accurate to judge whether the image frame reaches the preset definition.
[0061] Figure 4 FIG. 2 shows an example of calculating the intersection of face detection frames based on the coordinates and sizes of face detection frames according to an embodiment of the present disclosure. Figure 4 As shown, Figure 4The coordinate system above is a coordinate system established with the upper left corner of the image frame as the origin. The positive direction of the X axis is the direction extending along one side of the image frame, and the positive direction of the Y axis is the direction extending along the other side of the image frame. Figure 4 As shown, the parameter set [x 1 ,y 1 ,w 1 ,h 1 ] to represent the position and size of the face detection frame in the first image frame. 1 and 1 Indicates the coordinates of the upper left corner of the face detection frame, w 1 Indicates the length of the face detection frame along the X-axis, h 1 represents the length of the face detection frame along the Y axis. Below the coordinate system is the process of intersecting the face detection frame in the first image frame with the face detection frame in the second image frame. Figure 4 As shown, we can determine that the coordinates of the upper left corner of the intersection area are x min =max(x 1 ,x 2 ), y min =max(y 1 ,y 2 ), and the coordinates of the lower right corner of the intersection area can be determined to be x max =min(x 1 +w 1 ,x 2 +w 2 ), y max =min(y 1 +h 1 ,y 2 +h 2 ). The area of the intersection area can be calculated as S according to the coordinates of the upper left corner and the lower right corner of the intersection area. 12 =(x max -x min )*(y max -y min ).
[0062] According to an embodiment, determining whether faces represent the same object based on multiple groups of facial features includes determining the similarity between facial features in any two adjacent image frames in a first image frame sequence, and determining that the faces represent the same object when the determined similarities are greater than a third threshold; otherwise, determining that the faces represent different objects. In an embodiment of the present disclosure, facial features can be obtained by calling a facial feature extraction model. Different facial feature extraction models output feature vectors of different dimensions. For feature vectors, the similarity between facial features in any two adjacent image frames in a first image frame sequence can be determined by calculating the distance between feature vectors. According to an embodiment, the Euclidean distance can be used. Manhattan distance c = |m i -n i |, or Mahalanobis distance etc. to calculate the distance between feature vectors, where m i and n i All represent vectors. According to an embodiment, the setting of the third threshold value can be determined according to the database used by the adopted face feature extraction model. Different face feature extraction models will provide recognition accuracy and corresponding threshold settings. If it is determined through analysis and recognition that the faces in each image frame in the first image frame sequence belong to the same object, the specified image frame in the first image frame sequence can be used as the registration data registration object.
[0063] According to an embodiment, before saving the registration data, the registration data may be compared for similarity with the registration data previously saved in the face library. If the face has already been registered, the storage may not be overwritten.
[0064] According to an embodiment of the present disclosure, by using video for registration and analyzing the relative positions between face detection frames in multiple image frames, the clarity of the image frames can be determined without the user having to cooperate with operations such as blinking and opening the mouth, thereby simplifying the registration and authentication process and ensuring the reliability of the registration data.
[0065] Figure 5 FIG. 5 is a flowchart of a method 500 for identifying and unlocking a face in received video data based on a face library according to another embodiment of the present disclosure. Figure 5 As shown, the method 500 includes the following steps:
[0066] In step S510, input video frame data is received.
[0067] In step S520, a second image frame sequence is acquired from the received video data, and each image frame in the second image frame sequence includes a face detection frame including complete facial features.
[0068] In step S530, it is determined whether the image frame includes a living human face according to the relative positions between the face detection frames in each image frame.
[0069] In step S540, when it is determined that the image frame includes a living human face, facial features are extracted based on the face detection frame.
[0070] In step S550, it is determined whether the facial features match the registration data in the face database to identify the face.
[0071] In step S560 , unlocking is recognized.
[0072] Among them, the operations of steps S510, S520, S540 and S550 can be obtained by referring to steps S110, S120 and S140 in the method 100 for face registration based on video data in the aforementioned embodiment, and will not be repeated here.
[0073] According to an embodiment, determining whether an image frame includes a living face based on the relative positions between face detection frames in each image frame includes determining a face detection frame that meets an overlap condition among the face detection frames in each image frame, determining a third ratio of the number of face detection frames that meet the overlap condition relative to the number of all face detection frames in multiple face detection frames, and determining that the face is a non-living face when the third ratio is greater than or equal to a fourth threshold, and determining that the face is a living face when the third ratio is less than the fourth threshold.
[0074] According to an embodiment, determining a face detection frame that meets the overlap condition in the face detection frame in each image frame includes determining a fourth ratio of the area of an intersection region between any two face detection frames in the face detection frame relative to the area of each face detection frame in the any two face detection frames, and determining that any two face detection frames are face detection frames that meet the overlap condition when the determined fourth ratios are both greater than a fifth threshold value, and determining that any two face detection frames are face detection frames that do not meet the overlap condition when the determined fourth ratios are both less than the fifth threshold value.
[0075] In this embodiment, an intersection operation is performed between any two face detection frames in a plurality of face detection frames, and the ratio of the area of the obtained intersection area to the area of each face detection frame in the two face detection frames where the intersection operation is performed is calculated. The degree of overlap between the two face detection frames can be determined by the obtained ratio. According to the embodiment, a fifth threshold is set to measure the degree of overlap between the two face detection frames. If the fifth threshold is set to a higher value, the overlap can only be determined when the overlap degree of the two face detection frames is high. The overlap of the face detection frames indicates that the object is likely to have no action behavior in the time period between the two face detection frames, that is, the object can be considered to be stationary, and further, the object is considered to be not a living body. Therefore, if the fifth threshold is set to a higher value, the proportion of overlapping face detection frames in all face detection frames will be reduced, and the possibility of non-living bodies being identified as living bodies will be increased. On the contrary, if the fifth threshold is set to a lower value, more face detection frames will be determined to be overlapped, which will increase the possibility of living bodies being identified as non-living bodies. In the application, the fifth threshold can be set according to the occasion of the registration and authentication application. For example, for some occasions where the unlocking function is performed using the method of the embodiments of the present disclosure, the fifth threshold can be set higher, because in these occasions, it can be basically guaranteed that the object is a living body, thereby reducing the possibility of the living body being identified as a non-living body, and can fully ensure that the living object is correctly identified, thereby improving the user experience.
[0076] By analyzing the action of the object, it can be determined whether the face is a live face. That is, in the embodiment of the present disclosure, it is possible to determine whether the object is a live object simply by analyzing the relative positions between the face detection frames in multiple image frames, thereby effectively preventing the operation of unlocking based on a video of a non-live object, for example, avoiding the operation of unlocking using a photo of the object, thereby improving the security of the lock.
[0077] Figure 6 The process of determining a face detection frame that meets the overlap condition among multiple face detection frames according to an embodiment of the present disclosure is shown. Figure 6 As shown, after obtaining the face detection frame F in the first image frame 1 and the face detection frame F in the second image frame 2 After the intersection area between them, it is necessary to calculate the intersection area F 12 The area of the face detection box F 1 The fourth ratio F of the area 12 / F 1 , and calculate the intersection area F 12 The area of the face detection box F 2 The fourth ratio F of the area 12 / F 2 Here, we still use F 1 and F2 Indicates the face detection frame area F 1 and F 2 Then, we need to compare the fourth ratio F 12 / F 1 and F 12 / F 2 The relationship with the fifth threshold is only in the fourth ratio F 12 / F 1 and F 12 / F 2 When both are greater than the fifth threshold, the face detection frame F is determined. 1 and the face detection frame F 2 The overlap condition is met. Similarly, the face detection frame F is calculated 1 and the face detection frame F 7 The fourth ratio F 17 / F 1 and F 17 / F 7 , if the fourth ratio F 17 / F 1 and F 17 / F 7 are all less than or equal to the fifth threshold, then the face detection frame F is determined 1 and the face detection frame F 7 The overlap condition is not met.
[0078] According to an embodiment, if one of the second ratios is less than or equal to the second threshold and the other second ratio is greater than the second threshold, it is determined that the face is a non-living face. This is a special case, which is caused by the large difference in the sizes of the two face detection frames. Figure 6 As shown, each face detection frame (such as F 1 、F 2 、F 7 ) are shown as having the same size. In practice, the sizes of face detection frames may be different from each other, but the size difference between each other is not large. If the size of a face detection frame is significantly different from the sizes of other face detection frames, it means that the face contained in the face detection frame may have a large range of motion, or the face contained in the face detection frame may not belong to the same person as the faces contained in other face detection frames. Therefore, in this case, it is possible to directly determine that the object is non-living based on the comparison result of the fourth ratio and the fifth threshold, and no further determination is made as to whether other face detection frames meet the overlap condition.
[0079] Figure 7 FIG. 7 is a block diagram of an apparatus 700 for performing registration based on video data according to another embodiment of the present disclosure. Figure 7As shown, the device 700 includes a processor 701, a memory 702 and a camera 703. The memory 702 stores machine-readable instructions, and the processor 701 can execute these machine-readable instructions to implement the method 100 for face registration based on video data according to an embodiment of the present disclosure. The camera 703 can be configured to acquire video data, and the frame rate of the camera 703 can be in the range of 15 to 25 frames per second.
[0080] The memory 702 may be in the form of non-volatile or volatile memory, such as Electrically Erasable Programmable Read Only Memory (EEPROM), flash memory, or the like.
[0081] According to the embodiment of the present disclosure, various components inside the device 700 can be implemented by a variety of devices, including but not limited to: analog circuit devices, digital circuit devices, digital signal processing (DSP) circuits, programmable processors, application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), programmable logic devices (CPLDs), and the like.
[0082] Figure 8 FIG. 8 is a block diagram of an electronic whiteboard 800 according to another embodiment of the present disclosure. Figure 8 As shown, the electronic whiteboard 800 according to the embodiment of the present disclosure includes a display whiteboard 801 and a device 802 for performing face registration based on video data according to the embodiment of the present disclosure.
[0083] According to the electronic whiteboard of the embodiment of the present disclosure, a device for registration based on video data is installed, and no human interaction is required. Face registration is performed directly by intercepting the video stream, and registration by directly obtaining video frames is more convenient. According to the electronic whiteboard of the embodiment of the present disclosure, there is no need to manually turn it on and off. It can be directly unlocked and used by face information within a certain distance, and the confidentiality is good. In addition, only the fixed face registered by appointment can be unlocked, which effectively protects the information security of the appointment user during the use of the electronic whiteboard.
[0084] The above detailed description has described numerous embodiments by using schematic diagrams, flow charts and / or examples. Where such schematic diagrams, flow charts and / or examples include one or more functions and / or operations, it should be understood by those skilled in the art that each function and / or operation in such schematic diagrams, flow charts or examples can be implemented individually and / or together by various structures, hardware, software, firmware or substantially any combination thereof.
[0085] Although the present disclosure has been described with reference to several typical embodiments, it should be understood that the terms used are illustrative and exemplary, rather than restrictive. Since the present disclosure can be embodied in a variety of forms without departing from the spirit or substance of the disclosure, it should be understood that the above embodiments are not limited to any of the foregoing details, but should be interpreted broadly within the spirit and scope defined by the appended claims, so all changes and modifications falling within the scope of the claims or their equivalents should be covered by the appended claims.
Claims
1. A method for face registration based on video data, include: receiving video data; Acquire a first image frame sequence from the video data, wherein each image frame in the first image frame sequence comprises a face detection frame including complete facial features; Determining whether the image frame reaches a preset definition according to the relative positions between the face detection frames in each image frame; When it is determined that the image frame reaches a preset definition, extracting multiple groups of facial features based on image information of multiple face detection frames, and determining whether the faces represent the same object according to the multiple groups of facial features; In the case where it is determined that the faces represent the same object, registering the object according to the first sequence of image frames; The method further comprises: storing the registered data obtained by registering the object according to the first image frame sequence into a face database; Acquire a second image frame sequence from the received video data, each image frame in the second image frame sequence comprising a face detection frame including complete facial features; Determining whether the image frame includes a living human face according to the relative positions between the face detection frames in each image frame; When it is determined that the image frame includes a living human face, extracting facial features based on the face detection frame; Determining whether the facial features match the registration data in the face database to identify the face; Determining whether the image frame includes a living human face according to the relative positions of the face detection frames in each image frame comprises: Determine a face detection frame that meets the overlap condition in the face detection frame in each image frame; Determine a third ratio of the number of face detection frames that meet the overlap condition to the number of all face detection frames in the face detection frame; When the third ratio is greater than or equal to the fourth threshold, the face is determined to be a non-living face; when the third ratio is less than the fourth threshold, the face is determined to be a living face.
2. The method according to claim 1, in, Acquiring a first image frame sequence from the video data comprises: Acquire a plurality of image frames from the video data in the order in which the video was captured; Determining whether the image frame contains a face based on a face detection model; and In the case where it is determined that the image frame includes a human face, a face detection frame including the human face is determined in each of the multiple image frames.
3. The method according to claim 2, in, Acquiring a first image frame sequence from the video data further includes: Determining whether the acquired image frame contains complete facial features; When the image frame includes complete facial features, storing the image frame as a frame in a first image frame sequence; When the stored first image frame sequence includes a predetermined number of frames, the acquisition of image frames is terminated.
4. The method according to claim 3, in, Determining whether the acquired image frame contains complete facial features includes: Determining whether the face is a frontal face based on a face posture detection model; In the case where it is determined that the human face included in the image frame is a frontal face, determining whether the human face is blocked based on a face occlusion detection model; In the case where it is determined that the human face included in the image frame is not blocked, determining that the image frame includes complete facial features; and Otherwise, it is determined that the image frame does not contain complete facial features.
5. The method according to claim 1, in, Determining whether the image frame reaches a preset definition according to the relative positions of the face detection frames in each image frame comprises: Determining a first ratio of an area of an intersection region of face detection frames in two image frames in the first image frame sequence to an area of a union region of face detection frames in the two image frames; and When the determined first ratios are all greater than the first threshold, it is determined that the image frame reaches a preset definition.
6. The method according to claim 1, in, Determining whether the image frame reaches a preset definition according to the relative positions of the face detection frames in each image frame comprises: Determine a first ratio of an area of an intersection region of face detection frames in two image frames in the first image frame sequence to an area of a union region of face detection frames in the two image frames; determining a second ratio of the number of the first ratios that are greater than a first threshold relative to the total number of the first ratios; and When the second ratio is greater than or equal to a second threshold, it is determined that the image frame reaches a preset definition.
7. The method according to claim 5 or 6, in, Determining whether the faces represent the same object according to the multiple groups of face features comprises: Determining the similarity between facial features in any two adjacent image frames in the first image frame sequence; and When the determined similarities are all greater than a third threshold, it is determined that the faces represent the same object.
8. The method according to claim 7, in, The facial features include facial feature vectors, and wherein determining the similarity between facial features in any two adjacent image frames in the first image frame sequence includes: Determine the distance between the facial feature vectors in two adjacent image frames in the first image frame sequence.
9. The method according to claim 1, in, Registering the object according to the first sequence of image frames comprises: The object is registered using a designated image frame in the first image frame sequence as registration data.
10. The method according to claim 1, in, The face detection frames in each image frame that meet the overlap conditions are determined to include: Determine a fourth ratio of the area of an intersection region between any two face detection frames in the face detection frames to the area of each face detection frame in the any two face detection frames; In a case where the determined fourth ratios are all greater than the fifth threshold, determining that the two arbitrary face detection frames are face detection frames that meet the overlap condition; and When the determined fourth ratios are all smaller than the fifth threshold, it is determined that the two arbitrary face detection frames are face detection frames that do not meet the overlap condition.
11. The method according to claim 10, in, Determining whether the image frame includes a living human face according to the relative positions between the face detection frames in each image frame further includes: In a case where one of the determined fourth ratios is greater than the fifth threshold and the other fourth ratio is less than or equal to the fifth threshold, it is determined that the human face is a non-living human face.
12. A device for face registration based on video data, include: a memory configured to store instructions; as well as A processor configured to execute the instructions to perform the method according to any one of claims 1 to 11.
13. An electronic whiteboard comprising the device according to claim 12.
Citation Information
Patent Citations
Living body face safety certification device and living body face safety certification method
CN102789572A
Method and device for detecting face image
CN110276277A
Video processing method and device, video operation method and device, storage medium and equipment
CN111541943A