Moving image processing system, moving image processing method, and program
The video processing system uses a learning model to analyze feature point position changes in video data, effectively identifying two-dimensional faces without specialized hardware, addressing the need for hardware-dependent face forgery detection.
Patent Information
- Application Number
- PCT/JP2023/046496
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-07-03
AI Technical Summary
Conventional techniques for preventing face forgery in video data require specialized hardware such as 3D or thermal cameras, which are not commonly installed in general computers, making it difficult to determine if a two-dimensional face is shown without such equipment.
A video processing system that utilizes a learning model to analyze the position changes of feature points in video data, estimating whether a two-dimensional face is present without requiring special hardware by detecting and processing two-dimensional information from standard cameras.
Enables accurate identification of two-dimensional faces in video data using standard imaging units, enhancing security by preventing impersonation without the need for specialized hardware, and improving estimation accuracy through filtering processes.
Smart Images

Figure JP2023046496_03072025_PF_FP_ABST
Abstract
Description
Video processing system, video processing method, and program
[0001] The present disclosure relates to a video processing system, a video processing method, and a program.
[0002] Conventionally, technologies for preventing spoofing by malicious individuals have been studied. For example, Patent Literature 1 describes a technology for performing face authentication using a three-dimensional camera capable of detecting depth and a thermal camera capable of detecting temperature. Even if a malicious individual attempts to spoof someone by using a photograph with another person's face printed on it or a screen on which another person's face is displayed, it is believed that the use of the three-dimensional camera and thermal camera disclosed in Patent Literature 1 can prevent spoofing by malicious individuals.
[0003] Japanese Patent Application Laid-Open No. 2023-029968
[0004] However, in conventional techniques, in order to estimate whether a two-dimensional face is shown in video data, special hardware that is not installed in general computers is required, such as the three-dimensional camera and thermal camera disclosed in Patent Document 1. Therefore, there is a need for a method for estimating whether a two-dimensional face is shown in video data without requiring special hardware.
[0005] One of the objectives of the present disclosure is to estimate whether a two-dimensional face is shown in video data without requiring special hardware.
[0006] The video processing system according to the present disclosure includes a video data acquisition unit that acquires video data relating to a video showing a user's face, a feature point detection unit that detects a plurality of feature points relating to the face based on the video data, and an estimation unit that estimates whether a two-dimensional face is shown in the video data based on a learning model in which positional changes of each of the plurality of feature points in the video data and training data relating to positional changes of each of the plurality of feature points relating to a training face have been learned.
[0007] The present disclosure can estimate, for example, whether a two-dimensional face is shown in video data without requiring special hardware.
[0008] FIG. 1 is a diagram illustrating an example of the hardware configuration of a video processing system. FIG. 2 is a diagram illustrating an example of a screen displayed on a user terminal. FIG. 3 is a diagram illustrating an example of how a malicious user performs impersonation. FIG. 4 is a diagram illustrating an example of functions implemented in the video processing system. FIG. 5 is a diagram illustrating an example of a training database. FIG. 6 is a diagram illustrating an example of a user database. FIG. 7 is a diagram illustrating an example of processing for individual frames of video data. FIG. 8 is a diagram illustrating an example of processing for all video data. FIG. 9 is a diagram illustrating an example of processing executed in the video processing system. FIG. 10 is a diagram illustrating an example of functions implemented in a modified example. FIG. 11 is a diagram illustrating an example of a learning model of modified example 1.
[0009] [1. Hardware configuration of video processing system] An example of an embodiment of a video processing system, video processing method, and program according to the present disclosure will be described. Fig. 1 is a diagram showing an example of the hardware configuration of a video processing system. For example, the video processing system 1 includes a learning terminal 10, a server 20, and a user terminal 30. Each of the learning terminal 10, the server 20, and the user terminal 30 is connected to a network N such as the Internet or a LAN.
[0010] The learning terminal 10 is a computer that performs learning of the learning model described below. For example, the learning terminal 10 is a personal computer, tablet, or smartphone. For example, the learning terminal 10 includes a control unit 11, a memory unit 12, a communication unit 13, an operation unit 14, and a display unit 15. For example, the control unit 11 includes at least one processor. The memory unit 12 includes at least one of volatile memory such as RAM and non-volatile memory such as flash memory. The communication unit 13 includes at least one of a communication interface for wired communication and a communication interface for wireless communication. The operation unit 14 is an input device such as a touch panel or a mouse. The display unit 15 is a display such as an LCD or organic EL.
[0011] The server 20 is a server computer that stores a trained learning model. For example, the server 20 includes a control unit 21, a storage unit 22, and a communication unit 23. The hardware configurations of the control unit 21, the storage unit 22, and the communication unit 23 may be similar to those of the control unit 11, the storage unit 12, and the communication unit 13, respectively.
[0012] The user terminal 30 is a user's computer. For example, the user terminal 30 is a personal computer, a smartphone, a tablet, or a wearable terminal. For example, the user terminal 30 includes a control unit 31, a memory unit 32, a communication unit 33, an operation unit 34, a display unit 35, and a photographing unit 36. The hardware configurations of the control unit 31, the memory unit 32, the communication unit 33, the operation unit 34, and the display unit 35 may be similar to those of the control unit 11, the memory unit 12, the communication unit 13, the operation unit 14, and the display unit 15, respectively. The photographing unit 36 includes at least one camera.
[0013] The programs stored in the storage units 12, 22, 32 may be supplied to the learning terminal 10, the server 20, or the user terminal 30 via the network N. Also, at least one of a reading unit (e.g., a memory card slot) that reads a computer-readable information storage medium and an input / output unit (e.g., a USB port) for inputting and outputting data to and from an external device may be included in the learning terminal 10, the server 20, or the user terminal 30. For example, a program stored in an information storage medium may be supplied to the learning terminal 10, the server 20, or the user terminal 30 via at least one of the reading unit and the input / output unit.
[0014] Furthermore, the video processing system 1 only needs to include at least one computer. The computers included in the video processing system 1 are not limited to the example in FIG. 1. For example, the video processing system 1 may include only the learning terminal 10 and the server 20. In this case, the user terminal 30 exists outside the video processing system 1. The video processing system 1 may include only the server 20. In this case, the learning terminal 10 and the user terminal 30 exist outside the video processing system 1. The video processing system 1 may include only the server 20 and a computer not shown in FIG. 1.
[0015] [2. Overview of Video Processing System] In this embodiment, an example is taken in which a user's face is photographed by the photographing unit 36 to verify the user's identity. Identity verification is a process for verifying that the user is the person in question. For example, eKYC (electronic Know Your Customer) is a type of identity verification. Identity verification may be performed in any situation for any service. For example, identity verification is performed when a user uses a payment service, financial service, communication service, e-commerce service, government service, or other service, or when registering as a member.
[0016] For example, when a situation arises where identity verification is required, the user terminal 30 activates the image capture unit 36. The image capture unit 36 continuously captures images of the user's face in video mode. The user terminal 30 generates video data showing the user's face based on the results of the image capture by the image capture unit 36. The user terminal 30 transmits the video data to the server 20. The server 20 executes a process for identity verification based on the video data.
[0017] 2 is a diagram showing an example of a screen displayed on the user terminal 30. For example, the user terminal 30 displays a shooting screen SC showing the results of shooting by the shooting unit 36 on the display unit 35. The shooting screen SC displays a guide G for guiding the position of the face and a message MS indicating an action required of the user. In the example of FIG. 2, the user is requested to turn his / her face from the front to the left. The user adjusts the position and orientation of the user terminal 30 so that his / her face fits within the guide G, and performs the action indicated by the message MS.
[0018] The action requested of the user may be any action. The action requested of the user is not limited to the example of FIG. 2 . For example, the user may be requested to turn his / her face up from the front, turn his / her face to the right from the front, turn his / her face down from the front, blink, wink, or other action. As in a modified example described later, the action requested of the user may be randomly determined from a plurality of actions. These actions are performed to prevent impersonation by a malicious user.
[0019] For example, a malicious user may impersonate another person by taking a photo of a piece of paper on which the face of another person is printed using the user's own user terminal 30. A malicious user may impersonate another person by taking a photo of a first user terminal 30 on which the face of another person is displayed using a second user terminal 30. Even if the operation shown in FIG. 2 is requested, a malicious user may cleverly circumvent identity verification.
[0020] 3 is a diagram showing an example of how a malicious user may impersonate another person. For example, the malicious user holds a piece of paper P with another person's face printed on it in front of his or her own face, with the side with the other person's face printed on it facing the photographing unit 36. The malicious user causes the photographing unit 36 of his or her own user terminal 30 to photograph the paper P. When the malicious user is required to turn his or her face from the front to the left, the malicious user may turn the paper P from the front to the left in an attempt to evade identity verification.
[0021] For example, even when other actions are requested, a malicious user may attempt to bypass identity verification by, for example, holding the paper P facing up. It is believed that such impersonation can be prevented if the person in charge of identity verification checks the video data uploaded from the user terminal 30. However, when verifying the identities of a large number of users, it is not realistic for the person in charge to check all of the video data. Therefore, the video processing system 1 analyzes the video data to estimate whether a two-dimensional face is shown in the video data.
[0022] For example, if the image capturing unit 36 is a 3D camera, a stereo camera, or a thermal camera, the video processing system 1 can detect the surface shape or surface temperature of the subject to estimate whether a 2D face is shown in the video data. However, in this case, special hardware such as a 3D camera, a stereo camera, or a thermal camera is required. A typical user terminal 30, such as a smartphone, does not include such special hardware.
[0023] Therefore, the video processing system 1 of this embodiment is configured to estimate whether a two-dimensional face is shown in the video data by analyzing the video data without requiring special hardware. For example, the positional changes of feature points detected from each frame of the video data differ between when a three-dimensional face is shown in the video data and when a two-dimensional face is shown in the video data. The video processing system 1 estimates whether a two-dimensional face is shown in the video data based on a learning model in which these features have been learned. The video processing system 1 will be described in detail below.
[0024] 3. Functions Implemented by the Video Processing System FIG. 4 is a diagram showing an example of functions implemented by the video processing system 1. As shown in FIG.
[0025] [3-1. Functions Realized by the Learning Terminal] For example, the learning terminal 10 includes a data storage unit 100 and a learning unit 101. The data storage unit 100 is realized by the storage unit 12. The learning unit 101 is realized by the control unit 11.
[0026] [Data Storage Unit] The data storage unit 100 stores data necessary for learning the learning model M. For example, the data storage unit 100 stores actual data of the learning model M and a training database DB1. The learning model M is a model created using a machine learning technique. Machine learning may have various known definitions. Machine learning in this embodiment has a broad meaning including deep learning, reinforcement learning, and AI (Artificial Intelligence). For example, the learning model M may be any of supervised learning, semi-supervised learning, and unsupervised learning models. The learning model M in this embodiment is assumed to be a supervised learning model.
[0027] The learning model M can be any of various models used in the field of image analysis. For example, the learning model M may be a neural network, a vision transformer, a support vector machine, a random forest, or a model of another method. The actual data of the learning model M includes a program that indicates various processes such as calculation of an embedded representation, and parameters referenced by the program. For example, the parameters are a weighting coefficient and a bias. The program and parameters included in the actual data of the learning model M may be programs and parameters used in known methods.
[0028] 5 is a diagram showing an example of the training database DB1. The training database DB1 is a database that stores training data used in learning the learning model M. For example, the training data includes an input portion that is input to the learning model M during learning, and an output portion that indicates the content that should be output from the learning model M during learning. The output portion of the training data can also be considered the correct answer during learning. If the learning model M is an unsupervised learning model, annotation of the training data is not performed, and therefore the training data does not include an output portion.
[0029] For example, the input portion of the training data is position change data indicating the position changes of each of a plurality of feature points on a training face. The training face is a face shown in a video prepared for learning the learning model M. A feature point is a characteristic part of a face. For example, a feature point is at least one pixel indicating a part of a face. A feature point is also called a facial landmark. In this embodiment, an example is given in which the vertices constituting a mesh indicating the surface shape of the face correspond to the feature points.
[0030] The feature point detection method may be a known method. For example, feature points may be detected based on the Dlib library provided for Python, a type of programming language. Many libraries for feature point detection are also provided for other programming languages, so feature points may be detected based on a library of another programming language. Instead of a programming language library, feature points may be detected using a machine learning method such as deep learning (for example, a method called Deep3DFaceReconstruction or RingNet).
[0031] For example, the position change data corresponding to the input portion of the training data indicates the position change of a plurality of feature points detected from each of a plurality of frames of video data relating to a video showing a training face. The position change can also be referred to as a speed or a movement amount. In this embodiment, the position change data corresponding to the input portion of the training data is generated using a calculation method similar to that used by the estimation unit 203 described below. Details of this calculation method will be described later.
[0032] In this embodiment, the position change data is in vector format. For example, if the number of frames included in the video data is m (m is a natural number) and k (k is a natural number) feature points are detected from one frame, the position change data is a vector of at least m×k dimensions. Each element of the vector indicated by the position change data indicates the position change of a specific feature point corresponding to a specific frame.
[0033] For example, the first to k-th dimensions of the vector indicated by the position change data indicate the position change of each of the k feature points detected from the first frame. Each dimension indicates the position change of a feature point indicating a specific location on the forehead, the second dimension indicates the position change of the feature point to the right of that, and the third dimension indicates the position change of the feature point further to the right of that, and so on. The k+1th dimension and onwards indicate the position change from the first frame onwards. The position changes are arranged in a predetermined order up to the last dimension of the vector indicated by the position change data.
[0034] The position change data may be data in any format. The position change data may be data relating to the time-series position change of each of the multiple feature points of the video data. The position change data is not limited to vector format data. For example, the position change data may be multiple numerical values, a single numerical value, an array format, a matrix format, or other formats. Data in these formats may indicate the time-series position change of each of the multiple feature points of the video data.
[0035] For example, the output portion of the training data is a label indicating whether the input portion of the training data represents a two-dimensional face. A two-dimensional face means that it is not a three-dimensional face. For example, if a piece of paper with a photograph of another person's face printed on it or a computer displaying another person's face is shown in the training video data, the input portion of the training data indicates that it represents a two-dimensional face. Since the surface of the paper or the surface of a computer screen has a two-dimensional plane, the position change data (the input portion of the training data) obtained from the training video data representing these will represent a two-dimensional face.
[0036] The label is a classification of the input portion of the training data. A first value (e.g., 1) of the label indicates that the input portion of the training data is a two-dimensional face. A second value (e.g., 0) of the label indicates that the input portion of the training data is not a two-dimensional face. The value of the output portion of the training data may be specified by the creator of the learning model M or may be determined by a known annotation tool.
[0037] For example, position change data acquired from training video data showing a three-dimensional face and position change data acquired from training video data showing a two-dimensional face have different characteristics. The difference in these characteristics is thought to be correlated with whether or not a two-dimensional face is shown in the training video data. The training data can allow the learning model M to learn such correlation.
[0038] Taking the example of turning one's face from the front to the left, as shown in Figure 2, the positional change of feature points near the left eye or left cheek on a two-dimensional face is greater than the positional change of feature points near the left eye or left cheek on a three-dimensional face. This is because a two-dimensional face printed on paper P or displayed on a screen does not have any unevenness in the depth direction (the direction perpendicular to the paper P or screen). Such positional change is correlated with whether or not the two-dimensional face is shown in the training video data (i.e., the label value). Such positional change can be learned by the learning model M using training data. Not only the positional change of the left eye or left cheek in the direction the face is turned, but also the positional change of other parts, such as the right eye or right cheek, is correlated with whether or not the two-dimensional face is shown in the training video data. The learning model M can learn such correlations using training data.
[0039] For example, for other movements, such as turning one's head from front to up, there is a correlation between the positional changes of feature points and whether or not a two-dimensional face is shown in the training video data. Therefore, for other movements, the learning model M learns training data indicating the relationship between the positional change data acquired from the training video data, the positional changes of feature points, and whether or not a two-dimensional face is shown in the training video data, and thereby the learning model M can learn these correlations.
[0040] The training data is not limited to the example shown in FIG. 2 . The training data may be data indicating at least one of three-dimensional facial feature point characteristics and two-dimensional facial feature point characteristics. For example, the input portion of the training data may be video data itself relating to a video showing a training face. The input portion of the training data may indicate the positions of multiple feature points detected from each of multiple frames of the video data, rather than the positional changes of each of the multiple feature points. The output portion of the training data may be a score indicating suspicion of spoofing (e.g., a numerical value having an intermediate value, rather than a binary value like a label).
[0041] Furthermore, the data stored in the data storage unit 100 is not limited to the above example. The data storage unit 100 may store any data. For example, the data storage unit 100 may store a learning program that indicates a series of processes during learning. The learning program may be a program adopted in a known machine learning method. For example, the learning program may be a program created based on a known method such as backpropagation or gradient descent.
[0042] [Learning Unit] The learning unit 101 performs learning of the learning model M based on the training data stored in the training database DB1. For example, the learning unit 101 performs learning of the learning model M by adjusting the parameters of the learning model M based on a known learning program so that when an input portion of the training data is input to the learning model M, an output portion of the training data is output from the learning model M. When learning of the learning model M is completed, the learning unit 101 records the trained learning model M in the data storage unit 100. The learning model M before learning may be stored in the data storage unit 100 as data separate from the trained learning model M, or may be overwritten with the trained learning model M. The learning unit 101 transmits the trained learning model M to the server 20. The trained learning model M transmitted to the server 20 is made available for use by the user.
[0043] [3-2. Functions Realized by the Server] For example, the server 20 includes a data storage unit 200, a video data acquisition unit 201, a feature point detection unit 202, an estimation unit 203, and an identity verification unit 204. The data storage unit 200 is realized by the storage unit 22. Each of the video data acquisition unit 201, the feature point detection unit 202, the estimation unit 203, and the identity verification unit 204 is realized by the control unit 21.
[0044] [Data Storage Unit] The data storage unit 200 stores data necessary for estimation using the learning model M. For example, the data storage unit 200 stores a trained learning model M and a user database DB2. In this embodiment, since the learning of the learning model M is performed by the learning terminal 10, the data storage unit 200 stores the learning model M transmitted from the learning terminal 10. When the server 20 performs the learning of the learning model M, the server 20 has the same functions as the learning unit 101.
[0045] FIG. 6 is a diagram showing an example of the user database DB2. The user database DB2 is a database in which various data related to users is stored. For example, the user database DB2 stores a user ID, video data, and identity verification data. The user database DB2 may store any data related to users. The data stored in the user database DB2 is not limited to the example of FIG. 6. For example, if identity verification is performed when a user registers as a member of a service, data indicating the application status for the service may be stored in the user database DB2.
[0046] The user ID is an example of user identification information that can identify a user. For example, an email address or a phone number may be used as user identification information. The video data is data related to a video that shows the user identified by the user ID performing an action such as turning their face from the front to the left. The video data may be in any data format (e.g., MP4 or AVI). The user terminal 30 of the user to be identified transmits video data generated based on the shooting results of the shooting unit 36 to the server 20. The server 20 stores the video data received from the user terminal 30 in association with the user ID of the user. The identity verification data is data that indicates the result of identity verification. For example, the identity verification data may indicate the result of estimation by the estimation unit 203.
[0047] The data stored in the data storage unit 200 is not limited to the above example. The data storage unit 200 may store any data. For example, the data storage unit 200 may store an identity verification program that indicates processes other than the processes of the estimation unit 203 among a series of identity verification processes. The processes of the identity verification unit 204, which will be described later, are indicated in the identity verification program. In cases where a check of the user's identity verification document is performed in addition to a check of the user's face, the data storage unit 200 may store document data indicating the identity verification document uploaded by the user.
[0048] [Video Data Acquisition Unit] The video data acquisition unit 201 acquires video data related to a video showing a user's face performing a predetermined action. The predetermined action is a motion using the user's face. In other words, the predetermined action is a motion that changes the position of at least one feature point. In yet another way, the predetermined action is a motion in which the tendency of the position change of the feature point differs between when a three-dimensional face is shown in the motion data and when a two-dimensional face is shown in the video data. For example, the predetermined action may be a motion that changes the direction of the face, a motion that changes the position of the face, a motion that uses a part of the face (e.g., blinking or winking), a motion that changes facial expression, or any other motion.
[0049] In this embodiment, since the video data is stored in the user database DB2, the video data acquisition unit 201 acquires the video data from the user database DB2. If the video data is stored in a database other than the user database DB2, the video data acquisition unit 201 may acquire the video data from the other database. If the video data is stored in a computer other than the server 20 (e.g., the user terminal 30) or an external information storage medium, the video data acquisition unit 201 may acquire the video data from the other computer or the external information storage medium.
[0050] In this embodiment, video data is generated without using special hardware such as a 3D camera or a thermal camera. Therefore, the video data acquisition unit 201 acquires video data generated based on the results of imaging by a general imaging unit 36 (e.g., an RGB camera) that does not detect depth information or temperature information. For example, the general imaging unit 36 does not include a depth sensor or a temperature sensor. The video data does not include depth information or temperature information of the subject, but only includes two-dimensional information (planar information) of the subject. The video data indicates the color of each pixel in each frame. The video data does not include depth information or temperature information as information for each pixel.
[0051] [Feature Point Detection Unit] The feature point detection unit 202 detects a plurality of feature points related to a face based on video data. Detecting feature points means acquiring the positions of feature points from frames of video data. In this embodiment, an example is given in which the positions of feature points are expressed as two-dimensional information. However, since there are technologies that can estimate three-dimensional information from two-dimensional images, such technologies may be used to express the positions of feature points as three-dimensional information.
[0052] In this embodiment, a detection program for detecting feature points is stored in the data storage unit 200. As described above, the feature point detection method may be a known method. The detection program indicates the processing of the known detection method. The detection program may be a library such as the Dlib library described above, or may be a program of a machine learning method such as deep learning (for example, a method called Deep3DFaceReconstruction or RingNet).
[0053] 7 is a diagram showing an example of processing for each frame of video data. i (i is an integer) indicates the face of the user when he / she starts to turn his / her face from the front to the left. The feature point detection unit 202 detects individual frames x of the video data based on the detection program. i For example, the feature point detection unit 202 detects a plurality of feature points from the frame x iThe face mesh shown in is detected, and the vertices that make up the mesh are detected as feature points. i In the image, the feature points are indicated by dots. The feature point detection unit 202 detects the feature points of each frame x i Feature point data indicating a plurality of feature points detected from the image is recorded in the data storage unit 200.
[0054] For example, the feature point data indicates a feature point ID and the position of the feature point. The feature point ID is an ID that can identify the feature point. The feature point ID may be able to identify the facial feature that the feature point indicates. For example, the feature point ID may be able to identify whether the feature point indicates the tip of the nose or the edge of the right eye. The method for identifying the feature point may be a method that is commonly used for detecting feature points. For example, the feature point ID may not be particularly used, and the feature point data may indicate only the position of the feature point.
[0055] The position of the feature point is frame x i For example, the coordinates of the frame x i The position of the feature point is indicated by the coordinates of the coordinate axes with the upper left corner of the frame as the origin. By tracking the position associated with a certain feature point ID, it is possible to track the position change of the feature point indicated by that feature point ID. i indicates the position change of the feature point. The method for calculating the position change will be described later.
[0056] 8 is a diagram showing an example of processing for the entire video data. In the example of FIG. 8, the feature point detection unit 202 detects the feature points of frame x 1 ~x m (In FIG. 8, m is an integer equal to or greater than 21) for each frame, the feature point detection unit 202 executes the process described with reference to FIG. 7 to detect multiple feature points. 1 ~x m The feature point data can be considered as a time-series history of the feature points. 1 ~x m The position change corresponding to each of 1 ~h m It is indicated by the symbol y. 1 ~y nThe sign of is information calculated to complement the position change. The details of the complementation method will be described later.
[0057] The feature point detection unit 202 detects all frames x of the video data. i Feature points may be detected from a partial frame x i Alternatively, feature points may be detected from only the frame x constituting the video data. i The number of frames x that are the target of feature point detection may be any number. i The number of frames x may also be any number. 1 the previous frame x before 0 Similarly, there may be a frame x m frame x, which is a frame that is later in time than m+1 In the following description, the frame reference numerals will be omitted unless it is necessary to refer to Figures 7 and 8.
[0058] [Estimation Unit] The estimation unit 203 estimates whether a two-dimensional face is shown in the video data based on the positional changes of each of the multiple feature points in the video data and a learning model M in which training data related to the positional changes of each of the multiple feature points related to a training face has been learned. The estimation unit 203 calculates the positional changes of each of the multiple feature points detected by the feature point detection unit 202 to obtain positional change data. The estimation unit 203 inputs the obtained positional change data to the learning model M. The learning model M calculates an embedded representation of the positional change data based on parameters adjusted by learning. The learning model M outputs a label corresponding to the calculated embedded representation. The estimation unit 203 estimates whether a two-dimensional face is shown in the video data by obtaining the label output from the learning model M.
[0059] For example, for each calculation target frame (a frame for which a position change is to be calculated), the estimation unit 203 calculates the position change of the calculation target frame based on at least two of: a plurality of feature points detected from the calculation target frame; a plurality of feature points detected from a previous frame (a frame that precedes the calculation target frame); and a plurality of feature points detected from a subsequent frame (a frame that follows the calculation target frame). The estimation unit 203 may use at least one previous frame and at least one subsequent frame in calculating the position change of the calculation target frame. The estimation unit 203 may use one previous frame and one subsequent frame in calculating the position change of the calculation target frame, or may use multiple previous frames and multiple subsequent frames in calculating the position change of the calculation target frame.
[0060] In this embodiment, the estimation unit 203 performs a predetermined filtering process to suppress noise related to the video data and calculate the position change. For example, the estimation unit 203 performs filtering on a history of feature points detected from the video data. The estimation unit 203 may perform filtering on the video data itself, rather than on the history of feature points detected from the video data. The filtering may be a known filtering process that can be used to suppress noise. The filtering may be a known filtering process for smoothing features in an image. Note that the estimation unit 203 may calculate the position change without performing filtering. In this case, the estimation unit 203 does not have a filtering function. An estimation unit 203 that does not perform filtering is also included in the scope of the present disclosure.
[0061] For example, the estimation unit 203 performs filtering using a DoG (Derivative of Gaussian) filter. The Gaussian function used in the DoG filter may be a known function. Parameters specified by coefficients of the DoG filter may be any values. The coefficients adjust the Gaussian distribution. In addition to suppressing noise, the DoG filter may also be used to emphasize edges. The estimation unit 203 reduces noise generated in at least one of the multiple feature points by the DoG filter, and smooths the positional changes of each of the multiple feature points to calculate the positional changes of each of the multiple feature points.
[0062] The estimation unit 203 may perform filtering using a filter other than a DoG filter. For example, the estimation unit 203 may perform filtering using a Gaussian filter, a Sobel filter, a median filter, or a linear filter that are not classified as DoG filters. The estimation unit 203 may perform these filtering processes to calculate position changes of feature points. When a filter that is not capable of filtering feature point history is used, the estimation unit 203 may perform filtering using a Gaussian filter, a Sobel filter, or the like on each frame of the video data, and then cause the feature point detection unit 202 to detect each of the multiple feature points and calculate the position changes of each of the multiple feature points.
[0063] For example, the estimation unit 203 sets a local window W including a calculation target frame, which is a frame for which a position change is to be calculated, a previous frame, which is a frame before the calculation target frame, and a subsequent frame, which is a frame after the calculation target frame. In the example of FIG. 8 , the calculation target frame is frame x 4 The previous frame is frame x 1 ~x 3 The next frame is frame x 5 ~x 7 The number of previous frames and the number of subsequent frames may be any number. The number of previous frames and subsequent frames is not limited to three as shown in FIG. 8. For example, the number of previous frames and subsequent frames may be one, two, four or more.
[0064] For example, the estimation unit 203 calculates the position change of the calculation target frame based on the positions of each of the plurality of feature points detected from the previous frame and the positions of each of the plurality of feature points detected from the subsequent frame. In the example of FIG. 8 , the estimation unit 203 calculates the position change of the calculation target frame based on the positions of each of the plurality of feature points detected from the previous frame, 1 ~x 3 and the positions of the feature points detected from each of the frames x 5 ~x 7 and the position of each of the plurality of feature points detected from each of the frames x 4 Calculate the change in position of
[0065] For example, the estimation unit 203 estimates the previous frame x 1 ~x 3 and the average value of the positions of the plurality of feature points detected from each of the frames x 5 ~x 7 The difference between the average value of each position of the plurality of feature points detected from each of the frames x 4 The average value is calculated as a change in position of the frame. Even if a particular frame contains noise, the noise is suppressed by calculating the average value. The average value may be a simple average or a weighted average. When the average value is a weighted average, the weighting coefficient may be larger the closer the frame to be calculated.
[0066] In this embodiment, the frame x 4 The position change of the feature points is expressed by a vector with the same number of dimensions as the number of feature points. The estimation unit 203 performs the above-mentioned filtering process on the history of feature points in the previous and subsequent frames, rather than the average position value, to suppress noise and smooth the position change of the feature points, and then calculates the average position change of the frame x 4 For example, the estimation unit 203 may calculate the position change of the frame x within the local window W by DoG. 1 ~x 7 The time-series positional changes of the feature points are tracked and the target frame is calculated. 4The position change of the
[0067] For example, the estimation unit 203 moves the local window on the time axis of the video data (t axis in the example of FIG. 8 ) and calculates the position change of each of the multiple frames to be calculated one after another. In the example of FIG. 8 , the estimation unit 203 calculates the position change of each of the multiple frames to be calculated one after another. 4 When calculating the position change of the local window W, we move the local window W by one and calculate the position change of the frame x 5 is set as the frame to be calculated. In this case, the previous frame is frame x 2 ~x 4 The next frame is frame x 6 ~x 8 The estimation unit 203 estimates the previous frame x 2 ~x 4 and the positions of the feature points detected from each of the frames x 6 ~x 8 and the position of each of the plurality of feature points detected from each of the frames x 5 Calculate the change in position of
[0068] Thereafter, the estimation unit 203 moves the local window W one by one until the last frame x m After the calculation of the position change is completed, the estimation unit 203 acquires position change data of the video data to be estimated. The position change data indicates, in time series, the position change of each feature point calculated from each frame of the video data to be estimated. As described above, in this embodiment, the position change data is in vector format, but the position change data may be in a format other than vector format.
[0069] In the example of FIG. 8, frame x 1 Frame x before 0 etc., then frame x 1 ~x 3 To calculate the position change of the previous frame x 0 etc. may be used. 0If there is no local window W, the estimation unit 203 does not use the local window W but estimates the frame x 1 ~x 3 Similarly, the position change of frame x m Frame x after m+1 etc., then frame x m To calculate the position change, the following frame x m+1 etc. may be used. m+1 If there is no local window W, the estimation unit 203 does not use the local window W but estimates the frame x m The position change may be calculated as follows.
[0070] For example, the estimation unit 203 estimates frame x 1 The position change of frame x 1 and the positions of the feature points detected from the frame x 2 For example, the estimation unit 203 may calculate the feature points based on the positions of the feature points detected from the frame x 2 The position change of frame x 2 and the positions of the feature points detected from the frame x 3 For example, the estimation unit 203 may calculate the feature points based on the positions of the feature points detected from the frame x 3 The position change of frame x 3 and the positions of the feature points detected from the frame x 4 The calculation may be performed based on the positions of each of the feature points detected from the frame x m Similarly, for the last frames of the video data, such as the first frame, the position change of the last frames may be calculated.
[0071] In this embodiment, the estimation unit 203 calculates the positional change between at least one frame of video data based on the positional change between the frames and the positional change between the frames that follows the frame. The estimation unit 203 estimates whether a two-dimensional face is shown in the video data based on the positional change between the frames. For example, the estimation unit 203 calculates the positional change between frames based on the positional change between at least one frame of the video data, a first weighting factor associated with the frame, the positional change between the frame that follows the frame, and a second weighting factor associated with the frame that follows the frame.
[0072] In this embodiment, the estimation unit 203 estimates a frame x i The position change yi between frames is calculated based on the following equation 1. i , h i+1 are the frames x i , x i+1 is the position change of each of the multiple feature points detected from i is the first weighting factor. i is as shown in Equation 2. i+1 is the second weighting factor. i+1 is as shown in Equation 3. n in Equations 2 and 3 is the number of frames between which position changes are to be calculated. When position changes are interpolated between all frames, n may be m-1. n may be any value.
[0073]
[0074]
[0075]
[0076] The first weighting coefficient w i and the second weighting factor w i+1 is not limited to the examples of Equations 2 and 3. For example, the first weighting coefficient w i and the second weighting factor w i+1 may be calculated. iand the second weighting factor w i+1 The first weighting coefficient w may be a fixed value (e.g., 0.3, 0.7). i and the second weighting factor w i+1 The sum of the first weighting coefficient w i and the second weighting factor w i+1 Instead of using the above, the estimation unit 203 may calculate a simple average value of the position change of at least one frame of the video data and the position change of the frame following that frame as the position change between these frames.
[0077] For example, the estimation unit 203 acquires final position change data by inserting inter-frame position changes into the position changes based on each of the multiple feature points detected from each frame. Inserting inter-frame position changes increases the number of dimensions of the vector indicated by the position change data accordingly. Inter-frame position changes may also be inserted into the position change data that is the input portion of the training data. The number of dimensions of the position change data that is the input portion of the training data is assumed to be the same as the number of dimensions of the position change data acquired from the video data to be estimated. In other words, the number of dimensions of the position change data input to the learning model M during learning is assumed to be the same as the number of dimensions of the position change data input to the learning model M during estimation.
[0078] The number of dimensions of the position change data, which is the input portion of the training data, may differ from the number of dimensions of the position change data acquired from the video data to be estimated. The number of dimensions of the position change data input to the learning model M during learning may differ from the number of dimensions of the position change data input to the learning model M during estimation. In this case, the portion with insufficient dimensions may be treated as a missing value. Alternatively, for example, the learning model M may calculate an embedded representation to absorb the difference in the number of dimensions. For example, if the embedded representation is also expressed as a vector, the number of dimensions of the embedded representation may be predetermined. The learning model M may calculate an embedded representation with a predetermined number of dimensions, regardless of the number of dimensions of the position change data, which is the input portion of the training data, and the number of dimensions of the position change data acquired from the video data to be estimated. The number of dimensions of the embedded representation may not be predetermined and may be indefinite.
[0079] The estimation unit 203 may also determine whether the position change of each of multiple frames of the video data is within a reference range. The reference range is a range that serves as a reference for whether or not to complement the position change between frames. For example, if the user moves their face faster than expected, the position change may be too large and the estimation unit 203 may not be able to perform accurate estimation. Therefore, a position change equal to or greater than a predetermined threshold may be considered to be outside the reference range. The threshold may be arbitrarily determined by an administrator of the video processing system 1.
[0080] For example, the estimation unit 203 determines whether the position change for each of multiple frames of video data is equal to or greater than a threshold value, thereby determining whether the position change is within a reference range. The position changes that are the subject of the determination of the reference range may be the position changes of all feature points, or may be the position changes of some of the feature points. Furthermore, the position changes that are the subject of the determination of the reference range may be the position changes of all frames, or may be the position changes of some of the frames.
[0081] For example, when it is determined that the position change is outside the reference range, the estimation unit 203 calculates the position change between frames. Since the position changes of only some frames in the video data may fall outside the reference range, in this case the estimation unit 203 may calculate the position change between frames for some frames whose position changes fall outside the reference range, or may calculate the position change between frames for all frames of the video data.
[0082] For example, if the position change is not determined to be outside the reference range, the estimation unit 203 does not calculate the position change between frames. 1 ~y n is not calculated, and h 1 ~h m Only the positional changes of some frames in the video data may not be outside the reference range. In this case, the estimation unit 203 may not calculate the inter-frame positional changes of some frames whose positional changes are not determined to be outside the reference range, but may instead calculate the inter-frame positional changes of other frames whose positional changes are determined to be outside the reference range. Furthermore, in this case, the estimation unit 203 may not calculate the inter-frame positional changes of all frames in the video data. Furthermore, regardless of whether the positional changes are outside the reference range, the estimation unit 203 may perform estimation based on the positional changes calculated from the frames without calculating the inter-frame positional changes. In this case, the estimation unit 203 may not have a function for calculating the inter-frame positional changes.
[0083] [Personal Identity Verification Unit] The personal identity verification unit 204 performs user identity verification based on the estimation result of whether or not a two-dimensional face is shown in the video data. For example, if the personal identity verification unit 204 estimates that a two-dimensional face is shown in the video data, it determines that personal identity verification has failed. If the personal identity verification unit 204 estimates that a two-dimensional face is shown in the video data, it may request the user to perform a predetermined action again. The action requested again may be an action different from the action previously requested.
[0084] For example, if it is estimated that a two-dimensional face is not shown in the video data, the identity verification unit 204 determines that identity verification has been successful. If a next step (e.g., photographing an identification document such as a driver's license) is also prepared for identity verification, the identity verification unit 204 may proceed to the next step of identity verification. The steps required for identity verification may be the same as those of known identity verification. The information processing required for this step may also be processing employed in known identity verification. If it is estimated that a two-dimensional face is shown in the video data, the identity verification unit 204 may notify a person in charge of identity verification. In this case, the person in charge may visually check the video data.
[0085] [3-3. Functions Realized by User Terminal] For example, the user terminal 30 includes a data storage unit 300, a display control unit 301, and an operation reception unit 302. The data storage unit 300 is realized by the storage unit 32. The display control unit 301 and the operation reception unit 302 are each realized by the control unit 31.
[0086] [Data Storage Unit] The data storage unit 300 stores data necessary for uploading video data. For example, the data storage unit 300 stores video data generated based on the results of shooting by the shooting unit 36.
[0087] [Display Control Unit] The display control unit 301 causes the display unit 35 to display various screens in the moving image processing system 1. For example, the display control unit 301 causes the display unit 35 to display a shooting screen SC showing the shooting results of the shooting unit 36.
[0088] [Operation Acceptance Unit] The operation acceptance unit 302 accepts various operations in the video processing system 1. Data indicating the operation content accepted by the operation acceptance unit 302 is transmitted to the server 20 as appropriate.
[0089] 9 and 10 are diagrams showing an example of processing executed in the video processing system 1. The control units 11, 21, and 31 execute programs stored in the storage units 12, 22, and 32, respectively, to execute the processing in Fig. 9 and 10. In the examples of Fig. 9 and 10, it is assumed that identity verification is performed when a user applies for a service.
[0090] 9 , the learning terminal 10 executes learning of the learning model M based on the training data stored in the training database DB1 (S1). The learning terminal 10 transmits the trained learning model M to the server 20 (S2). The server 20 receives the trained learning model M from the learning terminal 10 (S3). In S3, the server 20 records the trained learning model M in the memory unit 22.
[0091] The user terminal 30 executes processing with the server 20 for the user to sign up for the service (S4). When the user proceeds to the step requiring identity verification, the server 20 transmits request data to the user terminal 30 requesting the user to perform a predetermined action (S5). The request data indicates the contents of a message MS. The user terminal 30 receives the request data (S6). The user terminal 30 activates the photographing unit 36 and displays the photographed screen SC on the display unit 35 (S7). The user performs the action guided by the message MS.
[0092] The user terminal 30 transmits video data to the server 20 (S8). The server 20 receives the video data from the user terminal 30 (S9). The server 20 detects multiple feature points related to the user's face based on the video data (S10). Moving to FIG. 10, the server 20 performs filtering using a DoG filter (S11). The server 20 calculates the position change for each frame (S12). The server 20 may perform the processes of S11 and S12 while moving the local window W. The server 20 calculates the position change between frames based on Equations 1 to 3 (S13).
[0093] The server 20 estimates whether a two-dimensional face is shown in the video data based on the position change for each frame, the position change between frames, and the trained learning model M (S14). If it is estimated in S14 that a two-dimensional face is shown in the video data (S14: Y), the server 20 executes processing with the user terminal 30 to display an error message (S15), and the processing ends. If it is not estimated in S14 that a two-dimensional face is shown in the video data (S14: N), the server 20 executes processing with the user terminal 30 to proceed to the next step of identity verification (S16), and the processing ends.
[0094] [5. Summary of the Embodiment] The video processing system 1 of the present embodiment estimates whether a two-dimensional face is shown in video data based on the positional changes of each of a plurality of feature points in the video data and a learning model M in which training data related to the positional changes of each of a plurality of feature points related to a training face is learned. This allows the video processing system 1 to estimate whether a two-dimensional face is shown in the video data without the need for special hardware. For example, it is possible to analyze the movements shown in the video data using optical flow without the need for special hardware, but optical flow may be weak against background movements and therefore not be accurate enough. In this regard, focusing on the positional changes of the facial feature points in the foreground can improve estimation accuracy. Furthermore, while performing estimation by analyzing mouth movements, etc., requires a high-frame-rate camera unit 36 and requires the user to move slowly, the video processing system 1 can estimate whether a two-dimensional face is shown in the video data even with a camera unit 36 that can only acquire two-dimensional information. The video processing system 1 can determine whether the person in question is actually present (liveness) by estimating whether a two-dimensional face is shown in the video data.
[0095] Furthermore, the video processing system 1 performs a predetermined filter process to suppress noise related to the video data and calculate the position change. By suppressing noise, the video processing system 1 can improve estimation accuracy. For example, the detection accuracy of feature points is not perfect, and the positions of feature points may shift even if the user does not move their face. Even if such position shifts occur, they can be suppressed as noise by filter process, improving estimation accuracy.
[0096] The video processing system 1 also performs filtering using a DoG filter. The video processing system 1 can effectively suppress noise and improve estimation accuracy by using the DoG filter. The video processing system 1 can also smooth out position changes indicated by feature points.
[0097] Furthermore, the video processing system 1 sets a local window W and calculates the position change of the frame to be calculated based on the positions of each of the multiple feature points detected from the previous frame and the positions of each of the multiple feature points detected from the subsequent frame. The video processing system 1 moves the local window on the time axis of the video data to sequentially calculate the position change of each of the multiple frames to be calculated. The video processing system 1 calculates the position change of the frame to be calculated based on the multiple feature points detected from both the frames before and after the frame to be calculated, rather than just one of the frames before and after the frame, thereby improving the accuracy of the position change calculation.
[0098] Furthermore, the video processing system 1 calculates the positional change between at least one frame of the video data based on the positional change between the frames and the positional change between the frames following the frame. The video processing system 1 estimates whether a two-dimensional face is shown in the video data based on the positional change between the frames. The video processing system 1 can compensate for the positional change between frames even if the speed at which a given movement is performed varies among users, thereby improving estimation accuracy. For example, even if a user takes only one second to turn left from a frontal position, rather than three seconds, the video processing system 1 compensates for the positional change between frames, thereby improving estimation accuracy.
[0099] The video processing system 1 also calculates a position change of at least one frame of the video data and a first weighting factor w associated with the frame. i the position change of the frame following the frame, and the second weighting factor w associated with the frame following the frame. i+1 and calculate the position change between frames based on the above equation. This allows the video processing system 1 to adjust which of the two frames to place more weight on when calculating the position change between frames.
[0100] The video processing system 1 also determines whether the position change of each of multiple frames of the video data is within a reference range. If the position change is determined to be outside the reference range, the video processing system 1 calculates the position change between frames. If the position change between frames needs to be complemented, the video processing system 1 can calculate the position change between frames. If the position change between frames does not need to be complemented, the video processing system 1 can avoid calculating the position change between frames.
[0101] Furthermore, the video processing system 1 performs user identity verification based on the estimation result of whether a two-dimensional face is shown in the video data, thereby preventing a malicious user from impersonating another user during identity verification.
[0102] [6. Modifications] The present disclosure is not limited to the above-described embodiments. The present disclosure can be modified as appropriate without departing from the spirit of the present disclosure.
[0103] Fig. 11 is a diagram showing an example of functions realized in the modified example. As shown in Fig. 11, the server 20 in the modified example includes an operation request unit 205 and a registration unit 206. Each of the operation request unit 205 and the registration unit 206 is realized by the control unit 21.
[0104] [6-1. Modification 1] For example, in the embodiment, a learning model M is exemplified that, when position change data is input, outputs a label indicating whether or not a two-dimensional face is shown in video data. The learning model M may be any model that has learned the position changes of each of a plurality of feature points related to a training face. The learning model M is not limited to the example of the embodiment.
[0105] 12 is a diagram showing an example of a learning model M of Modification 1. The learning unit 101 of Modification 1 performs learning of a first learning model M1 in which first training data relating to position changes of each of a plurality of feature points on a three-dimensional training face is learned, and a second learning model M2 in which second training data relating to position changes of each of a plurality of feature points on a two-dimensional training face is learned.
[0106] For example, the first training data is position change data calculated from training video data generated by capturing a training human face with a camera. The method of acquiring the position change data may be the same as in the embodiment. The second training data is position change data calculated from training video data generated by capturing a training medium with a camera. The training medium is paper on which a training human face is printed, or a computer on which a training human face is displayed. The first training data and the second training data are stored in a training database DB1. Note that labels do not need to be assigned to the first training data and the second training data.
[0107] For example, the learning unit 101 trains the first learning model M1 to learn features indicated by the first training data. In the case where the input portion of the first training data is position change data and the output portion of the first training data is a label, as in the embodiment, the learning unit 101 adjusts the parameters of the first learning model M1 so that, when the position change data, which is the input portion of the first training data, is input to the first learning model M1, the first learning model M1 outputs a label indicating that the data is not a two-dimensional face. For example, when no label is assigned to the first training data, the learning unit 101 performs clustering of multiple pieces of first training data and adjusts the parameters of the first learning model M1 so that first training data having similar features belong to the same cluster.
[0108] For example, the learning unit 101 trains the second learning model M2 to learn features indicated by the second training data. In the case where the input portion of the second training data is position change data and the output portion of the second training data is a label, as in the embodiment, the learning unit 101 adjusts the parameters of the second learning model M2 so that, when the position change data, which is the input portion of the second training data, is input to the second learning model M2, the second learning model M2 outputs a label indicating that the data is not a two-dimensional face. For example, when no label is assigned to the second training data, the learning unit 101 performs clustering of multiple pieces of second training data and adjusts the parameters of the second learning model M2 so that second training data having similar features belong to the same cluster.
[0109] The estimation unit 203 of the first modification acquires position change data relating to characteristics of position changes in a multidimensional space. The method of acquiring the position change data may be the same as that of the embodiment. Based on the first learning model M1, the estimation unit 203 performs encoding to reduce the dimensions of the position change data and decoding to restore the dimensions of the position change data to the original, and calculates a first movement amount, which is the movement amount of the position change data before and after the encoding and decoding.
[0110] Encoding is also called dimension reduction. The encoding method may be a known method. For example, the estimation unit 203 may reduce the dimension of the position change data based on a method such as principal component analysis or autoencoding. Decoding may also be a known method. For example, the estimation unit 203 may restore the dimension-reduced position change data to its original dimension based on a method such as principal component analysis or autoencoding. The encoding and decoding are performed based on parameters of the first learning model M1.
[0111] For example, the estimation unit 203 performs encoding and decoding based on the second learning model M2 and calculates a second movement amount, which is the movement amount of the position change data before and after the encoding and decoding. The meaning of encoding and decoding is the same as that of the first learning model M1. The encoding and decoding are performed based on the parameters of the second learning model M2.
[0112] For example, the estimation unit 203 estimates whether a two-dimensional face is shown in the video data based on the first movement amount and the second movement amount. If a three-dimensional face is shown in the video data to be estimated, the first movement amount before and after encoding and decoding will be smaller than the second movement amount. This is because the characteristics of the video data are similar to the characteristics of the first training data learned by the first learning model M1. If the first movement amount is less than the second movement amount, the estimation unit 203 does not estimate that a two-dimensional face is shown in the video data.
[0113] On the other hand, if a two-dimensional face is shown in the video data to be estimated, the first movement amount before and after encoding and decoding will be equal to or greater than the second movement amount. This is because the features of the video data are similar to the features of the second training data learned by the second learning model M2. If the first movement amount is equal to or greater than the second movement amount, the estimation unit 203 estimates that a two-dimensional face is shown in the video data.
[0114] The video processing system 1 of Modification 1 calculates a first movement amount based on a first learning model M1. The video processing system 1 calculates a second movement amount based on a second learning model M2. The video processing system 1 estimates whether a two-dimensional face is shown in the video data based on the first movement amount and the second movement amount. This allows the video processing system 1 to improve the accuracy of estimating whether a two-dimensional face is shown in the video data.
[0115] [6-2. Modification 2] For example, in the embodiment, an example is given in which the user is requested to turn his / her face from the front to the left, but the action requested of the user is not limited to the example in the embodiment. The video processing system 1 of Modification 2 includes an action request unit 205. The action request unit 205 requests the user to perform an action selected from a plurality of actions related to the face using a predetermined selection method.
[0116] The candidate motions requested by the user may be any motion related to the face. For example, the candidate motions may be a motion to change the direction of the face, a motion to move a part of the face, a motion to deform a part of the face, or other motions. The candidate motions may be motions other than the motions exemplified in the embodiment. Data indicating the candidate motions is assumed to be stored in advance in the data storage unit 200.
[0117] The predetermined selection method may be any method. In Variation 2, a case will be described in which the selection method is a random selection method. The method of randomly selecting from among the candidates may be a known method. For example, the action request unit 205 randomly selects one action from among a plurality of actions based on a random number. The action request unit 205 requests the selected action from the user. In the examples of Figures 2 and 3, the user is requested to perform an action by displaying a message MS, but the user may also be requested to perform an action by other methods such as outputting a voice.
[0118] The selection method is not limited to a random selection method. For example, the selection method may be a method of selecting each of a plurality of actions in a predetermined order. The selection method may be a method of selecting an action according to data acquired from the user terminal 30 or the characteristics of the user. The video data acquisition unit 201 acquires video data when the action selected by the action selection unit is performed by the user. Although the present embodiment differs from the embodiment in that a video is selected by the action selection unit, the method of acquiring the video data is the same as the embodiment.
[0119] The data storage unit 200 of the second modification is assumed to have a learning model M prepared for each action. The training database DB1 stored in the learning terminal 10 stores training data indicating each of a plurality of actions. The learning unit 101 executes learning of the learning model M corresponding to a certain action based on the training data indicating the action. Although this differs from the embodiment in that training data and a learning model M are prepared for each action, the method of creating the learning model M itself may be the same as the embodiment.
[0120] For example, assume that four candidate actions are a first action of turning left from the front, a second action of turning right from the front, a third action of turning up from the front, and a fourth action of blinking. In this case, the learning unit 101 executes learning of four learning models M, namely, a learning model M for the first action, a learning model M for the second action, a learning model M for the third action, and a learning model M for the fourth action. The learning terminal 10 transmits the four learned learning models M to the server 20. The server 20 records the four learned learning models M in the data storage unit 200.
[0121] The estimation unit 203 of the modification 2 estimates whether a two-dimensional face is shown in the video data based on the learning model M corresponding to the action selected by the action selection unit among the learning models M corresponding to each of the plurality of actions. Although the estimation differs from the embodiment in that the learning model M corresponding to the action selected by the action selection unit among the plurality of learning models M stored in the data storage unit 200 is used for estimation, the estimation itself based on the learning model M may be the same as the embodiment.
[0122] For example, assuming that four actions, such as the first action to the fourth action described above, are candidates, when the user performs the first action, the estimation unit 203 performs estimation based on the learning model M corresponding to the first action out of the four learning models M. When the user performs the second action, the estimation unit 203 performs estimation based on the learning model M corresponding to the second action out of the four learning models M. When the user performs the third action, the estimation unit 203 performs estimation based on the learning model M corresponding to the third action out of the four learning models M. When the user performs the fourth action, the estimation unit 203 performs estimation based on the learning model M corresponding to the fourth action out of the four learning models M.
[0123] The video processing system 1 of the second modification example requests a user to perform an action selected from a plurality of actions using a predetermined selection method. The video processing system 1 acquires video data when the selected action is performed by the user. The video processing system 1 estimates whether a two-dimensional face is shown in the video data based on the learning model M corresponding to the selected action among the plurality of actions. This allows the video processing system 1 to perform estimation using the learning model M appropriate for the action performed by the user. A malicious user would need to imagine various actions to circumvent identity verification, making it more difficult for them to successfully impersonate someone.
[0124] [6-3. Modification 3] For example, in the embodiment, the case where the estimation result by the estimation unit 203 is used for identity verification has been exemplified. The estimation result by the estimation unit 203 can be used for any purpose other than identity verification. In Modification 3, the estimation result by the estimation unit 203 is used when a user registers image data for face authentication. A malicious user may register a face photo of another person as image data for face authentication and then continue to use the face photo to impersonate that person. The video processing system 1 can also be used to prevent such impersonation.
[0125] In Modification 3, similar to Modification 2, videos are randomly requested. The video processing system 1 of Modification 3 includes a registration unit 206. Based on an estimation result of whether or not a two-dimensional face is shown in the video data, the registration unit 206 registers, from among multiple frames of the video data, a frame showing the user's face facing forward as image data for user face authentication. The registration unit 206 may determine whether or not the user's face is facing forward by image analysis (e.g., template matching, a method of determining based on the shape of the contour, or a machine learning method), or may determine a specific frame (e.g., the first frame) as the frame showing the user's face facing forward.
[0126] For example, if it is estimated that a two-dimensional face is shown in the video data, the registration unit 206 does not register image data for facial authentication of the user. In this case, a malicious user is attempting to register image data for impersonation using a facial photograph of another person, so the registration unit 206 does not register the image data. If it is not estimated that a two-dimensional face is shown in the video data, the registration unit 206 registers image data for facial authentication of the user. The image data for facial authentication may be registered in the user database DB2 or in another database.
[0127] For example, the server 20 or another computer performs face authentication based on the image data registered by the registration unit 206. The authentication terminal for face authentication includes a camera. The authentication terminal may be the user terminal 30 or a computer other than the user terminal 30. When the authentication terminal captures a picture of the user's face with the camera, it transmits the image data to the server 20 or another computer. The server 20 or another computer performs face authentication based on the image data received from the authentication terminal and the image data registered by the registration unit 206. The face authentication may be performed by a known method.
[0128] The video processing system 1 of the third modification registers, from among multiple frames of the video data, frames showing the user's face facing forward as image data for face authentication of the user, based on the estimation result of whether or not a two-dimensional face is shown in the video data. The video processing system 1 can prevent a malicious user from registering image data while impersonating another person and performing face authentication.
[0129] [6-4. Other Modifications] For example, the above modifications 1 to 3 may be combined.
[0130] For example, a function described as being realized by the server 20 may be realized by another computer such as the learning terminal 10. A function described as being realized by the server 20 may be shared among multiple computers. A function described as being realized by the learning terminal 10 may be realized by another computer such as the server 20.
[0131] [7. Supplementary Note] For example, the video processing system may also be configured as follows: (1) A video processing system including: a video data acquisition unit that acquires video data related to a video showing the face of a user performing a predetermined action; a feature point detection unit that detects a plurality of feature points related to the face based on the video data; and an estimation unit that estimates whether a two-dimensional face is shown in the video data based on positional changes of each of the plurality of feature points in the video data and a learning model in which training data related to positional changes of each of the plurality of feature points related to a training face is learned. (2) The video processing system described in (1), in which the estimation unit calculates the positional changes by performing a predetermined filter process to suppress noise related to the video data. (3) The video processing system described in (2), in which the estimation unit performs the filter process using a DoG (Derivative of Gaussian) filter. (4) The video processing system according to any one of (1) to (3), wherein the estimation unit sets a local window including a calculation target frame that is a frame to be calculated for the position change, a previous frame that is a frame that precedes the calculation target frame, and a subsequent frame that is a frame that follows the calculation target frame, calculates the position change of the calculation target frame based on the positions of each of the plurality of feature points detected from the previous frame and the positions of each of the plurality of feature points detected from the subsequent frame, and moves the local window on the time axis of the video data to successively calculate the position change of each of the plurality of calculation target frames. (5) The video processing system according to any one of (1) to (4), wherein the estimation unit calculates the position change between at least one frame of the video data based on the position change of the frame and the position change of a frame following the frame, and estimates whether the two-dimensional face is shown in the video data based further on the position change between the frames.(6) The moving image processing system according to (5), wherein the estimation unit calculates the position change between frames based on the position change of at least one frame of the moving image data, a first weighting factor associated with the frame, the position change of a frame subsequent to the frame, and a second weighting factor associated with the frame subsequent to the frame. (7) The moving image processing system according to (5) or (6), wherein the estimation unit determines whether the position change of each of a plurality of frames of the moving image data is within a reference range, and calculates the position change between the frames when it is determined that the position change is outside the reference range. (8) The video processing system according to any one of (1) to (7), wherein the estimation unit acquires position change data relating to the characteristics of the position change in multidimensional space, performs encoding to reduce the dimension of the position change data and decoding to restore the dimension of the position change data based on a first learning model in which first training data relating to the position change of each of a plurality of feature points on a three-dimensional training face has been learned, and calculates a first movement amount which is the movement amount of the position change data before and after the encoding and decoding, performs the encoding and the decoding based on a second learning model in which second training data relating to the position change of each of a plurality of feature points on a two-dimensional training face has been learned, and calculates a second movement amount which is the movement amount of the position change data before and after the encoding and decoding, and estimates whether the two-dimensional face is shown in the video data based on the first movement amount and the second movement amount. (9) The video processing system according to any one of (1) to (8), further including an action request unit that requests the user to perform an action selected from a plurality of actions related to the face by a predetermined selection method, the video data acquisition unit that acquires the video data when the selected action is performed by the user, and the estimation unit that estimates whether the two-dimensional face is shown in the video data based on the learning model corresponding to the selected action among the plurality of actions.(10) The video processing system according to any one of (1) to (9), further including an identity verification unit that performs identity verification of the user based on an estimation result of whether the two-dimensional face is shown in the video data. (11) The video processing system according to any one of (1) to (10), further including a registration unit that, based on an estimation result of whether the two-dimensional face is shown in the video data, registers, among a plurality of frames of the video data, a frame that shows a face of the user facing forward as image data for facial authentication of the user.
Claims
1. A video processing system comprising: a video data acquisition unit that acquires video data regarding a video showing a face of a user performing a predetermined operation; a feature point detection unit that detects a plurality of feature points regarding the face based on the video data; and an estimation unit that estimates whether a two-dimensional face is shown in the video data based on a change in position of each of the plurality of feature points in the video data and a learning model in which training data regarding a change in position of each of the plurality of feature points regarding a training face has been learned.
2. The video processing system according to claim 1, wherein the estimation unit calculates the change in position by performing predetermined filter processing to suppress noise regarding the video data.
3. The video processing system according to claim 2, wherein the estimation unit performs the filter processing using a DoG (Derivative of Gaussian) filter.
4. The estimation unit of the video processing system according to any one of claims 1 to 3 sets a local window including a calculation target frame that is a frame for which the change in position is to be calculated, a previous frame that is a frame before the calculation target frame, and a subsequent frame that is a frame after the calculation target frame, calculates the change in position of the calculation target frame based on the position of each of the plurality of feature points detected from the previous frame and the position of each of the plurality of feature points detected from the subsequent frame, and moves the local window on the time axis of the video data to sequentially calculate the change in position of each of the plurality of calculation target frames.
5. The estimation unit of the video processing system according to any one of claims 1 to 3 calculates the change in position between frames based on the change in position of at least one frame of the video data and the change in position of the frame next to the frame, and estimates whether the two-dimensional face is shown in the video data based on the change in position between the frames.
6. The estimation unit of the video processing system according to claim 5 calculates the change in position between the frames based on the change in position of at least one frame of the video data, a first weight coefficient associated with the frame, the change in position of the frame next to the frame, and a second weight coefficient associated with the next frame.
7. The estimation unit determines whether or not the position change of each of a plurality of frames of the video data is within a reference range, and calculates the position change between the frames when it is determined that the position change is outside the reference range. The video processing system according to claim 5.
8. The estimation unit acquires position change data regarding characteristics of the position change in a multi-dimensional space, and performs encoding for reducing the dimension of the position change data and decoding for restoring the dimension of the position change data based on a first learning model in which first training data regarding the position change of each of a plurality of feature points regarding a three-dimensional training face has been learned, and calculates a first movement amount which is the movement amount of the position change data before and after the encoding and the decoding. The estimation unit performs the encoding and the decoding based on a second learning model in which second training data regarding the position change of each of a plurality of feature points regarding a two-dimensional training face has been learned, and calculates a second movement amount which is the movement amount of the position change data before and after the encoding and the decoding, and estimates whether or not the two-dimensional face is shown in the video data based on the first movement amount and the second movement amount. The video processing system according to any one of claims 1 to 3.
9. The video processing system further includes an operation request unit that requests the user to perform an operation selected by a predetermined selection method from among a plurality of operations regarding the face. The video data acquisition unit acquires the video data when the selected operation is performed by the user. The estimation unit estimates whether or not the two-dimensional face is shown in the video data based on the learning model corresponding to the selected operation among the learning models corresponding to each of the plurality of operations. The video processing system according to any one of claims 1 to 3.
10. The video processing system further includes an identity verification unit that performs identity verification of the user based on an estimation result as to whether or not the two-dimensional face is shown in the video data. The video processing system according to any one of claims 1 to 3.
11. The video processing system according to any one of claims 1 to 3, further comprising a registration unit that registers, as image data for face authentication of the user, a frame showing the face of the user facing forward among a plurality of frames of the video data, based on an estimation result as to whether or not the two-dimensional face is shown in the video data.
12. A video processing method including: a video data acquisition step of acquiring video data related to a video showing the face of a user performing a predetermined operation; a feature point detection step of detecting a plurality of feature points related to the face based on the video data; and an estimation step of estimating whether or not a two-dimensional face is shown in the video data, based on a change in the position of each of the plurality of feature points in the video data and a learning model in which training data related to a change in the position of each of the plurality of feature points related to a training face has been learned.
13. A program for causing a computer to function as: a video data acquisition unit that acquires video data related to a video showing the face of a user performing a predetermined operation; a feature point detection unit that detects a plurality of feature points related to the face based on the video data; and an estimation unit that estimates whether or not a two-dimensional face is shown in the video data, based on a change in the position of each of the plurality of feature points in the video data and a learning model in which training data related to a change in the position of each of the plurality of feature points related to a training face has been learned.
Citation Information
Patent Citations
Solidity authenticating method, solidity authenticating apparatus, and solidity authenticating program
JP2012069133A
Method for distinguishing three-dimensional real objects from two-dimensional spoofs of real objects
JP2021522591A
Information processing device, information processing method, and recording medium
JP2023063314A
User image verification
US20190347388A1
Methods and systems for facial recognition using motion vector trained model
US20210182539A1