Image processing systems, image processing methods and software products
Patent Information
- Application Number
- TW113149104
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-12-25
- Filing Date
- 2024-12-17
- Publication Date
- 2026-08-11
- Estimated Expiration
- 2044-12-16
AI Technical Summary
Existing image processing systems require special hardware like 3D cameras or thermal imagers to determine if a 2D face is displayed in image data, which is not commonly available in general user terminals.
An image processing system that uses a learning model to estimate whether a 2D face is displayed in image data based on the positional changes of feature points detected in the image data, without requiring special hardware.
Enables accurate determination of a 2D face in image data using standard imaging devices, enhancing security by preventing impersonation without the need for specialized hardware.
Smart Images

Figure TWG2TB001905511_001 
Figure TWG2TB001905511_002 
Figure TWG2TB001905511_003
Abstract
Description
Technical Field
[0001] This disclosure relates to an image processing system, image processing method, and program product. Prior Technology
[0002] Previously, the industry has explored technologies to prevent impersonation by malicious actors. For example, Patent Document 1 describes a facial authentication technique using a 3D camera capable of detecting depth and a thermal imager capable of detecting temperature. It is believed that even if malicious actors use photographs or images displaying other people's faces to impersonate others, the 3D camera and thermal imager in Patent Document 1 can prevent such impersonation. [Previous Technical Documents] [Patent Literature]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-029968 Summary of the Invention
[0004] [The problem the invention aims to solve] However, in previous technologies, in order to determine whether a 2D face is displayed in image data, special hardware not typically found in computers is required, such as the 3D camera and thermal imager in Patent Document 1. Therefore, there is a need to determine whether a 2D face is displayed in image data without requiring special hardware.
[0005] One of the purposes of this disclosure is to presume whether a 2D face is displayed in image data without requiring special hardware. [Technical means to solve the problem]
[0006] The image processing system disclosed herein includes: an image data acquisition unit that acquires image data related to the display of an image of a user's face; a feature point detection unit that detects a plurality of feature points related to the aforementioned face based on the aforementioned image data; and an estimation unit that estimates whether a 2D face is displayed in the aforementioned image data based on the positional changes of each of the aforementioned plurality of feature points in the aforementioned image data and a learning model that has been learned in relation to the positional changes of each of the plurality of feature points of the aforementioned face used for training. [Effects of the Invention]
[0007] This disclosure, for example, can presume whether a 2D face is displayed in image data without the need for special hardware. Simple Explanation of the Diagram
[0008] Figure 1 shows an example of the hardware configuration of an image processing system. Figure 2 shows an example of a screen displayed on a user terminal. Figure 3 shows an example of a malicious user impersonating another user. Figure 4 shows one example of the functions implemented by the image processing system. Figure 5 shows an example of a training database. Figure 6 shows an example of a user database. Figure 7 shows an example of the processing of individual frames of image data. Figure 8 shows an example of the overall processing of image data. Figure 9 shows an example of processing performed by an image processing system. Figure 10 shows an example of processing performed by an image processing system. Figure 11 shows one example of the function implemented by the variation example. Figure 12 shows one example of the learning model of variation example 1. Implementation
[0009] [1. Hardware Components of an Image Processing System] This illustration shows an example of an implementation of the image processing system, image processing method, and program disclosed herein. Figure 1 is a diagram showing an example of the hardware configuration of the image processing system. For example, the image processing system 1 includes a learning terminal 10, a server 20, and a user terminal 30. The learning terminal 10, server 20, and user terminal 30 are each connected to a network N such as the Internet or a LAN.
[0010] Learning terminal 10 is a computer that performs learning using the learning model described later. For example, learning terminal 10 is a personal computer, tablet, or smartphone. For example, learning terminal 10 includes a control unit 11, a memory unit 12, a communication unit 13, an operation unit 14, and a display unit 15. For example, control unit 11 includes at least one processor. Memory unit 12 includes at least one of volatile memory such as RAM and non-volatile memory such as flash memory. Communication unit 13 includes at least one of a wired communication interface and a wireless communication interface. Operation unit 14 is an input device such as a touch panel or mouse. Display unit 15 is a display such as a liquid crystal display (LCD) or an organic EL display.
[0011] Server 20 is a server computer that stores a pre-trained learning model. For example, server 20 includes a control unit 21, a memory unit 22, and a communication unit 23. The hardware configuration of the control unit 21, memory unit 22, and communication unit 23 can be the same as that of the control unit 11, memory unit 12, and communication unit 13, respectively.
[0012] User terminal 30 is a user's computer. For example, user terminal 30 is a personal computer, smartphone, tablet, or wearable terminal. For example, user terminal 30 includes a control unit 31, a memory unit 32, a communication unit 33, an operation unit 34, a display unit 35, and a camera unit 36. The hardware configuration of control unit 31, memory unit 32, communication unit 33, operation unit 34, and display unit 35 can be the same as that of control unit 11, memory unit 12, communication unit 13, operation unit 14, and display unit 15, respectively. Camera unit 36 includes at least one camera.
[0013] Furthermore, the programs stored in memory units 12, 22, and 32 can be supplied to the learning terminal 10, server 20, or user terminal 30 via network N. Also, at least one of a reading unit (e.g., a memory card slot) for reading computer-readable information storage media and an input / output unit (e.g., a USB port) for inputting and outputting data to external machines can be included in the learning terminal 10, server 20, or user terminal 30. For example, the programs stored in information storage media can be supplied to the learning terminal 10, server 20, or user terminal 30 via at least one of the reading unit and the input / output unit.
[0014] Furthermore, the image processing system 1 only needs to include at least one computer. The computer included in the image processing system 1 is not limited to the example in Figure 1. For example, the image processing system 1 can consist of a learning terminal 10 and a server 20. In this case, the user terminal 30 is located outside the image processing system 1. The image processing system 1 can also consist of a server 20. In this case, the learning terminal 10 and the user terminal 30 are located outside the image processing system 1. The image processing system 1 can also consist of a server 20 and a computer not shown in Figure 1.
[0015] [2. Overview of Image Processing Systems] In this embodiment, an example is given where the camera unit 36 captures the user's face for user identity verification. Identity verification is a process used to confirm that the user is indeed the person they are looking for. For example, eKYC (electronic Know Your Customer) is a type of identity verification. Identity verification can be performed in any scenario of any service. For example, users may verify their identity when using payment services, financial services, communication services, e-commerce services, administrative services, or other services, or when logging in as a member.
[0016] For example, when a scenario requiring user confirmation occurs, user terminal 30 activates camera unit 36. Camera unit 36 continuously captures images of the user's face in image mode. Based on the images captured by camera unit 36, user terminal 30 generates image data displaying the user's face. User terminal 30 sends the image data to server 20. Server 20 performs processing for user confirmation based on the image data.
[0017] Figure 2 illustrates an example of a screen displayed on a user terminal 30. For instance, the user terminal 30 displays the captured image SC (screenshot) from the camera unit 36 on the display unit 35. The captured image SC displays a guide G for guiding the position of the face and a message MS for the user's requested action. In the example of Figure 2, the user is asked to turn their face from front to left. The user adjusts the position and orientation of the user terminal 30 by placing their face on the guide G and performs the action of displaying the message MS.
[0018] Furthermore, the actions requested from the user can be arbitrary. The actions requested from the user are not limited to the example in Figure 2. For example, the user may be asked to turn their face upwards, turn their face to the right, turn their face downwards, blink, glare, or perform other actions. As in the variations described later, the action requested from the user can be randomly determined from a plurality of actions. These actions are performed to prevent impersonation by malicious users.
[0019] For example, a malicious user could impersonate another person by taking a picture of a piece of paper with another person's face printed on it using their own user terminal 30. A malicious user could also impersonate another person by taking a picture of the first user terminal 30 displaying another person's face using a second user terminal 30. Even when asked to perform the actions shown in Figure 2, a malicious user could cleverly bypass the user's identity verification.
[0020] 3 is a diagram showing an example of a modality in which a user with malicious intent performs impersonation. For example, a user with malicious intent holds the paper P printed with the face of another person before his or her own face in such a way that the face of the printed face of another person is directed toward the square of the photographic department 36 . The malicious user causes the user terminal 30 of itself to take the photographic part 36 of the paper P . In situations where a malicious user is asked to move his face to the left from the front, he sometimes intends to resort to turning the paper P from the front to the left, and through his own confirmation.
[0021] For example, the same is true in the case of other action being requested, it is possible for a user with malice to intend to confirm by himself by recourse to the action of making the paper P face-up etc. It is considered that this impersonation can be prevented as long as the person acting as the person acting as a check confirmed by me checks the image data uploaded from the user terminal 30 . However, in cases where personal confirmations from a number of users were performed, the person in charge realistically did not examine all the image data. For this purpose, the image processing system 1 resorts to parsing the image data and determines whether a 2-dimensional face is displayed in the image data.
[0022] For example, it is considered that if the photography unit 36 is a three-dimensional camera, a stereo camera, or a thermal imager, the image processing system 1 can determine whether a 2-dimensional face is displayed in the image data by detecting the surface shape or surface temperature of the subject. However, in that case, special hardware such as a three-dimensional camera, a stereo camera, or a thermal imager must be available. General user terminals such as smartphones 30 do not contain such special hardware.
[0023] For this purpose, the image processing system 1 of the present embodiment has the configuration of whether or not a 2-dimensional face is presumed to be displayed in the image data by resorting to the resolution of the image data, without having to have special hardware. For example, the position changes of the feature points detected in each frame of the self-imaging data are different in the case of a 3-dimensional face in the image data, and in the case of a 2-dimensional face in the image data. The image processing system 1 determines whether a 2-dimensional face is displayed in the image data based on the learning model that has learned those features. Hereafter, the details of the image processing system 1 are explained.
[0024] [3.Functions implemented by image processing systems] 4 is a diagram showing an example of a function implemented by the image processing system 1 .
[0025] [3-1.Functions Implemented by Learning Terminal] For example, the learning terminal 10 contains a data memory section 100 and a learning section 101 . The data memory section 100 is implemented by the memory section 12 . The learning part 101 is implemented by the control part 11 .
[0026] [Data Memory Department] The data required for the learning of the data memory unit 100 and the memory learning model M. For example, the actual data of the data memory unit 100 and the training database DB1. The learning model M is a model created using machine learning methods. Machine learning can be defined in various well-known ways. In this embodiment, machine learning includes deep learning, reinforcement learning, and AI (Artificial Intelligence) in a broad sense. For example, the learning model M can be any model of supervised learning, semi-supervised learning, or unsupervised learning. It is assumed that the learning model M in this embodiment is a supervised learning model.
[0027] The learning model M can utilize various models used in the field of image analysis. For example, the learning model M can be a neural network, a visual converter, a support vector machine, a random forest, or other methods. The actual data of the learning model M includes: the programs for various processing such as the embedded computation, and the parameters referenced by those programs. For example, the parameters are weighting coefficients and biases. The programs and parameters contained in the actual data of the learning model M can be those used in well-known methods.
[0028] Figure 5 shows an example of a training database DB1. The training database DB1 stores the training data used by the learning model M. For example, the training data includes: the input portion that is fed into the learning model M during learning, and the output portion that displays what should be output from the learning model M during learning. The output portion of the training data can also be considered the correct solution during learning. In the case where the learning model M is an unsupervised learning model, no annotations are made to the training data; therefore, the training data does not include an output portion.
[0029] For example, the input portion of the training data displays positional change data of a plurality of feature points associated with the face used for training. The face used for training is an image of a face prepared for learning model M. Feature points are characteristic parts of the face. For example, a feature point is at least one pixel displaying a part of the face. Feature points are also sometimes referred to as landmarks of the face. In this embodiment, an example is given where the vertices of the mesh constituting the surface shape of the face are equivalent to feature points.
[0030] Furthermore, the methods for feature point detection can be well-known. For example, feature points can be detected using the Dlib library, which is available in Python, a programming language. Since numerous libraries for feature point detection are also available in other programming languages, feature point detection can be performed using those libraries. Alternatively, machine learning methods such as deep learning (e.g., methods known as deep 3D face reconstruction or ring network methods) can be used instead of programming language libraries to detect feature points.
[0031] For example, position change data, equivalent to the input portion of the training data, is displayed as the position changes of multiple feature points detected in each of multiple frames of image data related to the image of the face used for training. Position change can also be referred to as velocity or amount of movement. In this embodiment, position change data, equivalent to the input portion of the training data, is generated using the same calculation method as the estimation unit 203 described later. Details of this calculation method will be described later.
[0032] In this embodiment, an example is given where the position change data is in vector form. For instance, if the number of frames in the image data is set to m (m is a natural number), and k (k is a natural number) feature points are detected from one frame, the position change data is a vector of at least m×k dimensions. The vector of position change data displays the position change of a specific feature point corresponding to a specific frame.
[0033] For example, the vector displaying positional change data shows the positional changes of each of the k feature points detected from the first frame in dimensions 1 through k. For instance, if dimension 1 represents the positional changes of feature points in a specific area of the forehead, dimension 2 represents the positional changes of feature points to its right, and dimension 3 represents the positional changes of feature points further to its right, the positional changes of which feature point are displayed in each dimension are predetermined. Dimensions (k+1) onwards display positional changes from the first frame onwards. The positional changes are arranged in a prescribed order up to the last dimension of the vector displaying positional change data.
[0034] Furthermore, the positional change data can be in any form. Positional change data only needs to be data related to the temporal positional changes of multiple feature points in the image data. Positional change data is not limited to vector form. For example, positional change data can be multiple numerical values, single numerical values, arranged forms, matrix forms, or other forms. These forms of data can be used to display the temporal positional changes of multiple feature points in the image data.
[0035] For example, the output of the training data may display a label indicating whether the input of the training data displays a 2D face. A 2D face indicates a non-3D face. For instance, in the case of displaying a piece of paper with a photograph of another person's face printed on it in the training image data, or a computer displaying another person's face, the input of the training data will display a 2D face. Because the surface of the paper or the surface of the screen in the computer has a 2D plane, the positional change data obtained from displaying such training image data (the input of the training data) will display a 2D face.
[0036] The labels categorize the input portion of the training data. A label of the first value (e.g., 1) indicates that the input portion of the training data is a 2D face. A label of the second value (e.g., 0) indicates that the input portion of the training data is not a 2D face. The values of the output portion of the training data can be specified by the creator of the learning model M, or determined using well-known annotation tools.
[0037] For example, positional change data obtained from training image data displaying a 3D face has different characteristics from positional change data obtained from training image data displaying a 2D face. It is believed that these differences in characteristics are related to whether a 2D face is displayed in the training image data. This correlation can be learned by the training data to enable the learning model M to learn.
[0038] As shown in Figure 2, taking the action of turning the face from front to left as an example, the positional change of feature points near the left eye or left cheek in a 2D face is greater than that in a 3D face. This is because there is no depth direction (the vertical direction of the paper P or the image) in a 2D face printed on paper P or displayed on a screen. This positional change is related to whether the 2D face is displayed in the training image data (i.e., the label value). The learning model M can learn this positional change through training data. Not only the positional change of the left eye or left cheek in the direction the face is facing, but also the positional changes of other parts such as the right eye or right cheek are related to whether the 2D face is displayed in the training image data. The learning model M can learn this relationship through training data.
[0039] For example, for other actions such as turning the face upwards from the front, there is a correlation between the positional changes of feature points and whether a 2D face is displayed in the training image data. Therefore, the same applies to other actions. By training the model M to learn the relationship between positional changes obtained from the training image data, the positional changes of feature points, and whether a 2D face is displayed in the training image data, the learning model M can learn these relationships.
[0040] Furthermore, the training data is not limited to the example in Figure 2. Training data only needs to display at least one of the features of 3D facial feature points and the features of 2D facial feature points. For example, the input to the training data can be the image data itself associated with the image of the face used for training. The input to the training data can display the positions of the feature points, rather than the positional changes of multiple feature points detected from multiple frames in the image data. The output of the training data can be a score indicating suspicion of impersonation (e.g., a median value, rather than a binary value like the label).
[0041] Furthermore, the data stored in the data memory unit 100 is not limited to the examples mentioned above. The data memory unit 100 can store any data. For example, the data memory unit 100 can store a learning program that displays a series of processes during learning. The learning program can be a program used by well-known machine learning methods. For example, the learning program can be a program created based on well-known methods such as backpropagation or gradient descent.
[0042] [Study Department] Learning unit 101 performs learning of learning model M based on training data stored in training database DB1. For example, when learning unit 101 inputs the input portion of training data into learning model M based on a known learning program, it adjusts the parameters of learning model M by outputting the output portion of training data, thereby performing learning of learning model M. When learning model M completes learning, learning unit 101 records the pre-trained learning model M in data memory unit 100. The learning model M before learning can be stored in data memory unit 100 as data different from the pre-trained learning model M, or it can be overwritten by the pre-trained learning model M. Learning unit 101 sends the pre-trained learning model M to server 20. The pre-trained learning model M sent to server 20 is available for user use.
[0043] [3-2. Functionality Implemented by the Server] For example, server 20 includes: a data memory unit 200, an image data acquisition unit 201, a feature point detection unit 202, an estimation unit 203, and a person verification unit 204. The data memory unit 200 is implemented by the memory unit 22. The image data acquisition unit 201, the feature point detection unit 202, the estimation unit 203, and the person verification unit 204 are each implemented by the control unit 21.
[0044] [Data Memory Department] The data memory unit 200 stores the data required for the estimation of the learning model M. For example, the data memory unit 200 stores the pre-trained learning model M and the user database DB2. In this embodiment, since the learning model M is learned by the learning terminal 10, the data memory unit 200 stores the learning model M sent from the learning terminal 10. When the learning model M is learned by the server 20, the server 20 has the same function as the learning unit 101.
[0045] Figure 6 shows an example of a user database DB2. The user database DB2 stores various data related to the user. For example, the user database DB2 stores the user ID, image data, and user verification information. Any data related to the user can be stored in the user database DB2. The data stored in the user database DB2 is not limited to the example in Figure 6. For example, when a user confirms their identity during membership login for certain services, information displaying the service application status can be stored in the user database DB2.
[0046] A user ID is one example of user identification information that can identify a user. For example, an email address or phone number can be used as user identification information. Image data is data related to an image showing the user, represented by the user ID, performing actions such as turning their face from front to left. Image data can be in any data format (e.g., MP4 or AVI). The user terminal 30, which is the object of the user's confirmation, sends the image data generated based on the photography results of the camera unit 36 to the server 20. The server 20 establishes an association with the user's user ID and stores the image data received from the user terminal 30. User confirmation data is data that displays the result of user confirmation. For example, user confirmation data can display the estimation result of the estimation unit 203.
[0047] Furthermore, the data stored in the data memory unit 200 is not limited to the examples mentioned above. The data memory unit 200 can store any data. For example, the data memory unit 200 can store personal verification procedures that include processes other than the presumption unit 203 in the series of processes that display personal verification. The processing of the personal verification unit 204, described later, is displayed in the personal verification procedure. In the case of checking personal verification documents together with the user's face check, the data memory unit 200 can store and display the document data of the personal verification documents uploaded by the user.
[0048] [Image Data Acquisition Department] The image data acquisition unit 201 acquires image data related to the image of a user's face performing a prescribed action. The prescribed action is an action using the user's face. In other words, the prescribed action is an action involving a change in the position of at least one feature point. Further, the prescribed action is an action where the tendency of the change in the position of the feature point differs depending on whether the action data displays a 3D face or the image data displays a 2D face. Examples include actions that change the direction of the face, actions that change the position of the face, actions that use parts of the face (e.g., blinking or squinting), actions that change facial expressions, or other actions equivalent to the prescribed action.
[0049] In this embodiment, since image data is stored in the user database DB2, the image data acquisition unit 201 acquires the image data from the user database DB2. If image data is stored in a database other than the user database DB2, the image data acquisition unit 201 only needs to acquire the image data from that database. If image data is stored in a computer other than the server 20 (e.g., the user terminal 30) or an external information memory medium, the image data acquisition unit 201 only needs to acquire the image data from that computer or the external information memory medium.
[0050] In this embodiment, image data is not generated using special hardware such as a 3D camera or thermal imager. Therefore, the image data acquisition unit 201 acquires image data generated based on the photographic results of a general imaging unit 36 (e.g., an RGB camera) that does not detect depth and temperature information. For example, the general imaging unit 36 does not include a depth sensor or a temperature sensor. The image data only contains 2D information (planar information) of the subject, and does not include depth and temperature information of the subject. The image data displays the color of each pixel in each frame. The image data is information for each pixel and does not include depth and temperature information.
[0051] [Feature Point Detection Department] The feature point detection unit 202 detects a plurality of feature points related to the face based on image data. Feature point detection refers to obtaining the position of feature points from the frame of the image data. In this embodiment, the position of feature points is expressed using 2D information as an example. However, there are also techniques for inferring 3D information from 2D images. Therefore, by using this technique, the position of feature points can be expressed using 3D information.
[0052] In this embodiment, it is assumed that the detection program for feature point detection is stored in the data memory unit 200. As mentioned above, the feature point detection method can be a well-known method. The processing of the well-known detection method is displayed in the detection program. The detection program can be a library such as the aforementioned Dlib library, or it can be a program using machine learning methods such as deep learning (e.g., methods known as deep 3D face reconstruction or ring network methods).
[0053] Figure 7 illustrates an example of processing individual frames of image data. The upper frame xi (i is an integer) in Figure 7 shows the user's face when viewed from the front to the left. The feature point detection unit 202 detects a plurality of feature points from each frame xi of the image data based on a detection program. For example, the feature point detection unit 202 detects the grid of the face shown in frame xi based on the detection program, and detects the vertices constituting the grid as feature points. In the central frame xi of Figure 7, feature points are represented by dots. The feature point detection unit 202 records the feature point data of the plurality of feature points detected from each frame xi in the data memory unit 200.
[0054] For example, the feature point data can display the feature point ID and the location of the feature point. The feature point ID is an identifier that can recognize the feature point. The feature point ID can be used to identify the facial portion represented by the feature point. For example, based on the feature point ID, it is possible to identify a feature point representing the tip of the nose or a feature point representing the edge of the right eye. The feature point identification method can be any method known for feature point detection. For example, it is possible to display only the location of the feature point in the feature point data without utilizing the feature point ID.
[0055] The position of a feature point is set on the coordinate axis of the frame xi. For example, the position of a feature point is represented by the coordinate axis with the upper left of frame xi as the origin. If the position associated with a certain feature point ID is tracked, the position change of the feature point represented by that feature point ID can be tracked. The frame xi at the bottom of Figure 7 represents the position change of the feature point. The method for calculating the position change will be described later.
[0056] Figure 8 illustrates an example of overall image data processing. In the example of Figure 8, the feature point detection unit 202 performs the processing described with reference to Figure 7 on each of the frames x1 to xm (where m is an integer greater than or equal to 21 in Figure 8), detecting a plurality of feature points. The feature point detection unit 202 acquires feature point data for each of the frames x1 to xm. This feature point data can refer to the time sequence history of the feature points. In Figure 8, the symbols h1 to hm represent the positional changes corresponding to each of the frames x1 to xm. The symbols y1 to yn are information calculated for interpolation of positional changes. Details of the interpolation method will be described later.
[0057] Furthermore, the feature point detection unit 202 can detect feature points from all frames xi of the image data, or it can detect feature points from only a portion of the frames xi. Also, the number of frames xi constituting the image data can be arbitrary. Furthermore, the number of frames xi that are the objects of feature point detection can also be arbitrary. For example, before frame x1, there may be an earlier frame x0, etc. Similarly, after frame xm, there may be a later frame xm+1, etc. In the following description, the frame symbols will be omitted unless reference to Figures 7 and 8 is required.
[0058] [Presumed Section] The estimation unit 203 estimates whether a 2D face is displayed in the image data based on the positional changes of a plurality of feature points in the image data and a learning model M that has been learned and correlated with the positional changes of a plurality of feature points of a face used for training. The estimation unit 203 calculates the positional changes of the plurality of feature points detected by the feature point detection unit 202, obtaining positional change data. The estimation unit 203 inputs the obtained positional change data into the learning model M. The learning model M calculates the embedding of the positional change data based on parameters adjusted through learning. The learning model M outputs a label corresponding to the calculated embedding. The estimation unit 203 estimates whether a 2D face is displayed in the image data based on the labels obtained from the output of the learning model M.
[0059] For example, the estimation unit 203 becomes each of the frames of the calculation object that are being calculated in terms of position change, and calculates the position change of the calculation object frame based on at least two of the following: a plurality of feature points detected from the calculation object frame, a plurality of feature points detected from the frame preceding the calculation object frame, and a plurality of feature points detected from the frame following the calculation object frame. The estimation unit 203 only needs to use at least one preceding frame and at least one following frame in the calculation of the position change of the calculation object frame. The estimation unit 203 may use one preceding frame and one following frame in the calculation of the position change of the calculation object frame, or it may use a plurality of preceding frames and a plurality of following frames in the calculation of the position change of the calculation object frame.
[0060] In this embodiment, an example is given where the estimation unit 203 performs a prescribed filtering process to suppress noise related to the image data and calculate position changes. For instance, the estimation unit 203 performs filtering on the history of feature points detected from the image data. The estimation unit 203 may perform filtering on the image data itself, rather than on the history of feature points detected from the image data. The filtering process may be a well-known filtering process that can be used to suppress noise. The filtering process may be a well-known filtering process used to smooth features within an image. Alternatively, the estimation unit 203 may calculate position changes without performing filtering. In this case, the estimation unit 203 does not have the function of filtering. An estimation unit 203 that does not perform filtering is also included within the scope of this disclosure.
[0061] For example, the estimation unit 203 performs filtering processing using a DoG (Derivative of Gaussian) filter. The Gaussian function used by the DoG filter can be a well-known function. The parameters specified by the coefficients of the DoG filter can be arbitrary values. The Gaussian distribution is adjusted according to the coefficients. In addition to noise suppression, the DoG filter is sometimes used to emphasize edges. The estimation unit 203 reduces the noise generated at least one of the plurality of feature points by using the DoG filter, and smooths the positional changes of each of the plurality of feature points, calculating the positional changes of each of the plurality of feature points.
[0062] Furthermore, the estimation unit 203 can perform filtering processing using filters other than the DoG filter. For example, the estimation unit 203 can perform filtering processing using Gaussian filters, Sobel filters, median filters, or linear filters that are not classified as DoG filters. The estimation unit 203 can perform such filtering processing and calculate the position changes of feature points. In cases where filters that cannot perform filtering processing on the history of feature points are used, the estimation unit 203 can perform filtering processing using Gaussian filters or Sobel filters on each frame of the image data, and then have the feature point detection unit 202 detect each of a plurality of feature points and calculate the position changes of each of the plurality of feature points.
[0063] For example, the estimation unit 203 sets a region window W, which includes: a frame that is the calculation object whose position changes (i.e., the calculation object frame), a frame preceding the calculation object frame (i.e., the front frame), and a frame following the calculation object frame (i.e., the rear frame). In the example of Figure 8, the calculation object frame is frame x4. The front frames are frames x1 to x3. The rear frames are frames x5 to x7. The number of front and rear frames can be arbitrary. The number of front and rear frames is not limited to the three shown in Figure 8. For example, the number of front and rear frames can be one, two, or four or more.
[0064] For example, the estimation unit 203 calculates the position change of the target frame based on the positions of a plurality of feature points detected from the preceding frame and the positions of a plurality of feature points detected from the following frame. In the example of Figure 8, the estimation unit 203 calculates the position change of the target frame, i.e., frame x4, based on the positions of a plurality of feature points detected from the preceding frames (frames x1 to x3) and the positions of a plurality of feature points detected from the following frames (frames x5 to x7).
[0065] For example, the estimation unit 203 calculates the difference between the average position of each of the plurality of feature points detected in the preceding frames (frames x1 to x3) and the average position of each of the plurality of feature points detected in the following frames (frames x5 to x7), and uses this difference as the position change of the target frame (frame x4). Even if noise is present in a specific frame, the noise is suppressed by calculating the average value. The average value can be a simple average or a weighted average. When the average value is a weighted average, the weighting coefficient can be larger the closer it is to the target frame.
[0066] In this embodiment, the positional change of the target frame, i.e., frame x4, is expressed as a vector with the same dimension as the number of feature points. The estimation unit 203 performs the aforementioned filtering process by analyzing the history of feature points in the preceding and following frames, rather than the average position, to suppress noise, smooth the positional changes of feature points, and calculate the positional change of the target frame, i.e., frame x4. For example, the estimation unit 203 can calculate the positional change of the target frame, i.e., frame x4, by tracking the temporal sequence of positional changes of feature points in frames x1 to x7 within the DoG tracking region window W.
[0067] For example, the estimation unit 203 moves the region window along the time axis of the image data (t-axis in the example of Figure 8) and calculates the positional changes of each of the plurality of calculation object frames one by one. In the example of Figure 8, when calculating the positional change of frame x4, the estimation unit 203 moves the region window W by one position and sets frame x5 as the calculation object frame. In this case, the previous frames are frames x2 to x4, and the subsequent frames are frames x6 to x8. The estimation unit 203 calculates the positional change of the calculation object frame, i.e., frame x5, based on the positions of the plurality of feature points detected from the previous frames (frames x2 to x4) and the positions of the plurality of feature points detected from the subsequent frames (frames x6 to x8).
[0068] Subsequently, the estimation unit 203 moves the region window W one position at a time while performing calculations until the position change of the last frame xm, which is the object of the position change calculation, is reached. At the end of the position change calculation, the estimation unit 203 acquires the position change data of the image data that is the object of estimation. In this position change data, the position changes of each feature point calculated from each frame of the image data that is the object of estimation are displayed sequentially over time. As described above, in this embodiment, the position change data is in vector form as an example, but the position change data can be in other forms besides vector form.
[0069] Furthermore, in the example of Figure 8, when there are frames such as x0 preceding frame x1, the positional changes of frames x1 to x3 can be calculated using these preceding frames such as x0. When there are no preceding frames such as x0, the estimation unit 203 can calculate the positional changes of frames x1 to x3 without using the area window W. Similarly, when there are frames such as xm+1 following frame xm, the positional changes of frames such as xm can be calculated using these following frames such as xm+1. When there are no following frames such as xm+1, the estimation unit 203 can calculate the positional changes of frames such as xm without using the area window W.
[0070] For example, the estimation unit 203 can calculate the position change of frame x1 based on the positions of each of the plurality of feature points detected by frame x1 and the positions of each of the plurality of feature points detected by frame x2. Similarly, the estimation unit 203 can calculate the position change of frame x2 based on the positions of each of the plurality of feature points detected by frame x2 and the positions of each of the plurality of feature points detected by frame x3. Likewise, the estimation unit 203 can calculate the position change of frame x3 based on the positions of each of the plurality of feature points detected by frame x3 and the positions of each of the plurality of feature points detected by frame x4. The same applies to the frame at the very end of image data such as frame xm; position changes in the frame at the very end can be suppressed.
[0071] In this embodiment, an example is given whereby the estimation unit 203 calculates the positional change between at least one frame in the image data, and the positional change of the frame following that frame. The estimation unit 203 further estimates whether a 2D face is displayed in the image data based on the positional change between the frames. For example, the estimation unit 203 calculates the positional change between the frames based on the positional change of at least one frame in the image data, a first weighting coefficient associated with that frame, the positional change of the frame following that frame, and a second weighting coefficient associated with that next frame.
[0072] In this embodiment, the estimation unit 203 calculates the positional change yi between frames of a given frame xi based on the following formula 1. In formula 1, hi and hi+1 represent the positional changes of each of the plurality of feature points detected from frames xi and xi+1, respectively. wi in formula 1 is the first weighting coefficient. The first weighting coefficient wi is the same as in formula 2. wi+1 in formula 1 is the second weighting coefficient. The second weighting coefficient wi+1 is the same as in formula 3. n in formulas 2 and 3 represents the number of frames that are the objects of the positional change calculation. When interpolation of positional changes is performed between all frames, n can be m-1. n can be any value.
[0073] [Number 1]
[0074] [Number 2]
[0075] [Number 3]
[0076] Furthermore, the first weighting coefficient wi and the second weighting coefficient wi+1 are not limited to the examples in equations 2 and 3. For example, the first weighting coefficient wi and the second weighting coefficient wi+1 can be calculated without using the floor function and ceiling function calculation formulas as in equations 2 and 3. The first weighting coefficient wi and the second weighting coefficient wi+1 can be fixed values (e.g., 0.3, 0.7). It is assumed that the sum of the first weighting coefficient wi and the second weighting coefficient wi+1 is 1. For example, especially without using the first weighting coefficient wi and the second weighting coefficient wi+1, the estimation unit 203 can calculate a simple average of the positional change of at least one frame of the image data and the positional change of the next frame as the positional change between the frames.
[0077] For example, the estimation unit 203 obtains the final positional change data by inserting positional changes between frames based on the positional changes detected from each of the multiple feature points in each frame. By inserting positional changes between frames, the dimension of the vector representing the positional change data also increases accordingly. Positional changes between frames can also be inserted into the input portion of the training data, i.e., the positional change data. It is assumed that the dimension of the input portion of the training data, i.e., the positional change data, is the same as the dimension of the positional change data obtained from the image data that becomes the estimation object. That is, it is assumed that the dimension of the positional change data input into the learning model M during learning is the same as the dimension of the positional change data input into the learning model M during estimation.
[0078] Furthermore, the dimension of the input portion of the training data, i.e., the positional change data, and the dimension of the positional change data obtained from the image data that serves as the inference object can be different. The dimension of the positional change data input into the learning model M during learning can also be different from the dimension of the positional change data input into the learning model M during inference. In this case, the portion with insufficient dimension can be treated as an omission value. Furthermore, for example, the learning model M can also calculate the embedding by absorbing the difference in dimension. For example, when the embedding is also expressed as a vector, the dimension of the embedding can be preset. The learning model M can calculate the embedding of a specified dimension regardless of the dimension of the input portion of the training data, i.e., the positional change data, and the dimension of the positional change data obtained from the image data that serves as the inference object. The dimension of the embedding can be indefinite and not preset.
[0079] Furthermore, the estimation unit 203 can determine whether the positional changes of each of the plurality of frames in the image data are within the reference range. The reference range is the range that serves as the basis for whether to interpolate the positional changes between frames. For example, if the user moves their face faster than expected, the positional change may be too large, and the estimation unit 203 may be unable to make an accurate estimation. Therefore, the positional change exceeds a predetermined threshold, which is equivalent to being outside the reference range. The threshold can be arbitrarily determined by the administrator of the image processing system 1.
[0080] For example, the estimation unit 203 determines whether the positional change of each of the plurality of frames in the image data is above a threshold value, and thus determines whether the positional change is within a reference range. The positional change that becomes the determination object of the reference range can be the positional change of all feature points, or it can be the positional change of a portion of the feature points. Furthermore, the positional change that becomes the determination object of the reference range can be the positional change of all frames, or it can be the positional change of a portion of the frames.
[0081] For example, when the estimation unit 203 determines that the position change is outside the reference range, it calculates the position change between frames. Since sometimes only the position change of a portion of the frames in the image data is outside the reference range, in this case, the estimation unit 203 can calculate the position change between frames in the portion of the frame whose position change is outside the reference range, or it can calculate the position change between frames in all frames of the image data.
[0082] For example, when the estimation unit 203 does not determine that the positional change is outside the reference range, it does not calculate the positional change between frames. In this case, y1~yn in Figure 8 are not calculated, and only h1~hm are used in the estimation. Since sometimes only the positional change of a portion of the frames in the image data is outside the reference range, in this case, the estimation unit 203 may not calculate the positional change between frames in the portion of the frames that are not determined to be outside the reference range, but calculate the positional change between frames in other frames that are determined to be outside the reference range. Furthermore, in this case, the estimation unit 203 may not calculate the positional change between frames in all frames of the image data. Also, regardless of whether the positional change is outside the reference range, the estimation unit 203 may not calculate the positional change between frames, and make an estimation based on the positional change calculated from the frames. In this case, the estimation unit 203 may not have the function of calculating the positional change between frames.
[0083] [Personal Confirmation Department] The user verification unit 204 performs user verification based on the presumption that a 2D face is displayed in the image data. For example, if the user verification unit 204 presumes that a 2D face is displayed in the image data, it determines that user verification has failed. The user verification unit 204 may, when it presumes that a 2D face is displayed in the image data, request the user to perform a prescribed action again. The action requested again may be different from the action requested so far.
[0084] For example, when the identity verification unit 204 determines that identity verification is successful because a 2D face is not displayed in the image data, it determines that identity verification is successful. If the next step is being prepared for identity verification (e.g., photographing identity verification documents such as a driver's license), the identity verification unit 204 can proceed to the next step. The steps required for identity verification can be the same as those known for identity verification. The information processing required for this step can also be the same as that used in known identity verification. When the identity verification unit 204 determines that a 2D face is displayed in the image data, it can notify the person responsible for identity verification. In this case, the person responsible can visually verify the image data.
[0085] [3-3. Functions implemented by the user terminal] For example, the user terminal 30 includes a data memory unit 300, a display control unit 301, and an operation receiving unit 302. The data memory unit 300 is implemented by a memory unit 32. The display control unit 301 and the operation receiving unit 302 are each implemented by a control unit 31.
[0086] [Data Memory Department] The data storage unit 300 stores the data required for uploading image data. For example, the data storage unit 300 stores image data generated based on the photographic results of the photography unit 36.
[0087] [Display Control Department] The display control unit 301 displays various screens from the image processing system 1 on the display unit 35. For example, the display control unit 301 displays the photographic screen SC, which shows the photographic results of the photographic unit 36, on the display unit 35.
[0088] [Operations and Reception Department] The operation receiving unit 302 receives various operations from the image processing system 1. The data showing the operation content received by the operation receiving unit 302 is appropriately sent to the server 20.
[0089] [4. Processing performed by the image processing system] Figures 9 and 10 illustrate an example of the processing performed by the image processing system 1. Control units 11, 21, and 31 execute the processing shown in Figures 9 and 10 by executing programs stored in memory units 12, 22, and 32, respectively. In the example shown in Figures 9 and 10, it is assumed that user confirmation is performed when a user requests a service.
[0090] As shown in Figure 9, the learning terminal 10 performs learning of the learning model M based on the training data stored in the training database DB1 (S1). The learning terminal 10 sends the pre-trained learning model M to the server 20 (S2). The server 20 receives the pre-trained learning model M from the learning terminal 10 (S3). In S3, the server 20 records the pre-trained learning model M in the memory unit 22.
[0091] User terminal 30 performs processing for user service requests between itself and server 20 (S4). Upon confirmation that the user has proceeded to the required step, server 20 sends request data to user terminal 30, instructing the user to perform a specified action (S5). The request data displays the content of message MS. User terminal 30 receives the request data (S6). User terminal 30 activates camera unit 36, displaying the captured image SC on display unit 35 (S7). The user performs the action guided by message MS.
[0092] User terminal 30 sends image data to server 20 (S8). Server 20 receives image data from user terminal 30 (S9). Based on the image data, server 20 detects a plurality of feature points related to the user's face (S10). Moving to Figure 10, server 20 performs filtering processing using a DoG filter (S11). Server 20 calculates the positional changes of each frame (S12). Server 20 can perform processing S11 and S12 while moving the area window W. Server 20 calculates the positional changes between frames based on equations 1 to 3 (S13).
[0093] Based on the positional changes of each frame, the positional changes between frames, and the pre-trained learning model M, server 20 infers whether a 2D face is displayed in the image data (S14). If, in S14, it is determined that a 2D face is displayed in the image data (S14: Yes), server 20 performs processing to display an error message between itself and user terminal 30 (S15), and this processing ends. If, in S14, it is not determined that a 2D face is displayed in the image data (S14: No), server 20 performs processing between itself and user terminal 30 to proceed to the next step of user confirmation (S16), and this processing ends.
[0094] [5. Summary of Implementation Modes] The image processing system 1 of this embodiment estimates whether a 2D face is displayed in the image data based on the positional changes of multiple feature points in the image data and a learning model M that has been trained on training data related to the positional changes of multiple feature points of the face. Thus, the image processing system 1 can estimate whether a 2D face is displayed in the image data even without special hardware. For example, it is also considered that without special hardware, the motion displayed in the image data can be analyzed using optical flow, but optical flow is not sensitive to background movement, so the accuracy is insufficient. Regarding this point, the estimation accuracy can be improved by focusing on the positional changes of feature points in the foreground, i.e., the face. Furthermore, if the estimation is to be performed by analyzing the movement of the mouth, a high frame rate camera unit 36 is required, and slow user movements are also required. However, even with a camera unit 36 that can only acquire 2D information, the image processing system 1 can estimate whether a 2D face is displayed in the image data. The image processing system 1 can determine whether a person is present (Liveness) by estimating whether a 2D face is displayed in the image data.
[0095] Furthermore, the image processing system 1 performs prescribed filtering to suppress noise related to the image data and calculate positional changes. By suppressing noise, the image processing system 1 can improve estimation accuracy. For example, if the detection accuracy of feature points is not perfect, sometimes the position of feature points will shift even if the user does not move their face. Even if this positional shift occurs, it can be suppressed as noise through filtering, thus improving estimation accuracy.
[0096] Furthermore, the image processing system 1 performs filtering processing using a DoG filter. By using the DoG filter, the image processing system 1 can effectively suppress noise and improve estimation accuracy. The image processing system 1 can also ensure smooth changes in the position of displayed feature points.
[0097] Furthermore, the image processing system 1 sets a region window W and calculates the positional change of the object frame based on the positions of a plurality of feature points detected from the front frame and the positions of a plurality of feature points detected from the rear frame. The image processing system 1 moves the region window along the time axis of the image data and calculates the positional change of each of the plurality of object frames one by one. By calculating the positional change of the object frame based not only on a plurality of feature points detected not only from either the front or rear frame, but also from both the front and rear frames, the image processing system 1 improves the accuracy of positional change calculation.
[0098] Furthermore, the image processing system 1 calculates the positional changes between frames based on the positional changes of at least one frame in the image data and the positional changes of the next frame. The image processing system 1 further infers whether a 2D face is displayed in the image data based on these positional changes. Even if the speed of the prescribed action deviates due to user error, the image processing system 1 can compensate for the positional changes between frames, thus improving the estimation accuracy. For example, if a user spends 3 seconds moving from front to left, even if it only takes 1 second, the image processing system 1 improves the estimation accuracy by interpolating the positional changes between frames.
[0099] Furthermore, the image processing system 1 calculates the positional change between frames based on the positional change of at least one frame in the image data, a first weighting coefficient wi associated with that frame, the positional change of the next frame, and a second weighting coefficient wi+1 associated with the next frame. In this way, when calculating the positional change between frames, the image processing system 1 can adjust which of the two frames to apply the weight to.
[0100] Furthermore, the image processing system 1 determines whether the positional changes of each of the plurality of frames in the image data are within a reference range. When the image processing system 1 determines that the positional changes are outside the reference range, it calculates the positional changes between frames. The image processing system 1 can calculate the positional changes between frames when interpolation of positional changes between frames is necessary. The image processing system 1 can choose not to calculate the positional changes between frames when interpolation of positional changes between frames is not necessary.
[0101] Furthermore, the image processing system 1 performs user identity verification based on the presumption of whether a 2D face is displayed in the image data. In this way, the image processing system 1 can prevent impersonation during identity verification by malicious users.
[0102] [6. Examples of Variation] This disclosure is not limited to the implementation described above. This disclosure may be modified as appropriate without departing from its intent.
[0103] Figure 11 shows one example of the function implemented by the variation example. As shown in Figure 11, the server 20 of the variation example includes an action request unit 205 and a registration unit 206. Both the action request unit 205 and the registration unit 206 are implemented by the control unit 21.
[0104] [6-1. Variation Example 1] For example, in an implementation, consider a learning model M. When inputted with positional change data, this model M outputs a label indicating whether a 2D face is displayed in the image data. The learning model M only needs to be a model of the positional changes of a plurality of face-related feature points that have been learned and trained. The learning model M is not limited to this implementation example.
[0105] Figure 12 shows an example of the learning model M in Variation Example 1. The learning unit 101 of Variation Example 1 learns a first learning model M1 and a second learning model M2. The first learning model M1 has learned first training data related to the positional changes of each of the plurality of feature points of the face used for 3D training, and the second learning model M2 has learned second training data related to the positional changes of each of the plurality of feature points of the face used for 2D training.
[0106] For example, the first training data is positional change data calculated from training image data generated by photographing the face of the training subject. The method of obtaining the positional change data can be the same as in the implementation. The second training data is positional change data calculated from training image data generated by photographing training media. The training media is paper printed with the face of the training subject, or a computer displaying the face of the training subject. The first and second training data are stored in the training database DB1. Furthermore, the first and second training data may not be tagged.
[0107] For example, the learning unit 101 enables the first learning model M1 to learn the features displayed in the first training data. As in the embodiment, when the input of the first training data is positional change data and the output is labels, the learning unit 101 adjusts the parameters of the first learning model M1 so that when the input of the first training data, i.e., the positional change data, is input into the first learning model M1, the output of the first learning model M1 represents a label that is not 2D of a face. For example, when no labels are assigned to the first training data, the learning unit 101 performs clustering of multiple sets of the first training data, adjusting the parameters of the first learning model M1 so that first training data with similar features belong to the same cluster.
[0108] For example, the learning unit 101 enables the second learning model M2 to learn the features displayed in the second training data. As in the embodiment, when the input of the second training data is positional change data and the output is labels, the learning unit 101 adjusts the parameters of the second learning model M2 so that when the input of the second training data, i.e., the positional change data, is input into the second learning model M2, the output of the second learning model M2 represents a label that is not 2D of a face. For example, when no labels are assigned to the second training data, the learning unit 101 performs clustering of multiple sets of the second training data, adjusting the parameters of the second learning model M2 so that second training data with similar features belong to the same cluster.
[0109] In Variation Example 1, the estimation unit 203 acquires position change data related to the characteristics of position changes in multidimensional space. The method for acquiring the position change data can be the same as in the embodiment. Based on the first learning model M1, the estimation unit 203 performs encoding to reduce the dimensions of the position change data and decoding to restore the dimensions of the position change data, and calculates the amount of movement of the position change data before and after the encoding and decoding, i.e., the first movement amount.
[0110] Encoding is also known as dimensionality reduction. Encoding methods can be well-known. For example, the estimation unit 203 can reduce the dimensionality of positional variation data based on methods such as principal component analysis or autoencoding. Decoding can also be well-known. For example, the estimation unit 203 can restore the dimensionality-reduced positional variation data to its original dimensionality based on methods such as principal component analysis or autoencoding. Encoding and decoding are performed based on the parameters of the first learning model M1.
[0111] For example, the estimation unit 203 performs encoding and decoding based on the second learning model M2, and calculates the shift amount of the positional change data before and after the encoding and decoding, which is the second shift amount. The meaning of encoding and decoding is the same as that of the first learning model M1. Encoding and decoding are performed based on the parameters of the second learning model M2.
[0112] For example, the estimation unit 203 estimates whether a 2D face is displayed in the image data based on the first and second motion amounts. When a 3D face is displayed in the image data that is the object of estimation, the first motion amount before and after encoding and decoding is smaller than the second motion amount. This is because the features of the image data are similar to the features of the first training data learned by the first learning model M1. The estimation unit 203 does not estimate that a 2D face is displayed in the image data when the first motion amount does not reach the second motion amount.
[0113] On the other hand, when a 2D face is displayed in the image data that is the object of estimation, the first movement amount before and after encoding and decoding is greater than the second movement amount. This is because the features of the image data are similar to the features of the second training data learned by the second learning model M2. When the first movement amount is greater than the second movement amount, the estimation unit 203 estimates that a 2D face is displayed in the image data.
[0114] In Variation Example 1, image processing system 1 calculates a first motion amount based on a first learning model M1. Image processing system 1 calculates a second motion amount based on a second learning model M2. Based on the first and second motion amounts, image processing system 1 estimates whether a 2D face is displayed in the image data. In this way, image processing system 1 can improve the estimation accuracy of whether a 2D face is displayed in the image data.
[0115] [6-2. Variation Example 2] For example, in the implementation, we take the case where the user is asked to turn their face from the front to the left, but the actions requested by the user are not limited to the implementation example. The image processing system 1 of Variation Example 2 includes an action request unit 205. The action request unit 205 requests the user to select an action from a plurality of actions related to the face using a prescribed selection method.
[0116] The candidate actions required of the user can be any facial-related action. For example, candidate actions could include changing the direction of the face, moving parts of the face, deforming parts of the face, or other actions. Candidate actions can also be actions other than those exemplified in the implementation example. It is assumed that the data representing the candidate actions has been pre-memorized in the data memory unit 200.
[0117] The selection method can be any method. In Variation Example 2, the case where the selection method is random selection is illustrated. The method of random selection from candidates can be a well-known method. For example, the action request unit 205 randomly selects one action from a plurality of actions based on a random number. The action request unit 205 then requests the user to perform the selected action. In the examples of Figures 2 and 3, the user is requested to perform the action by displaying a message MS, but other methods, such as outputting sound, can also be used to request the user to perform the action.
[0118] Furthermore, the selection method is not limited to random selection. For example, the selection method could be to select multiple actions in a preset order. The selection method could also be to select actions based on data obtained from the user terminal 30 or user characteristics. The image data acquisition unit 201 acquires image data when the user performs an action selected by the action selection unit. While the point of image selection by the action selection unit differs from the implementation mode, the method of acquiring the image data is the same.
[0119] Assume that the data memory unit 200 of Variation Example 2 has prepared a learning model M for each action. The training database DB1 stored in the learning terminal 10 contains training data for displaying a plurality of actions. Based on the training data for displaying a particular action, the learning unit 101 performs learning of the learning model M corresponding to that action. While the preparation of training data and learning model M for each action differs from the implementation form, the method for creating the learning model M itself can be the same as in the implementation form.
[0120] For example, suppose we consider four actions as candidates: the first action of facing left, the second action of facing right, the third action of facing up, and the fourth action of blinking. In this case, the learning unit 101 executes the learning model M for the first action, the learning model M for the second action, the learning model M for the third action, and the learning model M for the fourth action. The learning terminal 10 sends the four pre-trained learning models M to the server 20. The server 20 records the four pre-trained learning models M in the data memory unit 200.
[0121] In Variation Example 2, the estimation unit 203 estimates whether a 2D face is displayed in the image data based on the learning model M corresponding to each of the plurality of actions and the learning model M corresponding to the action selected by the action selection unit. The point of use in the estimation of the plurality of learning models M in the data memory unit 200 and the learning model M corresponding to the action selected by the action selection unit is different from the implementation form, but the estimation based on the learning model M can be the same as the implementation form.
[0122] For example, assuming the aforementioned four actions (actions 1 through 4) are candidates, the estimation unit 203, when the user performs action 1, makes an estimation based on the learning model M corresponding to action 1 out of the four learning models M. When the user performs action 2, the estimation unit 203 makes an estimation based on the learning model M corresponding to action 2 out of the four learning models M. When the user performs action 3, the estimation unit 203 makes an estimation based on the learning model M corresponding to action 3 out of the four learning models M. When the user performs action 4, the estimation unit 203 makes an estimation based on the learning model M corresponding to action 4 out of the four learning models M.
[0123] In Variation Example 2, the image processing system 1 requires the user to select an action from a set of actions using a prescribed selection method. The image processing system 1 acquires image data when the user performs the selected action. Based on the learning model M corresponding to the selected action in the learning model M corresponding to each of the set of actions, the image processing system 1 infers whether a 2D face should be displayed in the image data. In this way, the image processing system 1 can perform an estimation of the appropriate learning model M for the user's actions. Because a malicious user must devise various actions to verify their identity, it is difficult to successfully impersonate them.
[0124] [6-3. Variation Example 3] For example, in one embodiment, the presumption result of the presumption unit 203 is used in the user verification process. The presumption result of the presumption unit 203 can be used for any purpose other than user verification. In Variation 3, the presumption result of the presumption unit 203 is used when a user registers image data for facial authentication. A malicious user may register another person's facial photo as image data for facial authentication and then continue to use that facial photo to impersonate that person. To prevent this impersonation, the image processing system 1 can also be used.
[0125] In Variation 3, assuming the same conditions as in Variation 2, images are randomly requested. The image processing system 1 of Variation 3 includes a registration unit 206. Based on the estimation result of whether a 2D face is displayed in the image data, the registration unit 206 registers frames showing the user's face facing forward among a plurality of frames in the image data as image data for user facial authentication. The registration unit 206 can determine whether the user's face is facing forward by image analysis (e.g., template matching, methods based on contour shape, or machine learning methods), and can determine a specific frame (e.g., the first frame) as the frame showing the user's face facing forward.
[0126] For example, when the registration unit 206 presumes that a 2D face is displayed in the image data, it does not register the user's facial recognition image data. In this case, since a malicious user intends to register image data for impersonation using someone else's facial photograph, the registration unit 206 does not register the image data. When the registration unit 206 does not presume that a 2D face is displayed in the image data, it registers the user's facial recognition image data. The facial recognition image data can be registered in the user database DB2 or in other databases.
[0127] For example, server 20 or another computer performs facial authentication based on image data registered by login unit 206. The authentication terminal for facial authentication includes a camera. The authentication terminal can be user terminal 30 or another computer besides user terminal 30. When the authentication terminal captures the user's face with the camera, it sends image data to server 20 or the other computer. Server 20 or the other computer performs facial authentication based on the image data received from the authentication terminal and the image data registered by login unit 206. Facial authentication can be performed using well-known methods.
[0128] In Variation Example 3, the image processing system 1, based on the estimation result of whether a 2D face is displayed in the image data, registers the frame displaying the user's face facing forward among multiple frames in the image data as image data for user facial authentication. The image processing system 1 can prevent malicious users from impersonating others to register image data to perform facial authentication, etc.
[0129] [6-4. Other examples of variations] For example, variations 1 to 3 above can be combined.
[0130] For example, the functions described as implemented by server 20 can be implemented by other computers such as learning terminal 10. The functions described as implemented by server 20 can be shared by multiple computers. The functions described as implemented by learning terminal 10 can be implemented by other computers such as server 20.
[0131] [7. Postscript] For example, an image processing system can also be configured as follows. (1) An image processing system comprising: The image data acquisition unit acquires and displays image data related to the image of the user's face while performing the prescribed action; The feature point detection unit, based on the aforementioned image data, detects a plurality of feature points related to the aforementioned face; and The estimation unit, based on the positional changes of the aforementioned plurality of feature points in the aforementioned image data and the learning model that has been learned in relation to the positional changes of the plurality of feature points of the face used for training, estimates whether a 2D face is displayed in the aforementioned image data. (2) In the image processing system of (1), the aforementioned estimation unit performs the prescribed filtering process to suppress noise related to the aforementioned image data and calculate the aforementioned position change. (3) The image processing system of (2) wherein the aforementioned estimation unit performs the aforementioned filtering process using a DoG (Derivative of Gaussian) filter. (4) Image processing systems such as any one of (1) to (3), wherein the aforementioned presumption unit Define a region window, which includes a frame that is the calculation object of the aforementioned position change (i.e., the calculation object frame), a frame preceding the calculation object frame (i.e., the front frame), and a frame following the calculation object frame (i.e., the back frame). Based on the positions of the aforementioned plurality of feature points detected from the preceding frame and the positions of the aforementioned plurality of feature points detected from the following frame, the aforementioned positional change of the aforementioned object frame is calculated. Move the aforementioned region window along the timeline of the aforementioned image data, and calculate the aforementioned positional changes of each of the plurality of aforementioned calculation object frames. (5) Image processing systems such as any one of (1) to (4), wherein the aforementioned presumption unit Based on the aforementioned positional changes of at least one frame in the aforementioned image data, and the aforementioned positional changes of the next frame, calculate the positional changes between these frames, and Based on the positional changes between the frames, it can be inferred whether the aforementioned 2D face is displayed in the aforementioned image data. (6) As in the image processing system of (5), the aforementioned estimation unit calculates the positional changes between the aforementioned frames based on the aforementioned positional changes of at least one frame of the aforementioned image data, the first weighting coefficient associated with the frame, the aforementioned positional changes of the next frame of the frame, and the second weighting coefficient associated with the next frame. (7) In an image processing system such as (5) or (6), the aforementioned estimation unit determines whether the aforementioned positional changes of each of the plurality of frames of the aforementioned image data are within the reference range. When it is determined that the aforementioned positional changes are outside the aforementioned reference range, the unit calculates the positional changes between the aforementioned frames. (8) Image processing systems such as any one of (1) to (7), wherein the aforementioned presumption unit Obtain position change data related to the characteristics of the aforementioned position changes in multidimensional space. Based on a first learning model that has been learned and correlated with the positional changes of multiple feature points of a face used for training in 3D, the model performs encoding to reduce the dimensionality of the aforementioned positional change data and decoding to restore the dimensionality of the aforementioned positional change data. The amount of movement of the aforementioned positional change data before and after the encoding and decoding is calculated, i.e., the first movement amount. Based on the second learning model, which is derived from the second training data relating to the positional changes of multiple feature points of a face used for training in 2D, the aforementioned encoding and decoding are performed. The amount of shift in the aforementioned positional change data before and after the encoding and decoding is calculated, i.e., the second shift amount. Based on the aforementioned first and second motion amounts, it is presumed whether the aforementioned 2D face is displayed in the aforementioned image data. (9) The image processing system of any one of (1) to (8), wherein the aforementioned image processing system further includes a motion request unit, which requests the user to select an action from a plurality of actions related to the aforementioned face using a prescribed selection method; and The aforementioned image data acquisition unit acquires the aforementioned image data when the aforementioned user performs the aforementioned selected action; The aforementioned estimation unit, based on the learning model corresponding to each of the aforementioned plurality of actions, and the learning model corresponding to the aforementioned selected actions, estimates whether the aforementioned 2D face is displayed in the aforementioned image data. (10) The image processing system of any one of (1) to (9) further includes a self-verification unit, which performs the self-verification of the user based on the inference result of whether the aforementioned 2D face is displayed in the aforementioned image data. (11) The image processing system of any one of (1) to (10), wherein the aforementioned image processing system further includes a registration unit, which, based on the inference result of whether the aforementioned 2D face is displayed in the aforementioned image data, registers the aforementioned frame that displays the aforementioned user's face facing forward among the plurality of frames in the aforementioned image data as image data for the aforementioned user's face authentication.
[0132] 1: Image Processing System 10: Learning Terminal 11, 21, 31: Control Department 12, 22, 32: Memory Department 13, 23, 33: Communications Department 14.34: Operations Department 15, 35: Display section 20: Server 30: User terminal 36: Photography Department 100: Data Memory Department 101: Academic Department 200: Data Memory Department 201: Image Data Acquisition Department 202: Feature Point Detection Department 203: Presumption Section 204: Personal Confirmation Department 205: Action Requirements Department 206: Login Department 300: Data Memory Department 301: Display Control Unit 302: Operations and Acceptance Department DB1: Training Data Repository DB2: User Database G: Guidance h1~hm: Positional changes M: Learning Model MS: Message N: Network P: Paper SC: Camera Footage t: axis W: Area Window x1~xm: Frame y1~yn: News
Claims
1. An image processing system comprising: an image data acquisition unit that acquires and displays image data related to an image of the face of a user performing a prescribed action; The feature point detection unit detects a plurality of feature points related to the aforementioned face based on the aforementioned image data; and the estimation unit acquires position change data related to the position changes of each of the aforementioned plurality of feature points in multidimensional space, performs encoding to reduce the dimension of the aforementioned position change data and decoding to restore the dimension of the aforementioned position change data based on a first learning model learned from a first training data related to the position changes of each of the plurality of feature points of a 3D training face, calculates the amount of movement of the aforementioned position change data before and after the encoding and decoding, i.e., the first movement amount, based on a second learning model learned from a second training data related to the position changes of each of the plurality of feature points of a 2D training face, performs the aforementioned encoding and decoding, calculates the amount of movement of the aforementioned position change data before and after the encoding and decoding, i.e., the second movement amount, and estimates whether a 2D face is displayed in the aforementioned image data based on the aforementioned first movement amount and the aforementioned second movement amount.
2. The image processing system of claim 1, wherein the aforementioned estimation unit performs the prescribed filtering process to suppress noise related to the aforementioned image data and calculate the aforementioned positional change.
3. The image processing system of claim 2, wherein the aforementioned presumption unit performs the aforementioned filtering process using a DoG (Derivative of Gaussian) filter.
4. An image processing system as described in any of claims 1 to 3, wherein the aforementioned estimation unit sets a region window, the region window including a frame that is the object of the aforementioned position change calculation (i.e., a calculation object frame), a frame preceding the calculation object frame (i.e., a front frame), and a frame following the calculation object frame (i.e., a rear frame), and calculates the aforementioned position change of the aforementioned calculation object frame based on the positions of the aforementioned plurality of feature points detected from the aforementioned front frame and the positions of the aforementioned plurality of feature points detected from the aforementioned rear frame, and moves the aforementioned region window on the time axis of the aforementioned image data to calculate the aforementioned position change of each of the plurality of aforementioned calculation object frames one by one.
5. The image processing system of any one of claims 1 to 3, wherein the aforementioned estimation unit calculates the positional change between the frames based on the aforementioned positional change of at least one frame of the aforementioned image data and the aforementioned positional change of the next frame of the frame, and further estimates whether the aforementioned 2D face is displayed in the aforementioned image data based on the positional change between the frames.
6. The image processing system of claim 5, wherein the aforementioned estimation unit calculates the positional change between the aforementioned frames based on the aforementioned positional change of at least one frame of the aforementioned image data, a first weighting coefficient associated with the frame, the aforementioned positional change of the next frame of the frame, and a second weighting coefficient associated with the next frame.
7. The image processing system of claim 5, wherein the aforementioned estimation unit determines whether the aforementioned positional change of each of the plurality of frames of the aforementioned image data is within the reference range, and when it is determined that the aforementioned positional change is outside the aforementioned reference range, calculates the positional change between the aforementioned frames.
8. The image processing system of any one of claims 1 to 3, wherein the image processing system further comprises: an action request unit that requests an action from a plurality of actions related to the face by a prescribed selection method; and an image data acquisition unit that acquires the image data when the user performs the selected action; and an estimation unit that estimates whether the 2D face is displayed in the image data based on the first learning model and the second learning model corresponding to each of the plurality of actions, and the first learning model and the second learning model corresponding to the selected action.
9. The image processing system of any one of claims 1 to 3, wherein the aforementioned image processing system further includes: a self-confirmation unit, which performs the aforementioned self-confirmation of the user based on the inference result of whether the aforementioned image data displays the aforementioned 2D face.
10. The image processing system of any one of claims 1 to 3, wherein the image processing system further comprises: a registration unit, which, based on the estimation result of whether the aforementioned 2D face is displayed in the aforementioned image data, registers the aforementioned frame displaying the aforementioned user's face facing forward among a plurality of frames in the aforementioned image data as image data for the aforementioned user's facial authentication.
11. An image processing method, wherein a computer performs the following steps: an image data acquisition step, wherein the acquisition is related to the image of the face of a user performing a prescribed action; The feature point detection step detects a plurality of feature points related to the aforementioned face based on the aforementioned image data; and the estimation step obtains position change data related to the position changes of each of the aforementioned plurality of feature points in multidimensional space, performs encoding to reduce the dimension of the aforementioned position change data and decoding to restore the dimension of the aforementioned position change data based on a first learning model that has been learned and is related to the position changes of each of the plurality of feature points of the 3D training face, calculates the amount of movement of the aforementioned position change data before and after the encoding and decoding, i.e., the first movement amount, based on a second learning model that has been learned and is related to the position changes of each of the plurality of feature points of the 2D training face, performs the aforementioned encoding and decoding, calculates the amount of movement of the aforementioned position change data before and after the encoding and decoding, i.e., the second movement amount, based on the aforementioned first movement amount and the aforementioned second movement amount, and estimates whether a 2D face is displayed in the aforementioned image data.
12. A program product for enabling a computer to function as: an image data acquisition unit that acquires and displays image data related to an image of the face of a user performing a prescribed action; The feature point detection unit detects a plurality of feature points related to the aforementioned face based on the aforementioned image data; and the estimation unit acquires position change data related to the position changes of each of the aforementioned plurality of feature points in multidimensional space, performs encoding to reduce the dimension of the aforementioned position change data and decoding to restore the dimension of the aforementioned position change data based on a first learning model learned from a first training data related to the position changes of each of the plurality of feature points of a 3D training face, calculates the amount of movement of the aforementioned position change data before and after the encoding and decoding, i.e., the first movement amount, based on a second learning model learned from a second training data related to the position changes of each of the plurality of feature points of a 2D training face, performs the aforementioned encoding and decoding, calculates the amount of movement of the aforementioned position change data before and after the encoding and decoding, i.e., the second movement amount, and estimates whether a 2D face is displayed in the aforementioned image data based on the aforementioned first movement amount and the aforementioned second movement amount.
Citation Information
Patent Citations
Solidity authenticating method, solidity authenticating apparatus, and solidity authenticating program
JP2012069133A
Moving image determination method
TW202318336A
User image verification
US20190347388A1