Method of operating an information processing device, information processing device, and program
The information processing apparatus quantitatively evaluates facial expressions by classifying changes into intensity levels and patterns, addressing the subjectivity of existing methods and providing precise analysis of facial changes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- KOSE HOLDINGS CORP
- Filing Date
- 2022-06-17
- Publication Date
- 2026-07-29
AI Technical Summary
Existing methods for evaluating facial expressions are subjective and lack quantitative analysis, making it difficult to accurately assess the intensity and patterns of facial changes.
An information processing apparatus and method that classifies facial expression changes into intensity levels and patterns by analyzing the distribution and combination of facial elements over time, using machine learning to derive quantitative evaluations.
Enables quantitative evaluation of facial expressions, allowing for precise classification of expression intensities and patterns, and even atmospheres created by facial changes.
Smart Images

Figure 0007897053000003 
Figure 0007897053000004 
Figure 0007897053000005
Abstract
Description
Technical Field
[0001] The present disclosure relates to an operation method of an information processing apparatus, an information processing apparatus, and a program.
Background Art
[0002] The impression that a person's facial appearance gives to others is influenced by various factors including expressions. In the field of beauty, various techniques for analyzing and evaluating expressions have been proposed. For example, Patent Document 1 discloses a method for specifying a physical quantity of skin change having a correlation with an impression obtained from a facial expression.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The impression obtained from a facial expression is merely a subjectively verbalized type. Even if factors having a correlation with such a type are specified, it is not always possible to quantitatively evaluate the expression or the change in the expression itself.
[0005] In view of the above, an operation method of an information processing apparatus that enables quantitative evaluation of changes in facial expressions will be disclosed below.
Means for Solving the Problems
[0006] To solve the above problems, an operation method of an information processing apparatus in the present disclosure includes a first step of classifying a distribution of time change amounts of a plurality of elements in each of a plurality of facial images into a plurality of expression intensities according to the magnitude of the time change amount for each distribution, and a second step of classifying a plurality of the distributions at each expression intensity into a plurality of expression patterns according to the mode of the combination of the element and the time change amount for each distribution.
[0007] Furthermore, the information processing device in this disclosure includes a control unit that classifies the distribution of the time-varying amounts of multiple elements in each of the multiple facial images into multiple facial intensity levels according to the magnitude of the time-varying amounts for each distribution, and classifies the multiple distributions in each facial intensity into multiple facial expression patterns according to the combination of the elements and the time-varying amounts for each distribution.
[0008] Furthermore, the program in this disclosure is executed by an information processing device, and the information processing device performs a first step of classifying the distribution of the time-varying amounts of multiple elements in each of multiple facial images into multiple facial intensity levels according to the magnitude of the time-varying amounts for each distribution. The second step is to classify the multiple distributions at each facial intensity into multiple facial expression patterns according to the combination of the elements and the amount of change over time for each distribution.
[0009] Furthermore, the operation method of another information processing device in this disclosure is an operation method in which the information processing device stores information for classifying the distribution of the time change amounts of multiple elements in each of a plurality of face images into a plurality of facial intensity levels according to the magnitude of the time change amounts for each distribution, and includes the steps of acquiring a target face image and outputting information on the facial intensity corresponding to the face image based on the time change amounts of the elements in the target face image.
[0010] Furthermore, an operation method for yet another information processing device in this disclosure is an operation method in which the information processing device classifies the distribution of time changes of a plurality of elements in each of a plurality of face images into a plurality of facial intensity levels according to the magnitude of the time changes for each distribution, stores information for classifying the plurality of distributions corresponding to each facial intensity into a plurality of facial patterns according to the manner of the combination of the elements and the time changes for each distribution, and includes the steps of acquiring a target face image and outputting information of a facial pattern corresponding to the face image based on the time changes of the elements in the target face image.
[0011] Furthermore, an operation method for yet another information processing device in this disclosure is an operation method that includes the steps of: the information processing device classifying the distribution of time changes of a plurality of elements in each of a plurality of face images into a plurality of facial intensity levels according to the magnitude of the time changes for each distribution; classifying the plurality of distributions corresponding to each facial intensity into a plurality of facial patterns according to the manner of the combination of the elements and the time changes for each distribution; and storing information for classifying the time distribution of the plurality of facial patterns corresponding to each of a plurality of face images into a plurality of atmosphere patterns according to the manner of the time distribution; acquiring a target face image; and outputting information for the atmosphere pattern corresponding to the face image based on the time changes of the elements in the target face image.
[0012] Furthermore, another information processing device in this disclosure includes a storage unit that stores information for classifying the distribution of the time-varying amounts of multiple elements in each of a plurality of face images into a plurality of facial intensity levels according to the magnitude of the time-varying amounts for each distribution, and a control unit that outputs information on facial intensity levels corresponding to an input face image based on the time-varying amounts of the elements in that face image.
[0013] Furthermore, yet another information processing device in this disclosure includes a storage unit that stores information for classifying the distribution of the time-varying amounts of multiple elements in each of a plurality of facial images into a plurality of facial intensity levels according to the magnitude of the time-varying amounts for each distribution, and for classifying the plurality of distributions corresponding to each facial intensity into a plurality of facial expression patterns according to the combination of the elements and the time-varying amounts for each distribution, and a control unit that outputs information of a facial expression pattern corresponding to an input facial image based on the time-varying amounts of the elements in the facial image.
[0014] Furthermore, yet another information processing device in this disclosure includes a storage unit that stores information for classifying the distribution of the time-varying amounts of multiple elements in each of a plurality of facial images into a plurality of facial intensity levels according to the magnitude of the time-varying amounts for each distribution, classifying the plurality of distributions corresponding to each facial intensity into a plurality of facial patterns according to the manner of the combination of the elements and the time-varying amounts for each distribution, and classifying the time-varying distribution of the plurality of facial patterns corresponding to each of the plurality of facial images into a plurality of atmosphere patterns according to the manner of the time-varying distribution, and a control unit that outputs information of the atmosphere pattern corresponding to the facial image based on the time-varying amounts of the elements in the input facial image, It holds.
[0015] Furthermore, another program in this disclosure is a program that causes an information processing device to store information for classifying the distribution of the time-varying amounts of multiple elements in each of multiple face images into multiple facial expression intensities according to the magnitude of the time-varying amounts for each distribution, and to perform the steps of acquiring a target face image and outputting information on the facial expression intensity corresponding to the face image based on the time-varying amounts of the elements in the target face image.
[0016] Furthermore, yet another program in this disclosure is a program that causes the information processing device to perform the following steps: acquire a target face image and output information about the facial expression pattern corresponding to the target face image based on the amount of change of the elements in the target face image. The program also includes information for classifying the distribution of the amount of change of time of multiple elements in each of multiple face images into multiple facial expression intensities according to the magnitude of the amount of change of time for each distribution, and for classifying the multiple distributions corresponding to each facial expression intensities into multiple facial expression patterns according to the combination of the elements and the amount of change of time for each distribution.
[0017] Furthermore, another program in the present disclosure is such that an information processing apparatus classifies the distributions of the amounts of temporal change of a plurality of elements in each of a plurality of face images into a plurality of expression intensities according to the magnitudes of the amounts of temporal change for each distribution, classifies the plurality of distributions corresponding to each expression intensity into a plurality of expression patterns according to the modes of combinations of the elements and the amounts of temporal change for each distribution, stores information for classifying the temporal distributions of the plurality of expression patterns corresponding to each of the plurality of face images into a plurality of atmosphere patterns according to the modes of the temporal distributions, a step of acquiring a target face image, and a step of outputting information on an atmosphere pattern corresponding to the face image based on the amount of temporal change of the elements in the target face image, and is a program for causing the information processing apparatus to execute the steps.
Advantages of the Invention
[0018] According to the operation method of the information processing apparatus in the present disclosure, etc., it becomes possible to quantitatively evaluate changes in facial expressions.
Brief Description of the Drawings
[0019] [Figure 1] It is a diagram showing a configuration example of an information processing system. [Figure 2] It is a flowchart showing an example of an operation procedure of a server device. [Figure 3] It is a flowchart showing an example of an operation procedure of a server device. [Figure 4A] It is a diagram for explaining expression elements. [Figure 4B] It is a diagram for explaining frontalization correction. [Figure 4C] It is a diagram for explaining a distance matrix. [Figure 4D] It is a diagram for explaining a distance matrix by time. [Figure 5] It is a diagram for explaining expression intensity. [Figure 6] It is a flowchart showing an example of an operation procedure of a server device. [Figure 7A] It is a diagram for explaining the principal components of a distance matrix by time. [Figure 7B]This figure illustrates the principal components of the time-dependent distance matrix. [Figure 7C] This figure illustrates the principal components of the time-dependent distance matrix. [Figure 7D] This figure illustrates the principal components of the time-dependent distance matrix. [Figure 7E] This is a diagram explaining facial expression patterns. [Figure 8] This is a flowchart illustrating an example of the operation procedure for a server device. [Figure 9A] This diagram illustrates the time distribution of facial expression patterns. [Figure 9B] This is a diagram illustrating atmosphere patterns. [Figure 10] This is a sequence diagram showing an example of the operation procedure of the information processing system in the embodiment. [Figure 11] This figure illustrates an example of an output image in the embodiment. [Figure 12] This is a sequence diagram showing an example of the operation procedure of the information processing system in the embodiment. [Figure 13] This figure illustrates an example of an output image in the embodiment. [Modes for carrying out the invention]
[0020] Embodiments of the present invention will be described below.
[0021] [System Configuration] Figure 1 shows an example configuration of one embodiment of the present invention. The information processing system 1 has a server device 10 and a terminal device 12 that are connected to each other via a network 11 so that they can communicate with each other. In the information processing system 1, the server device 10 performs machine learning using various information sent from the terminal device 12. The terminal device 12 is, for example, one or more personal computers. The terminal device 12 may include a tablet terminal device, a smartphone, etc. The server device 10 is, for example, one or more server computers. If the server device 10 is a single server computer, the server device 10 may be multiple server computers that coordinately execute the operations in this embodiment and provide cloud services. The network 11 is, for example, a LAN (Local Area Network), the Internet, an ad hoc network, a MAN (Metropolitan Area Network), a mobile communication network, or other networks, or any combination thereof.
[0022] The server device 10 acquires a face image, including the entire face of a person, obtained by capturing a person's face, from the terminal device 12. The face image is a moving image and has multiple frames that are continuous in time. As an "information processing device" in this embodiment, the server device 10 performs the following operations: The server device 10 performs a first step (hereinafter referred to as the "expression intensity classification step") which classifies the distribution of the amount of change over time (hereinafter simply referred to as the "amount of change distribution") of multiple elements (hereinafter referred to as "expression elements") in each of the multiple face images into multiple expression intensities according to the magnitude of the amount of change for each amount of change distribution, and a second step (hereinafter referred to as the "expression pattern classification step") which classifies the multiple amount of change distributions in each expression intensity into multiple expression patterns according to the manner of the combination of expression elements and amount of change for each amount of change distribution. Therefore, the server device 10 classifies changes in expression into expression intensity or expression pattern, making it possible to quantitatively evaluate changes in expression. Furthermore, the server device 10 performs a third step (hereinafter referred to as the "atmosphere classification step") in which it classifies the time distribution of multiple facial expression patterns corresponding to each of the multiple facial images (hereinafter referred to as the "facial expression distribution") into multiple atmosphere patterns according to the nature of the facial expression distribution. Therefore, since the server device 10 classifies the atmosphere caused by changes in facial expression, it becomes possible to quantitatively evaluate even phenomena such as atmospheres that are created by changes in facial expression and cannot be fully put into words.
[0023] Next, the configurations of the server device 10 and the terminal device 12 will be described.
[0024] The server device 10 includes a communication unit 101, a storage unit 102, a control unit 103, an input unit 105, and an output unit 106. These components are appropriately arranged in two or more server computers when the server device 10 is composed of two or more server computers.
[0025] The communication unit 101 includes one or more communication interfaces. These communication interfaces are, for example, LAN interfaces. The communication unit 101 receives information used in the operation of the server device 10 and transmits information obtained through the operation of the server device 10. The server device 10 is connected to the network 11 via the communication unit 101 and communicates information with the terminal device 12 via the network 11.
[0026] The storage unit 102 includes, for example, one or more semiconductor memories, one or more magnetic memories, one or more optical memories, or a combination of at least two of these, which function as main memory, auxiliary memory, or cache memory. The semiconductor memory is, for example, RAM (Random Access Memory) or ROM (Read Only Memory). The RAM is, for example, SRAM (Static RAM) or DRAM (Dynamic RAM). The ROM is, for example, EEPROM (Electrically Erasable Programmable ROM). The storage unit 102 stores information used in the operation of the control unit 103 and information obtained by the operation of the control unit 103.
[0027] The control unit 103 includes one or more processors, one or more dedicated circuits, or a combination thereof. The processors are, for example, general-purpose processors such as CPUs (Central Processing Units) or dedicated processors such as GPUs (Graphics Processing Units) specialized for specific processing. The dedicated circuits are, for example, FPGAs (Field-Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits). The control unit 103 controls each part of the server device 10 and performs information processing related to the operation of the server device 10.
[0028] The functions of the server device 10 are realized by the processor included in the control unit 103 executing a control program. The control program is a program that causes the processor to function as the control unit 103. In addition, some or all of the functions of the server device 10 may be realized by a dedicated circuit included in the control unit 103. Furthermore, the control program may be stored in a non-transient recording / storage medium readable by the control unit 103, and the control unit 103 may read it from the medium.
[0029] The input unit 105 includes one or more input interfaces. The input interfaces are, for example, physical keys, capacitive keys, pointing devices, touchscreens integrated with a display, or microphones that accept voice input. The input unit 105 accepts operations to input information used for the operation of the server device 10 and sends the input information to the control unit 103.
[0030] The output unit 106 includes one or more output interfaces. The output interfaces are, for example, a display or a speaker. The display is, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display. The output unit 106 outputs information obtained from the operation of the server device 10.
[0031] The terminal device 12 includes a communication unit 121, a storage unit 122, a control unit 123, an input unit 125, and an output unit 126.
[0032] The communication unit 121 includes a communication module compatible with wired or wireless LAN standards, a module compatible with mobile communication standards such as LTE, 4G, and 5G, etc. The terminal device 12 is connected to the network 11 via the communication unit 121 through a nearby router device or mobile communication base station, and communicates information with the server device 10, etc. via the network 11.
[0033] The storage unit 122 includes one or more semiconductor memories, one or more magnetic memories, one or more optical memories, or a combination of at least two of these. The semiconductor memory is, for example, RAM or ROM. The RAM is, for example, SRAM or DRAM. The ROM is, for example, EEPROM. The storage unit 122 functions, for example, as a main memory, auxiliary memory, or cache memory. The storage unit 122 stores information used in the operation of the control unit 123 and information obtained by the operation of the control unit 123.
[0034] The control unit 123 has, for example, one or more general-purpose processors such as a CPU or MPU (Micro Processing Unit), or one or more dedicated processors such as a GPU specialized for specific processing. Alternatively, the control unit 123 may have one or more dedicated circuits such as FPGAs or ASICs. The control unit 123 comprehensively controls the operation of the terminal device 12 by operating according to a control and processing program, or by operating according to an operating procedure implemented as a circuit. The control unit 123 then sends and receives various information with the server device 10, etc., via the communication unit 121 and executes the operations according to this embodiment.
[0035] The functions of the terminal device 12 are realized by the execution of a control program by the processor included in the control unit 123. The control program is a program that causes the processor to function as the control unit 123. In addition, some or all of the functions of the terminal device 12 may be realized by a dedicated circuit included in the control unit 123. Furthermore, the control program may be stored in a non-transient recording / storage medium readable by the control unit 123, and the control unit 123 may read it from the medium.
[0036] The input unit 125 includes one or more input interfaces. These input interfaces include, for example, physical keys, capacitive keys, a pointing device, and a touchscreen integrated with a display. The input interfaces may also include a microphone for receiving voice input and a camera for capturing images. Furthermore, the input interfaces may include a scanner or camera for scanning image codes and an IC card reader. The input unit 125 accepts operations to input information used in the operation of the control unit 123 and sends the input information to the control unit 123. The input unit 125 also sends images captured by the camera to the control unit 123.
[0037] The output unit 126 includes one or more output interfaces. The output interfaces include, for example, a display and a speaker. The display is, for example, an LCD or an organic EL display. The output unit 126 outputs information obtained by the operation of the control unit 123.
[0038] [Operation of Server Device 10] Figure 2 is a flowchart illustrating an example of the operation of the server device 10. Each step is performed by the control unit 103.
[0039] In step S20, the control unit 103 acquires a face image. The face image is a moving image generated by capturing a person's face. The face image includes, for example, 30 still images per second. For example, a face image captured by the terminal device 13, or a face image stored in the terminal device 13, is sent from the terminal device 13 to the server device 10. In the server device 10, the control unit 103 receives multiple face images via the communication unit 101 and stores them in the storage unit 102. Alternatively, the control unit 103 may acquire multiple face images from open data in response to an operation performed by the operator on the input unit 105, or an instruction sent from the terminal device 12 to the server device 10 in response to an operation performed by the operator on the input unit 125 of the terminal device 12.
[0040] In step S22, the control unit 103 performs a facial expression intensity classification step in the facial image. A detailed procedure for the facial expression intensity classification step is shown in Figure 3.
[0041] Figure 3 is a flowchart illustrating an example of the operation of the server device 10 in the facial expression intensity classification step. The procedure in Figure 3 is executed by the control unit 103.
[0042] The control unit 103 sequentially executes steps S32, S34, and S36 for all face images (step S30).
[0043] In step S32, the control unit 103 extracts feature points and corrects the feature points to be frontal view in each face image. Feature points include points on the contours connecting the eyebrows, eyes, and mouth, which move characteristically during changes in facial expression, as well as points on the contours of the bridge of the nose, the wings of the nose, and the contours of the chin. The control unit 103 extracts feature points for each frame image in the face image using arbitrary image recognition processing. For example, the control unit 103 extracts feature points Pn (where n is an arbitrary natural number) from the frame image 40 by landmark detection using a face feature extraction library, as shown in Figure 4A. The control unit 103 also corrects the arrangement of feature points Pn obtained when the face image is a face image viewed from an oblique direction to an arrangement corresponding to a face image viewed from the front by correcting the frontal view. For example, the control unit 103 uses landmark-based approaches to correct the arrangement example 41 of feature points Pn in a face image viewed from an oblique direction to an arrangement example 42 corresponding to a frontal view, as shown in Figure 4B. The storage unit 102 of the server device 10 stores information about the reference positions of feature points Pn in a front view face image, and the control unit 103 performs frontalization correction by bringing each feature point Pn closer to its respective reference position. Alternatively, the control unit 103 may perform frontalization correction using appearance-based frontalization or other algorithms.
[0044] In step S34, the control unit 103 derives an expression element for each face image. The expression element is the distance between a pair of feature points Pn. For example, as shown in Figure 4C, the control unit 103 generates a distance matrix 43 for the feature points Pn. The distance matrix 43 is an N×N matrix whose components are the distances between any two feature points obtained by exhaustive testing from the N feature points. In the distance matrix 43, the value of the component in row i and column j (where i and j are any natural numbers less than or equal to N) corresponds to the distance between feature points Pi and Pj, where the magnitude of the distance is represented by the grayscale intensity.
[0045] In step S36, the control unit 103 derives the distribution of changes in facial expression elements for each face image. The amount of change in facial expression elements is obtained, for example, as the difference in facial expression elements over a time difference of several tens of milliseconds. The control unit 103 derives the distance matrix 43 at different time points and calculates the difference between the two for each component, thereby deriving the time difference distance matrix 44 as shown in Figure 4D. The time difference distance matrix 44 is an N×N matrix, and the value of the component P(i,j) in row i and column j corresponds to the amount of change in distance over time between feature points Pi and Pj. That is, the time difference distance matrix 44 shows the distribution of changes for facial expression elements. Here, if the value of the component of the time difference distance matrix 44 at time t is LFMt, the amount of change in the distance between feature points Pi and Pj over a time difference T is expressed by the following equation. In TIFF0007897053000001.tif10159, the magnitude of the change is represented by the grayscale density. Note that even for a single face image obtained from capturing the same subject, different distributions of change will be derived at different time points.
[0046] The control unit 103 repeats the process for all face images until steps S32, S34, and S36 are completed (step S30).
[0047] In step S38, the control unit 103 classifies the distribution of changes in the facial expression elements in each of the multiple facial images into facial expression strengths according to the magnitude of the change for each change distribution.
[0048] First, the control unit 103 derives the sum of the time difference distance matrix 44 as the magnitude of the change for each change distribution, for example, as the facial expression metric (FEM) used to measure facial intensity, using the following formula. JPEG0007897053000002.jpg18157
[0049] Next, the control unit 103 classifies the probability distribution of the matrix sum FEMt into four expression intensities, such as strong expression, intermediate expression, weak expression I, and weak expression II. Any algorithm, reference values, etc., can be used for classification. Figure 5 shows an example of classifying the probability distribution obtained from approximately 10,000 time difference distance matrices 44 at 30-millisecond intervals into expression intensities using a box plot. Here, in the probability distribution where the upper outlier corresponds to the top 2.2%, the median to the top 50%, the first quartile to the top 25%, and the third quartile to the top 75%, the probability distribution 51 from the maximum value to the minimum value of the upper outlier is classified as a strong expression, the probability distribution 52 from the minimum value of the upper outlier to the first quartile is classified as an intermediate expression, the probability distribution 53 from the first quartile to the median is classified as a weak expression I, and the probability distribution 54 from the third quartile to the minimum value is classified as a weak expression II. The control unit 103 associates the type of facial expression intensity classified for each time difference distance matrix 44 and stores it in the storage unit 102. The control unit 103 also stores in the storage unit 102 the reference value used when classifying the matrix sum FEMt of each time difference distance matrix 44 into facial expression intensity.
[0050] As the control unit 103 executes steps S30 to S38 shown in Figure 3, the change amount distribution of multiple facial expression elements in each of the multiple facial images, i.e., the time difference distance matrix 44 of feature points Pn, is classified into multiple facial expression intensities according to the magnitude of the change, i.e., the matrix sum FEMt.
[0051] Returning to Figure 2, the control unit 103 completes step SS22 and proceeds to step S24, in which step S24 performs the facial expression pattern classification step. The detailed procedure for the facial expression pattern classification step is shown in Figure 6.
[0052] Figure 6 is a flowchart illustrating an example of the operation of the server device 10 in the facial expression pattern classification step. The procedure in Figure 6 is executed by the control unit 103.
[0053] The control unit 103 performs steps S62, S64, and S66 for each of the facial expression intensities: strong expression, intermediate expression, weak expression I, and weak expression II (step S60).
[0054] In step S62, the control unit 103 acquires the change amount distribution of the facial expression elements classified by the facial expression intensity to be processed. The control unit 103 reads one or more time difference distance matrices 44 from the storage unit 102 that are classified by the facial expression intensity to be processed from strong facial expression, intermediate facial expression, weak facial expression I, and weak facial expression II.
[0055] In step S62, the control unit 103 extracts the principal components of the change amount distribution of the facial expression elements. The principal components are those that contribute more to the magnitude of the time change for each change amount distribution than the others. In other words, the principal components of the time difference distance matrix 44 are components obtained by dimensionality reduction of component P(i,j) to components that have a relatively high contribution rate to the matrix sum FEMt.
[0056] In step S64, the control unit 103 classifies the principal components into facial expression patterns. The principal components represent the combinations of facial expression elements and change amounts in the time-difference distance matrix 44 and correspond to facial expression patterns. For example, the control unit 103 performs Gaussian Mixture Model (GMM) inference using collapsing Gibbs sampling on the principal components of the time-difference distance matrix 44 to classify the principal components into facial expression patterns.
[0057] Then, the control unit 103 performs steps S62, S64, and S66 for all facial intensity levels, and then terminates the procedure shown in Figure 6 (step S60).
[0058] Examples of facial expression patterns classified according to the procedure in Figure 6 are shown in Figures 7A to 7D. Figures 7A to 7D show examples of the distribution of principal components of the time-difference distance matrix 44 as facial expression patterns when approximately 70 feature points are extracted from the frame images 20 of the face image by landmark detection. Here, the magnitude of the principal components is represented by the grayscale density.
[0059] Figure 7A shows facial expression patterns 71-1, 71-2, and 71-3 in the time-difference distance matrix 44, which is classified as a strong facial expression. Facial expression pattern 71-1 is an example of the distribution of principal components with a contribution rate of 43.8%, and this principal component corresponds to the amount of change in facial expression elements over time associated with vertical movement of the mouth. Facial expression pattern 71-2 is an example of the distribution of principal components with a contribution rate of 21.3%, and this principal component corresponds to the amount of change in facial expression elements over time associated with nodding of the face and movement of the eyes. And facial expression pattern 71-3 is an example of the distribution of principal components with a contribution rate of 18.3%, and this principal component corresponds to the amount of change in facial expression elements over time associated with movement of the eyes and mouth.
[0060] Figure 7B shows facial expression patterns 72-1, 72-2, and 72-3 in the time difference distance matrix 44, which is classified as an intermediate facial expression. Facial expression pattern 72-1 is an example of the distribution of principal components with a contribution rate of 29.6%, and this principal component corresponds to the amount of change in facial expression elements over time associated with vertical movement of the mouth. Facial expression pattern 72-2 is an example of the distribution of principal components with a contribution rate of 26.5%, and this principal component corresponds to the amount of change in facial expression elements over time associated with nodding of the face and movement of the eyes. And facial expression pattern 72-3 is an example of the distribution of principal components with a contribution rate of 20.4%, and this principal component corresponds to the amount of change in facial expression elements over time associated with movement of the eyes and mouth.
[0061] Figure 7C shows facial expression patterns 73-1, 73-2, and 73-3 in the time difference distance matrix 44, which is classified as a weak expression I. Facial expression pattern 73-1 is an example of the distribution of principal components with a contribution rate of 26.9%, and this principal component corresponds to the amount of change in facial expression elements over time associated with vertical movement of the mouth. Facial expression pattern 73-2 is an example of the distribution of principal components with a contribution rate of 26.1%, and this principal component corresponds to the amount of change in facial expression elements over time associated with nodding of the face and movement of the eyes. And facial expression pattern 73-3 is an example of the distribution of principal components with a contribution rate of 18.6%, and this principal component corresponds to the amount of change in facial expression elements over time associated with movement of the eyes and mouth.
[0062] Figure 7D shows facial expression patterns 74-1, 74-2, and 74-3 in the time difference distance matrix 44, which is classified as weak expression II. Facial expression pattern 74-1 is an example of the distribution of principal components with a contribution rate of 21.9%, and this principal component corresponds to the amount of change in facial expression elements over time associated with vertical movement of the mouth. Facial expression pattern 74-2 is an example of the distribution of principal components with a contribution rate of 20.9%, and this principal component corresponds to the amount of change in facial expression elements over time associated with nodding of the face and movement of the eyes. And facial expression pattern 74-3 is an example of the distribution of principal components with a contribution rate of 16.7%, and this principal component corresponds to the amount of change in facial expression elements over time associated with movement of the eyes and mouth.
[0063] As shown in Figures 7A-7D, at different levels of facial intensity, facial patterns with common facial elements exhibit a degree of similarity in the distribution of their principal components. For example, the facial patterns 71-1, 72-1, 73-1, and 74-1, which correspond to vertical mouth movements, are similar to each other. Similarly, the facial patterns 71-2, 72-2, 73-2, and 74-2, which correspond to nodding and eye movements, are similar to each other. Furthermore, the facial patterns 71-3, 72-3, 73-3, and 74-3, which correspond to eye and mouth movements, are similar to each other.
[0064] In Figures 7A to 7D, three different facial expression patterns are shown for each facial expression intensity, but the number of facial expression patterns is not limited to the examples shown here. For example, if the probability distribution obtained from the approximately 10,000 time-difference distance matrices 44, each occurring every 30 milliseconds, shown in Figure 5, is classified into four facial expression intensities—strong, intermediate, weak I, and weak II—and GMM (Gaussian Mixture Model) inference using collapsing Gibbs sampling is performed on the principal components extracted from the time-difference distance matrices 44 for each facial expression intensity to classify them into facial expression patterns, then around 10 clusters were identified for each facial expression intensity. By first classifying the facial expression intensity and then classifying the facial expression patterns for each intensity, it becomes possible to suppress the influence of the matrix sum FEMt intensity on clustering accuracy, for example, in weak I and weak II.
[0065] As the control unit 103 operates according to the procedure shown in Figure 6, a classification model is generated for classifying the distribution of changes in facial expression elements classified as facial intensity into facial expression patterns. This classification model (hereinafter referred to as the facial expression pattern classification model) is stored in the memory unit 102.
[0066] Furthermore, when the control unit 103 classifies a face image into an expression pattern, it may output information of the expression pattern corresponding to the face image to the output unit 106. Figure 7E shows an example of a display image by the output unit 106. The display image 700 includes three face images 75-1, 75-2, and 75-3, each having two frame images showing changes in expression; time difference distance matrix images 76-1, 76-2, and 76-3 showing the expression pattern corresponding to each face image; and information 77-1, 77-2, and 77-3 on the expression intensity into which each expression pattern is classified. With such a display image 700, the operator can visually confirm the correspondence between face images, expression intensity, and expression patterns.
[0067] Alternatively, instead of step S66, the control unit 103 may display the distribution of principal components showing the combinations of facial expression elements and change amounts in the time difference distance matrix 44 using the output unit 106, and the operator may classify the distribution of principal components into facial expression patterns and input the classification results using the input unit 105. For example, the control unit 103 may display a display image 700 including a face image having frame images showing facial expression changes, as shown in Figure 7E, and an example of the distribution of principal components, and accept input of information for labeling each distribution example, and classify the distribution of principal components into facial expression patterns according to the labeling. In this case, the control unit 103 can perform machine learning using the labeled distribution of principal components as training data and generate a facial expression pattern classification model.
[0068] Returning to Figure 2, the control unit 103 completes step SS24 and proceeds to step S26, in which it performs atmospheric pattern classification. The detailed procedure for step S26 is shown in Figure 8.
[0069] Figure 8 is a flowchart illustrating an example of the operation of the server device 10 in the atmosphere pattern classification step. The procedure in Figure 8 is executed by the control unit 103.
[0070] The control unit 103 sequentially performs step S82 for all face images (step S80).
[0071] In step S82, the control unit 103 derives the facial expression distribution for each face image. Since the face images are moving images with a fixed playback time, when the facial expression intensity classification step and the facial expression pattern classification step are executed, a time difference distance matrix 44 is derived every predetermined time (e.g., 30 milliseconds) over the playback time of each face image, and the facial expression intensity corresponding to the sum FEMt of each matrix and the facial expression pattern corresponding to the distribution of the principal components at that facial expression intensity are derived. The control unit 103 derives the time distribution of the facial expression pattern over the playback time for each face image.
[0072] The control unit 103 repeats the process for all face images until step S82 is completed (step S80). An example of the resulting facial expression distribution is shown in Figure 9A. Figure 9A shows the time distribution of facial expression patterns classified into strong expression, intermediate expression, weak expression I, and weak expression II for 11 face images 901 to 911, using bar graphs normalized by the playback time of the face. Here, the bar graphs are represented by different hatching depending on the facial expression pattern.
[0073] In step S84, the control unit 103 classifies the facial expression distribution obtained from the multiple facial images into multiple atmosphere patterns according to the nature of the facial expression distribution. The control unit 103 classifies the facial expression distribution into multiple atmosphere patterns by unsupervised learning such as clustering. The control unit 103 may also classify the facial expression distribution into atmosphere patterns by assigning weights to the facial expression distribution according to the intensity of the facial expression. For example, stronger expressions that appear less frequently are assigned larger weights, and weaker expressions that appear more frequently are assigned smaller weights, thereby normalizing the facial expression distribution to some extent.
[0074] The control unit 103 operates according to the procedure shown in Figure 8, generating a classification model for classifying the facial expression distribution obtained from the face image into atmosphere patterns. This classification model (hereinafter referred to as the atmosphere pattern classification model) is stored in the memory unit 102.
[0075] Furthermore, when the control unit 103 classifies a face image into an atmosphere pattern, it may output information of the atmosphere pattern corresponding to the face image to the output unit 106. Figure 9B shows an example of a display image by the output unit 106. The display image 900 includes frame images 921 to 931 of 11 face images at any given time point, and bar graphs 901 to 911 showing the atmosphere pattern corresponding to each face image. This display image 900 allows the operator to visually confirm the correspondence between face images and atmosphere patterns.
[0076] When using machine learning with training data obtained by subjectively labeling facial images with the atmosphere they evoke, it is possible to classify according to the labeling used in the training data, but it is difficult to classify facial intensity, facial pattern, or atmosphere pattern that cannot be fully expressed in words. In this respect, according to this embodiment, changes in facial expression that cannot be fully expressed in words can be quantitatively classified into facial intensity, facial pattern, or atmosphere pattern, making it possible to quantitatively evaluate changes in facial expression.
[0077] [Example 1] Figure 10 is a sequence diagram illustrating an example of the operation of the information processing system 1 in an embodiment. The procedure in Figure 10 relates to the coordinated operation of a server device 10 and a terminal device 12, which have information for facial expression intensity classification and information for facial expression pattern classification (facial expression pattern classification model) obtained in at least the procedure of this embodiment. The terminal device 12 is used, for example, by a user to select a customer service representative who will provide various types of counseling. Such counseling may include, for example, advice on beauty. The counseling may be conducted in person or via online video call. The procedure in Figure 10 is performed, for example, when a user selects a customer service representative using the terminal device 12 in a store or at home prior to counseling.
[0078] In step S100, the terminal device 12 sends a request for a face image and facial expression information to the server device 10. The face image is a video image obtained by capturing the customer service staff member in advance and is stored in the server device 10. The facial expression information includes the facial expression intensity and facial expression pattern to which the face image is classified. In response to the user's operation on the input unit 125, the control unit 123 sends information for the request for the face image and facial expression information to the server device 10 via the communication unit 121. In the server device 10, the control unit 103 receives the information from the terminal device 12 via the communication unit 101.
[0079] In step S102, the server device 10 acquires a facial image. The control unit 103 reads the facial image of the customer service representative that has been previously stored in the storage unit 102.
[0080] In step S104, the server device 10 performs an expression intensity classification step and an expression pattern classification step for each face image. The control unit 103 derives a time difference distance matrix 44 at predetermined time intervals (e.g., 30 milliseconds) over the playback time of each face image. Furthermore, the control unit 103 derives the matrix sum FEMt of each time difference distance matrix 44 and derives the principal components. Furthermore, the control unit 103 identifies the expression intensity corresponding to each matrix sum FEMt. Then, the control unit 103 classifies each matrix sum FEMt into an expression pattern using an expression pattern classification model. The storage unit 102 may pre-store character information for describing the classified expression patterns.
[0081] In step S106, the server device 10 sends the output face image and facial expression information to the terminal device 12. The control unit 103 sends the face image and facial expression information to the terminal device 12 via the communication unit 101. In the terminal device 12, the control unit 113 receives the information from the server device 10 via the communication unit 121.
[0082] In step S108, the terminal device 12 displays a face image and facial expression information. The control unit 123 displays the face image and facial expression information indicating facial expression intensity and facial expression pattern via the output unit 126. The output unit 126 displays a display image 110, for example, as shown in Figure 11. The display image 110 includes images 111-1, 112-2, and 111-3 of two frames of the face images of each of the three customer service staff, images 110-1, 110-2, and 110-3 of time difference distance matrices representing the corresponding facial expression patterns, and strings 112-1, 112-2, and 112-3 indicating the respective facial expression patterns.
[0083] According to this embodiment, by viewing the displayed image 110, the user can recognize the face image of the customer service representative, the intensity of their facial expression, that is, the magnitude of facial movement associated with changes in expression, and the type of facial expression pattern, and thus be able to select a customer service representative that is more suitable to their preferences.
[0084] [Example 2] Figure 12 is a sequence diagram illustrating an example of the operation of the information processing system 1 in another embodiment. The procedure in Figure 12 relates to the coordinated operation of a server device 10 and a terminal device 12, which have information for facial intensity classification, information for facial pattern classification (classification model), and information for atmosphere pattern classification (classification model) obtained in the procedure of this embodiment. The terminal device 12 is used, for example, by a customer service representative or their supervisor who provides various counseling services. The procedure in Figure 12 is performed, for example, when a customer service representative or their supervisor evaluates the facial expressions of the customer service representative.
[0085] In step S120, the terminal device 12 captures an image of a subject and acquires a facial image. The control unit 123 of the terminal device 12 takes an image using the camera included in the input unit 125 in response to a user operation on the input unit 125. The subject is, for example, a customer service representative. As a result, the terminal device 12 acquires a facial image.
[0086] In step S122, the terminal device 12 sends a face image and a request for atmosphere information to the server device 10. The atmosphere information is information about the atmosphere pattern corresponding to the face image. In response to the user's operation on the input unit 125, the control unit 123 sends the face image and information for the atmosphere information request to the server device 10 via the communication unit 121. In the server device 10, the control unit 103 receives the information from the terminal device 12 via the communication unit 101.
[0087] In step S124, the server device 10 performs atmosphere pattern classification on the face images. The control unit 103 derives a time difference distance matrix 44 at predetermined time intervals (e.g., 30 milliseconds) over the playback time of each face image. Furthermore, the control unit 103 derives the matrix sum FEMt of each time difference distance matrix 44 and derives the principal components. Furthermore, the control unit 103 identifies the facial expression intensity corresponding to each matrix sum FEMt. Furthermore, the control unit 103 classifies the matrix sum FEMt into facial expression patterns for each facial expression intensity. Then, the control unit 103 classifies the facial expression distribution in each image into an atmosphere pattern using an atmosphere pattern classification model.
[0088] In step S126, the server device 10 sends the output face image and atmosphere information to the terminal device 12. The control unit 103 sends the face image and expression information to the terminal device 12 via the communication unit 101. In the terminal device 12, the control unit 113 receives the information from the server device 10 via the communication unit 121.
[0089] In step S128, the terminal device 12 displays a face image and atmosphere information. The control unit 113 displays the face image and expression information indicating the atmosphere pattern via the output unit 126. The output unit 126 displays a display image 130, for example, as shown in Figure 13. The display image 130 includes several frames of the customer service representative's face image 130-1, a bar graph 130-2 showing the expression distribution indicating the atmosphere pattern, and a string of characters 130-3 indicating the atmosphere pattern.
[0090] According to this embodiment, the customer service representative or their supervisor can recognize the customer service representative's facial image and the atmosphere patterns created by changes in their facial expressions by viewing the displayed image 120, making it possible to evaluate the customer service representative's atmosphere quantitatively and more objectively.
[0091] In this embodiment, instead of capturing a facial image in step S120, it is possible to acquire existing facial images from open source and evaluate the facial expression patterns, atmosphere patterns, etc., of the subject. By doing so, for example, it becomes possible to evaluate the facial expression patterns or atmosphere patterns of customer service staff from other companies and use this information to improve one's own facial expression patterns or atmosphere patterns.
[0092] As described above, this embodiment makes it possible to evaluate changes in facial expression as quantitative facial intensity and facial expression patterns. Furthermore, it makes it possible to quantitatively evaluate even the atmosphere patterns created by changes in facial expression that cannot be fully put into words.
[0093] In the above description, the server device 10 corresponds to the "information processing device". However, the server device 10 and the terminal device 12 may operate in conjunction to constitute the "information processing device", or the terminal device 12 may correspond to the "information processing device".
[0094] In the above-described embodiment, the processing and control program that defines the operation of the terminal device 12 may be stored in the storage unit 102 of the server device 10 or in the storage unit of another server device and downloaded to the terminal device 12 via the network 11, or it may be stored in a computer-readable non-transient recording and storage medium and read by the terminal device 12 from the medium.
[0095] As described above, embodiments have been explained based on various drawings and examples, but it should be noted that those skilled in the art will find it easy to make various modifications and alterations based on this disclosure. Therefore, it should be noted that these modifications and alterations are within the scope of this disclosure. For example, the functions, etc., included in each means, each step, etc., can be rearranged in a logically consistent manner, and multiple means, steps, etc., can be combined into one or divided. [Explanation of Symbols]
[0096] 10: Server equipment 11: Network 12: Terminal device 101, 121: Communications Department 102, 122: Storage section 103, 123: Control Unit 105, 125: Input section 106, 126: Output section 40: Frame image 43: Distance matrix 44: Time difference distance matrix Pn: Feature point
Claims
1. A method for operating an information processing device, The first step involves classifying the distribution of the time-dependent changes in multiple elements in each of multiple facial images into multiple facial expression intensities according to the magnitude of the time-dependent changes for each distribution. A second step involves classifying the multiple distributions at each facial intensity into multiple facial expression patterns according to the combination of the elements and the amount of change over time for each distribution. Includes, The second step further includes outputting information indicating the configuration of the combination of the elements and the amount of change over time for each distribution in each facial intensity, and receiving input information of a facial expression pattern corresponding to the configuration of the combination of the elements and the amount of change over time. How it works.
2. In claim 1, The first step includes correcting the amount of time change of the plurality of elements to correspond to a frontal view of the face in the corresponding face image. How it works.
3. A method for operating an information processing device, The first step involves classifying the distribution of the time-dependent changes in multiple elements in each of multiple facial images into multiple facial expression intensities according to the magnitude of the time-dependent changes for each distribution. A second step involves classifying the multiple distributions at each facial intensity into multiple facial expression patterns according to the combination of the elements and the amount of change over time for each distribution. The method further includes a third step of classifying the time distribution of the multiple facial expression patterns corresponding to each of the multiple facial images into multiple atmosphere patterns according to the nature of the time distribution. How it works.
4. In claim 3, In the third step, weights corresponding to the intensity of each facial expression are assigned to the multiple facial expression patterns. How it works.
5. In any of claims 1 to 4, The aforementioned element is the distance between pairs of feature points in the facial image. How it works.
6. A control unit that classifies the distribution of the time-varying amounts of multiple elements in each of multiple facial images into multiple facial intensity levels according to the magnitude of the time-varying amounts for each distribution, and classifies the multiple distributions in each facial intensity into multiple facial expression patterns according to the combination of the elements and the time-varying amounts for each distribution, An output unit that outputs information showing the combination of the elements and the amount of change over time for each distribution in each facial expression intensity, An input unit that receives input of information on facial expression patterns corresponding to the combination of the aforementioned elements and the amount of change over time, An information processing device having
7. In claim 6, When the control unit classifies the distribution of the time-varying amounts of the plurality of elements into the plurality of facial expression intensities, it corrects the time-varying amounts of the plurality of elements so that they correspond to a frontal view of the face in the corresponding face image. Information processing device.
8. In claim 6, An information processing device further having an output unit that outputs information of the facial expression pattern corresponding to each facial image.
9. The system includes a control unit that classifies the distribution of the time-varying amounts of multiple elements in each of the multiple facial images into multiple facial intensity levels according to the magnitude of the time-varying amounts for each distribution, and classifies the multiple distributions in each facial intensity into multiple facial expression patterns according to the combination of the elements and the time-varying amounts for each distribution, and the control unit further classifies the time distribution of the multiple facial expression patterns corresponding to each of the multiple facial images into multiple atmosphere patterns according to the nature of the time distribution, The system further includes an output unit that outputs information of the atmosphere pattern corresponding to each face image. Information processing device.
10. In claim 9, When the control unit classifies the plurality of atmosphere patterns, it assigns weights to the plurality of facial expression patterns according to the intensity of each expression. Information processing device.
11. In any of claims 6 to 10, The aforementioned element is the distance between pairs of feature points in the facial image. Information processing device.
12. By being executed by an information processing device, the information processing device will The first step involves classifying the distribution of the time-dependent changes in multiple elements in each of multiple facial images into multiple facial expression intensities according to the magnitude of the time-dependent changes for each distribution. A second step involves classifying the multiple distributions at each facial intensity into multiple facial expression patterns according to the combination of the elements and the amount of change over time for each distribution. Execute It is a program, The second step further includes outputting information indicating the configuration of the combination of the elements and the amount of change over time for each distribution in each facial intensity, and receiving input information of a facial expression pattern corresponding to the configuration of the combination of the elements and the amount of change over time. program.
13. In claim 12, The first step includes correcting the amount of time change of the plurality of elements to correspond to a frontal view of the face in the corresponding face image. program.
14. By being executed by an information processing device, the information processing device will The first step involves classifying the distribution of the time-dependent changes in multiple elements in each of multiple facial images into multiple facial expression intensities according to the magnitude of the time-dependent changes for each distribution. A second step involves classifying the multiple distributions at each facial intensity into multiple facial expression patterns according to the combination of the elements and the amount of change over time for each distribution. A third step involves classifying the time distribution of the multiple facial expressions corresponding to each of the multiple facial images into multiple atmosphere patterns according to the nature of the time distribution, Execute program.
15. In claim 14, In the third step, weights corresponding to the intensity of each facial expression are assigned to the multiple facial expression patterns. program.
16. In any of claims 12 to 15, The aforementioned element is the distance between pairs of feature points in the facial image. program.
17. A method for operating an information processing device, The information processing device stores information for classifying the distribution of the time-varying amounts of multiple elements in each of the multiple facial images into multiple facial intensity levels according to the magnitude of the time-varying amounts for each distribution, classifying the multiple distributions corresponding to each facial intensity into multiple facial patterns according to the combination of the elements and the time-varying amounts for each distribution, and classifying the time distribution of the multiple facial patterns corresponding to each of the multiple facial images into multiple atmosphere patterns according to the nature of the time distribution. Steps to acquire the target face image, The steps include outputting information about an atmosphere pattern corresponding to the face image based on the amount of time change of the elements in the target face image, A method of operation that includes this.
18. A storage unit that stores information for classifying the distribution of the time-varying amounts of multiple elements in each of multiple facial images into multiple facial intensity levels according to the magnitude of the time-varying amounts for each distribution, classifying the multiple distributions corresponding to each facial intensity into multiple facial patterns according to the combination of the elements and the time-varying amounts for each distribution, and classifying the time distribution of the multiple facial patterns corresponding to each of the multiple facial images into multiple atmosphere patterns according to the manner of the time distribution, A control unit that outputs information on an atmosphere pattern corresponding to an input facial image based on the amount of time change of the elements in the input facial image, An information processing device having
19. A program executed by an information processing device, The information processing device stores information for classifying the distribution of the time-varying amounts of multiple elements in each of the multiple facial images into multiple facial intensity levels according to the magnitude of the time-varying amounts for each distribution, classifying the multiple distributions corresponding to each facial intensity into multiple facial patterns according to the combination of the elements and the time-varying amounts for each distribution, and classifying the time distribution of the multiple facial patterns corresponding to each of the multiple facial images into multiple atmosphere patterns according to the nature of the time distribution. Steps to acquire the target face image, The steps include outputting information about an atmosphere pattern corresponding to the face image based on the amount of time change of the elements in the target face image, A program that causes the aforementioned information processing device to execute.