Information processing apparatus, supporting method, information processing method, and program
The information processing device classifies and extracts representative facial images to objectively reveal a person's unique expressions, aiding in self-discovery and cosmetic enhancement.
Patent Information
- Application Number
- JP2024082786
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-21
- Publication Date
- 2025-12-04
AI Technical Summary
Conventional technologies often extract facial images that reflect the subjective intentions of the device manufacturer, failing to discover a person's unique charms objectively.
An information processing device that classifies temporally consecutive facial images into groups based on similarity, extracts representative images from each group, and generates a display image to showcase various expressions, thereby uncovering the subject's inherent charms without bias.
The device enables the discovery of a person's unique and unnoticed facial expressions, enhancing self-awareness and improving the subject's impression through personalized cosmetic advice.
Smart Images

Figure 2025176551000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, a support method, an information processing method, and a program. [Background technology]
[0002] The inventors have been studying a method and a communication device for selecting "a person's typical, everyday, casual facial expression, an expression that one does not notice oneself making during conversation," with the aim of helping people discover new charms in themselves and improving their self-esteem and self-affirmation. However, when selecting a person's facial image, facial expressions that generally give a good impression are often extracted, and conventional technologies (see Patent Documents 1 and 2) have had the problem of easily extracting arbitrary facial images that reflect the intentions of the device manufacturer. Furthermore, a general cluster extraction method (see Patent Document 3) was insufficient in terms of efficiency.
[0003] Patent Document 1 describes a method for quantitatively evaluating changes in facial expressions that cannot be fully expressed in words by extracting feature points from multiple facial images of a person, and then using the distribution of changes in facial expression elements in the facial images after correcting the feature points to make them appear more forward-facing, and quantitatively classifying changes in facial expressions into expression intensity, expression pattern, or atmosphere pattern.
[0004] Patent Document 2 describes an information processing technology for estimating the position and orientation of an imaging device based on a captured image. Patent Document 2 describes a method for selecting feature points whose similarity with other nearby feature points is equal to or less than a threshold value, with the aim of enabling acquisition of feature points with a low risk of erroneous correspondence from among feature points extracted from a captured image.
[0005] Patent Document 3 describes an image search system for searching for images, which classifies image data in an image database into multiple clusters using normalized features of the image data, assigns a representative value to each cluster formed by clustering, and creates an index to be used for image search. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Japanese Patent Application Publication No. 2023-184309 [Patent Document 2] Japanese Patent Publication No. 2022-112168 [Patent Document 3] Japanese Patent Application Laid-Open No. 2010-250633 Summary of the Invention [Problem to be solved by the invention]
[0007] What is needed is to discover new charms of people that are free from the subjective or arbitrary intentions of the device maker.
[0008] An object of the present invention is to provide an information processing device, a support method, an information processing method and a program for discovering new charms of people without incorporating the arbitrary or subjective intentions of the creator. [Means for solving the problem]
[0009] An information processing device according to an aspect of the present invention includes a processing unit. The processing unit acquiring a plurality of temporally consecutive facial images of a subject; classifying the plurality of face images into a plurality of groups based on the similarity of a plurality of facial feature amounts extracted from each of the plurality of face images; extracting a plurality of groups at predetermined intervals from the plurality of groups in descending order of the number of face images belonging to each group as display groups; extracting a representative face image from the face images belonging to each group; A first display image is generated in which the representative face image extracted for each display group is displayed.
[0010] A support method according to one aspect of the present invention uses the information processing device to support an improvement in an impression of the subject.
[0011] An information processing method according to one aspect of the present invention is an information processing method executed by an information processing device, comprising: acquiring a plurality of temporally consecutively acquired facial images of a subject; classifying the plurality of face images into a plurality of groups based on similarities of a plurality of facial feature amounts extracted from each of the plurality of face images; extracting a plurality of groups at predetermined intervals from the plurality of groups in descending order of the number of face images belonging to each group as display groups; extracting, for each group, a representative face image from among the face images belonging to that group; and generating a first display image in which the representative face image extracted for each display group is displayed.
[0012] A program according to one aspect of the present invention includes: acquiring a plurality of temporally consecutively acquired facial images of a subject; classifying the plurality of face images into a plurality of groups based on similarities of a plurality of facial feature amounts extracted from each of the plurality of face images; extracting a plurality of groups at predetermined intervals from the plurality of groups in descending order of the number of face images belonging to each group as display groups; extracting, for each group, a representative face image from among the face images belonging to that group; generating a first display image in which the representative face image extracted for each display group is displayed; The information processing device executes the above. [Effects of the Invention]
[0013] According to the present invention, the arbitrary or subjective intentions of the creator are not included, and new charms of people can be discovered. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a diagram showing a state of beauty counseling using an information processing device according to a first embodiment of the present invention. [Figure 2] 1 is a functional block diagram of an information processing device (information processing system) according to a first embodiment of the present invention. [Figure 3] 10A and 10B are schematic diagrams illustrating generation of a first display image including facial images of a subject with various facial expressions, using a moving image of the facial image of the subject, by an information processing device according to an embodiment of the present invention. [Figure 4] 10A and 10B are schematic diagrams illustrating how a first display image including facial images of a target person with various facial expressions is generated from a moving image of facial images by an information processing device according to an embodiment of the present invention. [Figure 5] FIG. 10 is a flowchart illustrating a series of information processing steps involved in generating a first display image including facial images with various expressions and a second display image for counseling purposes, performed by an information processing device according to an embodiment of the present invention. [Figure 6] FIG. 10 is a flowchart illustrating the flow of a display group extraction process (process A) in the flow of the series of information processes. [Figure 7] FIG. 10 is a flowchart illustrating the flow of processing (process B) related to generating a second display image in the flow of the series of information processing described above. [Figure 8] 10A and 10B are examples of display images displayed on the display unit by the above series of information processing, where (A) shows a start image and (B) shows a setting image. [Figure 9] Examples of display images displayed on the display unit during the above series of information processing flows, where (A) shows a consent confirmation image, (B) shows an image for positioning the image for shooting, and (C) shows an image displayed during shooting. [Figure 10] (A) and (B) show the first display image displayed on the display unit in the flow of the above series of information processing. [Figure 11] The display images displayed on the display unit during the flow of the above series of information processing are (A) a first display image and (B) a second display image for confirmation. [Figure 12] (A) and (B) show the first display image displayed on the display unit in the flow of the above series of information processing. [Figure 13] The display images displayed on the display unit during the flow of the above series of information processing are (A) a first display image and (B) a second display image for confirmation. [Figure 14] (A) and (B) show the second display image displayed on the display unit in the flow of the above series of information processing. [Figure 15] FIG. 4 is a diagram illustrating a method for extracting a display group in the first embodiment. [Figure 16] FIG. 10 is a functional block diagram of an information processing system including an information processing device according to a second embodiment of the present invention. [Figure 17] (A) shows an example of three-dimensional facial feature points, and (B) shows an example of a limited portion of the face, such as the facial contour, eyebrows, eyes, nose, and mouth, extracted. [Figure 18] FIG. 10 is a diagram illustrating a method for extracting a display group in Modification 3. [Figure 19] 10(A) to 10(C) are diagrams illustrating methods for extracting display groups in Modifications 5 to 7, respectively. [Figure 20] 10(A) and 10(B) are diagrams illustrating a method for extracting a display group in Modifications 8 and 9, respectively. [Figure 21] FIG. 13 is a flowchart illustrating the flow of a display group extraction process (process A) in Modification 3. DETAILED DESCRIPTION OF THE INVENTION
[0015] <Summary of the Invention> An embodiment of the present invention will be described below with reference to the drawings. Here, an example is given in which a beauty consultant (beauty counselor) provides counseling to a counseling recipient (hereinafter simply referred to as the recipient) while referring to a facial image of the recipient as needed. The beauty consultant can provide the recipient with advice on, for example, selecting cosmetics (e.g., makeup products such as a makeup base, foundation, lipstick, blush, eye shadow, and mascara) that suit the recipient's facial structure, features, skin type, skin color, etc., and how to use the cosmetics to more attractively accentuate the recipient's facial features and expression, while referring to the facial image of the recipient displayed on a display as needed.
[0016] The information processing device of the present invention can discover new charms in a target person (person) without the arbitrary or subjective intention of the device manufacturer. Specifically, the information processing device can extract and present multiple facial images of a variety of expressions, including the target person's usual expressions, casual everyday expressions, rare expressions that appear during conversations that the target person does not notice, and expressions that reveal gestures that the target person does not notice, without the arbitrary or subjective intention of the device manufacturer. By presenting facial images of the target person with a variety of expressions, the target person can easily notice new expressions that make them look attractive, and beauty consultants can easily notice the target person's rare expressions that they tend to overlook even during face-to-face conversations, allowing them to provide advice on makeup and other aspects that further enhance the target person's charm. In this way, the information processing device can be used to help improve the impression made on the target person.
[0017] In this way, an information processing system including the information processing device of the present invention can present the subject and beauty consultant with a variety of casual expressions of the subject without any arbitrary or subjective intentions on the part of the system creator, making it possible to support the subject in becoming aware of charms that the subject was not aware of, and to provide counseling support to the beauty consultant by providing advice on the charms that the subject has become aware of as a new point of focus.
[0018] The following provides a detailed explanation. In the following first embodiment, an example is given in which a beauty consultant provides face-to-face counseling to a target person. In the second embodiment, an example is given in which a beauty consultant provides remote counseling to a target person via the Internet. Note that components already mentioned are given the same reference numerals, and explanations may be omitted.
[0019] First Embodiment
[0020] <<An example of face-to-face counseling>> As shown in FIG. 1, for example, a beauty consultant 9 can provide advice to a subject 8, while referring to the first facial image 101 and the second facial image 102 of the subject 8 on the second display image P2 displayed on the display unit 41 of the information processing device 1 as needed, on selecting cosmetics that suit the subject's skin type and skin color, how to use cosmetics to make the subject's facial expression more attractive, comparing the facial expression when using different products (fragrance, cream, etc.), and advice on gestures, hairstyles, facial expressions, etc. that suit an attractive expression even though the subject himself / herself is not aware of them.
[0021] The second display image P2, which includes the first face image 101 and the second face image 102, is a counseling image that the subject 8 and the beauty consultant 9 refer to during counseling. In this embodiment, an example is given in which an image showing the entire face is used as the facial image of the subject 8 included in the second display image P2. A facial image showing the entire face makes it easier to grasp the facial expression and overall mood of the subject. A variety of facial expressions can be created by combining the state of facial features such as the eyes, eyebrows, nose, and mouth, the distance between different facial features, the direction and tilt of the face, and the like.
[0022] In this embodiment, an example of using a full-face image showing the entire face of the subject as the "face image" to be referenced is given; however, a partial face image showing only a portion of the face may also be used. For example, facial parts such as the eyes and mouth are parts that reveal a person's facial expressions, and various facial expressions of the subject can be obtained from a partial face image including both eyes or a partial face image including the mouth. In this way, the "face image" to be referenced may be a full-face image, or a partial face image including parts of the face that are likely to show changes in the subject's facial expression. In actuality, when a person captures another person's facial expression, the use of a full-face image is preferable to the use of a partial face image, since the person grasps the whole picture.
[0023] In this embodiment, the first face image 101 and the second face image 102, which are the "face images" to be referenced, are face images selected by the subject 8 and the beauty consultant 9, respectively. In selecting the face image, the subject 8 and the beauty consultant 9 each select one face image from a plurality of face images of the subject. In the information processing device 1 of the present invention, an image including a plurality of face images of the subject with various facial expressions is generated as the first display image P1 used in selecting this face image. By presenting face images with various expressions, for example, the subject can learn expressions that he or she has not noticed before, and the beauty consultant can grasp attractive expressions of the subject that he or she might overlook even when directly interacting with the subject.
[0024] 1 shows an example of mutually orthogonal XYZ coordinate axes used to indicate the positions of three-dimensional facial feature points (described later). In this embodiment, the position of the imaging unit (camera) 5 is set as the origin, the Z axis (depth direction) is set as the distance direction between the position of the imaging unit 5 and the target person 8 (subject), the X axis is set as the horizontal direction in a plane perpendicular to the Z axis, and the Y axis is set as the vertical direction.
[0025] In this embodiment, as shown in FIG. 3 , the information processing device 1 generates a first display image P1 including a facial image 52 (hereinafter, sometimes referred to as a “representative facial image 52”) that represents various facial expressions of the subject 8 by using a captured video 50 of the subject's face and extracting a plurality of still facial images 51 (hereinafter, sometimes referred to as “facial images 51” (see FIG. 4 )) frame by frame from the captured video 50. The plurality of facial images 51 captured consecutively over time may include facial images with various facial expressions. As described above, various facial expressions can be created by combining various factors such as the facial direction, the state of facial features (eyes, eyebrows, nose, mouth, etc.), and changes in the distance between different facial features. For example, as shown in FIG. 3 , facial directions include facing forward, facing up, facing down, facing right, facing left, tilting (tilting the face due to tilting the head), or a combination of these directions (such as facing diagonally downward to the right), and facial expressions can change depending on the facial direction. Furthermore, facial expressions can also change depending on the magnitude of displacement of the face's orientation relative to when facing forward. The state of facial features, for example, in the case of the "eyes," includes the direction of gaze, the degree of eyelid opening, and the degree of eyelid widening, while in the case of the "mouth," includes the position of the corners of the mouth, the degree of mouth opening, and the degree of lip protrusion, and facial expressions can change depending on the state of these facial features. A change in the distance between each facial feature, for example, the distance between the "eye" feature and the "eyebrow" feature, can also change facial expressions. The information processing device 1 of the present invention is capable of extracting multiple such diverse facial expressions of the subject 8.
[0026] As shown in FIG. 4, when generating a first display image P1 including representative face images 52 of a subject 8 with various facial expressions, the information processing device 1 classifies the face images 51 into multiple (n) groups G, each group including face images 51 with similar features, based on facial feature amounts contained in the plurality of face images 51 acquired from a video 50 captured of the subject 8. The facial images 51 are grouped together based on facial feature amounts contained in the plurality of face images 51. In the example shown in FIG. 4, the face images 51 are classified into five groups G. The number of groups can be set arbitrarily in advance, and in this embodiment, the image data is classified into 100 groups. The information processing device 1 extracts a plurality of display groups from the 100 groups, extracts one representative face image 52 (representative face image 52) from the face images 51 belonging to each display group, and generates a first display image P1 including the representative face image 52.
[0027] In this embodiment, an example is given in which the feature amounts of three-dimensional facial feature points are used as the facial feature amounts. More specifically, an example is given in which the three-dimensional coordinate values of three-dimensional facial feature points are used as the facial feature amounts. Facial feature points are points that represent features related to the shape and expression of the face. Specifically, there are positions obtained by proportionally dividing the edge of a facial feature, the space between the edges of the same feature, and the position obtained by proportionally dividing a line connecting the edges of different features.
[0028] In this embodiment, the "plurality of facial images acquired continuously in time" is exemplified by an example in which a plurality of still facial images (facial images) are obtained by cutting out each frame from a video, but is not limited to this. For example, the "plurality of images acquired continuously in time" may be a plurality of still images (facial images) acquired by continuous shooting. Furthermore, the video used may be one acquired in real time during counseling, or a video that has been shot in advance.
[0029] <<Information processing equipment>> [Hardware configuration of information processing device] The information processing device 1 includes hardware necessary for a computer, such as a central processing unit (CPU), read-only memory (ROM), random access memory (RAM), a storage unit, and a communication unit. The ROM is a nonvolatile memory that stores firmware, such as an OS program and various parameters, to be executed by the CPU. The RAM is used as a working area for the CPU and temporarily stores the OS, various running applications, various data being processed, and the like. The storage unit is a nonvolatile memory, such as a hard disk drive (HDD), a solid state drive (SSD), or other solid-state memory. The storage unit stores the OS, various applications, various data, and the like. The communication unit is, for example, a network interface card (NIC) for Ethernet or various modules for wireless communication such as wireless LAN, and handles communication processing with other information processing devices. The CPU loads programs stored in the SSD or HDD into the RAM and executes them together with programs stored in the ROM as necessary, thereby controlling the operation of each functional block constituting the processing unit of the information processing device 1, which will be described later.
[0030] As the information processing device, for example, a PC (Personal Computer), a smartphone terminal, a tablet terminal, or the like is used, but any other computer may also be used.
[0031] The basic hardware configurations of the target person terminal 80, beauty consultant terminal 90, and server (information processing device) 1A included in an information processing system 100A of a second embodiment described later are approximately the same as the hardware configuration of the information processing device 1.
[0032] 2, the information processing system 100 includes an information processing device 1. The information processing device 1 includes a processing unit 2, an input unit 3, an output unit 4, an imaging unit 5, and a storage unit 6.
[0033] (Input section) The input unit 3 includes an operation device for accepting input operations from a user (in this embodiment, a beauty consultant or a subject) who uses the information processing device 1. The input unit 3 may also include an audio input unit for acquiring audio around the information processing device 1. The user interface of the input unit 3 is arbitrary, and may be, for example, a pointing device such as a mouse, a keyboard, a touch panel, or other input operation device. If the operation device of the input unit 3 is a touch panel, the touch panel may be integrated with the display unit 41 described later. The audio input unit is composed of a microphone and acquires audio around the user. The input unit 83 of the subject terminal 80 and the input unit 93 of the beauty consultant terminal 90 of the second embodiment described later are also similar to the input unit 3.
[0034] (output section) The output unit 4 includes a display unit 41 and an audio output unit. The display unit 41 has a function of outputting visual information such as images and text, and is equipped with a display device that presents visual information. The display device is, for example, a display using liquid crystal or organic EL (Electro-Luminescence). The display unit 41 displays the display image generated by the processing unit 2. The audio output unit is composed of a speaker and outputs audio. The output unit 84 and display unit 841 of the subject terminal 80 and the output unit 94 and display unit 941 of the beauty consultant terminal 90 of the second embodiment described below are also similar to the output unit 4 and display unit 41.
[0035] (imaging unit) The imaging unit 5 has a function of capturing images of the subject 8 and the surrounding environment of the information processing device 1, and generating still images or moving images. The imaging unit 5 includes an imaging element such as a CMOS (Complementary Metal Oxide Semiconductor) image sensor or a CCD (Charged Coupled Devices) image sensor for capturing images. The imaging unit 85 of the subject terminal 80 of the second embodiment, which will be described later, is similar to the imaging unit 5.
[0036] (Storage part) The storage unit 6 stores programs and the like for executing various processes performed by the processing unit 2. For example, the storage unit 6 stores programs that cause the information processing device 1 to execute an acquisition process, a feature extraction process, a classification process, a display group extraction process, a representative face image extraction process, and a display image generation process.
[0037] The program includes the steps of acquiring a plurality of facial images of a subject that are acquired consecutively over time; classifying the plurality of facial images into a plurality of groups based on the similarity of a plurality of facial features extracted from each of the plurality of facial images; extracting a plurality of groups at predetermined intervals from the plurality of groups as display groups in order of the number of facial images belonging to each group; extracting a representative facial image from the facial images belonging to each extracted display group for each of the extracted display groups; and generating a first display image in which the representative facial image extracted for each display group is displayed.
[0038] The storage unit 6 may also store a first display image P1 (see FIG. 3), a second display image P2 (see FIG. 1), etc. The storage unit 6 may also store, for each group G including a plurality of representative face images 52 (see FIG. 4) displayed in the first display image P1, a plurality of face images 51 belonging to the group G. These face images of the subject may be deleted after the counseling session is completed.
[0039] (Processing section) The processing unit 2 is configured using a CPU. In the information processing device 1, the CPU loads a program stored in the storage unit 6 into RAM and executes it, thereby performing processes related to generating a first display image including various facial images. The program may be stored in, for example, a non-transitory computer-readable recording medium and installed in the information processing device 1 using the recording medium. Alternatively, the program may be downloaded from a server on a network and installed in the information processing device 1. Examples of recording media for providing the program include magnetic disks such as hard disks and flexible disks, optical disks such as DVD-ROMs, CD-ROMs, and CD-Rs, and non-volatile memories such as USB memory and memory cards. By executing the program, the processing unit 2 functions as an acquisition unit 21, a feature extraction unit 22, a classification unit 23, a display group extraction unit 24, a representative facial image extraction unit 25, and a display image generation unit 26. The processing unit 2 also performs image display processing on the display unit 41, information storage processing in the storage unit 6, and drive control processing for the imaging unit 5. Note that FIG. 2 shows the main functional configuration of the processing unit 2 in a block diagram.
[0040] The processing unit 2 has an acquisition unit 21, a feature extraction unit 22, a classification unit 23, a display group extraction unit 24, a representative face image extraction unit 25, and a display image generation unit 26 as its functional blocks.
[0041] ((Acquisition Department)) The acquisition unit 21 acquires a video 50 of the face of the subject captured by the imaging unit 5. The video 50 of the face includes a plurality of facial images captured successively in time.
[0042] ((Feature Extraction Unit)) The feature extraction unit 22 extracts face images 51 for each frame from the captured video 50 acquired by the acquisition unit 21, and extracts facial features of the subject 8 for each face image 51 using the extracted face images 51. In this embodiment, the three-dimensional coordinate values (x, y, z coordinate values) of three-dimensional facial feature points are extracted as the facial features. z represents a depth value. In this embodiment, FaceMesh (a machine learning model) is used to obtain 468 three-dimensional key points (three-dimensional feature points) 55 for each face image, as shown in FIG. 17(A). The three-dimensional feature points 55 correspond to the vertices of a polygon mesh. The facial features are used by the classification unit 23 when classifying the faces into groups.
[0043] In this embodiment, the facial feature amounts are three-dimensional coordinate values (three-dimensional feature amounts) of a plurality of predetermined feature points of the face, including depth values. By using the three-dimensional coordinate values (position information) of the face, facial orientation information such as upward, downward, rightward, leftward, and diagonal orientations other than facing forward can be obtained, making it possible to acquire information on a variety of dynamic facial expressions. Note that the orientation information also includes information on the tilt of the face, i.e., tilt to the left or right (a state in which the center line of the face is tilted from the vertical direction).
[0044] Furthermore, in this embodiment, when extracting features, no processing such as correction of feature points to make them face forward is performed, so information on a variety of facial expressions with movement, including information on the direction of the face, can be obtained from multiple angles and in multiple aspects.
[0045] ((Classification Department)) The classification unit 23 classifies the plurality of face images 51 into n groups G (see FIG. 4 ) based on the facial feature amounts (three-dimensional coordinate values in this embodiment) included in each face image, with face images 51 with similar feature amounts grouped together. Each group G includes one or more face images 51. The number of groups n can be set arbitrarily in advance, and in this embodiment, an example is given in which the faces are classified into 100 groups. In this embodiment, the classification is performed using depth information of the face images, and each group G is classified based on differences in facial expressions that take into account face direction information.
[0046] Known methods can be used for group classification based on the similarity of feature quantities, which is executed by the classification unit 23. For example, K-means, hierarchical clustering, multidimensional scaling (distance matrix, correlation matrix), principal component analysis (summarizing features into a small number of dimensions, mapping facial images on the dimensions, and grouping similar images), etc. can be used. Examples of the K-means and hierarchical clustering include Euclidean distance, Manhattan distance, Mahalanobis distance, Ward's method, shortest distance method (nearest neighbor method), longest distance method (farthest neighbor method), centroid method, group average method, and median method. In this embodiment, the K-means method is used to simplify calculations when specifying the number of groups to be classified in advance.
[0047] ((Display group extraction part)) The display group extraction unit 24 extracts multiple groups as display groups DG at predetermined intervals from the multiple groups G classified by the classification unit 23 in descending or descending order of the number of facial images 51 belonging to the group G. The number of groups to be extracted is preferably two or more, more preferably four or more, from the viewpoint of the appearance of diverse facial expressions. Furthermore, from the viewpoint of ease of selection when selecting a facial image for counseling from the multiple representative facial images 52, the number is preferably 50 or less, more preferably 32 or less. In this embodiment, an example is given in which 30 display groups DG are extracted from 100 classified groups G. Note that in the case of extracting a number of display groups other than 30 (for example, 8, 16, 24, etc.), the numbers should be interpreted appropriately.
[0048] For the sake of convenience, FIG. 15 shows 100 groups G arranged from left to right in descending order of the number of facial images 51 belonging to each group G. In FIG. 15 and FIGS. 18 to 20 described below, groups extracted as display groups DG are shown filled in, and other groups (groups not extracted as display groups) are shown without filling in. In the example shown in FIG. 15, the display group extraction unit 24 extracts groups to become display groups DG at predetermined intervals from the multiple groups G arranged in descending order, extracting a total of 30 display groups DG. A specific example of the extraction of display groups DG will be described later. Note that the process of arranging the n groups G in descending order of the number of facial images 51 belonging to each group G may or may not actually be performed.
[0049] ((Representative face image extraction unit)) The representative face image extraction unit 25 extracts, for each display group DG extracted by the display group extraction unit 24, one representative face image (representative face image) 52 from among the face images 51 belonging to the display group DG.
[0050] There are no particular limitations on the method for extracting one representative face image 52 from the face images 51 belonging to each display group DG. When there is only one face image 51 belonging to each display group DG, the face image 51 is set as the representative face image 52. When there are multiple face images 51 belonging to each display group DG, the face image 51 to be the representative face image 52 may be extracted randomly, the face image 51 stored first in the group G to be classified may be extracted as the representative face image 52, or the face image 51 stored last in the group G to be classified may be extracted as the representative face image 52. Alternatively, the face image 51 to be the representative face image 52 may be extracted according to a specified numerical value. The "specified numerical value" may be, for example, the storage order of the face images within the group to be classified. For example, when the specified numerical value is 2, the second stored face image 51 is extracted as the representative face image 52.
[0051] In this embodiment, the representative face image 52 is extracted randomly from the viewpoint of being independent of the timing of face image acquisition, and from the viewpoint of more reliably ensuring objectivity by making it less likely that the creator's arbitrary or subjective intentions will be included.
[0052] ((Display image generation unit)) The display image generation unit 26 generates images to be displayed on the display unit 41. For example, the display image generation unit 26 generates a start image P10 (see FIG. 8(A)), a setting image P11 (see FIG. 8(B)), a photography consent confirmation image P12 (see FIG. 9(A)), a photography positioning image P13 (see FIG. 9(B)), an image P14 displayed during photography (see FIG. 9(C)), a first display image P1 (P1a to P1f) (see FIGS. 10(A) and (B) and FIGS. 11 to 13(A)) used as a face image selection image including a plurality of face images with various expressions, a second display image P2 (P2a to P2d) (see FIGS. 11(B), FIG. 13(B), FIG. 14(A) and (B)) used as a counseling image, and the like.
[0053] <<Information processing method>> A series of processes executed mainly by the processing unit 2 of the information processing device 1, from generating a first display image P1 including representative face images with various facial expressions to generating a second display image P2 using a face image selected from the representative face images included in the first display image P1, will be described with reference to Figures 5 to 14. In Figures 5 to 7, "ST" means "step."
[0054] The following will be explained according to the processing flow of Fig. 5, with reference to the processing flow of Fig. 6, the processing flow of Fig. 7, and examples of display images of Fig. 8 to Fig. 14. As shown in Fig. 8 to Fig. 14, the display content displayed on the display unit 41 changes according to the processing flow.
[0055] 10 to 13, the first display image in which a plurality of representative face images 52 are arranged is given the symbol P1. In some cases, a frame (61 or 62) surrounding the representative face image 52 is superimposed on the first display image in accordance with the processing flow, so that it is clear which representative face image 52 has been selected by the user. In FIGS. 10 to 13, images generated based on the first display image in which a plurality of representative face images 52 are arranged are given the symbols P1a to P1f, but when there is no particular need to distinguish between them, they are referred to as the first display image P1. Note that the first display image P1a shown in FIG. 10(A) corresponds to the base image in which a plurality of representative face images 52 are arranged.
[0056] 11, 13, and 14, the second display image in which the first face image 101 and the second face image 102, which are face images for counseling, are arranged is given the symbol P2. For example, a message or the like is superimposed on the second display image in accordance with the processing flow. In FIGS. 11, 13, and 14, the images generated based on the second display image in which the first face image 101 and the second face image 102 are arranged, and whose display content changes in accordance with the processing flow, are given the symbols P2a to P2d, but when there is no particular need to distinguish them, they are referred to as the second display image P2.
[0057] In this embodiment, an example will be described in which the display unit 41 of the information processing device 1 is a touch panel. In a series of processes from when the first display image P1 is generated until when the second display image P2 is generated, an input operation such as determining a face image on an image by the user can be performed by tapping a predetermined position on the screen with a finger. Note that the determination input operation may be, in addition to tapping the screen with a finger, key input on a keyboard (number, enter key), voice input, or eye gaze input, and is not limited thereto.
[0058] An information processing method will now be described. An application for performing the above series of processes is pre-installed in the information processing device 1. When the application is started by an input operation by a user of the information processing device 1, such as a beauty consultant 9 or a target person 8, the processing unit 2 displays a start image P10 on the display unit 41, as shown in FIG. 8(A) (ST1). The start image P10 displays a setting button and a start button, but is not limited to these, and any other buttons may be displayed.
[0059] When the setting button is selected by the user's input operation, the processing unit 2 displays a setting image P11 shown in Fig. 8(B) (ST2). The setting image P11 displays various condition setting items such as the subject's video shooting time (recording time), the number of clusters (the number of display groups to be extracted; this corresponds to the number of representative face images 52 displayed in the first display image P1), selection of still images or videos to be displayed on the display unit 41 during video shooting, items related to video playback, the color of the frame on the image and the option options, and items related to the questionnaire.
[0060] The video capture time is not particularly limited, but is preferably 20 seconds or more to capture a variety of facial expressions, and is preferably 120 seconds or less to capture the subject's natural facial expressions without boring them. The frame rate of the recorded video is not particularly limited, but is preferably at least one image per two seconds to capture a variety of facial expressions, i.e., at intervals of 2 seconds or less. To reduce image similarity, it is preferable to capture one image per 0.03 seconds or less, i.e., at intervals of 0.03 seconds or more. For example, in this embodiment, images are captured at intervals of 7 to 10 images per second, or approximately 0.15 to 0.1 seconds. The number of representative face images 52 displayed in the first display image is not particularly limited, but is preferably two or more, more preferably four or more, to capture a variety of facial expressions. To facilitate easy selection of face images for counseling from the multiple representative face images 52, the number of displayed images is preferably 50 or less, more preferably 32 or less. If the number of facial images displayed in the first display image is too small, it will be difficult to capture a variety of facial expressions. Conversely, if there are too many, similar images will be displayed across groups, or many images will be displayed that make it difficult to understand the user's facial expression (for example, images in which the face is turned extremely upside down or left or right, or is tilted significantly to the left or right and is out of the shooting area, making the selection more difficult and requiring the user to spend more time selecting an image.
[0061] When the user inputs various setting items and selects the OK button, the processing unit 2 acquires information related to the condition settings input on the setting image P11 (hereinafter referred to as "condition setting information") (ST3), stores the condition setting information in the storage unit 6, and displays the start image P10 shown in Fig. 8(A) again on the display unit 41 (ST4). Note that ST2 and ST3 do not have to be performed each time, and conditions that have been set in advance and stored in the storage unit 6 may be used.
[0062] When the start button on the start image P10 is selected by the user's input operation, the processing unit 2 displays the photography consent confirmation image P12 shown in Fig. 9(A) on the display unit 41 (ST5). Note that although an example in which the user selects the setting button immediately after starting the application has been given here, the user may also select the start button if, for example, there are no changes to the condition settings already stored. When the start button is selected, the process proceeds to ST5.
[0063] As shown in FIG. 9(A), the photographing consent confirmation image P12 may, for example, request permission to capture a facial image and display a notice that the image will be discarded at the end of the counseling session, but is not limited thereto. When the user selects the "Agree" button in the photographing consent confirmation image P12, the processing unit 2 activates the image capture unit 5 and generates a photographing positioning image P13 shown in FIG. 9(B) using images captured by the image capture unit 5, and displays the image on the display unit 41 (ST6). The photographing positioning image P13 is an image that includes an image in which a face positioning ellipse 53 is superimposed on a moving image of the subject 8 captured in real time by the image capture unit 5. The subject 8 can adjust their posture during video capture so that their face fits within the ellipse 53 displayed in the photographing positioning image P13. When the user selects the start button displayed in the photographing positioning image P13, the processing unit 2 starts recording a moving image including the face of the subject 8 using the image capture unit 5, and acquires the captured moving image 50 (recording information) (ST7). The recording time is set based on the video shooting time (recording time) set in the setting image P11. The face positioning ellipse 53 is not limited to an ellipse, and can be any shape such as a circle, square, or triangle. Even if the face is not completely within the ellipse, as long as 80% or more of the entire face is within the face positioning ellipse 53, the face image can be analyzed.
[0064] As shown in FIG. 9(C), during recording, the processing unit 2 does not display the facial image of the subject 8 on the display unit 41, but displays an image P14, such as a still image or a moving image, set in the setting image P11. By deliberately not displaying the facial image of the subject 8 during recording, it becomes easier to elicit a natural facial expression from the subject 8. While the facial image of the subject 8 being recorded may be displayed, the subject 8 may be conscious of posture and a good facial expression if his or her own face is displayed, making it difficult to elicit a natural facial expression from the subject 8. Therefore, it is preferable to display an image other than the facial image of the subject 8. In this embodiment, the image P14, such as a still image or a moving image, is selected and set in the setting image P11 as the image to be displayed on the display unit 41 during recording. However, instead of selecting and setting, it is also possible to configure the display unit 41 to display an image other than the image being captured, i.e., a predetermined still image or a moving image, that is determined by a program or preselected. Furthermore, during recording, the beauty consultant 9 may interact with the subject 8 to elicit a natural facial expression from the subject 8. Furthermore, the beauty consultant 9 may explain the product to the target person 8 based on the image displayed on the display unit 41 during recording, and may bring out the natural expressions that appear in casual conversations from the target person 8.
[0065] The processing unit 2 cuts out a facial image 51 for each frame from the recorded video in which the face of the subject is shown, and extracts a plurality of facial images 51 (ST8). The number of facial images to be extracted from the video is not particularly limited and may be set appropriately depending on the number of groups to be classified, but from the viewpoint of obtaining facial images with a variety of expressions, it is more preferable that the number be 140 or more, and from the viewpoint of reducing the discomfort of the subject by shortening the calculation processing, it is more preferable that the number be 3600 or less, and more preferably 840 or less.
[0066] Next, the processing unit 2 extracts facial features for each face image 51 (ST9). Specifically, in this embodiment, the processing unit 2 uses FaceMesh (machine learning model) to extract 468 three-dimensional key points (three-dimensional feature points) 55 from each face image 51, as shown in Fig. 17(A), and acquires the three-dimensional coordinate values of each three-dimensional feature point 55 as the three-dimensional feature amounts.
[0067] Next, the processing unit 2 classifies the plurality of face images 51 into, for example, 100 groups based on the similarity of the three-dimensional feature amounts of the faces included in each face image 51 (ST10). In detail, face images with high similarity of the three-dimensional feature amounts, i.e., similar three-dimensional feature amounts, are grouped together, and the plurality of face images are classified into 100 groups G. The groups are classified according to differences in facial expressions including facial orientation.
[0068] Next, the processing unit 2 extracts 30 display groups DG from the 100 groups (ST11). Specifically, a plurality of groups (30 in this embodiment) are extracted at predetermined intervals from the 100 groups in descending order of the number of facial images 51 belonging to the group G, and these extracted groups are designated as display groups DG. In this way, in this embodiment, a plurality of groups are extracted as display groups at predetermined intervals in descending order of the number of facial images 51 belonging to the group G. In other words, groups to be designated as display groups DG are intermittently extracted over the entire area or almost the entire area, from groups with a relatively large number of facial images 51 to groups with a relatively small number of facial images 51.
[0069] Here, it is considered that facial images belonging to a group with a relatively large number of facial images tend to be expressions that the subject 8 frequently shows, and are familiar to the subject 8 and people other than the subject 8. On the other hand, facial images belonging to a group with a relatively small number of facial images tend to be expressions that the subject 8 rarely shows, and are rare expressions that are difficult for the subject 8 and people other than the subject 8 to notice. Furthermore, facial images belonging to a group with a median or near the median number of facial images are considered to be expressions that appear with a moderate frequency. In other words, the display groups extracted in step 11 include not only groups of expressions that the subject 8 frequently shows, but also groups of expressions that the subject 8 rarely shows and groups of expressions that appear with a moderate frequency, and it is considered that in step 11, groups of diverse expressions of the subject can be selected as display groups.
[0070] The display group extraction process (process A) will be described in detail with reference to FIG. 15, following the flow of FIG.
[0071] For convenience of explanation, the processing unit 2 will be described as arranging the 100 groups G classified in step 10 in descending order of the number of facial images belonging to the group, as shown in FIG. 15. Here, an example of arranging in descending order will be given. In FIG. 15, the 100 (n) groups G are arranged from left to right so that the number of facial images belonging to the groups increases from "largest" to "smallest." As shown in FIG. 15, the processing unit 2 assigns group numbers 1, 2, 3, ..., 98 (n-2), 99 (n-1), and 100 (n) to each group in descending order of the number of facial images belonging to the group (ST20).
[0072] Next, the processing unit 2 extracts the group with group number 1 as the display start group number as one of the display groups (ST21).
[0073] Next, the processing unit 2 acquires display number information based on the condition setting information stored in the storage unit 6 (ST22). The display number information corresponds to the number of clusters set by the user's input operation on the setting image P11 acquired in step 3. Here, an example is given in which the number of representative face images (sometimes referred to as the display number) to be displayed in the first display image based on the display number information is m, the thinning number is a, b is an integer equal to or greater than 2, and the display number m is 30. The display number m is the number of representative face images 52 to be displayed in the first display image P1 and is the same as the number of display groups to be extracted. The thinning number is the thinning ratio. Specifically, a thinning number N means that one group is extracted from a group set consisting of N groups (N groups with consecutive group numbers in this embodiment).
[0074] Next, the processing unit 2 calculates the thinning number a using the following equation (1) (ST23). a=(n-2) / (m-1) (1)
[0075] Next, the processing unit 2 calculates the group numbers of the display groups DG (ST24, ST25). In this embodiment, an example is given in which a group G with group number 1 and a group G with group number 100 (n) are always extracted as the display groups DG. In other words, the group G with the largest number of facial images and the group G with the smallest number of facial images are always extracted as the display groups DG. When extracting the display groups, the group number of the group extracted first as the display group after the start of extraction is referred to as the "display start group number," and the group number used in determining whether or not to end the display group extraction process is referred to as the "display end group number." In this embodiment, the "display end group number" is the group number of the group extracted last. In this embodiment, the display start group number is "1," and the display end group number is "100," which is the same as the number of groups n. The display start group number and the display end group number can be set in advance. In this embodiment, the calculation of the group number of the group to be used as the display group DG using the above formula (1) and the formula (2) described below is the calculation of the group number of the group to be extracted as the display group after the group assigned group number 1.
[0076] Steps 24 and 25 will now be explained.
[0077] The processing unit 2 calculates the value to be extracted as the display group DG by performing so-called rounding processing, such as rounding down, rounding up, or rounding off, using the thinning number a calculated in step 23 according to the following equation (2). In this embodiment, an example is given in which decimal points are rounded down. Group number of display group = "Display start group number" + a × (b-1) (2)
[0078] In calculating the group number of the group to be designated as the display group DG, the processing unit 2 first substitutes 2 for b in the above formula (2) to calculate the group number of the group to be designated as the display group DG (ST24). Next, the processing unit 2 determines whether the number of the calculated group number exceeds the "display end group number," which in this embodiment is the number of groups n (ST25). If it is determined in step 24 that the number of the calculated group number exceeds the number of groups n (YES), the processing unit 2 ends the group number calculation process for the display group and proceeds to step 26. On the other hand, if it is determined in step 25 that the number of the calculated group number does not exceed the number of groups n (NO), the processing unit 2 returns to step 24, calculates the group number of the display group using the value obtained by adding 1 to the number used immediately before as the value to be substituted for b in the above formula (2), and proceeds to step 25. The processing unit 2 executes a loop process of repeating the processes of steps 24 and 25 until it is determined in step 25 that the number of the calculated group number exceeds the number of groups n (YES).
[0079] The group with the group number calculated immediately before the calculation in which it is determined in step 25 that the number of group numbers exceeds the number of groups becomes the group most recently extracted as a display group. In step 26, processing unit 2 extracts group G with group number 100 (n) as another display group, replacing group G last extracted as display group DG, terminates the display group extraction process (process A), and proceeds to step 12. In this embodiment, the group number of the group last extracted as a display group, calculated in step 24 using equations (1) and (2) above, is 99 (n-1), and in step 26, the group with group number 100 (n) is extracted as one of the display groups, replacing group G with group number 99.
[0080] In calculating the group numbers of the display groups DG in the display group extraction process (process A) (steps 24 and 25), the processing unit 2 calculates the group numbers of the display groups by substituting integers equal to or greater than 2 in ascending order, starting from 2, into formula (2), and continues calculating the group numbers of the display groups up to the point where the calculated display group number exceeds the value of n. As shown in Fig. 15, in this embodiment, the groups with group numbers 1, 4, 7, 11, 14, 17, 21, 24, 28, 31, 34, 38, 41, 44, 48, 51, 55, 58, 61, 65, 68, 71, 75, 78, 82, 85, 88, 92, 95, and 100 are extracted as display groups. In this manner, in this embodiment, the thinning number a is generally between 3 and 4 for group numbers 1 to n. In other words, one group is extracted as a display group for every 3 or 4 groups with consecutive group numbers, and the display groups are extracted at approximately equal intervals. Note that, although the display groups are extracted at approximately equal intervals here, depending on the number of groups to be classified and the value of the number of images to be displayed m, the thinning number a may always be the same and the display groups may be extracted at equal intervals.
[0081] As described above, in this embodiment, when extracting display groups from a plurality of groups, the groups with group numbers 1 and n are always extracted as display groups, and further, using the above formulas (1) and (2), display groups are extracted at approximately equal intervals from a plurality of groups arranged in descending order of the number of facial images 51 belonging thereto. In other words, groups to be used as display groups DG are extracted intermittently across the entire range (the entire range from group numbers 1 to 100), from groups with a relatively large number of belonging facial images 51 to groups with a relatively small number of belonging facial images 51. Note that the method of extracting display groups is not limited to this, and other extraction methods will be described later as modified examples 1 to 9.
[0082] In extracting display groups, when multiple groups are sorted in descending order of the number of facial images belonging to them, display groups whose number is less than the number of groups should be extracted at predetermined intervals across the entire area or almost the entire area from the group with the most to the group with the least number of facial images belonging to them.
[0083] In this specification, "at a predetermined interval" includes equal intervals, approximately equal intervals, and non-equal intervals. As an example of non-equal intervals, display groups may be extracted randomly using random numbers. Specifically, group numbers to be extracted as display groups are generated using random numbers based on the number of groups and the number of images to be displayed generated by the classification unit, and the group numbers of the groups to be extracted as display groups are determined by sorting the groups in descending order of the number of facial images belonging to the groups, and groups with group numbers determined by the generated random numbers are extracted.
[0084] In this embodiment, "extracting display groups at predetermined intervals" means "extracting display groups intermittently." Furthermore, in this specification, "extracting display groups from the entire area" means that when n groups are sorted in order of the number of facial images belonging to them, either in descending or descending order, and group numbers are assigned, groups with group numbers 1 and n are extracted as display groups, and display groups are extracted from the entire range of group numbers from 1 to n. Furthermore, in this specification, "extracting display groups from almost the entire area" refers to a range of groups with consecutive group numbers, which is equal to or greater than (n × 0.7) and less than n, when n groups are sorted in descending or descending order of the number of facial images belonging to each group and assigned group numbers. To give a specific example, extracting display groups including groups 11 and 80 from a range of groups with consecutive group numbers 11 to 80 (i.e., 70 groups with consecutive group numbers) out of 100 groups assigned group numbers 1 to 100 is "extracting display groups from almost the entire area." By setting the "display start group number" to 2 or greater and the "display end group number" to be smaller than the number n of groups to be classified, it is possible to "extract display groups from almost the entire area" using the above-described formulas (1) and (2). In this way, by intermittently extracting display groups from the entire range or almost the entire range, from the group with the most to the group with the fewest number of facial images, it becomes easier to obtain a variety of facial expressions, from those that the subject often makes to those that they rarely make.
[0085] In this embodiment, an example has been given in which decimal points are rounded down when calculating the group number using the above formula (2), but the group number can also be calculated by performing other rounding processes such as rounding up or rounding off the value obtained by the above formula (2), and the group number can be calculated in the same way.
[0086] When the display group extraction process (process A) is completed, the processing unit 2 selects, for each display group DG, one representative face image 52 from one or more face images belonging to the display group DG (ST12). In this embodiment, one representative face image 52 is selected from each of the 30 display groups, and a total of 30 representative face images 52 are selected.
[0087] Next, as shown in FIG. 10(A), the processing unit 2 generates a first display image P1 in which the selected 30 representative face images 52 are arranged in a matrix, and displays the first display image P1 on the display unit 41 (ST13). The first display image P1 is used to determine face images to be used in the second display image P2, which is a counseling image. The representative face images 52 displayed on the display unit 41 may be enlarged or reduced by a pinch operation by the user. Furthermore, all face images belonging to the display group DG including the selected representative face image 52 may be displayed by a predetermined operation, for example, a pinch-out operation after tapping.
[0088] Next, the processing unit 2 generates a second display image P2 in which a first face image 101 and a second face image 102 selected by users (the subject 8 and the beauty consultant 9) from among the plurality of representative face images 52 arranged in the first display image P1 are arranged (ST14). The second display image P2 arranges the first face image 101 selected by the subject 8, who is the first user (operator), and the second face image 102 selected by the beauty consultant 9, who is the second user (operator), and the first face image 101 and the second face image 102 are displayed simultaneously, allowing each user to easily compare the two.
[0089] Details of the second display image generation process (process B, step 14) will be described with reference to the flow in Fig. 7. Here, the two users (the subject and the beauty consultant) are referred to as the first user and the second user.
[0090] As shown in FIG. 10(A), the processing unit 2 displays the first display image P1a on the display unit 41. Four input operation buttons, "Retake," "Display Results," "First Selection," and "Second Selection," are displayed on the first display image P1a. When the processing unit 2 receives an input operation to select the "Retake" button, it returns to step 6 and repeats the process. In the example of the first display image P1a shown in FIG. 10(A), the processing unit 2 displays the "First Selection" button so that it flashes. The flashing of the "First Selection" button indicates that it is the first user's turn to select a face image.
[0091] When the first user taps and selects one of the representative face images 52 in the first display image P1a displayed on the display unit 41, the processing unit 2 acquires the selection input operation information from the first user (ST30), generates a first display image P1b based on the selection input operation information, and displays it on the display unit 41 as shown in FIG. 10(B) (ST31). The first display image P1b is a first display image in which a first frame 61 is displayed around the representative face image 52 selected by the first user, and a pop-up display including an enlarged image 63 of the selected representative face image 52 is displayed in the center of the screen. The display of the first frame 61 allows the first user to understand which representative face image 52 the enlarged image 63 is. The first frame 61 emphasizes the selected representative face image 52 so that it can be distinguished from the other unselected representative face images 52. Although the present embodiment provides an example of using a frame for emphasis, the present invention is not limited to this, and other techniques may be used for emphasis. For example, the selected representative face image 52 may be left unprocessed and displayed in color, while the other representative face images 52 may be processed and displayed in monochrome, or the selected face image may be left unprocessed and the other face images may be processed and displayed in a whitish overall manner, thereby highlighting the image. The same applies to the third frame 65 described below. Furthermore, by displaying the enlarged image 63, the first user can check the facial expression in more detail before selecting a face image for counseling. Note that if the selected representative face image 52 is in a position where it is hidden by the pop-up display, i.e., the enlarged image 63 is displayed overlaid, the position of the selected representative face image 52 can be shifted from the display position of the pop-up display to maintain the representative face image 52 selected by the user displayed in the first display image P1b on the display unit 41.
[0092] When the first user taps the "OK" button in the pop-up display to select and confirm, the processing unit 2 acquires information on the selection and confirmation input operation by the first user (ST32). Based on the selection and confirmation input operation, the processing unit 2 generates a first display image P1c in which a second frame 62 framing the selected and confirmed representative face image 52 is superimposed on an image in which a plurality of representative face images 52 are arranged, as shown in FIG. 11(A), and displays the first display image P1c on the display unit 41 (ST33). This first display image P1c is an image in which the selection result of the first user is reflected. By displaying the second frame 62 in the first display image P1c, the user can understand which representative face image the first user has selected and confirmed as the face image for counseling.
[0093] Next, when the user taps the "Display Results" button in the first display image P1c, the processing unit 2 acquires information about the user's operation of pressing the "Display Results" button (ST34), generates a second display image P2a based on the user's operation of pressing the "Display Results" button, and displays it on the display unit 41 as shown in FIG. 11(B) (ST35). The second display image P2a is a confirmation second display image reflecting the first user's selection result. The second display image P2a displays the first face image 101 for counseling, selected and confirmed by the first user, as "First Select," and a provisional second face image 102' that has not yet been selected and confirmed as "Second Select." The provisional second face image 102' is displayed in monochrome or in a whitish color to let the user know that it is undecided.
[0094] Next, when the user touches the "select" button on the First Select side in the second display image P2a and then the "OK" button, the processing unit 2 acquires OK button input operation information by the user (referred to as "OK input operation information") (ST36). Based on the OK input operation information, the processing unit 2 stores the representative face image selected and determined by the first user in the storage unit 6 as a face image for counseling set by the first user. Furthermore, as shown in FIG. 12(A), the processing unit 2 generates and displays a first display image P1d in which a second frame 62 framing the representative face image 52 selected and determined by the first user is superimposed on an image in which multiple representative face images 52 are arranged (ST37). In the example of the first display image P1d shown in FIG. 12(A), the processing unit 2 displays the "Second Selection" button so that it flashes. The flashing of the "Second Selection" button indicates that it is the second user's turn to select a face image.
[0095] Next, when the second user selects one representative face image 52 from the plurality of representative face images 52 in the first display image P1d, the processing unit 2 acquires the selection input operation information by the second user (ST38). The processing unit 2 generates a first display image P1e based on the selection input operation information and displays it on the display unit 41 as shown in Fig. 12(B) (ST39). The first display image P1e is a first display image in which the second frame 62 and a third frame 65 that frames the representative face image selected by the second user are displayed superimposed on each other, and a pop-up display of an enlarged image 64 of the representative face image 52 selected by the second user is displayed superimposed on the center of the screen.
[0096] Next, when the second user taps the "OK" button in the pop-up display to select and confirm, the processing unit 2 acquires selection and confirmation input operation information by the second user (ST40). Based on the selection and confirmation input operation, the processing unit 2 generates a first display image P1f in which a second frame 62 and a fourth frame 66 framing the representative face image 52 selected and confirmed by the second user are superimposed on an image in which a plurality of representative face images 52 are arranged, as shown in FIG. 13(A), and displays the first display image P1f on the display unit 41 (ST41). The first display image P1f is an image in which the selection results of the first and second users are reflected. The display of the fourth frame 66 in the first display image P1f allows the user to understand which face image the second user has selected and confirmed as the face image for counseling.
[0097] Next, when the user taps the "Display Results" button in the first display image P1f, the processing unit 2 acquires information about the user's operation to input the display results button (ST42), generates a second display image P2b based on the operation to input the display results button, and displays it on the display unit 41 as shown in FIG. 13(B) (ST43). The second display image P2b is a second display image for confirmation that reflects the selection results of the first and second users. The second display image P2b displays a first face image 101 for counseling that has been selected and determined by the first user as "First Select," and a second face image 102 for counseling that has been selected and determined by the second user as "Second Select."
[0098] Next, when the user touches the "select" button on the Second Select side in the second display image P2b and then the "OK" button, the processing unit 2 acquires OK input operation information by the user (ST44). Based on the OK input operation information, the processing unit 2 stores the representative face image selected and determined by the second user in the storage unit 6 as a face image for counseling set by the second user. Next, the processing unit 2 generates an image in which the first face image 101 and the second face image 102 determined by the first user and the second user, respectively, and set for counseling and stored in the storage unit 6 are arranged, and determines this as the second display image P2 for counseling to be referred to during counseling (ST45), thereby terminating the second display image generation process (process B). Subsequently, the process proceeds to step 15.
[0099] The processing unit 2 generates a second display image P2c by superimposing a message, for example, "We will suggest makeup that will enhance your attractiveness," on the second display image determined as the counseling display image in step 45, as shown in Fig. 14(A), and displays the generated second display image P2c on the display unit 41 (ST15). The beauty counseling is performed by the target person 8 and the beauty consultant 9, appropriately referring to the second display image P2c.
[0100] Typically, the first facial image 101 selected by the first user and the second facial image 102 selected by the second user, which are used in the second display image P2, are different images. In this way, two facial images with different facial expressions can be selected by different people selecting different facial images that they find attractive from among multiple facial images of a single person (subject) with various facial expressions. Therefore, the second display image P2 can display facial images with different facial expressions that reflect the subject's attractiveness from different perspectives. By providing beauty counseling while referring to such a second display image P2, for example, the subject can see the facial image selected by the beauty consultant and become aware of their own attractiveness that they had not noticed before. Furthermore, by referring to the two facial images for counseling with different facial expressions, the beauty consultant can more easily provide advice, such as on how to apply makeup to bring out the subject's attractiveness.
[0101] Note that, when the subject 8 and the beauty consultant 9 select the same representative face image, the processing unit 2 may generate the second display image P2 in which one representative face image selected by both of them is arranged. Also, the number of people selecting representative face images may be three or more in total, and the processing unit 2 may generate the second display image P2 in which one representative face image selected by each person is displayed. Also, in the present embodiment, an example has been given in which the subject and the beauty consultant each select one face image for counseling, but multiple images may be selected, and the number of face images selected by the subject and the beauty consultant may differ.
[0102] When the beauty consultation is completed and the user touches the "End" button displayed on the second display image P2c, the processing unit 2 generates a second display image P2d on which a message confirming the end is superimposed, as shown in FIG. 14(B), and displays the generated image on the display unit 41. When the user taps "Yes" on the message confirming the end to acknowledge the end, the processing unit 2 deletes the image of the subject based on the selection input operation (ST16) and ends the processing. Note that, as shown in FIG. 14(A), a "Survey" button may be displayed on the second display image P2c, and by selecting the "Survey" button, various surveys may be conducted for the subject. Furthermore, the facial image referenced by the subject may be output by printing, downloading, or the like.
[0103] <<Action and effect>> As described above, the processing unit of the information processing device of the present invention classifies a plurality of face images of a subject into a plurality of groups based on the similarity of facial features, and extracts a plurality of groups at predetermined intervals from the plurality of groups in descending order of the number of face images belonging to each group as display groups. Then, for each extracted display group, the processing unit extracts and displays a representative face image from the face images belonging to the display group.
[0104] According to this configuration, display groups are extracted at predetermined intervals from a plurality of groups in which facial images of a plurality of subjects are classified based on the similarity of facial features, in order of the number of facial images belonging thereto. This makes it possible to extract facial images of a variety of facial expressions of the subject, from expressions that the subject frequently displays to expressions that the subject rarely displays. The facial images of a variety of facial expressions extracted in this way are free of the arbitrary or subjective intentions of the device manufacturer, making it easier to capture the diverse facial expressions of the subject. Furthermore, by presenting facial images of a variety of facial expressions extracted in this way, the subject becomes more likely to notice attractive expressions that the subject normally does not notice, and beauty consultants become more likely to notice expressions that the subject rarely displays during conversations with the subject, which are easily overlooked.
[0105] <<Extracting groups for display>> As described above, the method of extracting a display group from a plurality of groups is not limited to the method described in the first embodiment. Modifications 1 to 9 will be described below as examples. In any of the modifications, display groups are extracted at predetermined intervals (extracted intermittently) over the entire area or almost the entire area, from the group with the largest number of facial images to the group with the smallest number, as in the first embodiment.
[0106] [Variation 1] In the first embodiment, an example was given in which the group with group number 1 (the group with the largest number of facial images) and the group with group number n (100 in the first embodiment) (the group with the smallest number of facial images) are always extracted as groups for display, but as in variant example 1, it is also possible to always extract the group with group number 1 as a group for display, and not extract the group with group number n as a group for display.
[0107] The display group extraction process in Modification 1 will be described in detail with reference to the processing flow in FIG. 6 as needed. In Modification 1, similar to the first embodiment, group numbers are assigned to a plurality of groups in descending order of the number of facial images belonging thereto (ST20), the group with group number 1 is designated as the display group (ST21), and the group numbers of the display groups are calculated using the above formulas (1) and (2) (ST23-25), but the process of step 26 is not performed. As a result, the group with the group number calculated in step 24 is extracted as the display group as is, the group with group number 1 is always extracted as the display group, and the group with group number n is not extracted as the display group. Also, the display start group number may be set to a value other than 1, for example, 3. Also, a display end group number may be determined as the last group number to be extracted, and when the result of rounding formula (2) reaches or exceeds the display end group number, a new calculation of formula (2) is terminated, i.e., the calculation of formula (2) is terminated. When the limit is exceeded, the group with the display end group number may be added as a display group, or the group number calculated just before the limit is exceeded may be the last display group. Also, as in the first embodiment, the group with the display end group number may be added to the display groups instead of the last display group calculated just before the limit is exceeded.
[0108] [Variation 2] In the first embodiment and the first modified example, group numbers are assigned to the groups in descending order of the number of facial images they contain, but group numbers may be assigned to the groups in ascending order of the number of facial images they contain. In this case, by extracting display groups in the same manner as in the first modified example, the group containing the fewest number of facial images (group number 1) is always extracted as a display group, and the group containing the most number of facial images (group number n) is prevented from being extracted as a display group.
[0109] Furthermore, a group number may be assigned to each group in order of the number of facial images belonging to it, and the display group may be extracted in the same manner as in the first embodiment described above. In this case, both the group with the fewest number of facial images belonging to it (group with group number 1) and the group with the most number of facial images belonging to it (group with group number n) will always be extracted as the display group.
[0110] [Variation 3] In the first embodiment, modified example 1, and modified example 2, when calculating the group number using the above formula (1) and formula (2), an example was given in which rounding processing such as rounding down, rounding up, or rounding to the nearest integer in the calculation of formula (2) was performed. In contrast, the thinning number a may be calculated by rounding down, rounding up, or rounding to the nearest integer in the calculation of formula (1), as will be described as modified example 3. Here, an example is given in which decimal points are rounded down, but other rounding processing such as rounding up or rounding to the nearest integer may also be used in the calculation.
[0111] In the first embodiment, modified example 1, and modified example 2, the calculation using formula (1) does not truncate decimal points, so the value calculated using formula (2) before the process of truncate decimal points can be an integer or a numeric value that includes a decimal point. For this reason, in the first embodiment, modified example 1, and modified example 2, there may be cases where the thinning number a is not constant, that is, where display groups are extracted at "almost equal intervals."
[0112] In contrast, in Modification 3, the value calculated by Equation (1) is truncated to the nearest integer, so the thinning number a becomes an integer, and the value calculated by Equation (2) also becomes an integer. Note that even if other rounding processes such as rounding up or rounding off the decimal point are performed on the value calculated by Equation (1), the thinning number a will also become an integer. In contrast, in Modification 3, by truncating, the thinning number a becomes constant, as shown in FIG. 18, and display groups are extracted at "equal intervals."
[0113] The display group extraction process in Modification 3 will be described in detail with reference to Fig. 21. In Fig. 21, the same steps as those shown in Fig. 6 are assigned the same step numbers. As shown in Fig. 21, in Modification 3, as in the first embodiment, the processing unit 2 assigns group numbers to a plurality of groups in descending order of number (ST20), sets the group with group number 1 as the display group (ST21), acquires display number information (ST22), and calculates the group numbers of the display groups using the above formulas (1) and (2). However, the calculation of the group numbers of the display groups differs from that in the first embodiment. Also, in Modification 3, the process of step 26 is not performed.
[0114] 21, the calculation of the group number of the display group in Modification 3 will be described. In Modification 3, in calculating the thinning number in step 23, the processing unit 2 calculates the thinning number a by using the above formula (1) and discarding the decimal point. Using this calculated thinning number a, the processing unit 2 calculates the group number of the group to be extracted as the display group according to the above formula (2) (ST24).
[0115] Next, the processing unit 2 determines whether the total number of extracted display groups is m (the value of the number of display images m of the representative face images) (ST27). "The total number of extracted display groups is m" means that extraction of the number of display groups to be extracted (the number of display images) has been completed. If the processing unit 2 determines in step 27 that the total number of display groups is not m (NO), the processing unit 2 returns to step 24, calculates the group number of the group to be used as the display group by adding 1 to the value used immediately before as the value to be substituted for b in the above formula (2), and proceeds to step 27. The processing unit 2 repeats the processes of steps 24 and 27 until it determines in step 27 that the number of display groups extracted is m (YES). If the processing unit 2 determines in step 27 that the total number of display groups extracted is m (YES), the processing unit 2 ends the display group extraction process (process A) and proceeds to step 12.
[0116] [Variation 4] In the first embodiment and Modifications 1 to 3, the thinning number a is determined using the above formula (1) and formula (2), and display groups are extracted based on the thinning number a. However, the thinning number a may be determined without using such a formula. For example, when the number of groups is set in advance as in the present embodiment, a table in which the number of images to be displayed and the thinning number a corresponding to the number of images to be displayed are associated may be prepared in advance and stored in the storage unit 6. The processing unit 2 may refer to the table stored in the storage unit 6 based on the number of images to be displayed (the number of clusters set in the setting image) input by the user, obtain the thinning number a associated with the same number of images to be displayed, and extract display groups using the thinning number a. Furthermore, when the number of groups is set by the user, a table in which the thinning number a corresponding to each combination of the number of groups and the number of images to be displayed is associated may be prepared. Note that the thinning number may be set so that display groups are extracted across the entire range or almost the entire range in the table, from the group with the most to the group with the fewest number of facial images. When display groups are extracted using such a table, display groups can be extracted at equal intervals.
[0117] [Variation 5] A group consisting of a plurality of groups having consecutive group numbers, including the group having the largest number of facial images and / or the group having the smallest number of facial images, may be extracted as a display group. The number of groups constituting a group may be set in advance or may be set by the user. For example, in the example shown in FIG. 19(A), the number of groups constituting a group is set to three, and three groups with consecutive group numbers, group number 1 (the group number of the group having the largest number of images), group number 2, and group number 3, form one group group. Furthermore, three groups with consecutive group numbers, group number 100 (the group number of the group having the largest number of images), group number 99, and group number 98, form another group group. Groups belonging to these two groups are extracted as display groups, and within the range of groups not belonging to the group groups (groups with group numbers 4 to 97 in the example shown in the figure), display groups may be extracted with a thinning number of, for example, 3. 19(A), display groups are extracted with a thinning number of 1 for group range (α) of group numbers 1 to 3 belonging to one group group and group range (α) of group numbers 98 to 100 belonging to another group group, and with a thinning number of 3 for group range (β) of group numbers 4 to 97. When extracting display groups within the range of group numbers 4 to 97, the group numbers may be calculated by applying the above formulas (1) and (2) mutatis mutandis.
[0118] In this way, a plurality of groups with consecutive group numbers may be extracted as display groups. Also, the number of thinning-out operations may be changed to extract display groups.
[0119] 19(A) to 19(C), "α" indicates a group range in which the thinning number is 1, and "β" indicates a group range in which the thinning number is an integer of 2 or more other than 1.
[0120] [Variation 6] In the fifth modification, an example was given in which the groups to be displayed are formed so as to include the group with the largest number of images and / or the group with the smallest number of images, but the groups may be formed so as not to include the group with the largest number of images and / or the group with the smallest number of images. For example, as in the sixth modification shown in FIG. 19(B), a plurality of groups with consecutive group numbers may be extracted from a plurality of groups with a relatively large number of images and used as a group, and a plurality of groups with consecutive group numbers may be extracted from a plurality of groups with a relatively small number of images and used as a group. In this case, the number of groups to be included in the group and the group numbers from which the groups are to be consecutively selected to form the group may be set in advance or may be set by the user.
[0121] In the example shown in FIG. 19(B), three groups with group numbers 3 to 5 form one group group, and three groups with group numbers 96 to 98 form another group group. Groups belonging to these two group groups are extracted as display groups. In a group range with consecutive group numbers that do not belong to a group group (in the example shown in the figure, group range (β) of group numbers 6 to 95), display groups may be extracted with a thinning number of 4, for example. In the example shown in FIG. 19(B), display groups are extracted with a thinning number of 1 for group range (α) of group numbers 3 to 5 that belong to one group group and group range (α) of group numbers 96 to 98 that belong to another group group, and with a thinning number of 4 for group range (β) of group numbers 6 to 95. When extracting display groups within the range of group numbers 6 to 95, the group numbers may be calculated by applying the above formulas (1) and (2) mutatis mutandis.
[0122] In this way, the group with the most number of images and the group with the least number of images may not be extracted as display groups. Alternatively, multiple groups with consecutive group numbers may be extracted as display groups. Alternatively, the number of thinning-out operations may be changed to extract display groups.
[0123] [Variation 7] In the fifth and sixth variants, examples were given in which a display group was extracted by setting the thinning number to 1 for a group range of a plurality of consecutive groups with a relatively large number of images and / or a group range of a plurality of consecutive groups with a relatively small number of images, and by setting the thinning number to a value other than 1 for a group range of a plurality of consecutive groups with a median number of images. In contrast to this, as in the seventh variant shown in Fig. 19(C), a display group may be extracted by setting the thinning number to a value other than 1 (the thinning number is 3 in Fig. 19(C)) for a group range (β) of a plurality of consecutive groups with a relatively large number of images and / or a group range (β) of a plurality of consecutive groups with a relatively small number of images, and by setting the thinning number to 1 for a group range (α) of a plurality of consecutive groups with a number close to the median.
[0124] In this way, a plurality of groups with consecutive group numbers that are near the median number of sheets may be extracted as display groups. Also, the number of thinning-out operations may be changed to extract display groups.
[0125] [Variation 8] In the fifth to seventh modified examples, a plurality of groups with consecutive group numbers are extracted as display groups from a plurality of groups arranged in descending order of the number of facial images belonging to them, and display groups are extracted intermittently in other portions, and the number of thinning-out operations is varied. In contrast to this, as in the eighth modified example shown in Fig. 20(A), display groups may be extracted by changing the number of thinning-out operations so that groups with consecutive group numbers are not extracted as display groups.
[0126] In FIGS. 20A and 20B, "γ" and "δ" indicate group ranges whose thinning number is other than 1, and the thinning number for the group range of γ is greater than the thinning number for the group range of δ.
[0127] As shown in FIG. 20A, for example, the multiple groups are divided into a group range (δ) of multiple consecutive groups with a relatively large number of images, a group range (γ) of multiple consecutive groups with a number of images near the median, and a group range (δ) of multiple consecutive groups with a relatively small number of images. The group numbers belonging to each range can be set arbitrarily. In the example shown in FIG. 20A, display groups are extracted using a thinning number of 4 for group ranges with a relatively large number of images and group ranges with a relatively small number of images, and display groups are extracted using a thinning number of 2 for group ranges with a number of images near the median. When extracting display groups for each range, the group numbers may be calculated using the above formulas (1) and (2). In this way, display groups may be extracted by changing the thinning number depending on the number of images included in the group, thereby obtaining representative face images containing a variety of facial expressions.
[0128] [Variation 9] In the eighth modification, an example was given in which display groups are extracted so that the thinning number for group ranges with a relatively large number of images and group ranges with a relatively small number of images is smaller than the thinning number for group ranges with a number near the median. In contrast, as shown in the ninth modification in Fig. 20(B), display groups may be extracted so that the thinning number for group ranges (γ) with a relatively large number of images and group ranges (γ) with a relatively small number of images (thinning number 4 in the illustrated example) is larger than the thinning number for group ranges (δ) with a number near the median (thinning number 2 in the illustrated example). In this way, display groups may be extracted by changing the thinning number, and representative face images containing a variety of expressions can be obtained.
[0129] <Regarding the placement of the representative face image in the first display image> In the first display image P1, the order in which the multiple (30 in the first embodiment) representative face images 52 are arranged is not particularly limited. For example, they can be arranged randomly, arranged in descending order of the number of face images belonging to the group including each representative face image 52, arranged in descending order of the number of face images belonging to the group including each representative face image 52, arranged in order of shooting time, arranged in descending order of similarity based on the feature amount of the representative image of the group with the largest number of face images, or arranged in descending order of similarity based on the feature amount of an average face image obtained by averaging all captured images.
[0130] The arrangement of the multiple representative face images 52 in the first display image P1 is not particularly limited. For example, all of the representative face images 52 may be arranged on one screen page, or the representative face images 52 may be arranged across multiple pages so that the representative face images can be viewed by scrolling. From the perspective of ease of selection of a representative face image, it is preferable that the size of each representative face image be large enough to intuitively recognize the differences between different representative face images, in other words, the differences in facial expressions. The number of face images arranged on one screen page may be adjusted depending on the size of the display unit 41. Furthermore, in the first display image P1, the multiple representative face images 52 may be arranged in a matrix as shown in FIG. 13(A), or may be arranged in a spiral or radial pattern, such as arranged from the center of the screen outward. Furthermore, the shape of the image group consisting of the multiple arranged representative face images may be any shape, such as a circle, a rectangle, or a triangle.
[0131] The representative face images 52 may be arranged, for example, from right to left on the screen, from top to bottom on the screen, or from the center of the screen to the outside.
[0132] In the first embodiment, in which an example of how the representative face images 52 are arranged is shown in FIG. 13A, the extracted representative face images 52 are arranged in a matrix in a predetermined order in descending order of the number of face images belonging to the display group to which the representative face images 52 belong. The "predetermined order" refers to an order from left to right, top to bottom, such that the representative face images are arranged from the top left or right, and after arranging in that row, the representative images are arranged from left to right in the next row. In this case, the representative face image arranged at the top and leftmost position is the representative face image of the display group with the largest number of face images, and the representative face image arranged at the bottom and rightmost position is the representative face image of the display group with the smallest number of face images. By using this arrangement, for example, when a beauty consultant 9 is unsure which face image to select, selecting the representative face image arranged at the top and leftmost position makes it easier to select a face image with an expression that is familiar to the subject 8 and frequently used by the subject 8. This allows the beauty consultant 9 to provide beauty counseling to the target person 8 by referring to the facial image of the target person's usual facial expression.
[0133] Second Embodiment In the second embodiment, an example is given in which the present invention is used for remote counseling, i.e., counseling where the counselor and the recipient are in different locations (remote counseling). In the first embodiment, an example of face-to-face counseling was given, but the present invention can also be applied to remote counseling as in this embodiment, and similar effects can be obtained. The information processing system of this embodiment will be described below with reference to FIG. 16.
[0134] 16, the information processing system 100A includes a target person terminal 80, a beauty consultant terminal 90, and a server 1A as an information processing device. In the information processing system 100A, beauty counseling can be performed without restrictions on the locations of the target person and the beauty consultant.
[0135] The server 1A as an information processing device is a web server that provides a service of presenting various facial expressions of a target person to a user. The server 1A is connected to a target person terminal 80 and a beauty consultant terminal 90 via the Internet. The target person terminal 80 and the beauty consultant terminal 90 may also be connected via the Internet.
[0136] The subject terminal 80 and the beauty consultant terminal 90 are information terminals used by the subject and the beauty consultant, respectively, and are, for example, a smartphone, a mobile phone, a tablet PC (Personal Computer), a notebook PC, a desktop PC, or the like.
[0137] [Target device] The subject terminal 80 includes a communication unit 81 , a control unit 82 , an input unit 83 , an output unit 84 , and an imaging unit 85 .
[0138] The communication unit 81 transmits and receives data to and from the server 1A. For example, the communication unit 81 receives from the server 1A various images including the first display image P1 and the second display image P2 generated by the server 1A, and audio information of the beauty consultant acquired by the server 1A from the beauty consultant terminal 90. The communication unit 81 also transmits to the server 1A images (captured video or still images) of the subject's face captured by the imaging unit 85 of the subject's terminal 80, input operation information performed by the subject on the subject's terminal 80 (e.g., input operation information related to determining a face image for counseling), audio information acquired by the microphone (audio input unit) of the subject's terminal 80, and the like. In this embodiment, an example is given in which audio information acquired by each of the subject's terminal 80 and the beauty consultant terminal 90 is transmitted to the other terminal via the server 1A. However, audio information may also be transmitted and received between the subject's terminal 80 and the beauty consultant terminal 90 directly or via various call services without going through the server 1A.
[0139] The input unit 83 includes an operation device for accepting input operations on the subject terminal 80 by the subject, and an audio input unit (microphone) for acquiring audio around the subject terminal 80. The output unit 84 includes a display unit 841 and an audio output unit (speaker). The imaging unit 85 is capable of acquiring a facial image of the subject. The control unit 82 controls the display unit 841, imaging unit 85, etc. for image display, and performs control calculations, etc.
[0140] [Beauty consultant terminal] The beauty consultant terminal 90 includes a communication unit 91, a control unit 92, an input unit 93, and an output unit 94. The beauty consultant terminal 90 may also include an imaging unit.
[0141] The communication unit 91 transmits and receives data to and from the server 1A. For example, the communication unit 91 receives from the server 1A various images including the first display image P1 and the second display image P2 generated by the server 1A, and voice information of the subject acquired by the server 1A from the subject's terminal 80. The communication unit 91 also transmits to the server 1A input operation information performed by the beauty consultant on the beauty consultant terminal 90 (for example, input operation information related to setting conditions, input operation information related to determining a face image for counseling, etc.), voice information acquired by the microphone (voice input unit) of the beauty consultant terminal 90, etc.
[0142] The input unit 93 includes an operation device for receiving input operations made by the beauty consultant on the beauty consultant terminal 90, and an audio input unit (microphone) for acquiring audio around the beauty consultant terminal 90. The output unit 94 includes a display unit 941 and an audio output unit (speaker). The control unit 92 controls the display unit 941 for image display, performs control calculations, etc.
[0143] [Server (information processing device)] The server 1A includes a communication unit 7 and a processing unit 2.
[0144] The communication unit 7 transmits and receives data between the subject terminal 80 and the beauty consultant terminal 90. The communication unit 7 transmits various images generated by the processing unit 2 to the subject terminal 80 and the beauty consultant terminal 90. The various images generated by the processing unit 2 may, for example, be transmitted with the same content to the subject terminal 80 and the beauty consultant terminal 90 at approximately the same time and displayed on the display units 841 and 941 of each terminal, or the subject and the beauty consultant may share the same images. The communication unit 7 receives input operation information at each terminal, audio information acquired at each terminal, and the like from the subject terminal 80 and the beauty consultant terminal 90. The communication unit 7 receives images (captured videos or still images) of the subject from the subject terminal 80.
[0145] The main processes executed by the processing unit 2 will now be described. The processing unit 2 acquires from the subject terminal 80 a plurality of facial images (photographed images, etc.) of the subject that are acquired successively in time. Next, the processing unit 2 classifies the multiple facial images into multiple groups using a method similar to that of the first embodiment, extracts display groups from the classified multiple groups, and generates a first display image P1 using the representative facial images 52 extracted for each display group. Next, the processing unit 2 transmits the generated first display image to the subject terminal 80 and the beauty consultant terminal 90 via the communication unit 7. The first display image is displayed on the display units 841 and 941, and the subject and the beauty consultant use the first display image to perform input operations related to the selection and determination of a face image for counseling.
[0146] Next, processing unit 2 acquires input operation information from each of target person terminal 80 and beauty consultant terminal 90, and generates second display image P2 using a total of two face images for counseling set by the target person and beauty consultant, respectively, based on the input operation information, and transmits the second display image P2 to target person terminal 80 and beauty consultant terminal 90. Second display image P2 is displayed on display units 841 and 941 of each terminal.
[0147] The second display image P2 may display facial images with different expressions selected by the subject and the beauty consultant. By referring to the second display image P2, the subject may become aware of and reaffirm attractive facial expressions that they had not previously noticed. By referring to the second display image P2, the beauty consultant can give the subject advice on makeup techniques that will bring out the subject's attractive features. Although the beauty consultant cannot see the subject's face directly because they are remote, they can provide beauty counseling while referring to the second display image P2. This allows the beauty consultant to carefully check the subject's skin color, bone structure, overall appearance, etc. in a still image, making it easier to provide accurate advice.
[0148] The beauty consultant terminal 90 may include the processing unit 2 and function as an information processing device that provides a service of presenting various facial expressions of the subject to the user, or the information processing system may be composed of the subject terminal and the beauty consultant terminal. Also, the subject terminal 80 may include the processing unit 2.
[0149] <Other variations> Although the embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the spirit and scope of the present invention. Other modifications will be described below.
[0150] [Variation 10] In the above-described embodiment, an example of beauty counseling by a beauty consultant (counselor) using an information processing device as a professional job has been given, but the present invention is not limited to this. For example, the present invention can also be applied to hairstyle counseling by a hairdresser (counselor), counseling by a non-professional information provider or acquaintance who is knowledgeable about beauty, presentations by a facial expression counselor, and facial expression counseling when serving customers.
[0151] The present invention can also be used in situations other than beauty counseling. For example, there are many situations in daily life where facial expressions and gestures are important, such as presentations and customer service in stores. In these situations, it is important to know what facial expressions and gestures give a good or bad impression. By applying the present invention, it is possible to extract facial expressions and gestures that are difficult to notice, thereby providing a new means for, for example, presentation rehearsals and customer service training. In other words, the "subject" of the present invention is not limited to the subject of beauty counseling, and does not necessarily have to be a different person. For example, the present invention can be used when filming a presentation practice and checking one's own facial expressions and gestures.
[0152] [Variation 11] In the above-described embodiment, an example has been given in which FaceMesh is used to obtain three-dimensional coordinate values of three-dimensional feature points, including depth information, from two-dimensional image data of a face captured by an imaging unit. However, the acquisition of depth information of feature points is not limited to this. For example, three-dimensional coordinate values of facial feature points may be obtained using a camera equipped with the functions of an image sensor (a color two-dimensional image capturing unit) and a distance sensor (a distance image capturing unit). The image sensor captures a color two-dimensional image of the subject, and the distance sensor captures a distance image of the subject. A known sensor can be used as the distance sensor; for example, a stereo camera may be used to obtain depth information and obtain three-dimensional coordinate values of facial feature points.
[0153] Furthermore, in the above-described embodiment, an example was given in which FaceMesh was used to acquire the three-dimensional coordinate values of three-dimensional facial feature points. However, an existing learning model may be used to extract facial feature points and acquire the three-dimensional coordinate values of the feature points. For example, as shown in FIG. 17(B), an existing learning model can extract 56 facial feature points, such as the facial contour, eyebrows, eyes, nose, and mouth. From the viewpoint of accuracy in similarity calculation, it is preferable to have 66 or more facial feature points, and from the viewpoint of reducing the subject's discomfort due to shortened calculation processing, it is preferable to have 1,000 or fewer facial feature points.
[0154] [Variation 12] In the above embodiment, an example has been given in which the three-dimensional coordinate values of three-dimensional feature points are used as face feature amounts used for group classification, but the present invention is not limited to this.
[0155] For example, instead of or in addition to the three-dimensional coordinate values (feature amounts) of the three-dimensional feature points, group classification may be performed using one or more feature amounts selected from the distance between three-dimensional feature points (the distance between two three-dimensional feature points), the area between three-dimensional feature points (for example, the area of a triangle with three three-dimensional feature points as vertices), the volume between three-dimensional feature points (the volume of a polyhedron with four or more three-dimensional feature points as vertices), facial skin brightness information, and facial skin color information. By appropriately combining these, the accuracy of similarity calculation increases, and face images with various expressions can be classified into groups with greater accuracy.
[0156] Furthermore, for example, instead of three-dimensional coordinate values, two-dimensional coordinate values of two-dimensional feature points, i.e., coordinate values on a two-dimensional image captured by an imaging device, may be used as feature values used for group classification. Furthermore, in addition to two-dimensional coordinate values, group classification may be performed using feature values that take into account one or more feature values selected from the distance between two-dimensional feature points, the area between two-dimensional feature points, facial skin luminance information, and facial skin color information. By appropriately combining these, the accuracy of similarity calculations can be improved, allowing for more accurate group classification of facial images with various expressions.
[0157] In the first embodiment, the face direction can be estimated based on three-dimensional coordinate values including depth information. However, the face direction may also be estimated using two-dimensional coordinate values of the face. For example, a first triangle is drawn with vertices at the two-dimensional feature points at the corners of the right eye, the left eye, and the tip of the nose. A second triangle is drawn with vertices at the two-dimensional feature points at the tip of the nose, the right corner of the mouth, and the left corner of the mouth. The face direction can be estimated based on changes in the shape and size of the first triangle and / or the second triangle. Group classification may be performed using the two-dimensional coordinate values of the facial feature points and the shape and size of the first triangle and / or the second triangle as facial feature quantities. Group classification can also be performed based on differences in facial expressions taking into account face direction information.
[0158] Furthermore, for example, brightness information and / or color information of facial skin may be used as the feature amount used for group classification.
[0159] [Variation 13] In the first and second embodiments, the subject and the beauty consultant review the same first display image P1 and select a face image for counseling. However, this is not limiting. For example, the first display image P1, which includes 30 representative face images and is generated by the information processing device 1 or 1A, may be first reviewed by the beauty consultant alone, and the beauty consultant may select some representative face images (e.g., 10) from the 30 representative face images. A newly generated display image including the 10 representative face images selected by the beauty consultant may then be presented to the subject. Then, the subject and the beauty consultant may each select one face image for counseling from the 10 representative face images included in the newly generated display image, and the second display image P2 may be generated.
[0160] [Variation 14] In each of the above-described embodiments, the processing unit 2 is configured by one information processing device 1 or 1A, but a configuration in which multiple information processing devices are integrated to perform the operations of the processing unit 2 may also be used. Furthermore, in the first embodiment, an example was given in which the generated display image is displayed on the display unit 41 of the information processing device 1, but the display image may also be displayed on a display device, such as a monitor or screen, separate from the information processing device 1 that includes the processing unit 2. Furthermore, in each of the above-described embodiments, an example was given in which the generated first display image P1 and second display image P2 are presented to the subject 8 and the beauty consultant 9 on a display unit (display), but they may also be presented on printed matter on which the first display image P1 and the second display image P2 are printed on paper.
[0161] [Variation 15] In the first and second embodiments, the subject and the beauty consultant review the same first display image P1 to select a facial image for counseling. However, this is not limiting. For example, the information processing device 1, 1A may display a first display image P1 containing 30 representative facial images, without prompting the user to select a representative image. The first display image P1 includes facial images that include a variety of natural facial expressions of the subject, without any arbitrary or subjective intentions of the device manufacturer. By reviewing such a first display image P1, the user (the subject or the beauty consultant) can easily notice new facial expressions that make the subject look attractive.
[0162] [Variation 16] In the above-described embodiment and modified example, facial images are classified into a plurality of groups, a plurality of display groups are extracted, and then a representative image to be used for display is extracted and selected. However, a display group may be extracted after extracting and selecting a representative image for each group. That is, the present invention may be implemented by extracting a representative facial image from the facial images belonging to each group, and then extracting a plurality of groups at predetermined intervals as display groups from the plurality of groups in descending or descending order of the number of facial images belonging to each group, and generating a first display image in which the representative facial image extracted for each display group is displayed. [Explanation of symbols]
[0163] 1, 1A...Information processing device 2...Processing section 8. Target 51...Facial image 52...Representative face image (representative face image) G...Group DG...Display group P1, P1a to P1f...First display image
Claims
1. acquiring a plurality of temporally consecutive facial images of a subject; classifying the plurality of face images into a plurality of groups based on the similarity of a plurality of facial feature amounts extracted from each of the plurality of face images; extracting a plurality of groups at predetermined intervals from the plurality of groups in descending order of the number of face images belonging to each group as display groups; extracting a representative face image from the face images belonging to each group; A first display image is generated in which the representative face image extracted for each display group is displayed. Processing section An information processing device comprising:
2. The processing unit acquiring information on the number of the representative face images to be displayed in the first display image; In the classification into the plurality of groups, the plurality of facial images are classified into n groups; In the extraction of the representative face image, one representative face image is extracted for each of the groups; In extracting the plurality of display groups, assigning group numbers of 1, 2, 3, . . . (n-2), (n-1), and (n) to the groups in order of the number of face images belonging to each group, extracting a group having a display start group number as one of the plurality of display groups; where m is the number of the representative face images to be displayed in the first display image based on the number information, a is the number of thinning-out operations, and b is an integer equal to or greater than 2, using the number of thinning-out operations a calculated by the following formula (1) and applying a rounding process to a value calculated by the following formula (2), thereby calculating the group number of the display group to be extracted; a=(n-2) / (m-1)...(1) Group number of display group=display start group number+a×(b−1) (2) In calculating the group number of the display group, integers of 2 or more are substituted into b in ascending order starting from 2 to find the group number of the display group, and the group numbers of the display groups are calculated until the calculated group number of the display group exceeds the value of the display end group number, and the group with the last calculated group number is extracted, or the group with the display end group number is extracted instead of the group with the last calculated group number. The information processing device according to claim 1 .
3. The processing unit acquiring information on the number of the representative face images to be displayed in the first display image; In the classification into the plurality of groups, the plurality of facial images are classified into n groups; In the extraction of the representative face image, one representative face image is extracted for each of the groups; In extracting the plurality of display groups, assigning group numbers of 1, 2, 3, . . . (n-2), (n-1), and (n) to the groups in order of the number of face images belonging to each group, extracting a group having a display start group number as one of the plurality of display groups; where m is the number of the representative face images to be displayed in the first display image based on the number information, a is the number of thinning-out images, and b is an integer equal to or greater than 2, using the number of thinning-out images a calculated by rounding a value calculated by the following formula (1), calculate the group number of the display group to be extracted by the following formula (2): a=(n-2) / (m-1)...(1) Group number of display group=display start group number+a×(b−1) (2) In calculating the group numbers of the display groups, integers of 2 or more are substituted into b in ascending order, starting from 2, to find the group numbers of the display groups, and the group numbers of the display groups are calculated until the total number of the extracted display groups reaches m. The information processing device according to claim 1 .
4. The feature amount is a feature amount of a three-dimensional feature point of a face.
3. The information processing device according to claim 1 or 2.
5. The feature amounts are three-dimensional coordinate values of three-dimensional feature points of the face. The information processing device according to claim 4 .
6. The classification into the plurality of groups based on the similarity of the feature quantities is performed using k-means.
3. The information processing system according to claim 1 or 2.
7. The processing unit The subject and the counselor who provides the counseling each select a face image from the plurality of representative face images included in the first display image, and set the selected face image as a face image for counseling, and generate a second display image including the face image for counseling.
3. The information processing device according to claim 1 or 2.
8. The processing unit In generating the first display image, the extracted representative face images are arranged in a predetermined order in descending order of the number of face images belonging to a display group including the extracted representative face images, to generate the first display image. The information processing device according to claim 7 .
9. A method for supporting improvement of an impression of a subject, using the information processing device according to claim 1 or 2.
10. An information processing method executed by an information processing device, acquiring a plurality of temporally consecutively acquired facial images of a subject; classifying the plurality of face images into a plurality of groups based on similarities of a plurality of facial feature amounts extracted from each of the plurality of face images; extracting a plurality of groups at predetermined intervals from the plurality of groups in descending order of the number of face images belonging to each group as display groups; extracting a representative face image from the face images belonging to each of the groups; generating a first display image in which the representative face image extracted for each display group is displayed. Information processing methods.
11. acquiring a plurality of temporally consecutively acquired facial images of a subject; classifying the plurality of face images into a plurality of groups based on similarities of a plurality of facial feature amounts extracted from each of the plurality of face images; extracting a plurality of groups at predetermined intervals from the plurality of groups in descending order of the number of face images belonging to each group as display groups; extracting a representative face image from the face images belonging to each group; generating a first display image in which the representative face image extracted for each display group is displayed; A program that causes an information processing device to execute the above.
Citation Information
Patent Citations
Image server, image retrieval system, image retrieval method, and index creation method
JP2010250633A
Information processing device, information processing method, and program
JP2022112168A
Operating method of information processing device, information processing device, and program
JP2023184309A