Inner surface state estimation device, inner surface state estimation method, and program

JPWO2024154292A5Pending Publication Date: 2025-09-16
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024571535
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2025-07-07
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing internal state estimation devices face accuracy issues due to variations in camera installation positions and differences in eye reflection patterns between training and estimation phases, leading to deteriorated estimation performance, especially in low frame rate facial images.

Method used

An internal state estimation device that calculates inner state feature amounts from face images and integrates estimation results from multiple models trained on different face orientations, using a face orientation feature amount to adjust for camera position variations and improve accuracy.

Benefits of technology

The device achieves high accuracy in estimating internal states by integrating estimation results from multiple models, effectively mitigating the impact of camera position changes and eye reflection pattern differences, thereby enhancing robustness and reliability.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

An inner state estimation device 1X mainly includes an inner state feature quantity acquisition means 21X, an inner state estimation means 22X, and an integration means 25X. The inner state feature quantity acquisition means 21X acquires an inner state feature quantity, which is a feature quantity calculated from a face image of a subject and used to estimate the inner state of the subject. The inner state estimation means 22X obtains a prescribed number of estimation results of the inner state on the basis of the inner state feature quantity and a prescribed number of inner state estimation models, each of which has been trained using face images of a face facing in a different direction. The integration means 25X generates an integrated estimation result of the inner state by integrating the prescribed number of estimation results of the inner state.
Need to check novelty before this filing date? Find Prior Art

Description

Inner surface state estimation device, inner surface state estimation method, and storage medium

[0001] The present disclosure relates to the technical fields of an inner surface state estimation device, an inner surface state estimation method, and a storage medium that perform processing related to estimation of an inner surface state.

[0002] There are known devices or systems that estimate the internal state of a subject based on a facial image of the subject. For example, Patent Literature 1 discloses a device that estimates the subject's drowsiness based on a facial image (facial video) of the subject captured by a camera. Also, Non-Patent Literature 1 discloses a technology that generates a separate image showing a different facial orientation from a facial image.

[0003] International Publication WO2019 / 123569

[0004] Hang Zhou el al., Rotate-and-Render: Unsupervised Photorealistic Face Rotation From Single-View Images, CVPR2020.

[0005] In Patent Literature 1, by capturing eyelid fluctuation, which is slower than blinking, as a characteristic of drowsiness, it is possible to estimate drowsiness with high accuracy even with facial images at a low frame rate. However, if the way the eyes are viewed differs between when the model is trained and when estimation is performed using the model due to factors such as the camera installation position, the estimation accuracy may deteriorate.

[0006] In view of the above-mentioned problems, one of the objects of the present disclosure is to provide an inner state estimation device, an inner state estimation method, and a storage medium that are capable of estimating the inner state of a subject with high accuracy.

[0007] One aspect of an inner state estimation device is an inner state estimation device having: inner state feature acquisition means for acquiring inner state feature amounts calculated from a facial image of a subject, the inner state feature amount being a feature amount used to estimate the inner state of the subject; inner state estimation means for acquiring the predetermined number of inner state estimation results based on the inner state feature amounts and a predetermined number of inner state estimation models each trained using facial images with different facial orientations; and integration means for generating an integrated estimation result by integrating the predetermined number of inner state estimation results.

[0008] Another aspect of the inner state estimating device is an inner state estimating device having: inner state feature acquisition means for acquiring inner state feature amounts calculated from a facial image of a subject, the inner state feature amount being a feature amount used to estimate the inner state of the subject; face direction feature acquisition means for acquiring face direction feature amounts calculated from the facial image, the face direction feature amount being a feature amount related to the facial direction of the subject; and inner state estimating means for estimating the inner state based on the inner state feature amounts and the face direction feature amounts.

[0009] One aspect of the inner state estimation method is a method in which a computer acquires inner state features calculated from a facial image of a subject, the inner state features being features used to estimate the inner state of the subject, acquires the predetermined number of inner state estimation results based on the inner state features and a predetermined number of inner state estimation models each trained using facial images with different facial orientations, and generates an integrated estimation result by integrating the predetermined number of inner state estimation results. Note that the term "computer" includes any electronic device (which may be a processor included in an electronic device) and may be composed of multiple electronic devices.

[0010] Another aspect of the inner state estimation method is a method in which a computer acquires inner state features, which are calculated from a facial image of a subject and are features used to estimate the inner state of the subject, acquires face direction features, which are calculated from the facial image and are features related to the facial direction of the subject, and estimates the inner state based on the inner state features and the face direction features.

[0011] One aspect of the storage medium is a storage medium storing a program that causes a computer to execute the following processes: acquire inner state features, which are features calculated from a facial image of a subject and used to estimate the inner state of the subject; acquire a predetermined number of inner state estimation results based on the inner state features and a predetermined number of inner state estimation models, each trained using facial images with different facial orientations; and generate an integrated estimation result by integrating the predetermined number of inner state estimation results.

[0012] Another aspect of the storage medium is a storage medium storing a program that causes a computer to execute a process of acquiring inner state features, which are features calculated from a facial image of a subject and used to estimate the inner state of the subject, acquiring face direction features, which are features related to the facial direction of the subject, calculated from the facial image, and estimating the inner state based on the inner state features and the face direction features.

[0013] The subject's internal state can be estimated with high accuracy.

[0014] 1 shows a schematic configuration of an inner state estimation system according to a first embodiment; FIG. 2 shows an example of the hardware configuration of an inner state estimation device common to all embodiments; (A) A face image of a subject facing forward relative to a camera taking the image; (B) A face image of a subject facing upward at a predetermined angle relative to a camera taking the image; (C) A face image of a subject facing rightward at a predetermined angle relative to a camera taking the image; FIG. 3 shows an example of functional blocks of an inner state estimation device related to training of an inner state estimation model according to a first embodiment; (A) A distribution of face orientations of original face images stored in a training data storage unit; (B) A distribution of face orientations of augmented face images generated from original face images; FIG. 4 shows an example of a flowchart related to training of an inner state estimation model executed by the inner state estimation device according to the first embodiment; (B) A flowchart related to inner state estimation executed by the inner state estimation device according to the first embodiment; (A) A diagram of the generation environment of an evaluation dataset observed from a direction in which the subject is present; (B) A diagram of the generation environment of an evaluation dataset observed from the side of the subject. 10A is a graph showing evaluation results of the disclosed method and the comparative method when a facial image generated by camera 5A is used as input to the inner state estimation model. (B) is a graph showing evaluation results of the disclosed method and the comparative method when a facial image generated by camera 5B is used as input to the inner state estimation model. (A) is a graph showing evaluation results of the disclosed method and the comparative method when a facial image generated by camera 5C is used as input to the inner state estimation model. (B) is a graph showing overall evaluation results of the disclosed method and the comparative method when facial images generated by cameras 5A to 5C are used as input to the inner state estimation model.

[0033] FIG. 10B is a graph showing an example of functional blocks of the inner state estimation device 1 related to training of the inner state estimation model in the second embodiment.

[0034] FIG. 10C is an example of a flowchart related to training of the inner state estimation model executed by the inner state estimation device in the second embodiment.

[0035] FIG. 10D is an example of functional blocks of the inner state estimation device 1 related to inner state estimation in the second embodiment.

[0036] FIG. 10E is an example of a flowchart related to inner state estimation executed by the inner state estimation device in the second embodiment. A schematic configuration of an inner state estimation system in a third embodiment is shown.Fig. 10 is a block diagram of an inner surface state estimating device according to a fourth embodiment; Fig. 11 is an example of a flowchart executed by the inner surface state estimating device according to the fourth embodiment; Fig. 12 is a block diagram of an inner surface state estimating device according to a fifth embodiment; Fig. 13 is an example of a flowchart executed by the inner surface state estimating device according to the fifth embodiment;

[0015] Hereinafter, embodiments of an inner surface state estimation device, an inner surface state estimation method, and a storage medium will be described with reference to the drawings.

[0016] <First Embodiment> (1) System Configuration Fig. 1 shows a schematic configuration of an inner state estimating system 100 according to the first embodiment. The inner state estimating system 100 is a system that estimates the inner state of a subject based on an image (facial image) of the subject's face, and mainly includes an inner state estimating device 1, an input device 2, a display device 3, a storage device 4, and a camera (image capturing device) 5. Note that the "subject" is a person who is the subject of inner state estimation, and may be an athlete or employee whose inner state is managed by an organization, or may be an individual user.

[0017] The internal state estimation device 1 estimates the internal state of a subject based on facial images of the subject generated by the camera 5 (including a video, which is a group of a predetermined number of images acquired in time series; the same applies hereinafter). The internal state estimation device 1 calculates an arbitrary index value (score) representing the subject's internal state as an estimation result of the subject's internal state. Examples of index values ​​representing the internal state include alertness, sleepiness, concentration, tension, health, and anxiety. In addition, in this embodiment, the internal state estimation device 1 trains a model to be used for estimating the internal state (also referred to as an "internal state estimation model") before executing the above-described internal state estimation. As described below, the internal state estimation model is a plurality of models each trained using facial images classified by facial orientation. By training such internal state estimation models and using the trained internal state estimation models to estimate the internal state, the internal state estimation device 1 can estimate the internal state without degrading estimation accuracy, even if the installation position of the camera 5 differs between when the model is trained and when the internal state is estimated. The inner surface state estimation model may be learned by a device separate from the inner surface state estimation device 1 before the inner surface state estimation device 1 estimates the inner surface state.

[0018] The inner state estimation device 1 communicates data with the input device 2, display device 3, and camera 5 via a communication network or by direct wireless or wired communication. For example, the inner state estimation device 1 receives an input signal "S1" from the input device 2. The inner state estimation device 1 also receives a facial image from the camera 5 that captures the subject's face. The inner state estimation device 1 also generates a display signal "S2" based on the estimation result of the subject's inner state and supplies the generated display signal S2 to the display device 3.

[0019] The input device 2 is an interface that accepts user input (manual input) of information about each subject. The user who inputs information using the input device 2 may be the subject himself / herself, or a person who manages or supervises the subject's activities. The input device 2 may be, for example, a variety of user input interfaces, such as a touch panel, buttons, a keyboard, a mouse, or a voice input device. The input device 2 supplies an input signal S1 generated based on the user's input to the internal state estimation device 1. The display device 3 displays predetermined information based on a display signal S2 supplied from the internal state estimation device 1. The display device 3 is, for example, a display or a projector.

[0020] The storage device 4 is a memory that stores various information necessary for estimating the inner state, etc. The storage device 4 may be an external storage device such as a hard disk connected to or built into the inner state estimating device 1, or may be a storage medium such as a flash memory. The storage device 4 may also be a server device that performs data communication with the inner state estimating device 1. The storage device 4 may also be composed of multiple devices.

[0021] The storage device 4 functionally includes an inner state estimation model storage unit 41, a training data storage unit 42, and a face direction weight calculation model storage unit 43. The inner state estimation model storage unit 41 stores parameters of the inner state estimation model. In this embodiment, the inner state estimation model is N models (where "N" is an integer equal to or greater than 2) trained using face images classified by subject's face direction (i.e., the direction of the subject's face relative to the camera that captured the face image). The training data storage unit 42 stores training data used for training the inner state estimation model. The face direction weight calculation model storage unit 43 stores parameters of the face direction weight calculation model. Here, the face direction weight calculation model is a model that calculates weights (also referred to as "face direction weights") for each estimation result to integrate the inner state estimation results output by the N inner state estimation models.

[0022] The information stored in the inner surface state estimation model storage unit 41, the training data storage unit 42, and the face direction weight calculation model storage unit 43 will be described in detail later.

[0023] The configuration of the inner state estimation system 100 shown in FIG. 1 is an example, and various modifications may be made to the configuration. For example, the input device 2 and the display device 3 may be configured as an integrated device. In this case, the input device 2 and the display device 3 may be configured as a tablet terminal that is integrated with or separate from the inner state estimation device 1. In this case, the inner state estimation device 1, the input device 2, the display device 3, and the camera 5 (and may also include the storage device 4) may be configured as a single smartphone or wearable device used by the subject. Furthermore, the inner state estimation device 1 may be configured as multiple devices. In this case, the multiple devices that make up the inner state estimation device 1 exchange information necessary to execute pre-assigned processes between these multiple devices.

[0024] (2) Hardware Configuration Fig. 2 shows the hardware configuration of the inner surface state estimation device 1. The inner surface state estimation device 1 includes, as hardware components, a processor 11, a memory 12, and an interface 13. The processor 11, the memory 12, and the interface 13 are connected via a data bus 90.

[0025] The processor 11 executes programs stored in the memory 12 to function as a controller (arithmetic unit) that performs overall control of the inner surface state estimation device 1. The processor 11 is, for example, a processor such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or a TPU (Tensor Processing Unit). The processor 11 may be composed of multiple processors. The processor 11 is an example of a computer.

[0026] The memory 12 is composed of various types of volatile and non-volatile memory, such as RAM (Random Access Memory), ROM (Read Only Memory), and flash memory. The memory 12 also stores programs for executing the processes performed by the inner state estimation device 1. Some of the information stored in the memory 12 may be stored in one or more external storage devices capable of communicating with the inner state estimation device 1, or in a storage medium that is detachable from the inner state estimation device 1. The memory 12 may also function as at least a part of the storage device 4. In this case, the memory 12 functions as at least one of an inner state estimation model storage unit 41, a training data storage unit 42, and a facial direction weight calculation model storage unit 43.

[0027] The interface 13 is an interface for electrically connecting the inner surface state estimation device 1 to other devices. These interfaces may be wireless interfaces such as network adapters for wirelessly transmitting and receiving data to and from other devices, or may be hardware interfaces for connecting to other devices via cables or the like.

[0028] The hardware configuration of the inner state estimating device 1 is not limited to the configuration shown in Fig. 2. For example, the inner state estimating device 1 may include at least one of the input device 2 and the display device 3. Furthermore, the inner state estimating device 1 may be connected to or built-in with a sound output device such as a speaker.

[0029] (3) Details of Data Next, we will explain the details of the data stored in the inner state estimation model storage unit 41, the training data storage unit 42, and the facial direction weight calculation model storage unit 43 of the storage device 4. Hereinafter, a "subject" refers to a person who is the subject of observation in generating the training data, and may be multiple people, and may or may not include the target person.

[0030] The inner state estimation model storage unit 41 stores parameters of the inner state estimation model (in other words, information necessary to configure the inner state estimation model). In this embodiment, the parameters of the inner state estimation model are learned by the inner state estimation device 1 before estimating the inner state of a subject.

[0031] The inner state estimation models are N models trained using face images classified into N patterns according to the subject's facial orientation. Hereinafter, the N inner state estimation models will be referred to as the "first inner state estimation model," ..., "Nth inner state estimation model," respectively, and the facial orientations corresponding to the "first inner state estimation model," ..., "Nth inner state estimation model," respectively, will be referred to as the "first orientation," ..., "Nth orientation." Here, the "first orientation," ..., "Nth orientation" respectively represent N different patterns of facial orientation, and are different from each other in at least one of the up-down direction or the left-right direction, for example.

[0032] When "n = 1, ..., N," the nth inner state estimation model is a model trained based on a face image in the nth orientation. Specifically, the nth inner state estimation model is a model trained on the relationship between features (also referred to as "inner state features") related to the inner state of a face image in which the subject's face is oriented in the nth orientation and the subject's inner state at the time the face image was generated. In other words, the nth inner state estimation model is trained in advance so that, when input with inner state features calculated based on a face image in the nth orientation, it outputs an estimation result of the inner state of the person depicted in the face image. Here, the inner state features are features used to estimate the inner state from the face image of the subject. For example, if the inner state to be estimated is drowsiness, the inner state features are values ​​indicating the degree to which the eyes are open. The inner state features are tensor-format data with a predetermined number of dimensions, which serves as an input format for the inner state estimation model.

[0033] 3(A) to 3(C) show examples of facial images with different facial orientations. The facial image shown in Fig. 3(A) is a facial image of a subject facing forward relative to the camera capturing the images. The facial image shown in Fig. 3(B) is a facial image of a subject facing upward at a predetermined angle (elevation angle) relative to the camera capturing the images. The facial image shown in Fig. 3(C) is a facial image of a subject facing rightward at a predetermined angle relative to the camera capturing the images. For example, if the facial orientations shown in Figs. 3(A) to 3(C) are designated as first to third orientations, respectively, parameters of a first inner state estimation model trained based on a facial image corresponding to the facial orientation shown in Fig. 3(A), a second inner state estimation model trained based on a facial image corresponding to the facial orientation shown in Fig. 3(B), and a third inner state estimation model trained based on a facial image corresponding to the facial orientation shown in Fig. 3(C) are stored in the inner state estimation model storage unit 41.

[0034] Each internal state estimation model may be any machine learning model (including a statistical model), such as a neural network or a support vector machine. For example, if the internal state estimation model is a model based on a neural network, such as a convolutional neural network, the internal state estimation model storage unit 41 stores information on various parameters, such as the layer structure, the neuron structure of each layer, the number and size of filters in each layer, and the weight of each element of each filter. Furthermore, each internal state estimation model may have a common architecture or may have different architectures.

[0035] Furthermore, the internal state estimation model may be a model trained by further classifying facial images according to predetermined attributes of the subject. In this case, parameters of each internal state estimation model trained based on facial images classified according to the predetermined attributes and facial orientation are stored in the internal state estimation model storage unit 41. Examples of the above-mentioned predetermined attributes include gender, occupation, race, age, height, weight, muscle mass, internal state tolerance, lifestyle habits, exercise habits, cognitive tendencies, and combinations thereof.

[0036] The training data storage unit 42 stores training data used to train the internal state estimation model. The training data includes facial images of the subject (e.g., time-series images of a predetermined duration) and correct-answer data indicating the correct internal state estimation result that the internal state estimation model should output when the facial images are input to the internal state estimation model. Here, the internal state estimation device 1 uses, as input data to the internal state estimation model during training, facial images of the subject (also referred to as "original facial images") stored in the training data storage unit 42, as well as facial images (also referred to as "augmented facial images") generated from the original facial images by data augmentation. Here, the augmented facial images are facial images of the subject in a different facial orientation from the subject's facial orientation in the original facial images, and are generated to include facial orientations that are missing in the original facial images. The original facial images are an example of a "first facial image," and the augmented facial images are an example of a "second facial image."

[0037] The face direction weight calculation model storage unit 43 stores parameters of the face direction weight calculation model (in other words, information necessary to configure the face direction weight calculation model). Here, the face direction weight calculation model calculates face direction weights so that the weight applied to the estimation result corresponding to a direction closer to the face direction indicated by the input face image is increased. Hereinafter, when "n = 1,...,N", the estimation result of the inner state output by the nth inner state estimation model will also be referred to as the "nth estimation result." The face direction weight calculation model is a model that estimates the relationship between the face image of a subject and a face direction weight according to the face direction of the subject.

[0038] In this embodiment, when a feature calculated based on a face image (also referred to as a "face direction feature") is input, the face direction weight calculation model outputs each face direction weight according to the face direction of the person indicated by the face image. The face direction feature is, for example, an angle representing the face direction, and may represent a pair of an angle in the up-down direction and an angle in the left-right direction of the face, or either one of the angles.

[0039] The face direction weight calculation model may be any model that calculates a face direction weight such that the weight assigned to an estimation result corresponding to a direction closer to the face direction indicated by the input face image is greater. For example, the face direction weight calculation model may store a Gaussian distribution of representative face direction feature amounts of each of the first to Nth estimation results, calculate a confidence interval to which the face direction feature amount input to the face direction weight calculation model belongs for each of the Gaussian distributions, and set a face direction weight according to the calculated confidence interval. In another example, the face direction weight calculation model may store representative face direction feature amounts of the first to Nth estimation results, and set a face direction weight according to the distance between the representative face direction feature amount of the first to Nth estimation results and the face direction feature amount input to the face direction weight calculation model.

[0040] In yet another example, the face direction weight calculation model may be a classification model that classifies the face direction of the original face image as one of the first to Nth directions based on the input face direction feature amount. In this case, for example, the confidence levels for the first to Nth directions output by the classification model when the face direction feature amount is input are set as face direction weights for the first to Nth estimation results. In this case, the classification model may be any machine learning model (including statistical models) such as a neural network or a support vector machine. For example, if the inner state estimation model is a model based on a neural network such as a convolutional neural network, the face direction weight calculation model storage unit 43 stores information on various parameters such as the layer structure, the neuron structure of each layer, the number and size of filters in each layer, and the weight of each element of each filter.

[0041] In addition to the various types of information described above, the storage device 4 also stores various types of information necessary for learning the inner surface state estimation model and estimating the inner surface state using the inner surface state estimation model.

[0042] For example, the storage device 4 stores parameters of an inner state feature calculation model that is a model for calculating inner state feature amounts from a facial image of a subject (in other words, information required to configure the inner state feature calculation model). Similarly, the storage device 4 stores parameters of a face direction feature calculation model that is a model for calculating facial direction feature amounts from a facial image of a subject (in other words, information required to configure the face direction feature calculation model).

[0043] Each feature calculation model used in this embodiment may be trained to extract features suitable for this embodiment. In this case, for example, each feature calculation model is trained using facial images prepared as training data as input data, and the parameters of each feature calculation model obtained by training are stored in the storage device 4 in advance (i.e., before estimating the subject's internal state). Various forms of feature calculation models (feature extractors) that use images as input have been proposed, and any of these forms may be adopted as each of the above-mentioned feature calculation models. For example, various deep learning models such as VGG16, VGG19, and MobileNet exist as such feature calculation models. For example, if each of the above-mentioned feature calculation models is a model based on a neural network, the storage device 4 stores in advance information on various parameters, such as the layer structure, the neuron structure of each layer, the number and size of filters in each layer, and the weight of each element of each filter.

[0044] (4) Learning of inner state estimation models Next, a process related to learning of the inner state estimation models will be described. In summary, the inner state estimation device 1 prepares face images corresponding to the first to Nth orientations, respectively, by generating extended face images from original face images stored in the training data storage unit 42, and learns the first to Nth inner state estimation models. In this way, data expansion of the training data is performed so that the amount is sufficient for learning the first to Nth inner state estimation models, and the first to Nth inner state estimation models that output highly accurate inner state estimation results are learned.

[0045] FIG. 4 shows an example of functional blocks of the inner state estimation device 1 related to learning of the inner state estimation model. The processor 11 of the inner state estimation device 1 functionally includes a data expansion unit 15, N inner state feature amount calculation units 16 (161 to 16N), and N learning units 17 (171 to 17N) related to learning of the inner state estimation model. Furthermore, the inner state estimation model storage unit 41 functionally includes a first inner state estimation model storage unit 411 to an Nth inner state estimation model storage unit 41N, which store parameters of the first inner state estimation model to the Nth inner state estimation model to be learned, respectively. While FIG. 4 shows blocks between which data is exchanged connected by solid lines, the combination of blocks between which data is exchanged is not limited to that shown. This also applies to other functional block diagrams described below.

[0046] The data expansion unit 15 acquires original face images stored in the training data storage unit 42 and generates expanded face images from the original face images by data augmentation, with the expanded face images having facial orientations different from those of the original face images. As a result, the data expansion unit 15 preferably generates face images necessary for training N number of first to Nth inner state estimation models corresponding to the first to Nth orientations. In this case, the data expansion unit 15 may convert the original face images into expanded face images based on any face orientation conversion technology that changes the facial orientation of a person in an image. Such face orientation conversion technology may be, for example, the technique described in Non-Patent Document 1. A specific example of the generation of expanded face images by the data expansion unit 15 will be described later. The data expansion unit 15 then supplies the face image corresponding to the nth orientation to the inner state feature amount calculation unit 16n. As a result, the face images corresponding to the first to Nth orientations are supplied to the inner state feature amount calculation units 161 to 16N, respectively. The data expansion unit 15 may use the original face images as they are without converting them. In this case, the original face image is supplied to one of inner state feature amount calculation units 161 to 16N depending on the facial orientation of the person in the image.

[0047] Each inner surface state feature calculation unit 16n (n = 1, ..., N) calculates an inner surface state feature from a face image corresponding to the nth direction. In this case, each inner surface state feature calculation unit 16n inputs the face image to an inner surface state feature calculation model to acquire an inner surface state feature output from the inner surface state feature calculation model. Each inner surface state feature calculation unit 16n supplies the calculated inner surface state feature to a corresponding learning unit 17n.

[0048] The learning unit 17n (n = 1, ..., N) learns the n-th inner state estimation model based on the inner state feature amounts obtained from the inner state feature amount calculation unit 16n and the correct answer data corresponding to the facial image used to calculate the inner state feature amounts. If the facial image used to calculate the inner state feature amounts is an original facial image, the correct answer data is the correct answer data stored in the training data storage unit 42 as the same record as the original facial image. If the facial image used to calculate the inner state feature amounts is an extended facial image, the correct answer data is the correct answer data stored in the training data storage unit 42 as the same record as the original facial image used to generate the extended facial image. The learning unit 17n determines the parameters of the n-th inner state estimation model so that the error (loss) between the inner state estimation result output by the n-th inner state estimation model when the inner state feature amounts obtained from the inner state feature amount calculation unit 16n are input to the n-th inner state estimation model is minimized and the correct answer indicated by the correct answer data. The algorithm for determining the parameters so as to minimize the loss may be any learning algorithm used in machine learning, such as gradient descent, backpropagation, etc. The learning unit 17 n then stores the learned parameters of the nth inner surface state estimation model in the nth inner surface state estimation model storage unit 41 n.

[0049] The components of the data expansion unit 15, the inner state feature amount calculation unit 16, and the learning unit 17 described in FIG. 4 can be realized, for example, by the processor 11 executing a program. Alternatively, the necessary programs may be recorded on any non-volatile storage medium and installed as needed to realize each component. At least some of these components may not necessarily be realized by software programs, but may be realized by any combination of hardware, firmware, and software. At least some of these components may be realized using a user-programmable integrated circuit, such as an FPGA (Field-Programmable Gate Array) or a microcontroller. In this case, a program consisting of the above components may be realized using this integrated circuit. Furthermore, at least a portion of each component may be configured by an ASSP (Application Specific Standard Product), an ASIC (Application Specific Integrated Circuit), or a quantum processor (quantum computer control chip). In this way, each component may be realized by various hardware. The same applies to other embodiments described below. Furthermore, each of these components may be realized by the cooperation of multiple computers, for example, using cloud computing technology.

[0050] Next, a supplementary explanation will be given of the expanded facial images generated by the data expansion unit 15. Fig. 5(A) shows a distribution of facial orientations of original facial images stored in the training data storage unit 42, and Fig. 5(B) shows a distribution of facial orientations of expanded facial images generated by the data expansion unit 15. Here, Figs. 5(A) and 5(B) show, as an example, a frequency distribution in which facial images are classified by facial orientation in the up-down direction. The facial orientation in the up-down direction is represented by a numerical value in which the forward direction is set to 0 degrees, the direction in which the elevation angle increases is a positive direction, and the direction in which the depression angle increases is a negative direction, and the frequency indicates the proportion of frequency when the total is set to 1.

[0051] Here, as an example, the data expansion unit 15 generates an expanded facial image so that the distribution has a different peak position from the distribution of the original facial image. Specifically, the distribution of the original facial image shown in Fig. 5(A) has an average value around 10 to 15 degrees, while the distribution of the expanded facial image shown in Fig. 5(B) has an average value around -10 to -5 degrees. In this case, for example, the data expansion unit 15 may generate an expanded facial image so that the expanded facial image has a Gaussian distribution with an average value and variance specified by user input.

[0052] Note that the data expansion unit 15 does not need to generate expanded facial images so as to have a Gaussian distribution, but may generate expanded facial images according to any rule so as to obtain the number of facial image samples required for training the N first to Nth inner state estimation models corresponding to the first to Nth orientations. Furthermore, in the examples of Figures 5(A) and 5(B), the facial orientations are classified by the facial orientation in the up-down direction, but this is not limiting, and the facial orientations may also be classified by the facial orientation in the left-right direction, or the facial orientations may be classified by a combination of the up-down direction and the left-right direction.

[0053] In this way, the data expansion unit 15 generates expanded facial images so as to increase the number of facial image samples for facial orientations for which the number of samples is insufficient with only the original facial image. This allows the data expansion unit 15 to secure the number of facial image samples required for training the N first to Nth inner inner state estimation models corresponding to the first to Nth orientations, respectively, and makes it possible to train the first to Nth inner inner state estimation models with high accuracy.

[0054] FIG. 6 is an example of a flowchart relating to the learning of the inner surface state estimation model executed by the inner surface state estimation device 1.

[0055] First, the inner state estimating device 1 generates extended face images in which the subject's facial orientation differs from the original face image (step S11) based on the original face image of the training data stored in the training data storage unit 42. In this way, the inner state estimating device 1 obtains the number of face image samples required for training N first to Nth inner state estimation models corresponding to the first to Nth orientations, respectively.

[0056] Next, the inner state estimation device 1 calculates inner state feature quantities of the facial image (step S12). In this case, the inner state estimation device 1 calculates inner state feature quantities to be input to the inner state estimation model in training for each facial image sample (e.g., one minute of video).

[0057] The inner state estimation device 1 then learns an inner state estimation model for each face direction based on the inner state feature amounts and the ground truth data (step S13). In this case, the inner state estimation device 1 updates the parameters of the n-th inner state estimation model based on the inner state feature amounts of the face image corresponding to the n-th direction and the ground truth data corresponding to the record of that face image (or the original face image in the case of an extended face image).

[0058] (5) Estimation of Inner State Next, a process for estimating an inner state using the trained inner state estimation model will be described. In summary, the inner state estimation device 1 acquires the estimation results of the first inner state estimation model to the Nth inner state estimation model and the face direction weights output by the face direction weight calculation model based on the facial images of the subject obtained from the camera 5, and integrates the above-mentioned estimation results using the face direction weights. This makes it possible to estimate the inner state without degrading the estimation accuracy, even if the installation position of the camera 5 is different between when the model is trained and when the inner state is estimated.

[0059] 7 shows an example of functional blocks of the inner state estimation device 1 related to estimating an inner state using an inner state estimation model. The processor 11 of the inner state estimation device 1, related to estimating an inner state using an inner state estimation model, functionally includes an inner state feature amount calculation unit 21, N inner state estimation units 22 (221 to 22N), a face direction feature amount calculation unit 23, a face direction weight calculation unit 24, and an integration unit 25. Furthermore, the inner state estimation model storage unit 41 functionally includes first inner state estimation model storage units 411 to Nth inner state estimation model storage units 41N that store trained parameters of the first to Nth inner state estimation models, respectively.

[0060] The inner state feature quantity calculation unit 21 acquires face images generated by the camera 5 via the interface 13 and calculates inner state feature quantities from the acquired face images. In this case, the inner state feature quantity calculation unit 21 calculates inner state feature quantities based on a predetermined number of time-series face images of the subject (e.g., one-minute video data) and an inner state feature quantity calculation model. The inner state feature quantity calculation model used by the inner state feature quantity calculation unit 21 is the same as the inner state feature quantity calculation model used by the inner state feature quantity calculation unit 16n. Then, the inner state feature quantity calculation unit 21 supplies the calculated inner state feature quantities to each of the inner state estimation units 221 to 22N.

[0061] The inner state estimation unit 22n (n = 1, ..., N) generates an nth estimation result regarding the inner state of the subject based on an nth inner state estimation model configured using parameters stored in the nth inner state estimation model storage unit 41n and the inner state feature amount. In this case, the inner state estimation unit 22n acquires, as the nth estimation result, the estimation result output by the nth inner state estimation model when the inner state feature amount is input to the nth inner state estimation model. The inner state estimation unit 22n supplies the generated nth estimation result to the integrating unit 25.

[0062] The face direction feature amount calculation unit 23 calculates a face direction feature amount based on the face image acquired by the inner state feature amount calculation unit 21 from the camera 5. In this case, the face direction feature amount calculation unit 23 acquires a face direction feature amount output by the face direction feature amount calculation model when the acquired face image is input to the face direction feature amount calculation model. The face direction feature amount calculation unit 23 supplies the calculated face direction feature amount to the face direction weight calculation unit 24.

[0063] The face direction weight calculation unit 24 calculates face direction weights for the first to Nth estimation results based on the face direction feature amount and a face direction weight calculation model configured using parameters stored in the face direction weight calculation model storage unit 43. The face direction weight calculation unit 24 supplies the face direction weights for the first to Nth estimation results to the integration unit 25.

[0064] The integrating unit 25 generates an integrated estimation result by integrating the first to Nth estimation results based on the first to Nth estimation results supplied from the inner state estimating unit 22 and the face direction weight supplied from the face direction weight calculating unit 24. In this case, the integrating unit 25 generates, for example, a weighted average of the first to Nth estimation results based on the face direction weights for the first to Nth estimation results as the integrated estimation result. The integrating unit 25 then generates a display signal S2 for displaying the generated integrated estimation result as the final estimation result of the subject's inner state, and supplies the generated display signal S2 to the display device 3. As a result, the display device 3 displays the integrated estimation result as the final estimation result of the subject's inner state.

[0065] The components of the inner state feature amount calculation unit 21, the inner state estimation unit 22, the face direction feature amount calculation unit 23, the face direction weight calculation unit 24, and the integration unit 25 described in FIG. 7 can be realized, for example, by the processor 11 executing a program. Alternatively, the necessary programs may be recorded in any non-volatile storage medium and installed as needed to realize the components. At least a portion of the components may not necessarily be realized by software programs, but may be realized by any combination of hardware, firmware, and software. At least a portion of the components may be realized using a user-programmable integrated circuit, such as an FPGA or a microcontroller. In this case, the integrated circuit may be used to realize a program consisting of the above components. At least a portion of the components may be configured using an ASSP, an ASIC, or a quantum processor. In this way, the components may be realized by various types of hardware. The same applies to other embodiments described below. Furthermore, the components may be realized by the cooperation of multiple computers, for example, using cloud computing technology.

[0066] FIG. 8 is an example of a flowchart executed by the inner surface state estimating device 1 regarding estimation of the inner surface state in the first embodiment.

[0067] First, the internal state estimation device 1 acquires face images generated by the camera 5 capturing an image of the subject (step S21). In this case, the internal state estimation device 1 acquires a predetermined number of face images (e.g., time-series images of a predetermined length) necessary for calculating the internal state feature amounts and the facial orientation feature amounts.

[0068] The inner state estimating device 1 then calculates inner state features and facial direction features from the facial image acquired in step S21 (step S22). In this case, the inner state estimating device 1 acquires inner state features output by the inner state feature calculation model when the above-mentioned facial image is input to the inner state feature calculation model, and acquires facial direction features output by the facial direction feature calculation model when the above-mentioned facial image is input to the facial direction feature calculation model.

[0069] Next, the inner state estimation device 1 generates first to Nth estimation results of the subject's inner state based on the first to Nth inner state estimation models constructed with reference to the inner state estimation model storage unit 41 and the inner state feature amounts calculated in step S22 (step S23). In this case, the inner state estimation device 1 inputs the inner state feature amounts into the first to Nth inner state estimation models, respectively, to obtain the first to Nth estimation results from the first to Nth inner state estimation models, respectively.

[0070] Next, the internal state estimating device 1 sets weights for the first to Nth estimation results based on the facial direction feature calculated in step S22 (step S24). In this case, the internal state estimating device 1 calculates facial direction weights for the first to Nth estimation results based on the facial direction feature and a facial direction weight calculation model configured with reference to the facial direction weight calculation model storage unit 43. Note that steps S23 and S24 may be performed in any order, may be performed in reverse order, or may be performed substantially simultaneously through parallel processing.

[0071] Next, the inner state estimation device 1 calculates an integrated estimation result by weighting and integrating the first to Nth estimation results using the face direction weight (step S25). The inner state estimation device 1 then outputs the calculated integrated estimation result (step S26). In this case, the inner state estimation device 1 may display the integrated estimation result on the display device 3, output it as audio using an audio output device (not shown), store it in the storage device 4, or transmit it to another device as the final estimation result of the subject's inner state.

[0072] (6) Example The applicant recorded facial images of a subject performing a calculation task using three cameras, and constructed an evaluation dataset consisting of the recorded facial images and correct data indicating the subject's correct internal state at the time of recording, which was generated based on questionnaire results of the subject or measurements using a sensor, and evaluated the internal state estimation method based on this embodiment (also referred to as the "disclosed method"). Here, the internal state estimated is arousal level. In addition, to verify the effectiveness of the disclosed method, an evaluation was also conducted on a method (also referred to as the "comparison method") that estimates the internal state using a single internal state estimation model trained regardless of face orientation. In the comparative method, the inner state estimation model was trained using original face images generated by a single camera and having the distribution shown in FIG. 5(A), while in the disclosed method, the first inner state estimation model and the second inner state estimation model were trained with N=2, where the first inner state estimation model was trained using original face images generated by a single camera and having the distribution shown in FIG. 5(A), and the second inner state estimation model was trained using augmented face images having the distribution shown in FIG. 5(B).

[0073] FIG. 9(A) is a diagram of the environment in which the evaluation dataset was generated, observed from the direction in which the subject was present, and FIG. 9(B) is a diagram of the environment in which the evaluation dataset was generated, observed from the side of the subject. As shown in FIG. 9(A), cameras 5A to 5C are installed at different heights, capturing images of the same subject at different facial angles. Also, as shown in FIGS. 9(A) and 9(B), camera 5C is installed at a position offset in the left-right and depth directions from cameras 5A and 5B. Here, there were a total of 27 subjects (24 men and 3 women). The subjects performed a calculation task displayed on a display for 15 minutes, and facial images of the subjects during the calculation task were generated from cameras 5A to 5C as the evaluation dataset. The length of each facial image per sample (i.e., the length of the facial image included in one record) was set to one minute. The evaluation data set was also classified into four groups according to whether the vertical tilt of the face belonged to "-30 to -15," "-15 to 0," "0 to 15," or "15 to 30," and then tabulated.

[0074] Fig. 10(A) is a graph showing the evaluation results of the disclosed method and the comparative method when a facial image generated by camera 5A is used as input to the inner state estimation model. Fig. 10(B) is a graph showing the evaluation results of the disclosed method and the comparative method when a facial image generated by camera 5B is used as input to the inner state estimation model. Fig. 11(A) is a graph showing the evaluation results of the disclosed method and the comparative method when a facial image generated by camera 5C is used as input to the inner state estimation model. Fig. 11(B) is a graph showing the overall evaluation results of the disclosed method and the comparative method when facial images generated by cameras 5A to 5C are used as input to the inner state estimation model. 10(A) to 11(B), the vertical axis represents the mean absolute error (a value normalized so that the range is 5.0 or less) between the correct arousal level indicated by the correct data and the estimated value of arousal level obtained by the disclosed method or the comparative method, and the horizontal axis represents the vertical tilt of the subject's face corresponding to "-30 to -15," "-15 to 0," "0 to 15," and "15 to 30." Also, the horizontal axis indicates the number of samples (i.e., the number of records) of face images obtained for each range of the subject's face orientation.

[0075] 10A to 11B, in the face images generated by any of cameras 5A to 5C, the disclosed method suppresses the increase in error when the face has a negative vertical tilt (i.e., when the face is captured from below) compared to the comparative method. In this way, the disclosed method is more robust against the position of the capturing camera 5.

[0076] Second Embodiment In the second embodiment, the inner state estimation device 1 differs from the first embodiment in that it uses, as an inner state estimation model, a model that receives as input a pair of inner state feature amounts and facial direction feature amounts obtained from a facial image, and outputs an estimation result of an inner state that takes into account the facial direction in the facial image. In other words, the inner state estimation model in the second embodiment is a model that has learned the relationship between the pair of inner state feature amounts and facial direction feature amounts calculated from a facial image and the inner state of the subject at the time the facial image was generated. Hereinafter, the same components as those in the first embodiment will be appropriately designated by the same reference numerals, and their description will be omitted.

[0077] 12 shows an example of functional blocks of the inner state estimation device 1 related to learning of the inner state estimation model in the second embodiment. The processor 11 of the inner state estimation device 1 in the second embodiment functionally includes a data expansion unit 15A, an inner state feature calculation unit 16Aa, a face direction feature calculation unit 16Ab, and a learning unit 17A, related to learning of the inner state estimation model.

[0078] The data expansion unit 15A acquires original face images stored in the training data storage unit 42 and generates expanded face images, each having a different facial orientation from the original face image, through data augmentation. This increases the variation in facial orientation of the face images used to train the inner state estimation model, improving the estimation accuracy of the trained inner state estimation model. The data expansion unit 15A supplies each sample of face image (original face image or expanded face image) to the inner state feature calculation unit 16Aa and the facial orientation feature calculation unit 16Ab.

[0079] The inner surface state feature amount calculation unit 16Aa calculates inner surface state feature amounts from the facial image supplied from the data expansion unit 15A. The inner surface state feature amount calculation unit 16Aa inputs the facial image into the inner surface state feature amount calculation model, thereby acquiring inner surface state feature amounts output by the inner surface state feature amount calculation model. The inner surface state feature amount calculation unit 16Aa supplies the calculated inner surface state feature amounts to the learning unit 17A.

[0080] The face direction feature amount calculation unit 16Ab calculates face direction feature amounts from the face image supplied from the data expansion unit 15A. The face direction feature amount calculation unit 16Ab inputs the face image to the face direction feature amount calculation model, thereby acquiring face direction feature amounts output by the face direction feature amount calculation model. The face direction feature amount calculation unit 16Ab supplies the calculated face direction feature amounts to the learning unit 17A. Note that, as in the first embodiment, parameters of the inner state feature amount calculation model and the face direction feature amount calculation model are stored in advance in, for example, the storage device 4.

[0081] Learning unit 17A learns the inner state estimation model based on the inner state feature amounts acquired from inner state feature amount calculation unit 16Aa, the facial direction feature amounts acquired from facial direction feature amount calculation unit 16Ab, and correct answer data corresponding to the facial image used to calculate the inner state feature amounts and the facial direction feature amounts. If the facial image used to calculate the inner state feature amounts is an original facial image, the correct answer data is the correct answer data stored in training data storage unit 42 as the same record as the original facial image. If the facial image used to calculate the inner state feature amounts is an extended facial image, the correct answer data is the correct answer data stored in training data storage unit 42 as the same record as the original facial image used to generate the extended facial image. Learning unit 17A determines the parameters of the inner state estimation model so that the error (loss) between the inner state estimation result output by the inner state estimation model when a set of inner state feature amounts and facial direction feature amounts is input to the inner state estimation model is minimized and the correct answer indicated by the correct answer data. Then, the learning unit 17A stores the learned parameters of the inner surface state estimation model in the inner surface state estimation model storage unit 41.

[0082] FIG. 13 is an example of a flowchart relating to learning of the inner surface state estimation model executed by the inner surface state estimation device 1 in the second embodiment.

[0083] First, the inner state estimating device 1 generates an expanded face image in which the face direction of the subject is changed based on the original face image of the training data stored in the training data storage unit 42 (step S31).

[0084] Next, the inner state estimating device 1 calculates inner state feature quantities and facial pose feature quantities of the facial image (step S32). In this case, the inner state estimating device 1 calculates inner state feature quantities and facial pose feature quantities to be input to the inner state estimation model for each sample of facial image (e.g., one minute of video).

[0085] The inner state estimation device 1 then learns an inner state estimation model based on the inner state feature amounts, facial direction feature amounts, and ground truth data (step S33). In this case, the inner state estimation device 1 updates the parameters of the inner state estimation model based on the pair of inner state feature amounts and facial direction feature amounts, and the ground truth data corresponding to the record of the face image (or the original face image in the case of an extended face image) used to calculate the inner state feature amounts and facial direction feature amounts.

[0086] 14 shows an example of functional blocks of the inner state estimation device 1 related to estimation of an inner state using an inner state estimation model in the second embodiment. The processor 11 of the inner state estimation device 1 in the second embodiment is related to estimation of an inner state using an inner state estimation model and functionally includes an inner state feature calculation unit 21A, a face direction feature calculation unit 23A, and an inner state estimation unit 22A.

[0087] The inner state feature quantity calculation unit 21A acquires face images generated by the camera 5 via the interface 13 and calculates inner state feature quantities from the acquired face images. In this case, the inner state feature quantity calculation unit 21A calculates inner state feature quantities based on an inner state feature quantity calculation model from a predetermined number of time-series face images of the subject (e.g., one-minute video data). The inner state feature quantity calculation unit 21A then supplies the calculated inner state feature quantities to the inner state estimation unit 22A.

[0088] Facial direction feature amount calculation unit 23A calculates facial direction feature amounts based on the facial image acquired by inner state feature amount calculation unit 21A from camera 5. In this case, facial direction feature amount calculation unit 23A acquires facial direction feature amounts output by the facial direction feature amount calculation model when a facial image is input to the facial direction feature amount calculation model. Facial direction feature amount calculation unit 23A supplies the calculated facial direction feature amounts to inner state estimation unit 22A.

[0089] The inner state estimation unit 22A generates an estimation result regarding the inner state of the subject based on an inner state estimation model configured using learned parameters stored in the inner state estimation model storage unit 41, the inner state feature amounts calculated by the inner state feature amount calculation unit 21A, and the facial direction feature amounts calculated by the facial direction feature amount calculation unit 23A. In this case, the inner state estimation unit 22A acquires the estimation result output by the inner state estimation model when a pair of the inner state feature amounts and the facial direction feature amounts is input to the inner state estimation model. The inner state estimation unit 22A then generates a display signal S2 for displaying the generated estimation result as a final estimation result of the inner state of the subject, and supplies the generated display signal S2 to the display device 3. As a result, the display device 3 displays the estimation result of the inner state of the subject.

[0090] FIG. 15 is an example of a flowchart executed by the inner surface state estimating device 1 according to the second embodiment in relation to estimating the inner surface state using the inner surface state estimation model.

[0091] First, the inner state estimating device 1 acquires face images generated by the camera 5 capturing an image of a subject (step S41). In this case, the inner state estimating device 1 acquires a predetermined number of face images necessary for calculating inner state feature amounts and facial orientation feature amounts.

[0092] The inner state estimating device 1 then calculates inner state feature amounts and facial direction feature amounts from the facial image acquired in step S41 (step S42). In this case, the inner state estimating device 1 acquires inner state feature amounts output by the inner state feature amount calculation model when the above-mentioned facial image is input to the inner state feature amount calculation model, and acquires facial direction feature amounts output by the facial direction feature amount calculation model when the above-mentioned facial image is input to the facial direction feature amount calculation model.

[0093] Next, the inner state estimation device 1 generates an estimation result of the inner state of the subject based on the inner state estimation model constructed with reference to the inner state estimation model storage unit 41 and the pair of inner state feature amounts and facial direction feature amounts calculated in step S42 (step S43). In this case, the inner state estimation device 1 inputs the pair of inner state feature amounts and facial direction feature amounts into the inner state estimation model, thereby acquiring the estimation result of the inner state from the inner state estimation model.

[0094] The inner state estimating device 1 then outputs the calculated estimation result (step S44). In this case, the inner state estimating device 1 may display the estimation result on the display device 3, output it as audio using an audio output device (not shown), store it in the storage device 4, or transmit it to another device.

[0095] Like the inner state estimation device 1 in the first embodiment, the inner state estimation device 1 in the second embodiment is capable of estimating the inner state without degrading the estimation accuracy even if the installation position of the camera 5 is different between when learning the model and when estimating the inner state.

[0096] 16 shows a schematic configuration of an inner state estimating system 100A according to a third embodiment. The inner state estimating system 100A according to the third embodiment includes an inner state estimating device 1A that performs the same processing as the inner state estimating device 1 according to the first or second embodiment, a storage device 4, and a terminal device 8 used by a subject. Hereinafter, the same components as those in the first embodiment will be appropriately designated by the same reference numerals, and their description will be omitted.

[0097] In the third embodiment, the inner surface state estimating device 1A functions as a server, and the terminal device 8 functions as a client. The inner surface state estimating device 1A and the terminal device 8 perform data communication via a network 9.

[0098] The terminal device 8 is a terminal used by a user who is a subject, and has input, display, communication, and imaging functions, and functions as the input device 2, display device 3, camera 5, etc. shown in Fig. 1. The terminal device 8 may be, for example, a personal computer, a tablet terminal such as a smartphone, or a PDA (Personal Digital Assistant). The terminal device 8 is electrically connected to a camera 5 such as a wearable sensor worn by the user, and transmits a facial image of the subject output by the camera 5 to the inner state estimation device 1A via a network 9.

[0099] The inner state estimation device 1A has the same hardware configuration as the inner state estimation device 1 shown in Fig. 2, and the processor 11 of the inner state estimation device 1A has the functional blocks described in the first or second embodiment. The inner state estimation device 1A receives a face image from a terminal device 8 via a network 9 and performs a process of estimating the inner state of a subject by referring to various information stored in the storage device 4. In addition, the inner state estimation device 1A transmits an output signal for outputting an inner state estimation result to the terminal device 8 via the network 9, based on a display request from the terminal device 8. In addition, the inner state estimation device 1A may perform a process related to learning of the inner state estimation model in the first or second embodiment, based on training data stored in the storage device 4.

[0100] In this way, the inner state estimation device 1A in the third embodiment estimates the inner state of a subject who is the user of the terminal device 8, and can present the estimation results to the subject via the terminal device 8 in a more suitable manner.

[0101] 17 is a block diagram of an inner surface state estimation device 1X according to a fourth embodiment. The inner surface state estimation device 1X mainly includes an inner surface state feature amount acquisition unit 21X, an inner surface state estimation unit 22X, and an integration unit 25X. Note that the inner surface state estimation device 1X may be configured by a plurality of units.

[0102] The inner state feature amount acquisition means 21X acquires inner state feature amounts, which are feature amounts used to estimate the inner state of the subject, calculated from a facial image of the subject. The facial image may be a single image, or a predetermined number of images equal to or greater than two. The inner state feature amount acquisition means 21X may acquire the inner state feature amounts by calculating the inner state feature amounts from the facial image of the subject, or may acquire the inner state feature amounts calculated from the facial image of the subject by a device other than the inner state estimation device 1X. The inner state feature amount acquisition means 21X may be, for example, the inner state feature amount calculation unit 21 in the first or third embodiment.

[0103] The inner state estimation means 22X acquires a predetermined number of inner state estimation results based on the inner state feature amounts and a predetermined number of inner state estimation models each trained using face images with different facial orientations. The inner state estimation means 22X may be, for example, the inner state estimation unit 22 in the first or third embodiment.

[0104] The integrating means 25X generates an integrated estimation result of the inner surface state by integrating a predetermined number of estimation results of the inner surface state. The integrating means 25X may use a weighted average of the predetermined number of estimation results of the inner surface state as the integrated estimation result, or a representative value such as an average of the predetermined number of estimation results of the inner surface state as the integrated estimation result, or may use a weighted average of the predetermined number of estimation results of the inner surface state as the integrated estimation result. The integrating means 25X may be, for example, the integrating unit 25 in the first or third embodiment.

[0105] 18 is an example of a flowchart executed by the inner state estimation device 1X in the fourth embodiment. First, the inner state feature acquisition means 21X acquires inner state feature values, which are calculated from a facial image of the subject and are used to estimate the subject's inner state (step S51). Next, the inner state estimation means 22X acquires a predetermined number of inner state estimation results based on the inner state feature values ​​and a predetermined number of inner state estimation models trained using facial images with different facial orientations (step S52). The integration means 25X generates an integrated inner state estimation result by integrating the predetermined number of inner state estimation results (step S53).

[0106] According to the fourth embodiment, the inner state estimating device 1X can estimate the inner state of a subject with high accuracy.

[0107] 19 is a block diagram of an inner state estimating device 1Y according to a fifth embodiment. The inner state estimating device 1Y mainly includes an inner state feature amount acquiring unit 21Y, a facial direction feature amount acquiring unit 23Y, and an inner state estimating unit 22Y. Note that the inner state estimating device 1Y may be configured with multiple units.

[0108] The inner state feature amount acquisition means 21Y acquires inner state feature amounts, which are feature amounts used to estimate the inner state of the subject, calculated from a facial image of the subject. The facial image may be a single image, or a predetermined number of images equal to or greater than two. The inner state feature amount acquisition means 21Y may acquire the inner state feature amounts by calculating the inner state feature amounts from the facial image of the subject, or may acquire inner state feature amounts calculated from the facial image of the subject by a device other than the inner state estimation device 1Y. The inner state feature amount acquisition means 21Y may be, for example, the inner state feature amount calculation unit 21A in the second or third embodiment.

[0109] The facial direction feature amount acquiring means 23Y acquires a facial direction feature amount, which is a feature amount related to the facial direction of the subject, calculated from the facial image of the subject. The facial direction feature amount acquiring means 23Y may acquire the facial direction feature amount by calculating the facial direction feature amount from the facial image of the subject, or may acquire the facial direction feature amount calculated from the facial image of the subject by a device other than the internal state estimating device 1Y. The facial direction feature amount acquiring means 23Y may be, for example, the facial direction feature amount calculating unit 23A in the second or third embodiment.

[0110] The inner state estimation means 22Y estimates the inner state of the subject based on the inner state feature amount and the facial direction feature amount. The inner state estimation means 22Y may be, for example, the inner state estimation unit 22A in the second or third embodiment.

[0111] FIG. 20 is an example of a flowchart executed by the inner surface state estimating device 1Y in the fifth embodiment.

[0112] The internal state feature acquisition means 21Y acquires internal state feature amounts calculated from the subject's facial image, which are feature amounts used to estimate the subject's internal state (step S61). The face direction feature acquisition means 23Y acquires face direction feature amounts calculated from the subject's facial image, which are feature amounts related to the subject's facial direction (step S62). The internal state estimation means 22Y estimates the subject's internal state based on the internal state feature amounts and the face direction feature amounts (step S63).

[0113] According to the fifth embodiment, the inner state estimating device 1Y can estimate the inner state of a subject with high accuracy.

[0114] In the above-described embodiments, the program can be stored using various types of non-transitory computer-readable media and supplied to a computer processor or the like. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic storage media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical storage media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)). The program may also be supplied to a computer by various types of transitory computer-readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. The transitory computer-readable media can be supplied to a computer via wired communication paths such as electric wires and optical fibers, or via wireless communication paths.

[0115] In addition, some or all of the above embodiments may be described as, but are not limited to, the following supplementary notes.

[0116] [Supplementary Note 1] An inner state estimation device comprising: inner state feature acquisition means for acquiring inner state feature amounts calculated from a facial image of a subject, the inner state feature amount being a feature amount used to estimate the inner state of the subject; inner state estimation means for acquiring a predetermined number of estimation results of the inner state based on the inner state feature amounts and a predetermined number of inner state estimation models each trained using facial images with different facial orientations; and integration means for generating an integrated estimation result by integrating the predetermined number of estimation results of the inner state. [Supplementary Note 2] The inner state estimation device according to Supplementary Note 1, wherein each of the predetermined number of inner state estimation models is a model that has learned the relationship between the inner state feature amount of the facial image with the facial orientation associated with that inner state estimation model and the inner state at the time of generation of the facial image. [Supplementary Note 3] The inner state estimation device according to Supplementary Note 1, wherein the predetermined number of inner state estimation models are trained using a first facial image that is a photograph of the subject, and a second facial image that has been transformed from the first facial image so that the second facial image has a facial orientation different from that of the subject in the first facial image. [Supplementary Note 4] The inner state estimation device according to Supplementary Note 1, further comprising weight determination means for determining a weight to be set for each of the predetermined number of inner state estimation results based on the face image, wherein the integrating means generates the integrated estimation result based on the weights and the predetermined number of inner state estimation results. [Supplementary Note 5] The inner state estimation device according to Supplementary Note 4, further comprising face direction feature acquisition means for acquiring face direction feature amounts calculated from the face image, the face direction feature amounts being feature amounts related to the face direction of the subject, wherein the weight determination means determines the weights based on the face direction feature amounts. [Supplementary Note 6] The inner state estimation device comprises: inner state feature acquisition means for acquiring inner state feature amounts calculated from the face image of the subject, the feature amounts being used to estimate the inner state of the subject; face direction feature acquisition means for acquiring face direction feature amounts calculated from the face image, the feature amounts related to the face direction of the subject; and inner state estimation means for estimating the inner state based on the inner state feature amounts and the face direction feature amounts.[Supplementary Note 7] The inner state estimation device according to Supplementary Note 6, wherein the inner state estimation means estimates the inner state based on the inner state feature amount, the facial direction feature amount, and an inner state estimation model, and the inner state estimation model is a model that learns a relationship between a set of the inner state feature amount and the facial direction feature amount calculated from the face image, and the inner state at the time of generating the face image. [Supplementary Note 8] An inner state estimation method, comprising: a computer; acquiring inner state feature amounts that are feature amounts used to estimate the inner state of the subject, calculated from a face image of the subject; acquiring the predetermined number of inner state estimation results based on the inner state feature amount and a predetermined number of inner state estimation models that are each trained using face images with different facial directions; and generating an integrated estimation result by integrating the predetermined number of inner state estimation results. [Supplementary Note 9] An inner state estimation method, comprising: a computer acquiring inner state features calculated from a facial image of a subject, the inner state features being features used to estimate the inner state of the subject; acquiring face direction features calculated from the facial image, the inner state features being features related to the facial direction of the subject; and estimating the inner state based on the inner state features and the face direction features. [Supplementary Note 10] A storage medium storing a program causing a computer to execute the following processes: acquiring inner state features calculated from a facial image of the subject, the inner state features being features used to estimate the inner state of the subject; acquiring the predetermined number of inner state estimation results based on the inner state features and a predetermined number of inner state estimation models each trained using facial images with different facial directions; and generating an integrated estimation result by integrating the predetermined number of inner state estimation results. [Supplementary Note 11] A storage medium storing a program that causes a computer to execute a process of acquiring inner state features, which are features used to estimate an inner state of a subject, calculated from a face image of the subject, acquiring face direction features, which are features related to the face direction of the subject, calculated from the face image, and estimating the inner state based on the inner state features and the face direction features.

[0117] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above embodiments. Various modifications within the scope of the present invention that would be understood by those skilled in the art can be made to the configuration and details of the present invention. In other words, the present invention naturally includes various modifications and alterations that would be possible for those skilled in the art based on the entire disclosure, including the claims, and the technical ideas. Furthermore, the disclosures of the above-cited patent and non-patent documents are incorporated herein by reference.

[0118] DESCRIPTION OF SYMBOLS 1, 1A, 1X, 1Y Inner surface state estimation device 2 Input device 3 Display device 4 Storage device 5, 5A to 5C Camera 8 Terminal device 9 Network 11 Processor 12 Memory 13 Interface 41 Inner surface state estimation model storage unit 42 Training data storage unit 43 Face direction weight calculation model storage unit 90 Data bus 100, 100A Inner surface state estimation system

Claims

1. an inner state feature amount acquiring means for acquiring inner state feature amounts calculated from a face image of a subject, the inner state feature amounts being feature amounts used to estimate the inner state of the subject; an inner state estimation means for acquiring a predetermined number of inner state estimation results based on the inner state feature amount and a predetermined number of inner state estimation models each trained using face images with different facial orientations; an integration means for generating an integrated estimation result by integrating the predetermined number of estimation results of the inner surface state; An inner surface state estimation device having the above structure.

2. 2. The inner state estimation device according to claim 1, wherein each of the predetermined number of inner state estimation models is a model that learns a relationship between the inner state feature of the face image corresponding to the face pose associated with the inner state estimation model and the inner state at the time of generation of the face image.

3. 2. The inner state estimating device according to claim 1, wherein the predetermined number of inner state estimation models are trained using a first face image that is a facial image of a subject captured, and a second face image obtained by transforming the first face image so that the subject's facial orientation is different from that of the subject in the first face image.

4. a weight determination means for determining a weight to be set for each of the predetermined number of estimation results of inner state based on the face image; The inner surface state estimation device according to claim 1 , wherein the integration means generates the integrated estimation result based on the weights and the predetermined number of estimation results of the inner surface state.

5. The method further includes a facial direction feature amount acquisition means for acquiring a facial direction feature amount, which is a feature amount relating to the facial direction of the subject, calculated from the facial image, The inner state estimating device according to claim 4 , wherein said weight determining means determines said weight based on said facial pose feature amount.

6. an inner state feature amount acquiring means for acquiring inner state feature amounts calculated from a face image of a subject, the inner state feature amounts being feature amounts used to estimate the inner state of the subject; a face direction feature amount acquiring means for acquiring a face direction feature amount, which is a feature amount relating to the face direction of the subject, calculated from the face image; an inner state estimation means for estimating the inner state based on the inner state feature amount and the facial direction feature amount; An inner surface state estimation device having the above structure.

7. The computer acquiring inner state features calculated from a face image of the subject, the inner state features being features used to estimate the inner state of the subject; acquiring a predetermined number of estimation results of inner states based on the inner state feature amounts and a predetermined number of inner state estimation models trained using face images with different facial orientations; generating an integrated estimation result by integrating the predetermined number of estimation results of the inner surface state; Internal state estimation method.

8. The computer acquiring inner state features calculated from a face image of the subject, the inner state features being features used to estimate the inner state of the subject; acquiring a facial direction feature amount calculated from the facial image, the facial direction feature amount being a feature amount relating to the facial direction of the subject; estimating the inner state based on the inner state feature amount and the facial direction feature amount; Internal state estimation method.

9. acquiring inner state features calculated from a face image of the subject, the inner state features being features used to estimate the inner state of the subject; acquiring a predetermined number of estimation results of inner states based on the inner state feature amounts and a predetermined number of inner state estimation models trained using face images with different facial orientations; a program that causes a computer to execute a process of generating an integrated estimation result by integrating the predetermined number of estimation results of the inner surface state;

10. acquiring inner state features calculated from a face image of the subject, the inner state features being features used to estimate the inner state of the subject; acquiring a facial direction feature amount calculated from the facial image, the facial direction feature amount being a feature amount relating to the facial direction of the subject; a program that causes a computer to execute a process of estimating the inner state based on the inner state feature amount and the face direction feature amount;