Information processing system and information processing method
The information processing system estimates user mental states by analyzing video and audio attributes through machine learning, addressing the inability of existing technologies to accurately assess mental states.
Patent Information
- Application Number
- PCT/JP2024/020991
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-10
- Publication Date
- 2025-12-18
Smart Images

Figure JP2024020991_18122025_PF_FP_ABST
Abstract
Description
Information processing system and information processing method
[0001] The present invention relates to an information processing system and an information processing method.
[0002] There is known a technique for analyzing the emotions felt by others in response to a speaker's statement (see, for example, Patent Document 1).
[0003] JP 2019-58625 A
[0004] However, although the technology of Patent Document 1 can analyze the emotions of a target person, it cannot estimate a specific mental state.
[0005] The present invention has been made in view of the above background, and aims to provide a technique that can estimate the mental state of a user.
[0006] The main invention of the present invention for solving the above problem is an information processing system comprising: a detection unit that analyzes video images of a user to detect the degree of the user's attributes; a calculation unit that calculates a second degree of variation of a first degree of variation related to the degree of the attributes in the video images; and an estimation unit that estimates the mental state of the user based on at least the second degree of variation.
[0007] Other problems and solutions disclosed in this application will be made clear in the section on preferred embodiments of the invention and the drawings.
[0008] According to the present invention, it is possible to estimate the mental state of a user.
[0009] It is a diagram showing an example of the overall configuration of an information processing system. It is a diagram showing an example of the hardware configuration of a management server 2. It is a diagram showing an example of the software configuration of a management server 2. It is a diagram explaining the operation of a management server 2.
[0010] <System Overview> An information processing system according to one embodiment of the present invention will be described below. The information processing system of this embodiment attempts to estimate a user's mental state (particularly the degree of depression) from a video of the user.
[0011] 1 is a diagram showing an example of the overall configuration of an information processing system. The information processing system of this embodiment is configured to include a management server 2. The management server 2 is communicably connected to a user terminal 1 via a communication network. The communication network is, for example, the Internet, and is constructed using a public telephone line network, a mobile phone line network, a wireless communication path, Ethernet (registered trademark), or the like.
[0012] The user terminal 1 is a computer operated by a user, and may be, for example, a smartphone, a tablet computer, or a personal computer.
[0013] The management server 2 is a computer that estimates the mental state of the user. The management server 2 may be a general-purpose computer such as a workstation or a personal computer, or may be logically realized by cloud computing.
[0014] <Management Server> FIG. 2 is a diagram illustrating an example of the hardware configuration of the management server 2. Note that the illustrated configuration is an example, and other configurations may also be used. The management server 2 includes a CPU 201, memory 202, storage device 203, communication interface 204, input device 205, and output device 206. The storage device 203 stores various data and programs, and is, for example, a hard disk drive, solid state drive, or flash memory. The communication interface 204 is an interface for connecting to a communication network, and is, for example, an adapter for connecting to Ethernet (registered trademark), a modem for connecting to a public telephone network, a wireless communication device for wireless communication, or a USB (Universal Serial Bus) connector or RS232C connector for serial communication. The input device 205 is, for example, a keyboard, mouse, touch panel, button, microphone, or the like for inputting data. The output device 206 is, for example, a display, printer, speaker, or the like for outputting data. Each functional unit of the management server 2 described below is realized by the CPU 201 reading a program stored in the storage device 203 into the memory 202 and executing it, and each storage unit of the management server 2 is realized as part of the storage area provided by the memory 202 and the storage device 203.
[0015] 3 is a diagram illustrating an example of the software configuration of the management server 2. The management server 2 includes a learning model storage unit 231, a detection unit 211, a calculation unit 212, an estimation unit 213, and an output unit 214.
[0016] <Memory unit> The learning model memory unit 231 stores a first learning model for detecting the degree of the attributes of a user appearing in a video image (hereinafter referred to as attribute degree), and a second learning model for estimating the mental state of the user based on the detected attribute degree.
[0017] The first and second learning models in this embodiment are created by machine learning. Machine learning can be broadly divided into supervised learning and unsupervised learning. Supervised learning is a method of training a model using input data and corresponding output data (supervisor data), and by adjusting model parameters based on the supervised data, a mapping from input data to output data is learned. On the other hand, unsupervised learning is a method of learning the structure and pattern of input data without using supervised data, and learning the density distribution and feature representation of the input data. As the first and second learning models in this embodiment, for example, neural networks, support vector machines, decision trees, random forests, and the like, which are types of supervised learning, can be used. Self-organizing maps and k-means methods, which are types of unsupervised learning, can also be used. These machine learning algorithms perform processes such as weighting and transformation on the input data to learn the relationship between the input data and output data, and optimize parameters to minimize the error with the output data. This makes it possible to learn the mapping from input data to output data.
[0018] The first learning model can be created by machine learning using features extracted from video and attribute degrees as training data. Features input to the first learning model include, for example, images of the face region extracted from each frame of the video, the positions and sizes of organs extracted from the face region, such as the eyes, nose, and mouth, and temporal changes in these organs. On the other hand, attribute degrees output by the first learning model include, for example, the number of blinks, the degree of eye opening, the degree of mouth opening, the position of the eyebrows, the direction of the face, and temporal changes therein. Specifically, the first learning model can input features such as images of the face region and the positions and sizes of organs, and output attribute degrees such as the number of blinks and the degree of eye opening. Attributes can include the number of blinks, eye offset (the angle of the eyes relative to the camera), gaze estimated from the eye offset, and facial expression. Facial expressions include anger, disgust, fear, happiness, sadness, surprise, neutral, negative / positive, etc., and it is possible to infer the number of times these expressions appear in a given period (such as one second), the average or median of their direction, and the degree of emotion (facial expression) such as anger.
[0019] The second learning model is a learning model that estimates a mental state (degree of depression) when given at least one of the degree of an attribute (value), a statistical value (such as a standard deviation) related to the degree of the attribute, and a statistical value of the statistical value (such as a standard deviation of standard deviations). In this embodiment, the feature values given to the second learning model include at least a statistical value of a statistical value (such as a standard deviation of standard deviations). The second learning model can be created by machine learning using at least one of the degree of an attribute, a statistical value related to the degree of the attribute, and a statistical value of the statistical value, as well as a mental state determined by an expert, as training data.
[0020] <Functional Unit> The detection unit 211 analyzes the video to detect the degree of a user's attribute. The detection unit 211 can detect the degree of the attribute for each section of a predetermined length. The detection unit 211 can estimate the degree of the attribute by providing the video to a first learning model.
[0021] The calculation unit 212 calculates a second degree of variation of a first degree of variation related to the degree of an attribute in the video. Specifically, the calculation unit 212 calculates a first degree of variation (e.g., a standard deviation of the degrees of the attribute) that is the degree of variation of the degree of the attribute, and further calculates a second degree of variation (e.g., a standard deviation of the standard deviations of the degrees of the attribute) that is the degree of variation of the first degree of variation. In this embodiment, the degree of variation is assumed to be a standard deviation, but it may also be a variance. For example, the calculation unit 212 can calculate a first degree of variation that is the standard deviation of the degree of anger, and a second degree of variation that is the standard deviation of the standard deviation.
[0022] The estimation unit 213 estimates the user's mental state based on at least the second degree of variation. The estimation unit 213 can estimate the mental state based on the second degree of variation and at least one of the first degree of variation and the attribute degree. The estimation unit 213 can estimate the mental state by applying at least the second degree of variation to a second learning model.
[0023] The output unit 214 outputs the estimated mental state.
[0024] <Operation> FIG. 4 is a diagram illustrating the operation of the management server 2.
[0025] The management server 2 acquires video images captured by the user (S301), estimates the degree of each attribute of the user for a predetermined period (e.g., one second) based on the acquired video images and the first learning model (S302), and calculates the degree of variation of the estimated degrees (S303). Here, standard deviation or variance can be used as the degree of variation. For example, when calculating the standard deviation of the attribute degrees, the calculation unit 212 calculates the standard deviation from the estimated attribute degrees. On the other hand, when calculating the variance of the attribute degrees, the calculation unit 212 calculates the variance from the estimated attribute degrees. Next, the management server 2 calculates the degree of variation of the calculated degree of variation (S304). For example, when calculating the standard deviation of the standard deviation of the attribute degrees, the calculation unit 212 calculates the standard deviation from the standard deviation of the attribute degrees. On the other hand, when calculating the variance of the variance of the attribute degree, the calculation unit 212 calculates the variance from the variance of the attribute degrees. Then, the management server 2 provides at least the degree of variation of the degree of variation (as well as the degree of each attribute and / or the degree of variation of the degree) to the second learning model to estimate the user's mental state (S305), and outputs the estimated mental state (S306).
[0026] As described above, the information processing system of this embodiment can estimate a user's mental state from video images of the user, allowing for easy estimation of the user's mental state without using tests such as QIDS. Furthermore, the information processing system of this embodiment can perform estimation using the standard deviation of the standard deviations of the attribute degrees as a feature. For example, a user whose mental state is deteriorating may overreact to certain topics while showing no interest in other topics. By evaluating the standard deviation of the standard deviations (the degree of variation of the degree of variation), it is possible to evaluate the degree of variation in facial expressions, etc. over time, which is expected to improve the accuracy of mental state estimation.
[0027] Although the present embodiment has been described above, the above embodiment is intended to facilitate understanding of the present invention and is not intended to limit the present invention. The present invention may be modified or improved without departing from the spirit thereof, and equivalents thereof are also included in the present invention.
[0028] For example, the processing by each of the functional units of the management server 2 described above may be performed by any of the functional units. Also, a different functional unit that performs part of the processing by each of the functional units described above may be added. Also, the functional units of the management server 2 may be distributed across multiple computers.
[0029] Furthermore, the information stored in each storage unit of the management server may be stored in any of the storage units. That is, the information stored in the above-mentioned multiple storage units may be stored in one storage unit, or part of the information stored in one storage unit may be stored in another storage unit.
[0030] <Modification 1>
[0031] In the above embodiment, the degree of a user's attributes is detected from a moving image, but this is not limiting. In addition to a moving image, the degree of a user's attributes may also be detected from audio. For example, the user's audio may be acquired using an audio input device such as a microphone, and the degree of attributes such as the user's voice volume, voice intonation, and speaking speed may be detected from the acquired audio. For example, the amplitude, sound pressure, and volume of the audio may be used to measure the degree of voice volume. For example, the amount of change in the fundamental frequency (pitch) of the audio may be used to measure the degree of voice intonation. For example, the number of syllables or the number of words per unit time may be used to measure the degree of speaking speed.
[0032] As in the above embodiment, a first degree of variation (such as a standard deviation or variance) and a second degree of variation (such as a standard deviation of standard deviations or a variance of variances) can be calculated for the attribute degrees based on the detected audio, and these values can be used to estimate the mental state. For example, the estimation unit 213 can estimate the user's mental state by using the attribute degrees detected from the audio, as well as the attribute degrees detected from the video, and their variances. This makes it possible to improve the accuracy of estimating the mental state by using information that cannot be obtained from the video alone.
[0033] <Modification 2>
[0034] In the above embodiment, the degree of one attribute (e.g., degree of smile, degree of eye opening, etc.) is used as the degree of an attribute, but a value obtained by combining multiple attributes may also be used as the degree of an attribute. For example, the detection unit 211 can detect the degree of smile and the degree of eye opening from a moving image and calculate a value obtained by combining these values as the degree of an attribute. For example, the combination of the smile degree and the eye opening degree can be a weighted sum of the smile degree and the eye opening degree, or a product of the smile degree and the eye opening degree.
[0035] As in the above embodiment, a first degree of variation (such as a standard deviation or variance) and a second degree of variation (such as a standard deviation of standard deviations or a variance of variances) can be calculated for the degree of combination of multiple attributes calculated in this way, and these values can be used to estimate the mental state. For example, the estimation unit 213 can estimate the user's mental state by using the degree of combination of the smile level and the eye opening level and the degree of variation thereof.
[0036] Furthermore, a value that combines three or more attributes may be used as the degree of an attribute. For example, a value that combines the degree of smile, the degree of eye opening, and the volume of the voice may be used as the degree of an attribute. This makes it possible to improve the accuracy of estimating mental states by using composite information that cannot be obtained from a single attribute alone.
[0037] <Modification 3>
[0038] In the above embodiment, the management server 2 acquires a video from the user terminal and detects the degree of an attribute from the video, but this is not limited to this. The user terminal 1 may detect the degree of an attribute from the video and transmit the detected degree of the attribute to the management server 2.
[0039] Specifically, the user terminal 1 captures an image of the user using an imaging device such as a camera, and detects the degree of the user's attribute from the captured video image in the same manner as in the above embodiment. The user terminal 1 then transmits the detected degree of the attribute to the management server 2. The degree of the attribute may be detected, for example, at a predetermined time interval (for example, every second), and the degree of the attribute for each detected time interval may be transmitted to the management server 2.
[0040] Based on the attribute degrees received from the user terminal 1, the management server 2 calculates, as in the above embodiment, a first degree of variation (such as a standard deviation or variance) of the attribute degrees and a second degree of variation (such as a standard deviation of standard deviations or a variance of variances) of the first degree of variation, and estimates the user's mental state using the calculated degree of variation.
[0041] Alternatively, the user terminal 1 may be provided with all of the functional units and storage units of the management server 2 without providing the management server 2, and the user terminal 1 may detect the degree of attributes and also detect the mental state. In this case, the management server 2 may be provided with the functions of managing the learning model and providing input data to the learning model, and the function of sending input data to the management server 2 and generating answers may be delegated to the management server 2.
[0042] <Modification 4>
[0043] In the above embodiment, the mental state of the current user is estimated based on the degree of the current user's attributes and the degree of variation thereof, but this is not limited to this. The mental state of the user in the future may also be estimated based on the degree of the current and past attributes and the degree of variation thereof.
[0044] Specifically, the management server 2 acquires the attribute levels and their variability over a predetermined period up to the present (e.g., the most recent week), and then estimates not only the current mental state but also the future mental state based on the acquired attribute levels and their variability.
[0045] One example of a method for estimating a future mental state is to predict the future attribute intensity and its variability from the current and past attribute intensities and their variability, and then estimate the future mental state based on the predicted future attribute intensity and its variability. Predicting the attribute intensity and its variability can be achieved using time-series data analysis techniques (e.g., ARIMA model, RNN, etc.).
[0046] Furthermore, future mental states may be predicted directly from changes in current and past mental states. For example, the estimation unit 213 may estimate changes in mental states over a predetermined period up to the present from the degree of an attribute and its variability over that period, and predict future mental states based on the estimated changes in mental states. Time-series data analysis techniques may also be used to predict future mental states from changes in mental states.
[0047] <Modification 5>
[0048] In the above embodiment, the degree of depression is estimated as the mental state, but this is not limited to this. Indicators other than the degree of depression, such as the degree of stress or concentration, may also be estimated as the mental state.
[0049] The stress level can be estimated based on the degree and the degree of variation of attributes such as the number and frequency of a user's blinks, the degree to which the eyes are open, the degree to which the mouth is open, the direction of the face, the volume and intonation of the voice, the speed of speech, etc. Generally, users who are highly stressed tend to blink frequently and frequently, have their eyes and mouths wide open, have an inconsistent direction of the face, have an unstable volume and intonation of the voice, and speak at a fast speed, so it is possible to estimate the stress level from the degree and the degree of variation of these attributes.
[0050] Concentration can be estimated based on the degree and variability of attributes such as the amount and speed of movement of the user's gaze, the degree of fixation of the gaze, the number and frequency of blinks, etc. Generally, users with high concentration tend to have a small amount and speed of movement of the gaze, a long time when the gaze is fixed on one point, and a small number and frequency of blinks, so it is possible to estimate concentration from the degree and variability of these attributes.
[0051] The estimation unit 213 can estimate the user's stress level and concentration level using a learning model that has learned the relationship between the degree of the above attributes, their variability, and the stress level and concentration level, thereby making it possible to understand the user's mental state from multiple angles.
[0052] <Disclosure> The present disclosure also includes the following configurations. [Item 1] An information processing system comprising: a detection unit that analyzes video images of a user to detect a degree of an attribute of the user; a calculation unit that calculates a second degree of variation in a first degree of variation related to the degree of the attribute in the video images; and an estimation unit that estimates a mental state of the user based on at least the second degree of variation. [Item 2] The information processing system according to item 1, wherein the detection unit detects the degree of the attribute for each section of a predetermined length. [Item 3] The information processing system according to item 1, wherein the first and second degrees of variation are expressed by a standard deviation. [Item 4] The information processing system according to item 1, wherein the estimation unit estimates the mental state based on the second degree of variation and at least one of the first degree of variation and the degree of the attribute. [Item 5] The information processing system according to Item 1, wherein the estimation unit estimates the mental state by applying at least the second degree of variation to a learning model created by machine learning using at least the second degree of variation and the mental state as training data. [Item 6] An information processing method, characterized in that a computer executes the following steps: analyzing a video of a user to detect a degree of an attribute of the user; calculating a second degree of variation of a first degree of variation related to the degree of the attribute in the video; and estimating the mental state of the user based on at least the second degree of variation.
[0053] 1 User terminal 2 Management server
Claims
1. An information processing system comprising: a detection unit that analyzes video images of a user to detect the degree of the user's attributes; a calculation unit that calculates a second degree of variation of a first degree of variation related to the degree of the attributes in the video images; and an estimation unit that estimates the mental state of the user based on at least the second degree of variation.
2. An information processing system according to claim 1, wherein the detection unit detects the degree of the attribute for each section of a predetermined length.
3. An information processing system according to claim 1, wherein the first and second degrees of variation are expressed by standard deviations.
4. An information processing system according to claim 1, characterized in that the estimation unit estimates the mental state based on the second degree of variation and at least one of the first degree of variation and the degree of the attribute.
5. An information processing system according to claim 1, wherein the estimation unit estimates the mental state by applying at least the second degree of variation to a learning model created by machine learning using at least the second degree of variation and the mental state as training data.
6. An information processing method characterized by the computer executing the following steps: analyzing video images of a user to detect the degree of the user's attributes; calculating a second degree of variation of a first degree of variation related to the degree of the attributes in the video images; and estimating the mental state of the user based on at least the second degree of variation.
Citation Information
Patent Citations
Stress evaluation device and method
JP2018057510A
Wakefulness estimation device, wakefulness estimation method and wakefulness estimation system
JP2018130342A
System and method for camera-based stress determination
US20210361208A1
System, method, and program for predicting diminished attentiveness state, and storage medium on which program is stored
WO2018097204A1