Facial expression recognition device, interactive robot, facial expression recognition system, facial expression recognition method, and program

The facial expression recognition device improves accuracy by using culturally or individually tailored models, addressing the variability in expression styles among individuals.

JP7790034B2Active Publication Date: 2025-12-23SINTOKOGIO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021091647
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-31
Publication Date
2025-12-23
Estimated Expiration
2041-05-31

AI Technical Summary

Technical Problem

Existing facial expression recognition technologies do not account for the variability in expression styles among individuals, leading to inaccurate recognition.

Method used

A facial expression recognition device that uses multiple recognition models tailored to specific cultural spheres or individuals, identified through facial images or voice analysis, to improve accuracy.

Benefits of technology

Enhances facial expression recognition accuracy by utilizing culturally or individually specialized models, accounting for diverse expression styles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007790034000001
    Figure 0007790034000001
  • Figure 0007790034000002
    Figure 0007790034000002
  • Figure 0007790034000003
    Figure 0007790034000003
Patent Text Reader

Abstract

To recognize expression with high accuracy according to a recognition object person.SOLUTION: An expression recognition device executes image acquisition processing (S104) for acquiring a face image and recognition result output processing (S103, S105 to S111) for outputting a recognition result of recognizing a human expression using a recognition model according to the human or an identification result of a human's attribute, among multiple recognition models that are different from each other, each of which outputs a recognition result of an expression with the face image as an input.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technique for recognizing facial expressions. [Background technology]

[0002] There are known techniques for recognizing facial expressions. For example, Patent Document 1 describes a technique in which images of characteristic parts such as eyebrows and eyes are extracted from a face image, and facial expression elements such as eyebrow movement and eye opening / closing are extracted and quantified from the characteristic part images, and emotions are determined by referring to the quantified facial expression elements. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 06-076058 Summary of the Invention [Problem to be solved by the invention]

[0004] Here, there is a possibility that the way of expressing facial expressions is not universal. For example, there is a possibility that the way of expressing facial expressions varies depending on individuals or their attributes. However, the technology described in Patent Document 1 does not take into consideration the possibility that the way of expressing facial expressions is not universal, and therefore, there are cases where accurate recognition is not possible depending on the person to be recognized.

[0005] An object of one aspect of the present invention is to realize a technology for recognizing facial expressions with higher accuracy depending on the person to be recognized. [Means for solving the problem]

[0006] In order to solve the above problems, a facial expression recognition device according to one aspect of the present invention includes one or more processors. The one or more processors execute an image acquisition process and a recognition result output process. Also, a facial expression recognition method according to one aspect of the present invention is a method executed by the one or more processors. The facial expression recognition method includes an image acquisition step and a recognition result output step.

[0007] In an image acquisition process (image acquisition step), the one or more processors acquire a facial image containing a human face as a subject. In a recognition result output process (recognition result output step), the one or more processors output a recognition result in which the facial expression of the person is recognized using a recognition model according to an identification result of the person or an attribute of the person, from among a plurality of mutually different recognition models, each of which takes a facial image as input and outputs a recognition result of facial expression.

[0008] In order to solve the above problems, a facial expression recognition system according to one aspect of the present invention includes a camera that captures an image of a person's face and generates a facial image, a facial expression recognition device that recognizes the person's facial expression by referring to the facial image, and an output device that outputs a recognition result by the facial expression recognition device. The facial expression recognition device executes an image acquisition process that acquires the facial image from the camera, and a recognition result output process that recognizes the person's facial expression using one of a plurality of different recognition models, each of which receives a facial image as an input and outputs a recognition result of the facial expression, according to an identification result of the person or an attribute of the person, and outputs the recognition result to the output device. [Effects of the Invention]

[0009] According to one aspect of the present invention, facial expressions can be recognized with higher accuracy depending on the person to be recognized. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a block diagram showing the configuration of a facial expression recognition device according to a first embodiment of the present invention. [Figure 2]FIG. 2 is a diagram illustrating an example of information indicating the correspondence between cultural areas and recognition models according to the first embodiment of the present invention. [Figure 3] 1 is a flowchart showing the flow of a facial expression recognition method according to a first embodiment of the present invention. [Figure 4] FIG. 4 is a diagram showing an example of a selection flag according to the first embodiment of the present invention. [Figure 5] FIG. 3 is a diagram showing an example of an identification flag according to the first embodiment of the present invention. [Figure 6] 1 is a flowchart showing the flow of an identification method according to a first embodiment of the present invention. [Figure 7] FIG. 4 is a diagram showing an example of a classification result in the first embodiment of the present invention. [Figure 8] FIG. 10 is a diagram illustrating an example of information indicating the correspondence between an individual and a recognition model according to the second embodiment of the present invention. [Figure 9] FIG. 10 is a block diagram showing the configuration of a conversational robot 1 according to a third embodiment of the present invention. [Figure 10] FIG. 11 is a diagram illustrating an example of information indicating the correspondence between dementia levels and recognition models according to the third embodiment of the present invention. [Figure 11] FIG. 10 is a block diagram showing the configuration of a facial expression recognition device according to a modified example of each embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] [Embodiment 1] A first embodiment of the present invention will be described below with reference to the drawings.

[0012] <Configuration of facial expression recognition device> The configuration of a facial expression recognition device 10 according to a first embodiment of the present invention will be described with reference to FIG. 1. FIG. 1 is a block diagram showing the configuration of the facial expression recognition device 10. The facial expression recognition device 10 is an example of a form realizing the "facial expression recognition device" set forth in the claims. As shown in FIG. 1, the facial expression recognition device 10 includes a processor 11, a primary memory 12, a secondary memory 13, and an input / output interface 14. The processor 11, the primary memory 12, the secondary memory 13, and the input / output interface 14 are interconnected via a bus. The facial expression recognition device 10 is also connected to a sensor 50, a camera 60, a microphone 70, and an output device 80 via the input / output interface 14.

[0013] The secondary memory 13 stores a program P1 and other information. The processor 11 executes each process included in a facial expression recognition method S1 and a classification method S2 (described later) in accordance with instructions included in the program P1. Details of the other information stored in the secondary memory 13 will be described later. Devices that can be used as the processor 11 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), or a combination thereof.

[0014] Furthermore, devices that can be used as the primary memory 12 include, but are not limited to, semiconductor RAM (Random Access Memory). Devices that can be used as the secondary memory 13 include, but are not limited to, flash memory, HDD (Hard Disk Drive), SSD (Solid State Drive), or a combination thereof.

[0015] Furthermore, the input / output interface 14 may be, for example, an interface such as a USB (Universal Serial Bus), but is not limited to this.

[0016] The sensor 50 outputs a detection signal indicating that a person has been detected to the facial expression recognition device 10. For example, the sensor 50 includes an infrared sensor or an ultrasonic sensor.

[0017] The camera 60 captures the surroundings to generate captured images, and outputs the generated captured images to the facial expression recognition device 10. The camera 60 may output the captured images as still images, or may sequentially output the captured images generated at a predetermined frame rate as moving images.

[0018] The microphone 70 detects surrounding sounds and outputs the detected sounds to the facial expression recognition device 10. The sounds input from the microphone 70 are stored in the secondary memory 13.

[0019] The output device 80 outputs the information generated by the facial expression recognition device 10. The output device 80 includes, for example, a display or a speaker.

[0020] The following describes the recognition models M1-0, M1-1, M1-2, ... stored in the secondary memory 13. The recognition models M1-0, M1-1, M1-2, ... are multiple mutually different recognition models. Hereinafter, when there is no need to particularly distinguish between these recognition models, each will also be simply referred to as recognition model M1.

[0021] The recognition model M1 is a model that receives a facial image as input and outputs a facial expression recognition result. A machine learning algorithm is used to generate the recognition model M1. Specific examples of the machine learning algorithm include neural networks such as a convolutional neural network (CNN) and a recurrent neural network (RNN), a support vector machine, and a random forest. However, the machine learning algorithm used to generate the recognition model M1 is not limited to these. The machine learning algorithm used to generate the recognition model M1 may be supervised learning or unsupervised learning. The machine learning algorithm used to generate each recognition model M1 may be the same as or different from the machine learning algorithm used to generate at least one other recognition model M1. Some or all of the multiple recognition models M1 may be generated by the facial expression recognition device 10 or by another device.

[0022] For example, the recognition model M1 outputs information indicating the classification of emotions, which may be, for example, but are not limited to, six basic emotions (anger, disgust, fear, joy, sadness, and surprise).

[0023] Each of the recognition models M1-1, M1-2, ... is generated by machine learning using facial images of people belonging to a specific cultural sphere as training data. Here, cultural sphere is an example of "person attributes" as defined in the claims. Such training data may include multiple facial images of at least one person belonging to the cultural sphere, but it is preferable that it includes facial images of each of multiple people belonging to the cultural sphere. Hereinafter, when there is no need to particularly distinguish between the recognition models M1-1, M1-2, ..., each will also be referred to as a "cultural sphere-specialized recognition model M1." Each of the multiple cultural sphere-specialized recognition models M1 is specialized for a different cultural sphere.

[0024] The recognition model M1-0 is a general-purpose model that is not specialized for a particular cultural sphere. For example, the recognition model M1-0 is generated by machine learning using facial images of a plurality of people belonging to various cultural spheres as training data. The cultural spheres to which the people who are the subjects of the facial images used as training data belong should be at least two types, but it is desirable to have more types.

[0025] Table T1 stored in the secondary memory 13 will be described. Table T1 stores information indicating the correspondence between cultural areas and recognition models M1. An example of table T1 will be described with reference to FIG. 2. FIG. 2 is a diagram illustrating an example of information indicating the correspondence between cultural areas and recognition models M1. In the example of FIG. 2, the cultural area "East Asia" is associated with a recognition model M1-1 having an ID of M1-1, and the cultural area "Europe" is associated with a recognition model M1-2 having an ID of M1-2.

[0026] Note that the cultural sphere is not limited to a unit spanning multiple countries, but may be a country, state, region, prefecture, city, town, or village. Furthermore, at least one of the multiple cultural spheres may not have the same granularity as at least one of the other cultural spheres. For example, a recognition model ID may be associated with each of East Asia, Europe, Eastern Japan, and Western Japan.

[0027] <Flow of facial expression recognition method S1> The flow of the facial expression recognition method S1 executed by the facial expression recognition device 10 configured as above will be described with reference to Fig. 3. Fig. 3 is a flow diagram showing the flow of the facial expression recognition method S1 according to this embodiment. As shown in Fig. 3, the facial expression recognition method S1 includes steps S101 to S112.

[0028] In step S101, the processor 11 refers to the detection signal input from the sensor 50 to detect a person.

[0029] In step S102, the processor 11 initializes a selection flag to start processing related to facial expression recognition for the person detected in step S101.

[0030] The selection flag will be described with reference to FIG. 4. FIG. 4 is a diagram showing an example of a selection flag used in the facial expression recognition method. The selection flag is information indicating whether one of a plurality of recognition models M1 is "selected" or "unselected," and its initial state is "unselected." In the example of FIG. 4, the selection flag indicates "selected." Therefore, the processor 11 initializes the selection flag by setting it to "unselected."

[0031] In step S103 (identification step), processor 11 starts executing an identification process to identify the cultural sphere of the person detected in step S101. Details of the identification process will be described later. Note that processor 11 may execute the process of the next step S104 even if the started identification process has not been completed. In other words, processor 11 executes the identification process and the processes from step S104 onwards in parallel until the identification process is completed.

[0032] In step S104 (image acquisition step), processor 11 executes image acquisition processing to acquire a face image by referring to the captured image input from camera 60. Specifically, processor 11 detects a human face area from the captured image and extracts the area as a face image. A known technique can be used to detect the face area. For example, processor 11 may detect the face area by dividing the captured image, extracting features indicating a face from each area, and determining whether each area is a face based on the extracted features. However, the technique for acquiring a face image by referring to the captured image is not limited to this.

[0033] In step S105, the processor 11 refers to the selection flag to determine whether or not any of the multiple recognition models M1 has been selected.

[0034] If the determination in step S105 is No, in step S106, the processor 11 refers to the identification flag and determines whether the cultural sphere of the person detected in step S101 has been identified.

[0035] The identification flag will be described with reference to FIG. 5. FIG. 5 is a diagram showing an example of an identification flag used in the facial expression recognition method. The identification flag is information indicating whether a cultural sphere is "identified" or "unidentified." In the example of FIG. 5, the identification flag indicates that identification has been performed. If the identification flag indicates "identified," the identification process started in step S103 has already ended. On the other hand, if the identification flag indicates "unidentified," the identification process started in step S103 has not yet ended.

[0036] If the determination in step S106 is Yes, in step S107, the processor 11 refers to the table T1 and determines whether or not there is a recognition model M1 specialized for the cultural sphere indicated by the identification result.

[0037] If the determination in step S107 is Yes, in step S108 (selection step), the processor 11 refers to table T1 and executes a selection process to select a recognition model M1 according to the identification result from among multiple recognition models M1. That is, the processor 11 selects a recognition model M1 specialized for the cultural sphere indicated by the identification result. The processor 11 also sets the selection flag to "selected."

[0038] If the determination in step S106 is No or the determination in step S107 is No, in step S109, processor 11 selects a general-purpose recognition model M1-0. Also, processor 11 sets the selection flag to "selected." Here, the determination in step S106 is No if, after a person is detected, the identification process for identifying the cultural sphere of the person has not been completed. The determination in step S107 is No if, after a person is detected, the identification process has been completed, but a recognition model M1 specialized for the cultural sphere indicated by the identification result has not been prepared.

[0039] In step S110 (recognition step), the processor 11 executes recognition processing to recognize the facial expression of the person detected in step S101 using the selected recognition model M1. Specifically, the face image acquired in step S104 is input to the selected recognition model M1, and the facial expression recognition result output from the recognition model M1 is obtained.

[0040] In step S111, processor 11 outputs the facial expression recognition result obtained in step S110 to output device 80. For example, processor 11 displays the facial expression recognition result on a display included in output device 80.

[0041] In step S112, the processor 11 refers to the detection signal from the sensor 50 and determines whether or not the person detected in step S101 is still being detected. For example, if the detection signal from the sensor 50 is still being received, the processor 11 determines that detection is still being performed. "Still being received" may mean, for example, that a period during which a detection signal cannot be received is within a predetermined length. However, the process of determining whether or not detection is still being performed is not limited to the above.

[0042] If the determination in step S112 is Yes, processor 11 repeats the process from step S104. In step S104, a face image different from the face image obtained in the previous step S104 is obtained for the person detected in step S101. In addition, in step S105, it is determined that recognition model M1 has been selected. Then, processor 11 executes the processes of steps S110 to S111 using the selected recognition model M1.

[0043] If the determination in step S112 is No, the processor 11 ends the facial expression recognition method S1.

[0044] As a result, the facial expression recognition device 10 uses a recognition model M1 specialized for the cultural sphere of the detected person to recognize and output the recognition results of the facial expression of the person in real time while detecting the person.

[0045] <Flow of identification method S2> Next, the flow of the classification method S2 for executing the classification process starting in step S103 will be described with reference to Fig. 6. Fig. 6 is a flow diagram showing the flow of the classification method S2 included in the facial expression recognition method S1 according to this embodiment. As shown in Fig. 6, the classification method S2 includes steps S201 to S206.

[0046] In step S201, the processor 11 initializes an identification flag to start the identification process. The identification flag has been described with reference to Fig. 5. If the identification flag indicates "identified," the processor 11 initializes the identification flag by setting it to "unidentified."

[0047] In step S202 (voice acquisition step), the processor 11 acquires voice input from the microphone 70. The voice to be acquired is voice including the human speech detected in step S101. For example, the processor 11 acquires a predetermined length of voice up to the present from the voice input from the microphone 70 and stored in the secondary memory 13.

[0048] In step S203 (identification step), processor 11 refers to the acquired voice and identifies the attribute (cultural area in this case) of the person detected in step S101. Specific examples 1 to 4 of the method for identifying cultural area by referring to voice will be described.

[0049] In specific example 1, processor 11 analyzes the acquired speech to determine the intonation, and identifies the cultural sphere according to the determined intonation as the identification result. For example, processor 11 may extract features indicating the intonation from the speech, and identify the cultural sphere by comparing the extracted features with intonation features specific to the cultural sphere registered in a database. Note that such a database may be stored in secondary memory 13 or in an external device.

[0050] In specific example 2, processor 11 extracts keywords contained in the acquired voice and identifies the cultural sphere corresponding to the extracted keywords. For example, processor 11 may identify the cultural sphere by comparing the extracted keywords with keywords specific to the cultural sphere registered in a database. Note that such a database may be stored in secondary memory 13 or in an external device.

[0051] In specific example 3, processor 11 determines the language of the utterance contained in the acquired voice and identifies the cultural area according to the language. For example, processor 11 may identify the cultural area by extracting features indicating the language from the voice and comparing the extracted features with features of languages ​​specific to the cultural area registered in a database. Note that such a database may be stored in secondary memory 13 or in an external device.

[0052] In specific example 4, the processor 11 identifies a cultural area using a machine-learned identification model that takes a voice as input and outputs a cultural area.

[0053] Furthermore, the processor 11 may recognize a cultural sphere by combining some or all of the specific examples 1 to 4.

[0054] In step S204, processor 11 determines whether or not the identification in step S203 was successful. Successful identification means that a cultural sphere was identified as the identification result.

[0055] If the determination in step S204 is No, the processor 11 repeats the process from step S202.

[0056] If the determination in step S204 is Yes, in step S205, the processor 11 sets the identification flag to "identified."

[0057] In step S206, the processor 11 sets the identification result. The identification result is stored in the primary memory 12 or the secondary memory 13. The identification result will be described with reference to FIG. 7. FIG. 7 is a diagram showing an example of the identification result. As shown in FIG. 7, the identification result indicates the cultural sphere identified in step S203. In this example, the identification result is "East Asia".

[0058] With this, the processor 11 ends the identification method S2.

[0059] <Effects of this embodiment> In this way, the present embodiment recognizes the facial expression of a person using a recognition model according to the person or the attributes of the person, thereby enabling the present embodiment to recognize the facial expression of the person with higher accuracy according to the person to be recognized.

[0060] Furthermore, this embodiment identifies the cultural sphere of the detected person and recognizes the facial expression of the detected person from the facial image of the person using a recognition model specialized for the identified cultural sphere. As a result, this embodiment takes into account the possibility that facial expressions may be expressed differently depending on the cultural sphere, and can recognize the facial expression of a person with higher accuracy according to the cultural sphere of the person to be recognized.

[0061] Furthermore, in this embodiment, a person or a person's attributes are identified by referring to a face image, which allows the present embodiment to realize an identification process for selecting an appropriate recognition model by referring to the face image.

[0062] [Embodiment 2] A second embodiment of the present invention will be described below. The second embodiment is a modification of the first embodiment. In the first embodiment, a person's cultural sphere is identified, and the facial expression of that person is recognized using an identification model specialized for the identified cultural sphere. This embodiment is a modification of this, in which a person is identified to identify an individual, and the facial expression of that person is recognized using an identification model specialized for the identified individual.

[0063] Below, differences from the first embodiment will be described, and explanation of similarities with the first embodiment will not be repeated.

[0064] In this embodiment, the recognition models M1-1, M1-2, ... are each generated by machine learning using a facial image of a specific person as training data. Such training data includes multiple facial images of the specific person. Hereinafter, when there is no need to particularly distinguish between the recognition models M1-1, M1-2, ..., each will also be referred to as an "individual-specific recognition model M1." Each of the multiple individual-specific recognition models M1 is specialized for a different person.

[0065] In this embodiment, the recognition model M1-0 is a general-purpose model that is not specialized for an individual. For example, the recognition model M1-0 is generated by machine learning using face images of multiple people as training data. Face images of at least two people are used as training data, but it is preferable to use face images of more people.

[0066] Furthermore, in this embodiment, table T1 stores information indicating the correspondence between individuals and recognition models M1. An example of table T1 will be described with reference to FIG. 8. FIG. 8 is a diagram illustrating an example of information indicating the correspondence between individuals and recognition models. In the example of FIG. 8, user ID "001" is associated with recognition model M1-1 having ID M1-1, and user ID "002" is associated with recognition model M1-2 having ID M1-2.

[0067] <Flow of facial expression recognition method S1> The facial expression recognition method S1 according to this embodiment is almost the same as the facial expression recognition method S1 described with reference to FIG. 3, but the processes in steps S103 and S106 to S108 are slightly different.

[0068] In step S103, the processor 11 starts executing an identification process to identify the person detected in step S101 and specify the user ID.

[0069] In step S106, the processor 11 refers to the identification flag to determine whether the user ID of the person detected in step S101 has been identified.

[0070] In step S107, the processor 11 refers to the table T1 and determines whether or not there is a recognition model M1 that is specialized for the individual indicated by the identified user ID.

[0071] In step S108, the processor 11 refers to the table T1 and executes a selection process to select a recognition model M1 specialized for the individual indicated by the identified user ID.

[0072] As a result, the facial expression recognition device 10 uses the recognition model M1 specialized for the detected person to recognize and output the recognition result of the facial expression of the person in real time while detecting the person.

[0073] <Flow of identification method S2> The identification method S2 according to this embodiment is almost the same as the identification method S2 described with reference to Fig. 6. However, the voice acquisition process in step S202 is omitted, and the process in S203 is slightly different.

[0074] In step S203, the processor 11 refers to the face image acquired in step S104, identifies the person detected in step S101, and specifies the user ID. A specific example of a method for identifying a person by referring to a face image will be described.

[0075] For example, the processor 11 may identify the user ID by comparing the facial image acquired in step S104 with a facial image of an individual registered in a database in association with the user ID.

[0076] <Effects of this embodiment> In this embodiment, a detected person is identified to specify a user ID, and a recognition model according to the individual indicated by the specified user ID is used to recognize the facial expression of the detected person from the facial image of the person. As a result, this embodiment takes into account the possibility that different individuals express their facial expressions differently, and can recognize the facial expressions of people with higher accuracy according to the individual.

[0077] [Embodiment 3] A third embodiment of the present invention will be described below. In the third embodiment, a facial expression recognition device 10, which is a modification of the first embodiment, is mounted on an interactive robot 1 that interacts with a care recipient. If the care recipient has dementia, it is possible that the way the care recipient expresses their facial expressions will differ depending on the level of dementia. Therefore, the interactive robot 1 according to this embodiment identifies the dementia level of the care recipient with whom they are interacting, and recognizes the facial expressions of the care recipient using a recognition model specialized for the identified dementia level. The dementia level is an example of a "person attribute" as defined in the claims.

[0078] Below, differences from the first embodiment will be described, and explanation of similarities with the first embodiment will not be repeated.

[0079] <Configuration of conversational robot 1> The configuration of the interactive robot 1 will be described with reference to Fig. 9. Fig. 9 is a block diagram showing the configuration of the interactive robot 1. As shown in Fig. 9, the interactive robot 1 includes a facial expression recognition device 10. A processor 11 of the facial expression recognition device 10 mounted on the interactive robot 1 recognizes the facial expression of a care recipient, who is the conversation partner, and executes output processing to output information according to the recognition result.

[0080] The facial expression recognition device 10 mounted on the interactive robot 1 is configured in almost the same manner as in the first embodiment, except for the following points.

[0081] In this embodiment, the recognition models M1-1, M1-2, ... are each generated by machine learning using training data that is a facial image of a person diagnosed with a particular level of dementia. Such training data may include a plurality of facial images of at least one person diagnosed with that level of dementia, but it is preferable that the training data include facial images of each of a plurality of people diagnosed with that level of dementia. Hereinafter, when there is no need to particularly distinguish between the recognition models M1-1, M1-2, ..., each will also be referred to as a "recognition model M1 specialized for dementia level." Each of the plurality of dementia level-specialized recognition models M1 is specialized for a different dementia level.

[0082] Furthermore, in this embodiment, the recognition model M1-0 is a general-purpose model that is not specialized for dementia levels. For example, the recognition model M1-0 is generated by machine learning using facial images of multiple people who have been certified as having multiple levels of dementia as training data. The number of dementia levels diagnosed for the multiple people who are the subjects of the facial images used as training data should be at least two, but it is desirable to have more levels.

[0083] Furthermore, in this embodiment, table T1 stores information indicating the correspondence between dementia levels and recognition models M1. An example of table T1 will be described with reference to Fig. 10. Fig. 10 is a diagram illustrating an example of information indicating the correspondence between dementia levels and recognition models M1. In the example of Fig. 10, dementia level "I" is associated with recognition model M1-1 having M1-1 as its ID, and dementia level "II" is associated with recognition model M1-2 having M1-2 as its ID.

[0084] In this embodiment, in addition to the table T1, the secondary memory 13 also stores a user information table. The user information table stores the user IDs of the care recipients in association with their dementia levels.

[0085] <Flow of facial expression recognition method S1 and classification method S2> The facial expression recognition method S1 and classification method S2 according to this embodiment can be explained in the same manner as in the explanation of these methods in embodiment 1 with reference to Figures 3 and 6, by replacing "cultural area" with "dementia level." However, the difference is that the processing in step S202 is omitted and the details of the operations in S203 and S111 are different.

[0086] In step S203, the processor 11 refers to the face image acquired in step S104 to identify the person detected in step S101 and specify the user ID. A specific example of the method for identifying a person by referring to a face image is as described in embodiment 2. The processor 11 also refers to the user information table and determines the dementia level associated with the specified user ID as the identification result.

[0087] In step S111, processor 11 outputs information corresponding to the facial expression recognition result obtained in step S110 to output device 80. For example, processor 11 outputs, to a speaker included in output device 80, a speech corresponding to the facial expression recognition result.

[0088] <Effects of this embodiment> In this embodiment, the interactive robot can recognize the facial expression of the conversation partner with higher accuracy, and therefore can present the conversation partner with information that is more suited to the conversation partner's facial expression.

[0089] In this embodiment, as shown in Fig. 9, the interactive robot 1 identifies the dementia level of the care recipient by referring to the speech of the care recipient. The interactive robot 1 also recognizes the facial expression of the care recipient by inputting a facial image of the care recipient into a recognition model M1 specialized for the identified dementia level. In step S111, the interactive robot 1 also outputs a speech sound according to the recognition result of the facial expression.

[0090] As a result, the conversational robot 1 according to this embodiment can utter a more appropriate response in conversation with a care recipient according to the dementia level of the care recipient.

[0091] [Variation 1] The facial expression recognition device 10 according to each of the above-described embodiments can be modified into a facial expression recognition system 2 including multiple processors. The facial expression recognition system 2 is an example of a configuration in which the "facial expression recognition device" recited in the claims is realized as a device including multiple processors. The facial expression recognition system 2 will be described with reference to FIG. 11. FIG. 11 is a block diagram showing the configuration of the facial expression recognition system 2. As shown in FIG. 11, the facial expression recognition system 2 includes a facial expression recognition device 10A and a server 20.

[0092] <Configuration of Server 20 and Facial Expression Recognition Device 10A> 11, the server 20 includes a processor 21, a primary memory 22, a secondary memory 23, and a communication interface 25. The processor 21, the primary memory 22, the secondary memory 23, and the communication interface 25 are connected to one another via a bus. The server 20 is also connected to the facial expression recognition device 10A via the communication interface 25 so as to be able to communicate with the facial expression recognition device 10A.

[0093] The secondary memory 23 stores a program P2, a table T1, and recognition models M1-0, M1-1, M1-2, .... The processor 21 executes at least a portion of the processes included in the facial expression recognition method S1 and the classification method S2 in accordance with instructions included in the program P2. Details of other information stored in the secondary memory 23 are as described in the first embodiment. Details of the processor 21, the primary memory 22, and the secondary memory 23 are the same as those of the processor 11, the primary memory 12, and the secondary memory 13 described in the first embodiment.

[0094] At least the facial expression recognition device 10A is connected to the communication interface 25 via a network. Examples of the communication interface 25 include, but are not limited to, interfaces such as Ethernet (registered trademark) and Wi-Fi (registered trademark). Available networks include, but are not limited to, a personal area network (PAN), a local area network (LAN), a campus area network (CAN), a metropolitan area network (MAN), a wide area network (WAN), a global area network (GAN), or an internetwork including these networks.

[0095] The facial expression recognition device 10A includes a communication interface 15 in addition to the same configuration as the facial expression recognition device 10 according to the first embodiment. The facial expression recognition device 10A is connected to at least a server 20 via the communication interface 15. Details of the communication interface 15 are the same as those of the communication interface 25. The secondary memory 13 stores a program P1A in place of "program P1, table T1, recognition models M1-0, M1-1, M1-2, . . . ". The processor 11 executes at least a part of the processes included in the facial expression recognition method S1 and the classification method S2 in accordance with instructions included in the program P1A. When the processor 11 needs to refer to "table T1, recognition models M1-0, M1-1, M1-2, . . . " stored in the server 20, the processor 11 references these by communicating with the server 20.

[0096] <Flow of facial expression recognition method S1 and classification method S2> The processor 11 and the processor 21 execute the facial expression recognition method S1 and the classification method S2 by transmitting and receiving necessary information to each other. For example, in the facial expression recognition method S1, step S110 (recognition processing for recognizing facial expressions) may be executed by the processor 21 of the server 20, and the other steps may be executed by the processor 11 of the facial expression recognition device 10A. Also, in the classification method S2, step S203 (classification processing for identifying a person or a person's attributes) may be executed by the processor 21 of the server 20, and the other steps may be executed by the processor 11 of the facial expression recognition device 10A. However, the steps executed by the server 20 and the facial expression recognition device 10A, respectively, are not limited to the examples described above. Details of the facial expression recognition method S1 and the classification method S2 are as described in each embodiment.

[0097] <Effects of this modified example> This modification can recognize human facial expressions with higher accuracy by referring to the recognition model M1 specialized for an individual or their attributes stored in the server 20. With this configuration, for example, the recognition model M1 specialized for an individual or their attributes can be shared with other information processing devices configured similarly to the facial expression recognition device 10A. As an example, when the third embodiment is modified in this way, the interactive robots 1 installed in multiple care facilities can share the recognition model M1 stored in the server 20 and utter responses appropriate to the dementia level of the care recipients.

[0098] [Variation 2] In each of the above-described embodiments, an example has been described in which the processor 11 detects a person by referring to a detection signal input from the sensor 50. The process of detecting a person can be modified as follows. For example, the processor 11 may detect a person by referring to a captured image input from the camera 60. As an example, the processor 11 determines that a person has been detected when it detects that an area showing a person is included in the captured image input from the camera 60. Furthermore, for example, the processor 11 may detect a person by referring to audio input from the microphone 70. As an example, the processor 11 determines that a person has been detected when it detects that the audio input from the microphone 70 includes a person's speech.

[0099] [Variation 3] In the first embodiment described above, the processor 11 identifies the cultural area by referring to the voice input from the microphone 70. The identification process for identifying the cultural area can be modified as follows.

[0100] Processor 11 may identify a person's attributes (e.g., cultural area) by referring to a facial image. For example, processor 11 may estimate a person's race by comparing facial features extracted from the facial image with facial features of races registered in a database, and may identify the cultural area corresponding to the estimated race as an identification result. Processor 11 may also identify the cultural area using an identification model trained by machine learning that receives a facial image as input and outputs a cultural area.

[0101] The processor 11 may also identify the cultural sphere by referring to the installation location of the facial expression recognition device (10, 10A). For example, the processor 11 may store information associating geographical areas with cultural spheres, and may identify the cultural sphere associated with the area including the installation location as the identification result.

[0102] The processor 11 may also identify the cultural sphere based on a user input. For example, the user may input a voice specifying the cultural sphere (e.g., "East Asia") into the microphone 70. In this case, the processor 11 performs a voice recognition process on the input voice to obtain information indicating the cultural sphere (in this example, "East Asia") and sets this as the identification result. The facial expression recognition device (10, 10A) may also include physical buttons corresponding to multiple cultural spheres, and the user may input the cultural sphere by operating one of the operation buttons. In this case, the processor 11 sets the cultural sphere corresponding to the physical button that received the operation as the identification result.

[0103] [Variation 4] In the above-described second embodiment, the processor 11 identifies an individual by referring to a face image. The identification process for identifying an individual can be modified as follows.

[0104] For example, the processor 11 may identify an individual by referring to a voice input from the microphone 70. As an example, the processor 11 may identify an individual by referring to a voice. Specifically, the processor 11 may extract features from the acquired voice and identify the user ID by comparing the extracted features with features of the individual's voice registered in a database in association with the user ID. Alternatively, the processor 11 may identify an individual by using an identification model trained by machine learning to input a voice and output a user ID.

[0105] Furthermore, the processor 11 may identify an individual through a user input. For example, the user may input information for identifying the individual (for example, a name) into the microphone 70. In this case, the processor 11 may perform a voice recognition process on the input voice to acquire the information for identifying the individual (for example, a name), and specify a user ID associated with the acquired information.

[0106] In this modification, the identification process for selecting an appropriate recognition model can be realized by referring to the speech.

[0107] [Variation 5] In addition, in each of the above-described embodiments, an example has been described in which multiple recognition models M1 specialized for individuals or attributes are included. However, in each embodiment, the number of recognition models M1 specialized for individuals or attributes may be one. In this case, in the selection process, if there is no recognition model M1 specialized for the identification result, the processor 11 may select a general-purpose recognition model M1-0.

[0108] [Variation 6] In the above-described first and third embodiments, examples have been described in which cultural background and dementia level are applied as attributes of a person. Each embodiment can be modified to apply other attributes instead. For example, other specific examples of attributes of a person include, but are not limited to, gender, level of care required, and age group.

[0109] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. [Explanation of symbols]

[0110] 1. Conversational robot 2. Facial Expression Recognition System 10, 10A facial expression recognition device 11, 21 processors 12, 22 Primary memory 13, 23 Secondary memory 14 Input / Output Interface 15, 25 Communication Interface 20 servers 50 sensors 60 cameras 70 Mike 80 Output Device

Claims

1. 1. A facial expression recognition device including one or more processors, the one or more processors: a voice acquisition process for acquiring a voice including a speech of the detected person by referring to a detection signal indicating that a person has been detected; an image acquisition process for acquiring a face image including the detected person's face as a subject; a recognition result output process for outputting a recognition result obtained by recognizing the facial expression of the detected person using a recognition model corresponding to an identification result of an attribute of the detected person from among a plurality of different recognition models, each of which receives a face image as an input and outputs a recognition result of a facial expression; Run The detected person is judged by referring to the detection signal whether or not the person is being continuously detected; the identification result of the attribute of the detected person is a cultural sphere identified by referring to the voice acquired in the voice acquisition process; The cultural area is determined by any one of the intonation of the speech, keywords contained in the speech, languages ​​contained in the speech, and a machine-learned discrimination model that takes the speech as input and outputs a cultural area. Facial expression recognition device.

2. the one or more processors: Among the plurality of recognition models, a recognition model generated by machine learning using facial images of people belonging to the cultural sphere as training data is used. The facial expression recognition device according to claim 1 .

3. A conversational robot including the facial expression recognition device according to claim 1 or 2, comprising one or more of the processors, The processor of the interactive robot includes: An interactive robot that at least executes the recognition result output process.

4. A microphone that detects audio including speech from a detected person by referring to a detection signal indicating that a person has been detected; a camera that captures an image of the detected person's face and generates a facial image; a facial expression recognition device that recognizes a facial expression of the detected person by referring to the voice and the facial image; an output device that outputs a recognition result by the facial expression recognition device, The facial expression recognition device a voice acquisition process for acquiring the voice from the microphone; an image acquisition process for acquiring the face image from the camera; a recognition result output process for outputting to the output device a recognition result obtained by recognizing the facial expression of the detected person using a recognition model corresponding to an identification result of an attribute of the detected person from among a plurality of different recognition models, each of which receives a face image as an input and outputs a recognition result of a facial expression; Run The detected person is judged by referring to the detection signal whether or not the person is being continuously detected; the identification result of the attribute of the detected person is a cultural sphere identified by referring to the voice acquired in the voice acquisition process; The cultural area is determined by any one of the intonation of the speech, keywords contained in the speech, languages ​​contained in the speech, and a machine-learned discrimination model that takes the speech as input and outputs a cultural area. Facial expression recognition system.

5. 3. A program for operating the facial expression recognition device according to claim 1, wherein the program causes the one or more processors to execute the processes.

6. 1. A method for facial expression recognition executed by one or more processors, comprising: a voice acquisition step of acquiring a voice including an utterance of the detected person by referring to a detection signal indicating that a person has been detected; an image acquisition step of acquiring a face image including the detected person's face as a subject; a recognition result output step of outputting a recognition result obtained by recognizing the facial expression of the detected person using a recognition model according to an identification result of an attribute of the detected person from among a plurality of different recognition models, each of which receives a face image as an input and outputs a recognition result of a facial expression; Including, The detected person is judged by referring to the detection signal whether or not the person is being continuously detected; the identification result of the attribute of the detected person is a cultural sphere identified by referring to the voice acquired in the voice acquisition step; The cultural area is determined by any one of the intonation of the speech, keywords contained in the speech, languages ​​contained in the speech, and a machine-learned discrimination model that takes the speech as input and outputs a cultural area. Facial expression recognition method.

Citation Information

Patent Citations

  • Expression encoding device and emotion discriminating device

    JP1994076058A

  • Nationality decision device and method, and program

    JP2010191530A

  • Display control device, method, and program

    JP2017054241A

  • Device, program and method for identifying state in specific object of predetermined object

    JP2018045350A

  • Electronic device and method of obtaining emotion information

    US20200104670A1