Method and computer program for calculating lip movements of a humanoid agent's facial model.

JP7923515B2Active Publication Date: 2026-09-18ATR ADVANCED TELECOMM RES INST INT
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2022030582
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-01
Publication Date
2026-09-18
Estimated Expiration
2042-03-01

Smart Images

  • Figure 0007923515000001
    Figure 0007923515000001
  • Figure 0007923515000002
    Figure 0007923515000002
  • Figure 0007923515000003
    Figure 0007923515000003
Patent Text Reader

Abstract

To provide a method for generating lip motions of a humanoid agent that allows a humanoid agent to provide a natural impression even when uttering a voice emotionally, and to provide a computer program.SOLUTION: A method for generating lip motions for an expression model of a humanoid agent includes: steps 306, 308, and 310 in which a computer generates lip motions associated with utterance; and steps 312 and 314 in which the computer modulates the lip motions with a parameter expressing a specified emotional state.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[[Technical Field]]

[0001] The present invention relates to a method for generating lip motion in consideration of emotions and expressions in a humanoid agent having a human appearance, such as an android robot or a CG avatar. [[Background Art]]

[0002] A humanoid agent has a three-dimensional or two-dimensional head imitating a human head, and has a face resembling a human face on the front side. In the case of a humanoid agent, it can interact with humans by having a speech function and a speech recognition function. As shown in FIG. 1, the humanoid agent can have facial expressions in the same manner as a human by means of a model 30 representing the shape of a face with many degrees of control freedom that imitates the muscles of a human face.

[0003] Furthermore, the humanoid agent can also speak. When the humanoid agent speaks, it is also necessary to change the lip shape. Many methods for moving lips in synchronization with speech have been proposed so far. For example, in MPEG (Moving Picture Expert Group)-4, visemes that define the corresponding lip motion for each phoneme are specified. A viseme is the minimum unit of lip shape during speech. By sequentially specifying visemes in accordance with the phoneme sequence of speech, the lip shape of the humanoid agent can be changed in accordance with speech. [[Prior Art Literature]] [[Non-Patent Literature]]

[0004] [[Non-Patent Literature 1]] Shogo Takeuchi et al., "Emotion Generation Model for Sensuous Conversation Robot Based on Interlocutor's Favorability", Journal of the Robotics Society of Japan, 2007, Vol. 25, No. 7, p. 1125-1133 [[Summary of the Invention]] [[Problem to be Solved by the Invention]]

[0005] Humans speak with a variety of emotions. When a person speaks with some emotion, their facial expression differs from when they speak in a neutral state. Even in the case of humanoid agents, research is being conducted to give them emotions. Non-patent document 1 discloses a method for generating a robot's emotions based on the feelings the robot has towards its conversation partner (likeability towards the conversation partner).

[0006] In the case of humanoid agents, it is preferable to give them facial expressions that correspond to their emotions. However, in the case of Non-Patent Document 1 mentioned above, the facial expressions of the robot are not disclosed at all. Even when a humanoid agent speaks according to its emotional state, the mouthpieces used are those used for normal speech. Therefore, when a humanoid agent is made to speak with emotions, it can sound unnatural. In order to make humanoid agents more relatable, it is necessary to eliminate this unnaturalness.

[0007] Therefore, the objective of this invention is to provide a method and computer program for generating lip movements of a humanoid agent that give a natural impression even when the humanoid agent is given emotions and made to speak. [Means for solving the problem]

[0008] A method for generating lip movements according to the first aspect of this invention is a method for generating lip movements in a facial expression model of a humanoid agent, comprising the steps of: a computer generating lip movements related to speech; and a computer modulating lip movements with parameters representing a specified emotional state.

[0009] Preferably, the parameters representing the emotional state include arousal level and emotional valence, and the modulation step includes the computer modifying the lip movements during speech as a function of arousal level and emotional valence.

[0010] More preferably, the step of generating lip movements includes the step of a computer generating lip movements according to a sequence of mouthformes corresponding to a sequence of phonemes that constitute an utterance. For each mouthforme, a function is predetermined whose value is determined by arousal level, emotional valence, or both. The modification step includes the step of a computer modifying the lip movements corresponding to each mouthforme forming the sequence by applying the function corresponding to the mouthforme to the value of arousal level, emotional valence, or both.

[0011] More preferably, each mouth element includes positional information of a plurality of parts that define the shape of the lips of the facial expression model, and the function includes a separate function for each of the plurality of parts of each mouth element, and each of the separate functions is a function of arousal level or emotional valence or both.

[0012] Preferably, the multiple parts in the facial expression model include the positions of control points on the cheeks, the positions of control points on the corners of the mouth, the positions of control points representing the amount of mouth opening, and the positions of control points representing the amount of lip pursing, and at least one of the functions relating to the positions of the cheeks, the positions relating to the corners of the mouth, and the functions relating to the control points representing the amount of lip pursing is a function of emotional valence only.

[0013] More preferably, at least one function representing the position of the control point representing the opening is a function of both emotional valence and arousal level.

[0014] More preferably, the modulation step further includes the step of the computer dynamically modifying the lip movements as a function of arousal level and time.

[0015] Preferably, the function of arousal level and time is a periodic function with respect to time.

[0016] More preferably, at least one of the period and amplitude of the periodic function changes as an increasing function of arousal.

[0017] More preferably, the dynamic modification step includes the computer changing the position of each part of the lip movement of the facial expression model according to a function of arousal level and time.

[0018] Preferably, the modulation step further includes the step in which the computer resolves the conflict between the lip movements at speech, modified as a function of arousal and emotional valence, and the dynamically modified lip movements, in a way that differs depending on the direction of movement of each part.

[0019] More preferably, the resolution steps include: the computer resolving the conflict by prioritizing the lip movements during speech over dynamically modified lip movements for the vertical movement of each part of the facial expression model associated with speech; and the computer resolving the conflict by using a weighted sum of the lip movements during speech and dynamically modified lip movements for the horizontal movement of each part of the facial expression model associated with speech.

[0020] More preferably, the step of resolving the conflict by prioritizing lip movements includes resolving the conflict by prioritizing the lip movements during speech over dynamically modified lip movements such that the proportion of vertical movement of the lips of the facial expression model is 80% or more.

[0021] Preferably, the step of resolving the conflict by prioritizing lip movements includes the step of the computer resolving the conflict by employing only the lip movements during speech.

[0022] A computer program according to the second aspect of this invention causes a computer to function to perform each step of any of the methods described above. The above and other objects, features, aspects and advantages of this invention will become apparent from the following detailed description relating to this invention, which will be understood in conjunction with the accompanying drawings. [Brief explanation of the drawing]

[0023] [Figure 1]FIG. 1 is a diagram for explaining a facial expression generation method using an actuator for generating facial expressions of a humanoid robot. [Figure 2] FIG. 2 is a diagram showing an emotion model. [Figure 3] FIG. 3 is a diagram for explaining mapping from an emotion model to facial expressions. [Figure 4] FIG. 4 is a diagram showing a list of visemes used in the embodiment in a table format. [Figure 5] FIG. 5 is a block diagram of a robot control system according to an embodiment of the present invention. [Figure 6] FIG. 6 is a diagram showing an example of a scenario executed by the robot control system shown in FIG. 4. [Figure 7] FIG. 7 is a flowchart showing a control structure of a program for generating robot lip motion executed by the robot control system shown in FIG. 5. [Figure 8] FIG. 8 is a diagram showing, in a table format, a correction method that takes facial expressions into account for each viseme used in an embodiment of the present invention. [Figure 9] FIG. 9 is a graph showing the concept of correction of visemes during non-utterance, which is used in an embodiment of the present invention. [Figure 10] FIG. 10 is a diagram for explaining conflict and resolution regarding deformation of each part of the face when uttering with a facial expression. [Figure 11] FIG. 11 is a block diagram showing an example of a hardware configuration of a computer for realizing the robot control system shown in FIG. 5. DETAILED DESCRIPTION OF THE INVENTION

[0024] In the following description and the drawings, the same components are denoted by the same reference numerals. Therefore, detailed description thereof will not be repeated.

[0025] First First Embodiment 1. Mapping from emotion model to facial expression When a robot is made to speak while displaying facial expressions, the resulting lip movements will differ from those generated from neutral facial expressions, such as those in Model 30 shown in Figure 1. Referring to Figure 1, the robot's head is equipped with multiple actuators that move each control point of Model 30 in directions determined based on physiological knowledge in order to generate facial expressions. These actuators are instructed to move each part of the face in a predetermined direction by a predetermined distance within a specified time. As a result, the robot's face can be made to display facial expressions, and its lip shape can be changed in synchronization with speech.

[0026] As background to this, there is a technology that involves preparing a model like Model 30 shown in Figure 1, which represents the shape of a robot's face, in advance, moving each control point of this model according to the facial expression, and giving commands to actuators in accordance with that movement to change the shape of the robot's face. In this specification, a model that represents the shape of a robot's face in this way is called a "face model." Here, the face model in a neutral state is called the "basic face model," and the face model when a certain expression is given is called the "basic expression." In the following explanation, the expression "moving each part of the face model" means moving the control point corresponding to that part.

[0027] To make the impression robots give when interacting with humans more natural, research is being conducted on giving robots emotions. One model used in this research is the PAD (Pleasure, Arousal, Dominance) model, as shown in Figure 2. This model posits that human emotions can be distinguished and represented by points in a three-dimensional space defined by three axes: Pleasure or Valence, Arousal, and Dominance.

[0028] Referring to Figure 2, for example, a happy emotion is represented by a high Pleasure (Valence) value and relatively low Arousal and Dominance values.

[0029] Furthermore, as shown in Figure 3, rules for mapping this PAD model to the model 30 shown in Figure 1 have already been proposed. In the embodiments described below, based on Russell's annular model, it is assumed that emotions are expressed by the emotional valence and arousal level among the three axes mentioned above, and a mapping is performed from these two values ​​to the face model.

[0030] Figure 4 shows examples of phonemes in tabular form. The phonemes shown here are defined in MPEG-4 and are primarily related to the English language. In Figure 4, the leftmost column represents the phoneme name. The second column shows examples of phonemes corresponding to the leftmost phoneme name. The rightmost column shows examples of words containing the phonemes from the second column.

[0031] In the example in Figure 4, 15 types of mouth shapes are defined. For example, the mouth shape sil represents the mouth shape when not speaking. The mouth shape PP represents the mouth shape corresponding to the phonemes p, b, and m. The mouth shape FF represents the mouth shape corresponding to the phonemes f and v.

[0032] The phonemes shown here, as mentioned earlier, are related to English. However, they can also be used for Japanese speech. Note that the set of phonemes differs depending on the language. Therefore, different sets of phonemes may be prepared depending on the language used for the robot's speech.

[0033] As will be described later with reference to Figure 8, in this embodiment, these 15 types of mouth shapes are used to integrate lip movements from facial expressions and lip movements from speech in order to naturally express emotions. That is, the lip movements in facial expressions expressed by the face model are modulated by parameters that represent the emotional state.

[0034] 2. Structure A. Overall configuration Figure 5 shows a block diagram of the robot control system 50 according to this embodiment. Figure 5 shows only the parts of the robot control system 50 that are relevant to this invention. There are other parts, such as those for controlling the robot's hands, feet, waist, etc., but these are not shown in Figure 5 in order to simplify the diagram and explanation.

[0035] Referring to Figure 5, the robot control system 50 includes a scenario storage unit 60 that stores a program for controlling the robot's movements as a scenario 70, an expression generation unit 62 that reads the scenario 70 from the scenario storage unit 60, executes it, and outputs commands to each actuator of the robot's head, and a robot control unit 64 that responds to commands from the expression generation unit 62 and controls each actuator of the robot's head.

[0036] The facial expression generation unit 62 includes an emotion information extraction unit 80 for interpreting the scenario 70 and extracting emotion information 82 that specifies the emotions of the robot as defined by the scenario. In this embodiment, arousal level and emotional valence are used as parameters representing the emotional state.

[0037] The facial expression generation unit 62 further includes a speech and speech text extraction unit 84 that extracts the robot's speech voice and speech text specified by the scenario, determines the prosodic features of the speech voice and the mouth shapeme sequence determined from the speech text, and outputs it as a speech command 86, and a facial expression information extraction unit 88 that extracts facial expression information that specifies the robot's facial expression from the scenario, deforms the face model according to the specified facial expression based on that facial expression information, and generates and outputs a basic facial expression 90. The determination of the mouth shapeme sequence is performed by extracting a phoneme sequence from the speech text and determining the mouth shapeme corresponding to each phoneme that makes up this phoneme sequence. In the speech command 86, one mouth shapeme is specified for each phoneme.

[0038] The facial expression generation unit 62 further receives emotion information 82 from the emotion information extraction unit 80, speech commands 86 from the speech voice and speech text extraction unit 84, and basic facial expressions 90 from the facial expression information extraction unit 88, and modifies the commands related to mouth shapes in the speech commands 86 using a function of the arousal level and emotional valence contained in the emotion information 82 and integrates them with the basic facial expression 90 (i.e., the lip movements of the basic facial expression 90) Arousal level It includes a lip motion generation unit 92 that modulates according to emotional valence and outputs commands related to lip movements to the robot control unit 64.

[0039] The robot control unit 64, upon receiving command values ​​for each part of the robot from the facial expression generation unit 62, has the function of controlling the actuators of each part in parallel with the movements of the facial expression generation unit 62, in order to realize movements in each part according to those command values.

[0040] B. program Figure 6 shows the structure of Scenario 150, a simplified example of Scenario 70. Referring to Figure 6, Scenario 150 is defined using a graph format that includes nodes 160, 162, 164, 166, 168, 170, and 172 representing the state of the robot, and edges representing the transitions between these states. Each of these nodes 160, 162, 164, 166, 168, 170, and 172 stores at least information specifying the robot's current processing, movement, speech, emotions, and the next node it will move to.

[0041] In the example shown in Figure 6, the robot is activated at node 160, and then at node 162, the robot speaks to the other party. Further, at node 164, the robot asks a question to the other party, and once the question is finished, the state moves to node 166. Assuming that the other party will respond to this question, the robot performs speech recognition on the other party's utterance at node 166. At node 166, the robot further determines the content of the other party's answer to the question based on the recognition result and decides what action the robot should take next. Based on this decision, the robot performs the execution of node 168 and node 170 in parallel, using different information.

[0042] For example, in node 168, the answer to the question is a robot but The robot will make different facial expressions and utter different statements depending on whether the response is expected or unexpected. Simultaneously, node 170 will perform pre-prepared actions according to the content of the response. Once the execution of nodes 168 and 170 is complete and control reaches node 172, the execution of this scenario ends.

[0043] The actual processing performed by the robot becomes more complex. In this embodiment, the content of that processing is basically prepared in advance in the form of a scenario as shown in Figure 6.

[0044] C. Lip movements that match emotions Figure 7 shows the control structure of the program executed by the robot control system 50 shown in Figure 5 to generate lip movements that correspond to emotions. This program is executed repeatedly, moving through the scenario nodes according to the scenario in accordance with the cycle that controls the robot's movements.

[0045] Referring to Figure 7, this program includes the steps of: obtaining the arousal level and emotional valence specified in the node being processed (step 300); and calculating a basic facial expression F (corresponding to basic facial expression 90 in Figure 5) according to the facial expression information specified by the node being processed. 302, and step 304, which branches the control flow depending on whether the robot is speaking or not. Whether the robot is speaking or not is determined by whether an utterance is specified in the scenario in which the robot is running and whether the end time of that utterance has been reached.

[0046] The program further includes step 306, which is executed in response to the determination in step 304 being affirmative, and which acquires the robot's spoken voice and spoken text from the scenario; step 308, which acquires and outputs the prosodic features and mouth shape of the robot's spoken voice based on the spoken voice and spoken text acquired in step 306; and step 310, which calculates the speech lip shape and eyelid / eyebrow position M according to emotion.

[0047] The program further includes, following step 310, a step 310 which determines an overlap ratio x for superimposing the basic facial expression F calculated in step 302 with the speaking lip shape and eyelid and eyebrow positions M calculated in step 310, according to the dialogue situation, and a step 314 which calculates the control amount C of the robot's face and lips by superimposing (adding in a weighted form) the basic facial expression F and the speaking lip shape and eyelid and eyebrow positions M using the determined overlap ratio x.

[0048] The program further includes step 318, which calculates a modification amount A1 of the lip shape of the robot's basic facial expression in response to the determination in step 304 being negative, and step 320, which calculates a control amount C of the robot's face and lips by adding the basic facial expression F and the modification amount A1 calculated in step 318.

[0049] The program further includes step 316, which is executed after step 314 and after step 320, and controls the robot according to the control variable C calculated in step 314 or step 320 to terminate the processing of this cycle.

[0050] The following describes the processes performed by this program, specifically the processes during speech execution from step 306 to step 314, and the processes during non-speech execution from step 318 to step 320.

[0051] D. Processing during speech In the following explanation, Cheek is a control variable representing the vertical movement of the cheeks. LipCorner is a control variable representing the vertical movement of the corners of the mouth. MouthOpen is a control variable representing the amount of mouth opening. MouthCornerStickout is a control variable representing the amount of lip pursing. EyeLidOpen is a control variable representing the position of the eyelids and eyebrows.

[0052] Figure 10(A) shows the facial parts related to lip shape during speech. Referring to Figure 10(A), the part of the face model 500 indicated by reference numeral 510 is the part controlled by EyeLidOpen (eyelid and eyebrow position). The part indicated by reference numeral 512 is the part controlled by Cheek, LipCorner, and MouthCornerStickout (cheek, corner of the mouth, and amount of lip pursing). The part indicated by reference numeral 514 is the part controlled by MouthOpen (amount of mouth opening). In Figure 10(A), each arrow indicates the direction of movement of each part. As is clear from Figure 10(A), EyeLidOpen and MouthOpen only include vertical movement. In contrast, Cheek, LipCorner, and MouthCornerStickout also include horizontal movement. In particular, MouthCornerStickout does not include vertical movement.

[0053] For Cheek, LipCorner, and MouthOpen, their range of motion is restricted to the range [0,1] due to mechanical constraints. In this embodiment, as the value of Cheek increases, the cheeks lift. As the value of LipCorner increases, the corners of the mouth lift. As the value of MouthOpen increases, the mouth opens. For MouthCornerStickout, its range of motion is restricted to the range [-1,1]. As the value of MouthCornerStickout decreases, the lips are pursed. When MouthCornerStickout is 0, the amount of lip pursing is neutral. As MouthCornerStickout increases, the lips are pulled to the sides. EyeLidOpen will be described later.

[0054] The calculation of the lip shape and eyelid and eyebrow positions M during speech step 310 is performed as follows. Figure 8 shows in tabular form the calculation method for the corresponding control variable values ​​for each part of each mouth element. That is, for each of the multiple parts of each mouth element, a separate function is defined that determines the value of the control variable.

[0055] The leftmost column of Figure 8 contains the names of the mouth shapes. The second column contains the formula for Cheek. The third column contains the formula for LipCorner. The fourth column contains the formula for MouthOpen. The fifth column shows the formula for MouthCornerStickout. The sixth column contains the information necessary for calculating EyeLidOpen. In the position labeled "open" in the sixth column, a value indicating the timing of raising the upper eyelid and eyebrows will be substituted, as will be explained later. In the position labeled "normal," the timing of returning the raised upper eyelid and eyebrows to their original positions will be substituted.

[0056] For example, in the case of mouth element FF, Cheek is calculated as 0*V (emotional valence). LipCorner is calculated as 0*V. MouthOpen is 0. MouthCornerStickout is calculated as 0.75 + 0.1*V. In the case of mouth element DD, Cheek and LipCorner are the same as for mouth element PP. MouthOpen is calculated as 0.5 + 0.2*V + 0.2*A (arousal level).

[0057] Examining Figures 10(A) and 8, it can be seen that Cheek, LipCorner, and MouthCornerStickout (in reference numerals 510 and 512) are all functions of emotional valence V, except when their values ​​are specified as 0. However, MouthOpen (reference numeral 514) is a function of both emotional valence V and arousal level A, except when its value is 0.

[0058] The behavior of EyeLidOpen is determined as follows. As shown in Figure 8, EyeLidOpen has two possible values, normal and open, depending on the corresponding morpheme.

[0059] a) The value placed at the position marked "open" is used to raise and hold the upper eyelid and eyebrows by the value of upperLidAdditionalValue, determined by the following formula, at the timing determined by "open" when the level of arousal is positive.

[0060] upperLidAdditionalValue = (Maximum value - Minimum value) * Awakening level + Minimum value (Maximum value = 1.0, minimum value = 0.3) As a result, when the level of arousal is 0, the eyelids will be lowered to a minimum of 0.3, and when the level of arousal is 1, the eyelids will be raised to their maximum. In humans, when there is a large vertical movement of the mouth, the face tends to naturally widen. By calculating the movement of the upper eyelids and eyebrows using the above formula, human-like movements can be achieved.

[0061] b) The value arranged at the position labeled "normal" is used to move the opened upper eyelid and eyebrow back by an amount corresponding to upperLidAdditionalValue over approximately 150 msec starting from the timing defined by the normal value.

[0062] E. Processing when not speaking Even when the robot is not speaking, it is unnatural to maintain a static facial expression without changing the robot's facial expression at all. Therefore, in this embodiment, while any emotion is set, the actuator for facial expression generation is constantly moved in the direction of a predetermined movement axis by the following method. Through this processing, a natural facial expression corresponding to the emotion is generated.

[0063] That is, assuming that the robot breathes similarly to a human, the mouth is moved in the axial direction that opens the mouth along with inhalation and in the axial direction that pulls the mouth laterally according to the breathing cycle, and moved in the axial direction that closes the mouth along with exhalation and in the axial direction that purses the mouth. This state is shown in FIG. 9.

[0064] Referring to FIG. 9, this graph 400 shows a change in a correction amount m that corrects the movement amount of each part of the mouth according to emotion over time. In this example, the graph 400 is substantially sinusoidal, and periodically changes over time with a period T indicated by reference numeral 402 and an amplitude A indicated by reference numeral 404. Furthermore, the period T is shortened as the arousal level is higher, and lengthened as the arousal level is lower. The amplitude A is changed such that the amplitude A is larger as the arousal level is higher, and smaller as the arousal level is lower. As described above, the correction amount m is a periodic function whose period and amplitude change according to the value of the arousal level. By applying such a correction amount m to the face model, a human-like facial expression can be expressed even when the robot is in a static state.

[0065] Note that when the value of the arousal level is represented by arousal, arousal changes within the range of -1 < arousal < 1. If the breathing period at rest is T0 and the parameter for changing the period is b, since the average breathing rate of an adult is 16 to 18 [times / minute], T0 is approximately 4000 msec and b is approximately 0.2.

[0066] Therefore, in this embodiment, the respiratory cycle T and correction amount m(t) that take emotion into account are calculated using the following formula, where t represents time and a0 is a parameter for changing the amplitude of respiration according to the level of arousal.

[0067] T = T0 × (1 + b × (arousal)) m(t)=(1+a0×(arousal))*sin(T*t) Using the correction amount m(t) obtained in this way, the Cheek stretch correction amount, LipCorner stretch correction amount, MouthOpen correction amount, and MouthCornerStickout correction amount are calculated according to the following formulas. However, in the following formulas, a1, a2, a3, and a4 are parameters that depend on the facial area.

[0068] Cheek tension correction amount (t) = a1 × m(t) LipCorner tension correction amount (t) = a² × m(t) MouthOpen correction amount (t)=a3×m(t) MouthCornerStickout correction amount (t)=-a4×m(t) Furthermore, when superimposing the basic facial expression F calculated in step 302 with the speech lip shape and eyelid and eyebrow positions M calculated in step 310, there may be conflicts between them. For example, control points located close to each other may be moved in opposite directions between the basic facial expression F and the speech lip shape. This could potentially cause mechanical damage to the robot's face. This could happen, for example, when an expression such as pursing the lips while pulling back the corners of the mouth is requested. Therefore, it is best to avoid independently controlling the basic facial expression and the mouth movements due to speech.

[0069] On the other hand, if you don't change your facial expression to avoid conflict, it will seem unnatural to the other person. However, simply adding the two together does not guarantee that the resulting expression will be natural and appropriate to the situation. In particular, in the case of robots and CG (Computer Graphics) avatars, there may be no degree of freedom in the cheeks, and in such cases, whether or not someone is smiling must be expressed solely through mouth movements. Therefore, it is necessary to resolve the conflict between facial expressions and lip movements caused by speech and to realize appropriate lip movements.

[0070] First, the following degrees of freedom can be considered as competing options.

[0071] a) Vertical movement: Opening of the mouth This is the movement shown by reference numeral 514 in Figure 10(A). The vertical direction is sometimes also referred to as the up and down direction.

[0072] b) Lateral movements: pulling the lips sideways (puckering), corners of the mouth.

[0073] This is the movement indicated by the arrow within the range shown by reference numeral 512 in Figure 10(B).

[0074] In this embodiment, natural superposition is achieved by changing the superposition rules when superimposing basic facial expressions and speech movements in the vertical and horizontal movements of the lips. In this embodiment, the following rules are adopted for both.

[0075] a) Rules regarding vertical movement In the case of vertical movement, basically the command value for facial expression generation is ignored, and utterance movement is given complete priority. That is, for vertical movement of the lips, the command value for utterance movement is completely followed. This is because if lip movement is not generated with priority given to utterance movement, the robot will not appear to be speaking properly. However, it is not absolutely necessary to give complete priority to utterance movement; facial expression generation may be included as long as it is superimposed at a ratio of 80% for utterance movement and 20% for facial expression generation. Note that "priority" in this specification generally refers to making the weight of one of two types of lip movement greater than 50% when the two types are weighted added and superimposed.

[0076] b) Rules for horizontal movement The movement of an actuator that conflicts between facial expression and utterance movement is added (superimposed) at a ratio of x:1-x (where 0<x<1).

[0077] For example, assume that the movable range of a conflicting lip pull axis is 0 to 100. Suppose that the command value for lip pulling to produce a surprised facial expression is 20, and at the same time the command value for forming the lip shape of the vowel / i / in utterance movement is 80. If x=0.4, the command value for the lip pull axis is 20*0.4 + 80*0.6 = 56. This value (56) is used as the command value for the lip pull axis.

[0078] The reason for superimposing the two in this way is as follows. First, if complete priority is given to facial expression, there will be more facial expression changes. More facial expression changes lead to a larger difference from the facial expression in a static state. As a result, there is a possibility that the robot's facial expression will give the other party the impression of being unnatural. Conversely, if complete priority is given to utterance movement, the robot will speak without altering facial expressions such as a smile, which is highly likely to give the other party the impression of a forced smile. There is also a possibility that the other party will get the impression that it is a doll talking. Furthermore, even though the robot is uttering speech, there is no change in facial expression, which gives the impression that it is different from a human being. The above is the reason for adding the two with weighting as done in the present case.

[0079] Note that x described above is a value that defines the superimposition ratio. This value may be a constant, but it is preferable to determine it based on some rule. This rule needs to be formulated based on some criteria when creating the scenario 70. In this embodiment, x is the superimposition ratio of facial expressions in superimposition. One criterion may be the idea that if the counterpart desires many facial expression changes in the robot (or avatar), they will prefer a value of x close to 1, and otherwise, they will prefer a value of x close to 0.

[0080] In this embodiment, in the scenario, a value close to x=1 is set for nodes where it is considered that there is no problem even if facial expression changes are large, such as when the robot is chatting with another person. When the conversation partner is a superior or a work partner, a value close to x+=0 is set. When an intermediate situation described in the second case is assumed, a value close to x=0.5 is set, that is, a value that considers both facial expressions and utterances is set.

[0081] If the conversational partner of the robot is not known in advance, this value may be dynamically determined based on the content of the partner's speech. For example, it is also conceivable to initially set x to a value close to 0, and add or subtract a predetermined value to / from x within the range of 0<x<1 when a specific word / phrase is found in the partner's utterance. In this case, a predetermined score may be assigned to each word / phrase in advance, and the score assigned to the word / phrase may be used for addition and subtraction every time such a word / phrase is found.

[0082] 3. Operation Referring to FIG. 5, a scenario 70 that defines the motion of the robot is created in advance and stored in the scenario storage unit 60. This scenario 70 is in a graph format as shown in FIG. 6. Each node holds information identifying the robot's motion, utterance content, facial expression, emotion, and the like at that node. It is also assumed that the value of the superimposition ratio x at that time is also specified for each node. Of course, the value of the superimposition ratio x may be changed by processing executed within the node.

[0083] When the robot starts moving, the program shown in Figure 7, which has a control structure, is launched and execution begins. The robot executes the processes specified by each node. At that time, the emotion information extraction unit 80 of the facial expression generation unit 62 interprets the scenario 70, extracts emotion information (emotional valence and arousal level) 82 that identifies the robot's emotions specified by the scenario, and outputs the emotion information 82. The speech voice and speech text extraction unit 84 extracts the robot's speech voice and speech text specified by the scenario, determines the prosodic features of the speech voice and the mouthpieces determined from the speech text, and outputs them as speech commands 86. The facial expression information extraction unit 88 extracts facial expression information that specifies the robot's facial expression from the scenario, and generates and outputs a basic facial expression 90 based on that facial expression information.

[0084] Furthermore, the lip movement generation unit 92 receives emotion information 82 from the emotion information extraction unit 80, speech commands 86 from the speech voice and speech text extraction unit 84, and basic facial expressions 90 from the facial expression information extraction unit 88. The lip movement generation unit 92 further modifies the commands related to mouth elements in the speech commands 86 by referring to the emotion information 82 and integrates them with the basic facial expressions 90. The lip movement generation unit 92 outputs the obtained values ​​to the robot control unit 64 as commands related to lip movements. When the robot control unit 64 receives command values ​​for each part of the robot from the facial expression generation unit 62, it operates separately from (in parallel with) the movements of the facial expression generation unit 62 to control the actuators of each part so that movements according to those command values ​​are realized in each part.

[0085] More specifically, the emotion information extraction unit 80 of the facial expression generation unit 62 shown in Figure 5 first acquires the level of arousal and emotional valence (step 300 in Figure 7). Furthermore, it generates a basic facial expression based on the facial expression information obtained from the scenario 70 (step 302 in Figure 7).

[0086] The lip movement generation unit 92 examines the information of the node being processed to determine whether the robot is speaking or not (step 304 in Figure 7). If the robot is speaking (YES in step 304), the speech voice and speech text extraction unit 84 shown in Figure 5 extracts the speech voice and speech text from the information of the node being executed, outputs the prosodic features of the speech voice and the mouth shapes corresponding to the speech as a speech command 86 (step 306 in Figure 7), and the lip movement generation unit 92 receives the speech command 86 (step 308 in Figure 7).

[0087] The lip movement generation unit 92 further calculates the speech lip shape and eyelid and eyebrow positions M according to the emotion, based on the arousal level and emotional valence received in step 300 (step 310 in Figure 7). The calculation formula shown in Figure 8 is used in this calculation. By performing such calculations, emotional modifications are added to the lip movements during speech, enabling lip movements appropriate to the emotion. The lip movement generation unit 92 also determines the overlap ratio x between the basic facial expression and the speech lip movements (step 312 in Figure 7). This overlap ratio x may be a constant, as mentioned above, or it may be determined according to the dialogue situation. In this example, the overlap ratio x is determined according to the dialogue situation.

[0088] Finally, the lip motion generation unit 92 superimposes the control amount C of the robot's face and lips with the basic facial expression F, the speech lip shape and eyelid / eyebrow position M according to the superposition ratio x, and outputs it as a command value to the robot control unit 64 shown in Figure 5 (step 316 in Figure 7). Based on the given command value, the robot control unit 64 controls each actuator of the robot to realize the facial movement according to the command value.

[0089] If the determination in step 304 is negative, the lip motion generation unit 92 calculates a modification amount A1 for the lip shape of the basic facial expression, which is the specified facial expression (step 318 in Figure 7). This modification amount A1 has a periodicity as shown in Figure 9. Its period T is shorter the higher the level of arousal and longer the lower the level of arousal. Also, the amplitude A is larger the higher the level of arousal and smaller the lower the level of arousal. Next, the lip motion generation unit 92 calculates the control amount C for the robot's face and lips by adding the basic facial expression F and the calculated modification amount A1 (step 320 in Figure 7). After this, this control amount is output to the robot control unit 64 shown in Figure 5, and the robot control unit 64 controls the actuators of each part of the robot according to that control amount.

[0090] When processing for a node is completed, the facial expression generation unit 62 moves the processing to the next node and executes the above-mentioned processing from the beginning. When the execution reaches the completion node, the robot also stops moving.

[0091] 4. Effects As described above, this embodiment allows emotions to be reflected in the lip movements of the robot when it speaks, once emotions are set for the robot. As a result, the robot's facial expressions become closer to those of a human, and the interaction with the robot feels more natural to the person interacting with it. Furthermore, even when the robot is not speaking, its lip movements can be given according to the set emotions. These movements are periodic, following the breathing pattern assumed for the robot. Therefore, to the observer, the robot's facial expressions appear more human-like and natural, even when it is not speaking. It also has the effect of effectively conveying the robot's emotions to the person interacting with it.

[0092] Furthermore, the conflict between the movement of the robot's actuators when facial expressions are set and the movement of the actuators that realize the lip movements when the robot speaks can be resolved by superimposing the two movements with a certain overlap ratio x. This prevents the robot's facial expressions from becoming mechanically distorted and unnatural. As a result, it is possible to provide a method for generating the robot's lip movements that gives a natural impression even when the robot is given emotions and speaks.

[0093] 5. Implementation by computer Figure 11 is a hardware block diagram of a computer system 950 that operates as, for example, the facial expression generation unit 62 shown in Figure 5. It should be noted that this is merely one example of a system for realizing the facial expression generation unit 62, and the configuration of the computer system realizing the facial expression generation unit 62 is not limited to that shown in Figure 11.

[0094] Referring to Figure 11, this computer system 950 includes a computer 970 having a DVD (Digital Versatile Disc) drive 1002, and a keyboard 974, a mouse 976, and a monitor 972, all connected to the computer 970, for user interaction. Of course, these are just examples of configurations for when user interaction is required, and any general hardware and software available for user interaction (e.g., touch panels, voice input, pointing devices in general) can be used. Here, "user" refers to the operator who operates the robot control system 50.

[0095] Referring to Figure 11, the computer 970 includes, in addition to the DVD drive 1002, a CPU (Central Processing Unit) 990, a GPU (Graphics Processing Unit) 992, and a bus 1010 connected to the CPU 990, GPU 992, and DVD drive 1002. The computer 970 further includes a ROM (Read-Only Memory) 996 connected to the bus 1010 for storing the computer 970's boot-up program, etc., and a RAM (Random Access Memory) 998 connected to the bus 1010 for storing program instructions, system programs, and work data, etc.

[0096] Computer 970 further includes an SSD (Solid State Drive) 1000, which is non-volatile memory connected to bus 1010. The SSD 1000 is for storing scenarios and programs to be executed by CPU 990 and GPU 992, face models, and data used by those scenarios and programs.

[0097] Computer 970 further includes a network interface 1008 that provides a connection to a network 986 enabling communication with other terminals, and a USB port 1006 that allows a USB memory 984 to be attached and detached, and provides communication between the USB memory 984 and various parts within Computer 970.

[0098] The computer 970 further includes an input / output interface 1004 connected to a bus 1010 and a microphone and speaker (not shown), as well as a control device (not shown) that controls actuators for various parts of the robot. The interface 1004 has the function of reading audio signals, video signals, and text data generated by the CPU 990 and stored in RAM 998 or SSD 1000 according to the instructions of the CPU 990, performing analog conversion and amplification processing to drive the speaker, digitizing analog audio signals from the microphone and storing them in RAM 998 or SSD 1000 at any address specified by the CPU 990, receiving input from sensors equipped on the robot, and transmitting commands to each actuator control device.

[0099] In the above embodiment, the scenario 70 (Figure 5) for realizing the functions of the robot control system 50 and the specific programs therefor are all stored in a storage medium of an external device (not shown) connected via network I / F 1008 and network 986, for example, as shown in Figure 11, such as SSD 1000, RAM 998, DVD 978, or USB memory 984. Typically, this data and parameters are written to the SSD 1000 from an external source and loaded into the RAM 998 when the computer 970 is running.

[0100] The computer program for operating this computer system to realize the functions of the robot control system 50 and its components shown in Figure 5 is stored on a DVD 978 inserted into the DVD drive 1002 and transferred from the DVD drive 1002 to the SSD 1000. Alternatively, these programs are stored on a USB memory 984, the USB memory 984 is inserted into the USB port 1006, and the programs are transferred to the SSD 1000. Alternatively, the programs may be transmitted to a computer 970 via the network 986 and stored in the SSD 1000. The programs are loaded into RAM 998 when executed.

[0101] Of course, the source program may be input using the keyboard 974, monitor 972, and mouse 976, and the compiled object program may be stored in the SSD 1000. If the scenario 70 used in the above embodiment is created as a script, the script input using the keyboard 974, etc., may be stored in the SSD 1000. In the case of a program that runs on a virtual machine, the program that functions as a virtual machine must be installed on the computer 970 in advance. Neural networks are used for speech recognition and speech synthesis, etc. A pre-trained network may be used, or training may be performed on a robot equipped with a robot control system 50.

[0102] The CPU990 reads the program from RAM998 according to the address indicated by an internal register called the program counter (not shown), interprets the instructions, reads the data necessary for executing the instructions from RAM998, SSD1000, or other devices according to the address specified by the instructions, and executes the processing specified by the instructions. The CPU990 stores the execution result data at an address specified by the program, such as RAM998, SSD1000, or a register within the CPU990. At this time, the value of the program counter is also updated by the program. Computer programs may be loaded directly into RAM998 from DVD978, USB memory 984, or via a network. Note that some tasks (mainly numerical calculations) within the program executed by the CPU990 are dispatched to GPU992 according to the instructions included in the program or according to the analysis results when the CPU990 executes the instructions.

[0103] The program that implements the functions of each part according to each embodiment described above using the computer 970 includes a plurality of instructions written and arranged to operate the computer 970 to implement those functions. Some of the basic functions necessary to execute these instructions are provided by an operating system (OS) running on the computer 970, a third-party program, or modules of various toolkits installed on the computer 970. Therefore, this program does not necessarily have to include all the functions necessary to implement the system and method of this embodiment. This program only needs to include instructions that execute the robot control system 50 and its components by statically linking appropriate functions or functions of the "programming toolkit" in a controlled manner to obtain the desired result, or by dynamically linking them during program execution. The method of operating the computer 970 for this purpose is well known and will not be repeated here.

[0104] Furthermore, the GPU992 is capable of parallel processing, allowing it to execute large amounts of calculations associated with machine learning concurrently, in parallel, or in a pipelined manner. For example, parallel computation elements discovered in the program during compilation, or during program execution, are dispatched from the CPU990 to the GPU992 as needed, executed, and the results are returned to the CPU990 either directly or via a predetermined address in RAM998, and assigned to a predetermined variable in the program.

[0105] Second variation The above embodiment assumes a case where emotions are set for a robot. However, this invention is not limited to such embodiments. It may be an avatar or agent realized by computer graphics instead of a robot. Furthermore, in the case of avatars and agents realized by computer graphics, the model that generates their facial expressions may be three-dimensional or two-dimensional. In this embodiment, such an entity that has a facial structure similar to that of a human is called a humanoid agent.

[0106] Furthermore, in the above embodiment, emotional valence and arousal level are used as information for setting emotions. However, the information for setting emotions is not limited to these two. The terminology and their meanings will differ depending on the model used, but if they can be used to set emotions, such terminology and their meanings may be used to set the emotions of the humanoid agent. For example, in the PAD model mentioned above, in addition to emotional valence and arousal level, there is Dominance, and this Dominance may also be added as information for setting emotions.

[0107] The embodiments disclosed herein are illustrative and not limited to those embodiments. The scope of the present invention is defined by the claims, with reference to the detailed description of the invention, and includes all modifications within the meaning and scope equivalent to the wording contained herein. [Explanation of symbols]

[0108] 30 Models 50 Robot Control Systems 60 Scenario Memory Unit 62 Facial expression generation section 64 Robot Control Unit 70 Scenarios 80 Emotion information extraction part 82 Emotional information 84. Speech and Speech Text Extraction Unit 86 Speech commands 88 Facial expression information extraction section 90 Basic facial expressions 92 Lip motion generation section 970 Computer

Claims

1. A method for calculating the lip movements of a humanoid agent's face model, The steps include: when the humanoid agent speaks, the computer determines the sequence of mouth shapes corresponding to the sequence of phonemes that make up the speech and outputs it as a speech command; The computer outputs a basic facial expression by deforming the face model according to the facial expression specified by the scenario, The computer receives the speech command and the basic facial expression, modifies the positions of the multiple parts constituting each mouth element in the mouth element sequence included in the speech command as a function of parameters including arousal level and emotional valence, and further integrates this with the basic facial expression to calculate positional information for the multiple parts of the face model, The function includes a plurality of individual functions for the positional information of each of the plurality of parts of each of the mouthforms, and each of the individual functions is a function of the arousal level, the emotional valence, or both. A method for calculating lip movements, comprising the steps of: a computer calculating the positional information of a plurality of parts that define the shape of the lips of the face model corresponding to the lip shape of the lip shape of the face model corresponding to the lip shape of the lip shape of the face model by applying the individual functions corresponding to each lip shape that make up the lip shape sequence to the arousal level or the emotional valence or both; and further superimposing this information with the basic facial expression using an overlap ratio determined in accordance with the dialogue situation.

2. The aforementioned multiple parts include, in the face model, the positions of the control points for the cheeks, the positions of the control points for the corners of the mouth, the positions of the control points representing the amount of mouth opening, and the positions of the control points representing the amount of lip pursing. The method for calculating lip movement according to claim 1, wherein at least one of the individual functions relating to the position of the cheek control point, the individual function relating to the position of the corner of the mouth control point, and the individual function relating to the control point representing the amount of lip pursing is a function of the emotional valence only.

3. The method for calculating lip movement according to claim 2, wherein at least one of the individual functions representing the position of the control point representing the amount of mouth opening is a function of both the emotional valence and the arousal level.

4. The method for calculating lip movements according to claim 1, further comprising the step of the computer repeatedly performing the steps of outputting the basic facial expression and calculating the basic facial expression using a value calculated as a predetermined function of time when the humanoid agent is not speaking.

5. The method for calculating lip movements according to claim 4, wherein the predetermined function is a periodic function with respect to time.

6. The method for calculating lip movements according to claim 5, wherein at least one of the period and amplitude of the periodic function changes as an increasing function of the arousal level.

7. The lip movement calculation method according to any one of claims 4 to 6, wherein the step of calculating the basic facial expression includes the step of a computer adding the value of the predetermined function to each of the positions of the plurality of parts relating to the lip movement of the face model.

8. The method for calculating lip movements according to claim 1, further comprising the step of a computer resolving a conflict between the positions of the plurality of parts modified by the function of the parameter and the positions of the plurality of parts relating to the basic facial expression in different ways according to the direction of movement of the plurality of parts.

9. The aforementioned steps to resolve the issue are: The first step is for the computer to resolve the conflict by prioritizing the change in the position of the multiple parts of the face model, which is modified by the function of the parameter, over the change in the position of the multiple parts due to the basic facial expression, with respect to the vertical movement of the multiple parts of the face model associated with speech. A method for calculating lip movements according to claim 8, wherein the computer resolves the conflict by weighting the change in the position of the plurality of parts, which has been modified by the function of the parameter, and the change in the position of the plurality of parts, which has been modified by the basic facial expression, with respect to the amount of lateral movement of each part of the face model associated with speech.

10. The method for calculating lip movements according to claim 9, wherein the first step includes resolving the conflict by prioritizing the change in the position of each part modified by the function of the parameter over the change in the position of each part due to the basic facial expression, such that the ratio of vertical movement of the positions of the plurality of parts of the face model is 80% or more.

11. The method for calculating lip movement according to claim 9, wherein the first step includes the step of a computer resolving the conflict by adopting only the changes in the position of each part modified by the function of the parameter.

12. A computer program that causes a computer to function in order to perform the method described in claim 1.

Citation Information

Patent Citations

  • JP2007

  • Image processor, image processing method, and program

    JP2011070623A

  • Communication device, communication robot, and communication control program

    JP2019000937A

  • Gesture control device and gesture control program

    JP2020037155A

  • Periocular and audio synthesis of full face image

    JP2021114324A