Control device, control program and control method

The control device and method improve user interaction by systematically enhancing the avatar's non-verbal cues like facial expression and voice tone, addressing the limitations of operator-dependent engagement in existing systems.

JP2025108820APending Publication Date: 2025-07-24AVITA INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024002246
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-11
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing information processing systems rely heavily on the operator's own responses to enliven conversations, lacking a systematic approach to enhance non-verbal interactions with users.

Method used

A control device and method that utilizes an avatar to control non-verbal information such as the clarity of the avatar's facial expression, tone of voice, and nodding, which are progressively enhanced throughout the conversation, based on operator inputs captured by a microphone and camera.

Benefits of technology

Enhances the engagement and liveliness of user interactions by naturally increasing the avatar's non-verbal cues, making the conversation more engaging and user-friendly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025108820000001_ABST
    Figure 2025108820000001_ABST
Patent Text Reader

Abstract

To enhance the atmosphere of an interaction with a user.SOLUTION: An operator terminal (16) is used by an operator who interacts with a user through an interaction using an avatar with the user who uses a user terminal (12). The action and voice of the avatar are controlled by the operator, and, from the start of the interaction toward the end of the interaction, the degrees of the avatar's voice tone, the clarity of its smile, and the magnitude of its nodding are increased.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a control device, a control program, and a control method, and more particularly, for example, to a control device, a control program, and a control method for controlling an avatar of an operator who responds to a user.

Background Art

[0002] An example of this type of background art is disclosed in Patent Document 1. In the information processing system disclosed in Patent Document 1, when a user wants to consult about a product or how to use a shopping site, the user can call an operator for consultation by pressing a call button. When the user terminal and the operator terminal are connected, the web site displayed on the user terminal is displayed on the operator terminal in the current display mode. Further, on the user terminal, an image of the operator or an avatar image synchronized with this is displayed on the web site. Therefore, the operator serves customers while using gestures and body language towards the user.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] According to the above background art, the operator serves customers by speaking while using gestures and body language towards the user, but enlivening the conversation scene with the customer largely depends on the operator's own way of responding, and there is room for improvement.

[0005] Therefore, the main object of this invention is to provide a novel information processing system, information processing device, control program, and control method.

[0006] Another object of the present invention is to provide an information processing system, an information processing apparatus, a control program, and a control method capable of enlivening the conversation scene.

Means for Solving the Problems

[0007] A first invention is a control device for an operator who responds to a user using an avatar, comprising avatar control means for controlling the operation and speech of the avatar, and the avatar control means increases the degree of at least one piece of non-verbal information of the avatar when the operator responds to the user from the start to the end of the conversation with the user. It is a control device that gets higher as it goes from the start to the end of the conversation.

[0008] A second invention is subordinate to the first invention, and the non-verbal information includes the clarity of the expression on the avatar's face, the tone of the avatar's voice, and the nodding of the avatar.

[0009] A third invention is subordinate to the second invention, and the expression on the avatar's face is a smiling face.

[0010] A fourth invention is subordinate to the second invention, further comprising a microphone for detecting the operator's voice, and the avatar control means generates voice data for controlling the avatar's speech based on the operator's voice detected by the microphone, and increases the tone of the avatar's voice as it goes from the start to the end of the conversation.

[0011] A fifth invention is subordinate to the first or second invention, further comprising a camera for photographing the operator, and the avatar control means generates motion data for controlling an operation including the expression on the avatar's face based on the image of the operator photographed by the camera, and increases the clarity of the avatar's smiling face as it goes from the start to the end of the conversation.

[0012] The sixth invention is subordinate to the second invention, further comprising a camera for photographing an operator, and the avatar control means generates operation data for controlling the operation of the avatar based on the video of the operator photographed by the camera, and the avatar control means increases the size of the avatar's nodding from the start to the end of the dialogue.

[0013] The seventh invention is subordinate to the second invention, and the avatar control means increases the rate of increasing the degree of each of the tone of the avatar's voice, the clarity of the avatar's facial expression, and the size of the avatar's nodding in the end stage of the dialogue compared to the beginning stage of the dialogue.

[0014] The eighth invention is subordinate to the first invention, and the operator is a program using a human or artificial intelligence.

[0015] The ninth invention is a control program executed by a control device of an operator who responds to a user using an avatar, and the control program causes a processor of the control device to control the operation and speech of the avatar, and during the period from the start to the end of the dialogue with the user, increases the degree of at least one piece of information of the non-verbal information of the avatar when the operator responds to the user from the start to the end of the dialogue.

[0016] The tenth invention is a control method for a control device of an operator who responds to a user using an avatar, and a processor of the control device controls the operation and speech of the avatar, and during the period from the start to the end of the dialogue with the user, increases the degree of at least one piece of information of the non-verbal information of the avatar when the operator responds to the user from the start to the end of the dialogue.

Advantages of the Invention

[0017] According to this invention, the atmosphere of the dialogue can be enlivened.

[0018] The above objects, other objects, features and advantages of this invention will become more apparent from the following detailed description of the embodiments with reference to the drawings.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

DETAILED DESCRIPTION OF THE INVENTION

[0020] <First Embodiment> Referring to FIG. 1, the information processing system 10 of the first embodiment includes a user terminal 12, and the user terminal 12 is communicably connected to an operator terminal 16 via a network 14.

[0021] As an example, by the technology of Web Real-Time Communication (WebRTC), the user terminal 12 and the operator terminal 16 can perform P2P (Peer to Peer) communication and transmit and receive video and / or audio in real time via a web browser.

[0022] Although the illustration is omitted, WebRTC can be used in combination with multiple servers. Specifically, multiple servers for using WebRTC include a signaling server, a STUN (Session Traversal Utilities for NAT) server, a TURN (Traversal Using Relay around NAT) server, and the like. Although detailed description is omitted, the signaling server is a server for acquiring information about a communication partner by WebRTC. Also, the STUN server and the TURN server are servers for performing so-called NAT (Network Address Translation) traversal when the communication partner exists in a different network.

[0023] In this first embodiment, one user terminal 12 and one operator terminal 16 are shown. Actually, a plurality of user terminals 12 and a plurality of operator terminals 16 are connected to the network 14, and among them, the connection state of one user terminal 12 and one operator terminal 16 that are set to perform a video call or a Web conference is established.

[0024] Note that when the operator terminal 16 connected to the user terminal 12 is predetermined, the user terminal 12 has the connection information of the operator terminal 16, and the operator terminal 16 has the connection information of the user terminal 12, the connection state between the user terminal 12 and the operator terminal 16 can be established without using the above servers.

[0025] The user terminal 12 is installed in a store where the information processing system 10 is applied, or is owned by the user and used by the user. The operator terminal 16 is used by a human operator. In this first embodiment, an avatar image (hereinafter, "avatar image") corresponding to the avatar of the operator is displayed on the display device 30 of the user terminal 12, and the user and the operator interact through the avatar. The operator operates the avatar, that is, controls the actions and utterances of the avatar to interact with the user. That is, the user is a person who uses the dialogue service using the avatar provided by the operator of the operator terminal 16.

[0026] The user terminal 12 is an information processing device, which is a general-purpose notebook PC or desktop PC, but other general-purpose terminals such as a smartphone or a tablet PC can also be used. However, when the user terminal 12 is set in a store and a notebook PC, a smartphone or a tablet PC is used, a display device with a relatively large size for displaying an avatar image may be separately prepared and connected to the notebook PC, the smartphone or the tablet PC.

[0027] The network 14 is composed of an IP network (or an IP network) including the Internet and an access network (or an access network) for accessing this IP network. As the access network, a public telephone network, a mobile phone network, a wired LAN, a wireless LAN, a CATV (Cable Television), etc. can be used.

[0028] The operator terminal 16 is another information processing device different from the avatar control device or the user terminal 12. As an example, it is a general-purpose notebook PC or desktop PC, but other general-purpose terminals such as a smartphone or a tablet PC can also be used.

[0029] FIG. 2 is a block diagram showing an example of the electrical configuration of the user terminal 12 shown in FIG. 1. As shown in FIG. 2, the user terminal 12 includes a CPU 20, and the CPU 20 is connected to a RAM 22, a communication interface (hereinafter referred to as "communication I / F") 24, and an input / output interface (hereinafter referred to as "input / output I / F") 26 via an internal bus.

[0030] The CPU 20 controls the overall operation of the user terminal 12. However, instead of the CPU 20, an SoC (System-on-a-chip) including a plurality of functions such as a CPU function and a GPU (Graphics Processing Unit) function may be provided.

[0031] The RAM 22 is the main memory device and is used as the working area or buffer area of the CPU 20. Although not shown in the figure, the user terminal 12 is provided with an HDD and a ROM as auxiliary storage devices. However, instead of or in addition to the HDD, a non-volatile memory such as an SSD may be used.

[0032] The communication I / F 24 is a wired interface for transmitting and receiving control signals and data between the CPU 20 and an external computer such as the operator terminal 16 via the network 14 under the control of the CPU 20. However, as the communication I / F 24, a wireless interface such as a wireless LAN or Bluetooth (registered trademark) can also be used.

[0033] An input device 28, a display device 30, a microphone 32, a speaker 34, and a camera 36 are connected to the input / output I / F 26.

[0034] The input device 28 is a keyboard and a computer mouse. Further, a touch panel may be provided. However, when a smartphone or a tablet PC is used as the operator terminal 12, the input device 28 is a touch panel and hardware buttons.

[0035] As an example, the display device 30 is a liquid crystal display device. The microphone 32 is a general-purpose sound pickup microphone. The speaker 34 is a general-purpose stereo speaker. The cameras 36 are each a general-purpose CCD camera.

[0036] Also, the input / output I / F 26 outputs the operation data (or operation information) input from the input device 28 to the CPU 20, and outputs the image data generated by the CPU 20 to the display device 30 to display a screen or an image corresponding to the image data on the display device 30. However, there may be a case where the image data received from an external computer (for example, the operator terminal 16) is output by the CPU 20.

[0037] In addition, the input / output I / F 26 converts the user's voice detected by the microphone 32 into digital voice data (hereinafter referred to as "user voice data") and outputs it to the CPU 20, or converts the voice data output by the CPU 20 into an analog voice signal and outputs it from the speaker 34, or outputs the video data of the video including the user captured (detected) by the camera 36 (hereinafter referred to as "user video data") to the CPU 20. The user video data is, for example, a moving image.

[0038] However, in this first embodiment, the voice data output from the CPU 20 is the voice data of the operator received from the operator terminal 16 or the voice data obtained by converting the operator's voice into the voice of the avatar. Hereinafter, these voice data will be referred to as operator voice data.

[0039] Note that the electrical configuration of the user terminal 12 shown in FIG. 2 is an example and does not need to be limited.

[0040] Also, when the user terminal 12 is a smartphone, it includes a call circuit for making a call via a mobile phone communication network or a mobile phone network and a public telephone network. However, in this first embodiment, such a call is not made, so the illustration is omitted.

[0041] FIG. 3 is a block diagram showing the electrical configuration of the operator terminal 16 shown in FIG. 1. As shown in FIG. 3, the operator terminal 16 includes a CPU 50, and the CPU 50 is connected to a RAM 52, a communication I / F 54, and an input / output I / F 56 via an internal bus.

[0042] The CPU 50 controls the overall operation of the operator terminal 16. However, instead of the CPU 50, an SoC including a plurality of functions such as a CPU function and a GPU function may be provided.

[0043] The RAM 52 is the main memory device and is used as the working area or buffer area of the CPU 50. Although not shown in the figure, the operator terminal 16 is provided with an HDD and a ROM as auxiliary storage devices. However, instead of or in addition to the HDD, a non-volatile memory such as an SSD may be used.

[0044] The communication I / F 54 is a wired interface for transmitting and receiving control signals and data between the CPU 50 and an external computer such as the operator terminal 16 via the network 14 under the control of the CPU 50. However, as the communication I / F 54, a wireless interface such as a wireless LAN or Bluetooth (registered trademark) can also be used.

[0045] An input device 58, a display device 60, a microphone 62, a speaker 64, and a camera 66 are connected to the input / output I / F 56.

[0046] The input device 58 is a keyboard and a computer mouse. Furthermore, a touch panel may be provided. However, when a smartphone or a tablet PC is used as the operator terminal 16, the input device 58 is a touch panel and hardware buttons.

[0047] As an example, the display device 60 is a liquid crystal display device. The microphone 62 is a general-purpose sound pickup microphone. The speaker 64 is a general-purpose stereo speaker. The camera 66 is a general-purpose CCD camera.

[0048] Also, the input / output I / F 56 outputs the operation data (or operation information) input from the input device 58 to the CPU 50, and outputs the image data generated by the CPU 50 to the display device 60 to display a screen corresponding to the image data on the display device 60. However, there may be a case where the image data received from an external computer (for example, the user terminal 12) is output by the CPU 50. The image data received from the user terminal 12 is user video data.

[0049] In addition, the input / output I / F 56 converts the voice of the operator detected by the microphone 62 into digital voice data and outputs it to the CPU 50, or converts the voice data output by the CPU 50 into an analog voice signal and outputs it from the speaker 64, or outputs the data of the video of the operator captured (detected) by the camera 66 (hereinafter referred to as "operator video data") to the CPU 50. However, in this first embodiment, the voice data output from the CPU 50 is the user voice data received from the user terminal 12. Also, the video of the operator is a moving image.

[0050] Note that the electrical configuration of the operator terminal 16 shown in FIG. 3 is an example and does not need to be limited.

[0051] Also, when the operator terminal 16 is a smartphone, it is provided with a call circuit for making a call via a mobile phone communication network or a mobile phone network and a public phone network. However, in this first embodiment, since such a call is not made, the illustration is omitted.

[0052] Hereinafter, the operation of the information processing system 10 of this first embodiment will be described. In such an information processing system 10, a user using the user terminal 12 and an operator using the operator terminal 16 perform a video call or a Web conference (or an online conference) using this user terminal 12 and this operator terminal 16.

[0053] As an example, when the user is accessing a desired EC (Electronic Commerce) site with the user terminal 12, the user requests an interaction with an avatar (or an operator), or the user terminal 12 and the operator terminal 16 access a predetermined URL notified in advance from a server or the like, whereby the user terminal 12 and the operator terminal 16 perform (or start) a video call or a Web conference.

[0054] FIG. 4 shows an example of the user-side interaction screen 100 displayed on the display device 30 of the user terminal 12 when performing a video call or a Web conference.

[0055] The dialogue screen 100 includes an image 102 of the avatar screen and is provided with a button image 104. The image 102 includes an avatar image (hereinafter referred to as the "avatar image") 110. The button image 104 is an icon for ending the dialogue.

[0056] Note that in the image 102 shown in FIG. 4, the background other than the avatar image 110 is omitted, but a single color or another background may be displayed.

[0057] In this first embodiment, the operations (including facial expressions) and voices (i.e., utterances) of the avatar corresponding to the avatar image 110 displayed on the display device 30 are controlled based on the operations (including facial expressions) and voices of the operator of the operator terminal 16. The operations include gestures, facial orientations, and facial expressions. The operations of the operator are detected from the video of the operator captured by the camera 66 of the operator terminal 16. The voice of the operator is detected by the microphone 62.

[0058] At the operator terminal 16, motion data for controlling the operations of the avatar is generated based on the operations of the operator, and voice data for controlling the utterances (voice output) of the avatar is generated based on the voice of the operator. The motion data is information (data) on the positions and orientations of each part for controlling the gestures, facial orientations, and facial expressions of the avatar. Specifically, the motion data includes parameters for each part of the face, namely the eyebrows (position, shape), the wrinkles between the eyebrows (degree of approximation), the eyes (size of the pupils, degree of raising and lowering of the outer corners of the eyes), the eyelids (degree of opening), the nasolabial folds (degree of approximation), and the mouth (position, shape, and degree of opening), the positions (coordinate data) of the head, hands, arms, shoulders, waist, knees, and ankles, and information on the orientations and degrees of bending (i.e., angles) of the neck (or face), wrists, and ankles.

[0059] However, as will be described later, in this first embodiment, a plurality of types of parameters for each part of the face are prepared according to the facial expression and its clarity, and the parameters corresponding to the estimated facial expression and its clarity are used for generating the motion data.

[0060] Also, when the lower part from the thigh of the avatar is not displayed, information about the knees and ankles may not be included in the motion data. Also, the voice data is data of the operator's voice or the voice obtained by converting the operator's voice into the avatar's voice.

[0061] Note that the motion data for controlling the motion of the avatar based on the operator's motion can be generated using, for example, an image processing library such as MediaPipe Holistic. In other examples, the method disclosed in JP-A-2021-56940 can also be used.

[0062] In the user terminal 12, the avatar image 110 is changed according to the motion data. Therefore, the avatar moves. Also, the user terminal 12 outputs the received operator voice data to the speaker 34. That is, the user terminal 12 controls the avatar image 110 (avatar) displayed on the display device 30 according to the instruction from the operator terminal 16. However, the instruction from the operator terminal 16 is only the motion data or the motion data and the operator voice data.

[0063] As described above, the motion and voice of the avatar are controlled based on the motion and voice of the operator. Also, in this first embodiment, the degree of each of the tone of the avatar's voice, the clarity of the avatar's happy expression (i.e., smiling face), and the size of the avatar's nodding increases (or becomes larger) from the start of the user's conversation (or start of response) to the end of the conversation (or end of response) as it goes from the start of the conversation to the end of the conversation.

[0064] In this way, by changing the tone of the avatar's voice, the clarity of the avatar's smiling face, and the size of the avatar's nodding, the user who is talking to the avatar naturally feels more involved in the conversation. As a result, it is more likely that the user will agree with the opinions of the operator of the avatar. For this reason, it is considered that the operator can smoothly proceed with the user's response.

[0065] In addition, in this first embodiment, for expressions representing emotions other than joy, they are expressed by the operator's emotion and its clarity (or intensity), and the clarity of expressions representing emotions other than joy is not forcibly increased from the start to the end of the dialogue.

[0066] However, expressions that may cause discomfort to the user, for example, expressions representing disgust and anger, may be suppressed.

[0067] The tone of the avatar's voice can be set in multiple levels from the lowest to the highest. When the voice tone is the lowest, the minimum value (for example, the frequency corresponding to the reference do sound (C4 or mid2C)) is set. When the voice tone is medium, the medium value (for example, the frequency corresponding to the reference fa sound (F4 or mid2F)) is set. When the voice tone is the highest, the maximum value (for example, the frequency corresponding to the sound one octave higher than the reference do (C5 or hiC)) is set. However, predetermined numerical values are also set for other levels of the voice tone.

[0068] FIG. 5(A) is a diagram showing the difference in the clarity of the avatar's smiling face, FIG. 5(B) is a diagram showing the difference in the size of the avatar's nodding, and FIG. 5(C) is a graph showing the time change of the target values of the avatar's voice tone, the clarity of the avatar's smiling face, and the size of the avatar's nodding. However, in FIG. 5(B), a simplified avatar is shown to clearly show the degree of inclination of the head or the degree of bending of the neck.

[0069] The clarity of the smiling face can also be set in multiple levels from the lowest to the highest. When the clarity of the smiling face is the lowest, the parameters of each part of the face image showing a neutral state without emotion are set. When the clarity of the smiling face is medium, the parameters of each part of the face image with a medium intensity of joy are set. When the clarity of the smiling face is the highest, the parameters of each part of the face image with the maximum intensity of joy are set. However, the so-called neutral state of expression without emotion means a so-called expressionless face.

[0070] In FIG. 5(A), an example of the face image of the avatar when the clarity of the emotion of joy changes from a neutral state without emotion to medium and maximum is shown. As described above, the facial expression representing the emotion of joy is a smiling face.

[0071] In the avatar in the neutral state, there are no wrinkles and the face is set to be symmetric. By designing in this way, it becomes difficult to judge the gender, and the avatar becomes acceptable to users with preferences biased towards men or women. Also, by making the face without wrinkles and without the characteristics of being symmetrically set, by simply creating a few wrinkles, facial expressions such as a smiling face or a frowning face can be easily expressed, and the clarity of the facial expression can also be easily controlled.

[0072] Also, a smiling face is determined by parameters for each part of the eyebrows (position, shape), the wrinkles between the eyebrows (degree of approach), the eyes (size of the black eyes, degree of raising and lowering of the eye corners), the eyelids (degree of opening), the nasolabial folds (degree of approach), and the mouth (position, shape, and degree of opening) with respect to the emotion of joy of the operator and the clarity of the emotion of joy.

[0073] However, for other facial expressions representing emotions other than the emotion of joy, similar to the smiling face, with respect to the clarity of each emotion, it is determined by the parameters for each of the above parts. Emotions other than the emotion of joy are the emotions of fear, sadness, disgust, and anger.

[0074] The size of the avatar's nodding can also be set in multiple stages from the lowest to the highest. As shown in FIG. 5(B), when the size of the avatar's nodding is the lowest, the angle at which the avatar bends its neck forward is set to the minimum value (for example, 0 degrees). When the size of the avatar's nodding is medium, the angle at which the avatar bends its neck forward is set to medium (for example, 22.5 degrees). When the size of the avatar's nodding is the largest, the angle at which the avatar bends its neck forward is set to the maximum value (for example, 45 degrees). However, the size of the nodding, that is, the angle at which the neck is bent forward, is also set to predetermined numerical values for other stages.

[0075] In this first embodiment, the time for the operator to respond to the user, that is, the time for conversation (hereinafter referred to as "conversation time"), is predetermined to be a first predetermined time (for example, 10 minutes). At the first predetermined time, as shown in FIG. 5(C), the tone of the avatar's voice, the clarity of the expression, and the amplitude of nodding are controlled to change. That is, FIG. 5(C) shows the time change of the target values (hereinafter sometimes simply referred to as "target values") of the tone of the avatar's voice, the clarity of the expression, and the amplitude of nodding.

[0076] As can be seen from FIG. 5(C), the time change of the target value is shown as a smooth curve. In this first embodiment, the initial value of the target value is set such that the tone of the voice, the clarity of the expression, and the amplitude of nodding are at a medium level. However, this is just an example, and the initial value of the target value may be set to be smaller than the medium level for the tone of the voice, the clarity of the expression, and the amplitude of nodding, or in some cases, may be set to be larger. In this case, the curve shown in FIG. 5(C) is translated parallel to the vertical axis. For example, the tone of the operator's voice when starting the conversation can be detected to set the initial value of the target value.

[0077] Also, as can be seen from FIG. 5(C), the time change of the target value at the end of the conversation when the conversation time is after 360 seconds (6 minutes) is set to be larger than the time change of the target value at the beginning of the conversation when the conversation time is from 0 to 240 seconds (4 minutes). This is to make it easier for the user to understand that the conversation is becoming lively at the end of the conversation.

[0078] The curve shown in FIG. 5(C) is just an example and does not need to be limited. Two or more line segments with different slopes may be connected. In this case, the slope of the line segment at the end of the conversation is set to be larger than the slope of the line segment at the beginning of the conversation. When there are three or more line segments, the slope of the line segment in the middle of the conversation is also set to be slightly larger than the slope of the line segment at the beginning of the conversation.

[0079] Basically, operation data is generated so that from the start of the dialogue, the tone of voice, clarity of expression, and amplitude of nodding are expressed according to the target values, and the tone of voice of the operator is changed.

[0080] In this first embodiment, the tone of voice, clarity of expression, and amplitude of nodding of the operator are each detected, and it is determined whether each of the detected tone of voice, clarity of expression, and amplitude of nodding exceeds the target value.

[0081] The tone of voice of the operator is extracted or detected by performing processes such as known noise removal and Fourier transform on the detected operator voice data.

[0082] The clarity of expression of the operator is estimated or detected based on the emotion and its intensity estimated based on the data of the face image among the detected operator video data. That is, the emotion and its intensity (or clarity) of the operator are estimated based on the face image of the operator, and the estimated emotion and its intensity are set as the facial expression and its clarity expressed by the avatar image 110. Since a method for estimating the emotion of a human like the operator using a face image is already known, the description of that method is omitted.

[0083] As a technique for estimating human emotion from a face image, known techniques can be used. For example, the techniques disclosed in "Hiroshi Kobayashi, Fumio Hara: Basic Facial Expression Recognition by Neural Network, Transactions of the Institute of Measurement and Control, Vol. 29, No. 1, 112 / 118 (1993)", "Yosuke Kotani, Tsuyoshi Homma, Masao Sakai, Kenichi Abe: Facial Expression Recognition Using Neural Network, Tohoku University School of Medicine and Health Sciences Annals 13(1):23 - 32, 2004", and "Daiki Nishiaki, Satoshi Endo, Akira Tokuma, Koji Yamada, Yuki Akamine: Acquisition of Facial Expression and Analysis of Facial Feature Amounts Using Convolutional Neural Network, Journal of the Japanese Society for Artificial Intelligence, Vol. 32, No. 5 FZ (2017)" can be used.

[0084] In addition, in other known technologies, as a method for estimating human emotions based on feature points extracted from a face image, the technology disclosed in Japanese Patent Application Laid-Open No. 2020-163660 can also be used.

[0085] Regarding the estimation of the intensity (or clarity) of emotions, by having a neural network learn face images of a plurality of expressions with different intensities for each emotion, it is possible to estimate not only the type of emotion but also the intensity of the emotion. Also, based on the difference in the output of the neural network when emotions are estimated, the intensity of the emotion can be estimated. For example, the intensity of the emotion is estimated based on the difference between the output for a face image in a neutral (expressionless) state and the output for the face image of the estimated emotion.

[0086] In addition, when using the method of Japanese Patent Application Laid-Open No. 2020-163660 described above, the intensity of the emotion can also be estimated based on the difference (distance) between the feature points extracted from the face image. For example, the distance for each feature point extracted from a face image in a neutral (expressionless) state is calculated with respect to each feature point extracted from the user's face image used for emotion estimation, and the intensity of the emotion is estimated based on the calculated distance. Since the distance is calculated for each feature point, the intensity of the emotion is estimated based on, for example, the average value, maximum value, or variance of the plurality of calculated distances.

[0087] However, when using the above technology as a method for estimating emotions and their intensities based on a face image, the necessary programs, circuit components, and data for this purpose may be provided in a server or the like on the network. Also, an apparatus for estimating emotions and their intensities based on a face image (hereinafter referred to as an "estimation apparatus") may be provided in the cloud, a face image may be transmitted to the estimation apparatus, and the estimation results of the emotions and their intensities may be received from the estimation apparatus.

[0088] In this first embodiment, nodding, i.e., the movement of the head and its magnitude, are detected based on the orientation of the operator's face. In such a case, based on the orientation of the face image of the operator facing the camera 66, the current face orientation is calculated based on the current face image, and the nodding operation of the operator and its magnitude are estimated based on the current face orientation. However, the face orientation can be detected by the movement of a plurality of facial feature points extracted from the face image.

[0089] When each of the tone of voice, the clarity of the expression, and the magnitude of nodding is less than the target value, it is set to reach the target value.

[0090] On the other hand, when each of the detected tone of voice, the clarity of the expression, and the magnitude of nodding exceeds the target value, it is further determined whether it exceeds the upper limit value. The upper limit value is a value obtained by multiplying the target value by a first predetermined ratio (for example, 1.05). However, the first predetermined ratio may be set larger in the latter stage of the dialogue than in the initial stage. Also, regarding the clarity of a smiling face, an upper limit value corresponding to the first predetermined ratio is set using a numerical value when the minimum value to the maximum value is digitized.

[0091] When each of the detected tone of voice, the clarity of the expression, and the magnitude of nodding exceeds the upper limit value, it is set to the upper limit value. On the other hand, when each of the tone of voice, the clarity of the expression, and the magnitude of nodding does not exceed the upper limit value, it is left as its original value.

[0092] Also, the target value and the upper limit value are updated every second predetermined time (for example, 5 seconds) from the start of the dialogue. Therefore, the data of a table (hereinafter referred to as "setting table") in which the target value and the upper limit value from 0 seconds to 600 seconds are described every 5 seconds are stored, and the target value and the upper limit value are set with reference to this setting table.

[0093] Figure 6 shows an example of a setting table. As shown in Figure 6, corresponding to the numbers, target values and upper limit values of the tone of voice, clarity of expression, and size of nodding are described. In this first embodiment, since the operator interacts with the user for 10 minutes and the target values and upper limit values are updated every 5 seconds, 120 target values and upper limit values of the tone of voice, clarity of expression, and size of nodding are provided respectively. Therefore, the numbers are from 1 to 120.

[0094] Also, the target value and upper limit value of the tone of voice are frequencies. The clarity of expression is the clarity of a smile, and the target value and upper limit value are parameters for each part of the face. The size of nodding is the angle of the neck, with the state where the neck is not bent being 0 degrees, and the angle of the neck when the head is tilted forward is described.

[0095] Note that in the setting table, specific numerical values and the like of the target value and upper limit value are not described, and only the target value and upper limit value are described, but they are different values respectively.

[0096] As described above, the upper limit value is set because when the conversation actually gets lively, each of the tone of voice, clarity of expression, and size of nodding may become excessively higher than the target value.

[0097] However, since it is not necessary to forcibly suppress the fact that the conversation is actually getting lively, when the button image 156 is turned on at the discretion of the operator, regardless of the actual time from the start of the conversation, the number of the setting table to be referred to is advanced to 72 (360 seconds from the start of the conversation), and thereafter, the target value and upper limit value are set every 5 seconds so that the number of the setting table to be referred to increases by 1 each time. In this case, after the target value and upper limit value corresponding to the number 120 (600 seconds from the start of the conversation) of the setting table (that is, the last target value and upper limit value) are set, the target value and upper limit value are not updated until the end of the conversation. However, by setting the target value and upper limit value for 600 seconds or more in the setting table, it is also possible to update the target value and upper limit value.

[0098] On one hand, a video including the user is captured by the camera 36, and video data corresponding to the video captured by the camera 36, that is, user video data, is transmitted from the user terminal 12 to the operator terminal 16. Also, the user's voice is detected by the microphone 32, and voice data corresponding to the voice detected by the microphone 32, that is, user voice data, is transmitted from the user terminal 12 to the operator terminal 16. The user video data is output to the display device 60 of the operator terminal 16, and the user voice data is output from the speaker 64 of the operator terminal 16. Therefore, the operator can control the avatar and respond to the user while seeing and hearing the user's state.

[0099] FIG. 7 shows an example of the operator-side dialogue screen 150 displayed on the display device 60 of the operator terminal 16. As shown in FIG. 7, on the operator-side dialogue screen 150, a display area 152 is provided slightly above the center of the screen, and button images 154 and 156 are provided below the display area 152.

[0100] The display area 152 is an area for outputting user video data. The video of the user with whom the operator responds is displayed.

[0101] The button image 154 is an icon for setting the tone height of the voice, the clarity of the facial expression, and the size of the nod to the target value and the upper limit value at the end of the dialogue according to the operator's intention.

[0102] The button image 156 is an icon for ending the user's response. When the button image 158 is turned on, a response end notification is transmitted to the user terminal 12, and the video call with the user terminal 12 is ended.

[0103] FIG. 8 shows an example of the memory map 300 of the RAM 22 built in the user terminal 12. As shown in FIG. 8, the RAM 22 includes a program storage area 302 and a data storage area 304. An information processing program executed by the user terminal 12 of this first embodiment is stored in the program storage area 302.

[0104] The information processing program executed on the user terminal 12 includes a main processing program 302a, an operation detection program 302b, a communication program 302c, an image generation program 302d, an image output program 302e, a shooting program 302f, an avatar control program 302g, a voice detection program 302h, and a voice output program 302i, etc.

[0105] The main processing program 302a is a program for executing the main routine of the information processing of the user terminal 12 in this first embodiment.

[0106] The operation detection program 302b is a program for detecting operation data 304a input from the input device 28 according to the user's operation and storing it in the data storage area 304.

[0107] The communication program 302c is a program for communicating (transmitting and receiving data, etc.) with an external device, which is the operator terminal 16 in this first embodiment.

[0108] The image generation program 302d is a program for generating image data corresponding to all or part of a screen (such as 100) to be displayed on the display device 30 using the image generation data 304b.

[0109] The image output program 302e is a program for outputting the image data generated according to the image generation program 302d to the display device 30. Therefore, the screen corresponding to the image data is displayed on the display device 30.

[0110] The shooting program 302f is a program for causing the camera 36 to execute a shooting process and storing the user video data 304c input from the camera 36 in the data storage area 304.

[0111] The avatar control program 302g is a program for operating the avatar using the motion data 304e.

[0112] The voice detection program 302h is a program for detecting the user's voice input from the microphone 32 and storing the user voice data 304d corresponding to the detected voice in the data storage area 304.

[0113] The voice output program 302i is a program for outputting the operator voice data 304f received from the operator terminal 16 to the speaker 34. Therefore, the voice corresponding to the operator voice data 304f or the voice obtained by converting this voice into the voice of the avatar comes out from the speaker 34.

[0114] Although not shown, in addition to the operating system and middleware of the user terminal 12, a browser and other application programs other than the information processing program of the present application are also stored in the program storage area 302.

[0115] The data storage area 304 stores operation data 304a, image generation data 304b, user video data 304c, user voice data 304d, motion data 304e, operator voice data 304f, and the like.

[0116] The operation data 304a is data of the user's operation detected according to the operation detection program 302b. When the operation data 304a is used for the processing of the CPU 20, it is deleted from the data storage area 304.

[0117] The image generation data 304b is data for generating the screen displayed on the display device 30 of the user terminal 12 and includes data for generating the avatar image 110.

[0118] The user video data 304c is data of the video captured by the camera 36. Basically, this user video data 304c is data of a video including the user, but when the user leaves or goes out of the shooting range of the camera 36, the video may not include the user's image. When the user video data 304c is transmitted by the CPU 20 to the operator terminal 16, it is deleted from the data storage area 304.

[0119] The user voice data 304d is the data of the user's voice detected by the microphone 32. When the user voice data 304d is transmitted to the operator terminal 16 by the CPU 20, it is erased from the data storage area 304.

[0120] The operation data 304e is the operation data received from the operator terminal 16. When the operation data 304e is used for the processing of the CPU 20, it is erased from the data storage area 304.

[0121] The operator voice data 304f is the operator voice data received from the operator terminal 16. When the operator voice data 304f is output to the speaker 34 by the CPU 20, it is erased from the data storage area 304.

[0122] Although not shown, other data necessary for executing information processing, a timer (counter) and a flag necessary for executing information processing are stored in the data storage area 304.

[0123] FIG. 9 shows an example of the memory map 400 of the RAM 52 built in the operator terminal 16. As shown in FIG. 9, the RAM 52 includes a program storage area 402 and a data storage area 404. A control program executed by the operator terminal 16 of this first embodiment is stored in the program storage area 402.

[0124] The control program executed by the operator terminal 16 includes a main processing program 402a, an operation detection program 402b, a communication program 402c, an image generation program 402d, an image output program 402e, a photographing program 402f, an expression estimation program 402g, a nodding detection program 402h, an operation data generation program 402i, a voice detection program 402j, a tone detection program 402k, a tone setting program 402m, a voice output program 402n, and a setting program 402p, etc.

[0125] The main processing program 402a is a program for executing the main routine of the control processing of the operator terminal 16 in this first embodiment.

[0126] The operation detection program 402b is a program for detecting operation data 404a input from the input device 58 according to the operation of the operator and storing it in the data storage area 404.

[0127] The communication program 402c is a program for communicating with an external device, which is the user terminal 12 in this first embodiment.

[0128] The image generation program 402d is a program for generating image data corresponding to the screen to be displayed on the display device 60 using the image generation data 404b. However, when generating image data corresponding to the operator-side interaction screen 150 as shown in FIG. 7, the user video data 404m is also used.

[0129] The image output program 402e is a program for outputting the image data generated according to the image generation program 402d to the display device 60. Therefore, the screen corresponding to the image data is displayed on the display device 60.

[0130] The photographing program 402f is a program for causing the camera 66 to execute a photographing process and storing the operator video data 404c input from the camera 66 in the data storage area 404.

[0131] The facial expression estimation program 402g is a program for estimating or detecting the facial expression of the operator and its clarity from the operator video data 404c photographed according to the photographing program 402f.

[0132] The nodding detection program 402h is a program for determining whether the operator has nodded from the operator video data 404c photographed according to the photographing program 402f, and detecting the angle if it is determined that the operator has nodded.

[0133] The operation data generation program 402i is a program for detecting the operation of an operator from the operator video data 404c and generating operation data 404h for controlling the operation of an avatar. However, the operation data generation program 402i generates the operation data 404h based on the clarity of the detected or set facial expression and the magnitude of nodding.

[0134] The voice detection program 402j is a program for detecting the voice of an operator input from the microphone 62 and storing the operator voice data 404d corresponding to the detected voice in the data storage area 404.

[0135] The tone detection program 402k is a program for extracting or detecting the tone of the operator's voice from the operator voice data 404d corresponding to the operator's voice detected according to the voice detection program 402j.

[0136] The tone setting program 402m is a program for setting or changing to a target value when the tone of the operator's voice does not exceed the target value, or setting or changing to an upper limit value when the tone of the operator's voice exceeds the upper limit value.

[0137] The voice output program 402n is a program for outputting the user voice data 404n received from the user terminal 12 to the speaker 64. Therefore, the voice corresponding to the user voice data 404n comes out from the speaker 64.

[0138] Although illustration is omitted, in addition to the operating system and middleware of the operator terminal 16, a browser and other application programs other than the control program of the present application are also stored in the program storage area 402.

[0139] Figure 10 shows an example of the specific content of the data storage area 404 shown in Figure 9. The data storage area 404 stores operation data 404a, image generation data 404b, operator video data 404c, operator voice data 404d, tone data 404e, expression data 404f, nodding data 404g, motion data 404h, setting table data 404i, target value data 404j, upper limit value data 404k, user video data 404m, user voice data 404n, and the like.

[0140] The operation data 404a is data on the operations of the operator detected according to the operation detection program 402b. When the operation data 404a is used for the processing of the CPU 50, it is deleted from the data storage area 404.

[0141] The image generation data 404b is data used to generate the screen displayed on the display device 60 of the operator terminal 16.

[0142] The operator video data 404c is data on the video captured by the camera 66. This operator video data 404c is basically data on the video including the operator, but when the operator leaves the seat or goes out of the shooting range of the camera 66, the video may not include the operator. When the operator video data 404c is used for the processing of the CPU 50, it is deleted from the data storage area 404.

[0143] The operator voice data 404d is data on the voice of the operator detected by the microphone 62 or the voice of the operator who sets (changes) the tone to the target value or the upper limit value. When the operator voice data 404d is transmitted to the user terminal 12 by the CPU 50, it is deleted from the data storage area 404.

[0144] The tone data 404e is data on the tone of the voice of the operator extracted or detected according to the tone detection program 402k.

[0145] The facial expression data 404f is data about the operator's facial expression estimated or detected according to the facial expression estimation program 402g and the clarity thereof.

[0146] The nodding data 404g is data about the angle detected when it is determined that the operator has nodded according to the nodding detection program 402h. Therefore, when it is not determined that the operator has nodded according to the nodding detection program 402h, the nodding data 404g is not stored.

[0147] The motion data 404h is data for controlling the motion of the avatar. In this first embodiment, it is information (data) about the position and orientation of each part for controlling the body language, face orientation, and facial expression of the avatar. Specifically, as described above, the motion data 404h is the position of each part of the face, namely the eyebrows, eyes (upper eyelids, lower eyelids), nose, and mouth (upper lip, lower lip), the position of the head, hands, arms, shoulders, waist, knees, and ankles, and information about the orientation and degree of bending (i.e., angle) of the neck (or face), wrists, and ankles. When the motion data 404h is transmitted to the user terminal 12 by the CPU 50, it is deleted from the data storage area 404.

[0148] The setting table data 404i is data of a setting table about the target values of each of the tone of voice, clarity of facial expression, and magnitude of nodding for each predetermined time from the start to the end of the dialogue.

[0149] The target value data 404j is data about the current target values of each of the tone of voice, clarity of facial expression, and magnitude of nodding.

[0150] The upper limit value data 404k is data about the current upper limit values of each of the tone of voice, clarity of facial expression, and magnitude of nodding. However, the current upper limit value is determined according to the current target value.

[0151] The user video data 404m is the user video data received from the user terminal 12. When the user video data 404m is used for the processing by the CPU 50, it is deleted from the data storage area 404.

[0152] The user audio data 404n is the user audio data received from the user terminal 12. When the user audio data 404n is output to the speaker 64 by the CPU 50, it is deleted from the data storage area 404.

[0153] Although illustration is omitted, other data necessary for executing control processing is stored in the data storage area 404, or a timer (counter) and a flag necessary for executing control processing are provided.

[0154] FIG. 11 is a flowchart showing the information processing of the CPU 20 of the user terminal 12. Although illustration is omitted, the CPU 20 executes a process of detecting operation data and storing the operation data, a process of causing the camera 36 to execute a shooting process and storing the user video data, a process of executing a voice detection process and storing the user audio data, or a process of receiving and storing data transmitted from the operator terminal 16 in parallel with the information processing. Hereinafter, the information processing of the CPU 20 will be described with reference to FIG. 11, but duplicate contents will be briefly described.

[0155] As shown in FIG. 11, when starting the information processing, the CPU 20 of the user terminal 12 establishes a connection state with the operator terminal 16 in step S1, and in step S3, displays the user-side dialogue screen 100 as shown in FIG. 4 on the display device 30. At this time, since the user terminal 12 has not received the operation data 304e, as an example, the dialogue screen 100 including the upright avatar image 110 is displayed. Also, at this time, the avatar may perform an operation (and speech) of greeting the user.

[0156] In the next step S5, it is determined whether the conversation has ended. Here, the CPU 20 determines whether there is a notification of the end of the conversation from the operator terminal 16. If it is "YES" in step S5, that is, if the conversation has ended, the information processing is terminated. On the other hand, if it is "NO" in step S5, that is, if the conversation has not ended, in step S7, it is determined whether data has been received from the operator terminal 16. Here, the CPU 20 determines whether the operation data 304e, or the operation data 304e and the operator voice data 304f have been received.

[0157] If it is "NO" in step S7, that is, if data has not been received from the operator terminal 16, the process proceeds to step S17. On the other hand, if it is "YES" in step S7, that is, if data has been received from the operator terminal 16, in step S9, it is determined whether there is operator voice data 304f.

[0158] If it is "YES" in step S9, that is, if there is operator voice data 304f, in step S11, the operator voice data 304f is output to the speaker 34, and in step S13, in accordance with the output of the operator voice data 304f, image data of an avatar that performs an operation according to the operation data 304e is generated and output to the display device 30, and then the process proceeds to step S17. Therefore, the avatar image 110 is updated, and the avatar displayed on the user-side conversation screen 100 talks to the user while performing gestures or changing the face direction. At this time, the avatar image 110 is lip-synced to the voice output from the speaker 34. Hereinafter, the same applies to the case where the avatar is operated in accordance with the voice data.

[0159] On the other hand, if it is "NO" in step S9, that is, if there is no operator voice data 304f, in step S15, image data of an avatar that performs an operation according to the operation data 304e is generated and output to the display device 30, and then the process proceeds to step S17. Therefore, the avatar image 110 is updated, and the avatar displayed on the user-side conversation screen 100 performs gestures or changes the face direction.

[0160] In step S17, it is determined whether there is voice input. If it is "YES" in step S17, that is, if there is voice input, in step S19, the user video data 304c and the user voice data 304d are transmitted to the operator terminal 16, and the process returns to step S5. On the other hand, if it is "NO" in step S17, that is, if there is no voice input, in step S21, the user video data 304c is transmitted to the operator terminal 16, and the process returns to step S5.

[0161] Figures 12 - 15 are flowcharts showing the control processing of the CPU 50 of the operator terminal 16. Although the illustration is omitted, the CPU 50 executes a detection process of operation data and stores the operation data in parallel with the control process, executes a shooting process on the camera 66 to store the operator video data, executes a detection process of voice to store the operator voice data, or executes a process of receiving and storing data transmitted from the user terminal 12. Hereinafter, the control processing of the CPU 50 of the operator terminal 16 will be described, but duplicate content will be briefly described.

[0162] As shown in Figure 12, when the CPU 50 starts the control process, in step S101, the connection state with the user terminal 12 is established. At this time, a timer that counts the first predetermined time is reset and started. In the subsequent step S103, the operator-side dialogue screen 150 as shown in Figure 5 is displayed on the display device 60. However, at the beginning of starting the control process, since the user video data has not been received, the user's video is not displayed in the display area 152.

[0163] In the next step S105, the target values and upper limit values of the tone of the avatar's voice, the clarity of the expression, and the size of the nod are set to the initial values. Here, the CPU 50 refers to the setting table data 404i and sets the target values of the tone of the avatar's voice, the clarity of the expression, and the size of the nod at the start of the dialogue, and also sets the upper limit values corresponding to the respective target values. At this time, a timer that counts the second predetermined time is reset and started.

[0164] Subsequently, in step S107, it is determined whether the response has ended. Here, the CPU 50 determines whether the first predetermined time (for example, 600 seconds) has elapsed since the start of the dialogue or whether the button image 158 has been turned on. If it is "YES" in step S107, that is, if the response has ended, the control process ends. At this time, the CPU 50 sends a notification of the end of the response to the user terminal 12.

[0165] On the other hand, if it is "NO" in step S107, that is, if the response has not ended, in step S109 shown in FIG. 13, the expression of the operator and its clarity are estimated or detected, and in step S111, the size of the operator's nod is detected. However, in step S109, the CPU 50 estimates the expression of the operator and its clarity by estimating the emotion of the operator and its intensity.

[0166] In the subsequent step S113, it is determined whether the expression of the operator is a smiling face. If it is "NO" in step S113, that is, if the expression of the operator is not a smiling face, the process proceeds to step S123.

[0167] On the other hand, if it is "YES" in step S113, that is, if the expression of the operator is a smiling face, in step S115, it is determined whether the clarity of the smiling face exceeds the target value. If it is "NO" in step S115, that is, if the clarity of the smiling face does not exceed the target value, in step S117, the clarity of the smiling face is set to the target value, and the process proceeds to step S123.

[0168] On the other hand, if it is "YES" in step S115, that is, if the clarity of the smiling face exceeds the target value, in step S119, it is determined whether the clarity of the smiling face exceeds the upper limit value. If it is "NO" in step S119, that is, if the clarity of the smiling face does not exceed the upper limit value, the process proceeds to step S123.

[0169] On the one hand, if it is "YES" in step S119, that is, if the clarity of the smiling face exceeds the upper limit value, then in step S121, the clarity of the smiling face is set to the upper limit value, and the process proceeds to step S123.

[0170] In step S123, it is determined whether the size of the nodding exceeds the target value. If it is "NO" in step S123, that is, if the size of the nodding does not exceed the target value, then in step S125, the size of the nodding is set to the target value, and the process proceeds to step S131 shown in FIG. 14.

[0171] On the other hand, if it is "YES" in step S123, that is, if the size of the nodding exceeds the target value, then in step S127, it is determined whether the size of the nodding exceeds the upper limit value. If it is "NO" in step S127, that is, if the size of the nodding does not exceed the upper limit value, then the process proceeds to step S131.

[0172] On the other hand, if it is "YES" in step S127, that is, if the size of the nodding exceeds the upper limit value, then in step S129, the size of the nodding is set to the upper limit value, and the process proceeds to step S131.

[0173] As shown in FIG. 14, in step S131, it is determined whether the operator's expression is a smiling face. If it is "YES" in step S131, then in step S133, operation data is generated from the operator video data with the detected or set clarity of the smiling face and the size of the nodding, and the process proceeds to step S137.

[0174] If it is "NO" in step S131, that is, if the operator's expression is not a smiling face, then in step S135, operation data is generated from the operator video data with the detected expression and the detected or set size of the nodding, and the process proceeds to step S137.

[0175] In step S137, it is determined whether there is voice input. If it is "NO" in step S137, that is, if there is no voice input, in step S139, the operation data 404h generated in step S133 or step S135 is transmitted to the user terminal 12, and the process proceeds to step S153 shown in FIG. 15.

[0176] On the other hand, if it is "YES" in step S137, that is, if there is voice input, in step S141, the tone of the operator's voice is detected, and in step S143, it is determined whether the voice tone exceeds the target value.

[0177] If it is "NO" in step S143, that is, if the voice tone does not exceed the target value, in step S145, the voice tone is set to the target value, and the process proceeds to step S151.

[0178] On the other hand, if it is "YES" in step S143, that is, if the voice tone exceeds the target value, in step S147, it is determined whether the voice tone exceeds the upper limit value. If it is "NO" in step S147, that is, if the voice tone does not exceed the upper limit value, the process proceeds to step S151.

[0179] On the other hand, if it is "YES" in step S147, that is, if the voice tone exceeds the upper limit value, in step S149, the voice tone is set to the upper limit value, and the process proceeds to step S151.

[0180] In step S151, the operation data 404h generated in step S133 or step S135 and the operator voice data 404d of the detected or set tone are transmitted to the user terminal 12, and the process proceeds to step S153.

[0181] As shown in FIG. 15, in step S153, it is determined whether data has been received from the user terminal 12. If the result in step S153 is "NO", that is, if data has not been received from the user terminal 12, the process proceeds to step S161. On the other hand, if the result in step S153 is "YES", that is, if data has been received from the user terminal 12, then in step S155, it is determined whether there is user voice data 404n.

[0182] If the result in step S155 is "YES", that is, if there is user voice data 404n, then in step S157, the user video data 404m and the user voice data 404n are output, and the process proceeds to step S161. On the other hand, if the result in step S155 is "NO", that is, if there is no user voice data 404n, then in step S159, the user video data 404m is output, and the process proceeds to step S161.

[0183] In step S161, after updating (or setting) the target value and the upper limit value, it is determined whether the second predetermined time (for example, 5 seconds) has elapsed. If the result in step S161 is "NO", that is, if the second predetermined time has not elapsed after updating the target value and the upper limit value, the process returns to step S107.

[0184] On the other hand, if the result in step S161 is "YES", that is, if the second predetermined time has elapsed after updating the target value and the upper limit value, then in step S163, the target value and the upper limit value are set (i.e., updated) to the target value and the upper limit value described in the next number of the setting table, and the process returns to step S107. At this time, the timer that counts the second predetermined time is reset and started.

[0185] According to the first embodiment, when interacting with a user using an avatar, the degree of the tone of the avatar's voice, the clarity of the smiling face, and the size of the nodding are gradually increased from the start to the end of the interaction, so that the atmosphere of the interaction can be enlivened. Also, in this first embodiment, since the degree of clarity of the avatar's smiling face is gradually increased, the user interacting with the avatar naturally feels excited about the interaction. As a result, the user is more likely to agree with the opinions of the avatar operator. For this reason, when gradually increasing the degree of clarity of the smiling face, the interaction can be enlivened, and thus it is considered that the operator can smoothly proceed with the response to the user.

[0186] As described above, in this first embodiment, the degrees of the tone of the avatar's voice, the clarity of the smiling face, and the size of the nodding are increased from the start to the end of the interaction. However, it is considered that the atmosphere of the interaction can be enlivened if at least one of the degrees is increased from the start to the end of the interaction. Also, in this first embodiment, since the degree of clarity of the avatar's smiling face is gradually increased, the user interacting with the avatar naturally feels excited about the interaction. As a result, the user is more likely to agree with the opinions of the avatar operator. For this reason, when gradually increasing the degree of clarity of the smiling face, the interaction can be enlivened, and thus it is considered that the operator can smoothly proceed with the response to the user.

[0187] However, it is also possible to increase the degrees of any two of the tone of the avatar's voice, the clarity of the smiling face, and the size of the nodding from the start to the end of the interaction. That is, the degrees of the tone of the avatar's voice and the clarity of the smiling face may be increased from the start to the end of the interaction, or the degrees of the tone of the avatar's voice and the size of the nodding may be increased from the start to the end of the interaction, or the degrees of the clarity of the smiling face and the size of the nodding may be increased from the start to the end of the interaction.

[0188] However, similar to not enhancing the clarity of expressions other than the smiling expression, when the operator's emotion is other than joy, the degrees of the tone of voice and the amplitude of nodding may not be increased.

[0189] Also, in this first embodiment, the degrees of the non-verbal information such as the tone of the avatar's voice, the smiling expression, and nodding are increased from the start to the end of the dialogue. However, it is not necessarily limited. Non-verbal information also includes body movements such as gestures other than nodding. That is, in this first embodiment, by increasing the degree of at least one piece of non-verbal information of the avatar when the operator responds to the user from the start to the end of the dialogue, the atmosphere of the dialogue can be enlivened.

[0190] Therefore, for body movements such as gestures, the degree of the movement of the body and the movement (angle) of the wrist is increased (made larger) from the start to the end of the dialogue.

[0191] Also, instead of or together with the tone of voice, the degree of the volume of the voice may be increased (made larger) from the start to the end of the dialogue. Similarly, instead of or together with the amplitude of nodding, the frequency of nodding may be increased from the start to the end of the dialogue.

[0192] Note that the size of the avatar's gestures, the volume of the voice, and the frequency of nodding can also be set in a plurality of stages from the minimum to the maximum, and numerical values are set for each stage. When the operator responds to the user, the size of the avatar's gestures, the volume of the voice, and the frequency of nodding are controlled to change according to the curve shown in FIG. 5(C).

[0193] Also, in this first embodiment, the clarity degree of the avatar's smiling face is gradually increased, but it is not necessary to be limited to this. In order to liven up the conversation scene, the clarity of facial expressions other than smiling may also be gradually increased. Therefore, for example, when the emotion of anger intensifies, the conversation scene may become lively. That is, not only positive emotions but also negative emotions can liven up the conversation scene. However, facial expressions other than smiling are, for example, expressions that represent emotions of fear, sadness, disgust, and anger.

[0194] Also, in this first embodiment, the degrees of the avatar's voice tone, the clarity of the smiling face, and the size of nodding are increased from the start to the end of the conversation according to one curve. However, separate curves with different changing manners may be prepared, and they may be increased from the start to the end of the conversation according to the corresponding curves. This also applies to the case where the degrees of other non-verbal information are increased from the start to the end of the conversation.

[0195] Furthermore, in this first embodiment, when the operator operates the button image, the setting table to be referred to is moved all at once to the start position of the final stage of the conversation, but it is not necessary to be limited to this. The number of the setting table to be referred to may be skipped by a predetermined number (for example, one or two) until the number of the setting table to be referred to becomes the start position of the final stage.

[0196] Alternatively, instead of the operator operating the button image, the setting table to be referred to may be automatically moved all at once to the start position of the end stage of the dialogue. The emotion and its intensity of the user are estimated from the face image of the user included in the user video data, or the tone of the user's voice is detected from the user voice data, and the degree of excitement of the dialogue scene is judged. When the degree of excitement of the dialogue scene is equal to or higher than a certain level, the setting table to be referred to may be moved all at once to the start position of the end stage of the dialogue. Also in such a case, the setting table to be referred to may be referred to skipping a predetermined number (for example, one or two) until the number of the setting table becomes the start position of the end stage. The method for estimating the emotion and its intensity of the user and the method for detecting the tone of the user's voice are the same as the method for estimating the emotion and its intensity of the operator and the method for detecting the tone of the operator's voice described in the first embodiment, respectively. Also, the degree of excitement of the dialogue scene can be judged by the intensity of the user's emotion or the height of the tone of the user's voice. In another example, the intensity of the user's emotion or the height of the tone of the user's voice at the start of the dialogue is memorized, and when the intensity of the current user's emotion becomes higher than the intensity of the user's emotion at the start of the dialogue by a certain degree or more, or when the height of the tone of the current user's voice becomes higher than the height of the tone of the user's voice at the start of the dialogue by a certain degree or more, it can also be judged that the degree of excitement of the dialogue scene is equal to or higher than a certain level.

[0197] Furthermore, in this first embodiment, based on the operator's image and voice, the degrees of the tone of the avatar's voice, the clarity of the smiling face, and the size of the nodding are increased from the start to the end of the dialogue according to one curve. However, based on the user's video and voice, the degrees of the tone of the avatar's voice, the clarity of the smiling face, and the size of the nodding can also be increased from the start to the end of the dialogue. In this case, the tone of the user's voice is detected from the user voice data, and the tone of the avatar's voice with a tone slightly higher than that of the user's voice is set. Also, the user's emotion and its intensity are estimated from the user's face image included in the user video data, and the facial expression of the avatar with a slightly higher degree than the estimated user's emotion and its intensity is set. Furthermore, the size of the user's nodding is detected from the user video data, and the size of the nodding with a slightly higher degree than the detected size of the nodding is set. At the start of the dialogue, the degrees of the tone of the avatar's voice, the clarity of the smiling face, and the size of the nodding are increased by multiplying the detected or estimated value by a first coefficient (for example, 1.02). From the start of the dialogue to the start position of the end stage of the dialogue, the first coefficient is increased by a second predetermined ratio (for example, 10%) every predetermined time (for example, 5 seconds). Also, at the start position of the end stage of the dialogue, the degrees of the tone of the avatar's voice, the clarity of the smiling face, and the size of the nodding are increased by multiplying the detected or estimated value by a second coefficient (for example, 1.1). After the start position of the end stage of the dialogue, the second coefficient is increased by a third predetermined ratio (for example, 30%) every predetermined time (for example, 5 seconds).

[0198] In addition, in this first embodiment, the operation data of the avatar is generated by the operator terminal, and the operation data is transmitted to the user terminal to generate the avatar image by the user terminal. However, the avatar image may be generated by the operator terminal and the avatar image may be transmitted to the user terminal.

[0199] Also, in this first embodiment, the operation of the avatar is controlled using the motion data generated from the video data of the operator being photographed, but it is not necessary to be limited to this. Motion data corresponding to a series of operations to be executed by the avatar may be prepared in advance, and the motion data corresponding to the series of operations instructed by the operator through button operations or the like may be transmitted to the user terminal to control the operation of the avatar. Examples of a series of operations include an operation of bowing and returning to the original posture (for example, an upright state), an operation of turning the face left and right several times to look around and then returning to the original posture, an operation of waving the right hand or / and the left hand in front of the body and then returning to the original posture, an operation of raising the right hand or / and the left hand and then returning to the original posture, and the like.

[0200] Also, in this first embodiment, the avatar is made to speak based on the voice of the operator, but voice data of the content to be spoken may be prepared in advance, and the voice data corresponding to the content to be spoken instructed by the operator through button operations or the like may be transmitted to the user terminal to control the speech of the avatar. However, the voice data of the content to be spoken may be set to one button together with the motion data corresponding to a series of operations of the avatar, and the motion data and the voice data may be transmitted to the user terminal at once.

[0201] In this way, even when performing button operations, the tone of voice, the clarity of the smiling face, and the size of the nodding are appropriately set using the setting table.

[0202] Furthermore, in this first embodiment, the video of the user is displayed on the display device of the operator terminal, but an avatar image may also be displayed on the operator terminal.

[0203] <Second Embodiment> The information processing system 10 of the second embodiment is the same as that of the first embodiment except that the avatar is controlled using artificial intelligence. Therefore, different contents will be described, and the description of overlapping contents will be omitted.

[0204] Therefore, in this specification, the avatar includes not only those operated by remote control by a human, but also those operated by automatic control using artificial intelligence.

[0205] Also, in the information processing system 10 of the second embodiment, the operator who controls the avatar is a program using artificial intelligence for controlling the avatar (hereinafter referred to as the "dialogue program"). When a human operator using the operator terminal 16 completely automatically controls the avatar without checking the excitement of the conversation scene, on the operator terminal 16, it is not necessary to display the operator-side dialogue screen 150, and it is not necessary to output the user's voice from the speaker 64. Also, in this case, since it is not necessary to transmit the user's video to the operator terminal 16, the camera 36 of the user terminal 12 is turned off. In some cases, the camera 36 may not be provided.

[0206] In the second embodiment, the case where the operator-side dialogue screen 150 is not displayed on the operator terminal 16 and the user's voice is not output from the speaker 64 will be described, but it is also possible to display the operator-side dialogue screen 150 and output the user's voice from the speaker 64. In this way, a human operator using the operator terminal 16 can check the excitement of the conversation scene. Also, in this case, the voice of the avatar's utterance content is output from the speaker 64, or the text of the avatar's utterance content is displayed on the operator-side dialogue screen 150.

[0207] In addition, in the operator terminal 16, an interaction unit for the avatar to automatically interact with the user is connected to circuit components such as the CPU 50 via an internal bus. The interaction unit automatically controls the avatar according to an interaction program using artificial intelligence (AI) such as a chatbot. Further, the interaction unit includes a voice recognition unit that recognizes the user's voice, a purpose interpretation unit that interprets the user's purpose based on the content indicated by the voice recognized by the voice recognition unit, a voice generation unit that generates speech content, that is, the voice of the avatar, according to a talk script from the interpreted purpose, and an action generation unit that generates the action of the avatar according to the generated voice (speech content) or the situation where the voice has not been generated (in this embodiment, selects the action data from pre-generated interaction action data). Therefore, when the avatar is completely automatically controlled by the interaction unit without controlling the avatar according to the operation of a human operator, the microphone 62 and the camera 66 do not have to be provided in the operator terminal 16.

[0208] Note that when the human operator and the interaction unit alternately control the avatar, when the human operator controls the avatar, the microphone 62 and the camera 66 are used as described in the first embodiment.

[0209] In the second embodiment, the case where the interaction unit is provided in the operator terminal 16 will be described. However, the interaction unit may be provided on the user terminal 12 or the network 14 (cloud).

[0210] In automatic control, the avatar either converses according to a talk script or operates according to operation data corresponding to actions during pre-generated conversations (hereinafter referred to as "conversation operation data"). As an example, the talk script and conversation operation data of the avatar are generated by machine learning the speech and actions of a human operator in response to the user's utterance. However, the conversation operation data includes operation data corresponding to multiple actions of the avatar. The text data of the avatar's speech content according to the talk script is converted into the avatar's voice data (hereinafter referred to as "avatar voice data") and transmitted to the user terminal 12. That is, in the second embodiment, instead of the operator voice data, the avatar voice data is transmitted to the user terminal 12.

[0211] In the second embodiment, the case where the avatar converses with the user according to a talk script will be described. However, it is also possible to automatically generate the avatar's speech content from the user's voice using a large language model such as ChatGPT, which is a type of generative AI. Since a large language model requires a high-performance computer, the computer incorporating the large language model is provided on the cloud. However, by improving the functions of the user terminal 12 and the operator terminal 16, it is also possible to incorporate the large language model into the user terminal 12 or the operator terminal 16.

[0212] Therefore, in the second embodiment, the user video data 304c is not stored in the RAM 22. Also, in the second embodiment, the operator video data 404c and the user video data 404m are not stored in the RAM 52. Also, in the second embodiment, instead of the operator voice data 404d, the avatar voice data is stored.

[0213] In the second embodiment, the tone of the avatar's voice, the clarity of the avatar's smile, and the size of the avatar's nod generated by the dialogue unit may be set to target values or upper limit values using a setting table.

[0214] However, since the tone of the avatar's voice is usually set to be constant, in this second embodiment, it is set to a target value using a setting table. Also, by machine learning the tone of the human operator's speech in response to the user's speech, when the dialogue unit converts the text data of the avatar's speech content according to the talk script into the avatar's voice data, that is, when generating the avatar's voice data, it can also be set to the machine-learned tone. In such a case, the tone of the avatar's voice data generated by the dialogue unit may be detected and set to a target value or an upper limit value using the setting table.

[0215] Specifically, the CPU 50 of the operator terminal 16 executes the control process shown in FIGS. 16-19. Hereinafter, the control process of the second embodiment will be described, but the same content as that described in the first embodiment will be briefly described. Note that the information processing of the CPU 20 of the user terminal 12 is the same as that described in the first embodiment except that the process of step S21 is deleted and the user voice data 304d is transmitted to the operator terminal 16 in step S19, and thus the illustration and description are omitted.

[0216] FIGS. 16-19 are flowcharts showing the control process of the CPU 50 of the operator terminal 16. Although the illustration is omitted, the CPU 50 executes a process of receiving and storing data and the like transmitted from the user terminal 12 in parallel with the control process.

[0217] As shown in FIG. 16, when starting the control process, the CPU 50 establishes a connection state with the user terminal 12 in step S201. In the subsequent step S203, the target values and upper limit values of the tone of the avatar's voice, the clarity of the expression, and the magnitude of the nodding are set to initial values.

[0218] Subsequently, in step S205, it is determined whether the response has ended. Here, the CPU 50 determines whether the first predetermined time (for example, 600 seconds) has elapsed since the start of the dialogue or the dialogue unit has determined to end the dialogue.

[0219] If it is “YES” in step S205, that is, if the response has ended, the control process ends. At this time, the CPU 50 sends a notification of the end of the response to the user terminal 12. On the other hand, if it is “NO” in step S205, that is, if the response has not ended, in step S207, operation data is selected by the dialogue unit, and in step S209, avatar voice is generated according to the dialogue script by the dialogue unit, and the process proceeds to step S211 shown in FIG. 17.

[0220] However, in the first time, since the user voice has not been received yet, operation data of the avatar's greeting operation to the user is selected, and greeting avatar voice data is generated. Also, even if the user voice has been received, there may be cases where operation data is not selected and / or avatar voice data is not generated.

[0221] As shown in FIG. 17, in step S211, the expression and clarity of the avatar based on the operation data selected in step S207 are estimated or detected, and in step S213, the size of the avatar's nodding based on the operation data selected in step S207 is detected. However, in step S211, the CPU 50 estimates or detects the facial expression and its clarity based on the parameters of each part of the avatar's face included in the operation data. Also, in step S213, the CPU 50 detects the size of the nodding based on the degree of bending of the avatar's neck included in the operation data.

[0222] In the subsequent step S215, it is determined whether the expression of the avatar is a smiling face. If it is “NO” in step S215, that is, if the expression of the avatar is not a smiling face, the process proceeds to step S225.

[0223] On the other hand, if it is “YES” in step S215, that is, if the expression of the avatar is a smiling face, in step S217, it is determined whether the clarity of the smiling face exceeds the target value. If it is “NO” in step S217, in step S219, the clarity of the smiling face is set to the target value, and the process proceeds to step S225.

[0224] On the one hand, if it is "YES" in step S217, in step S221, it is determined whether the clarity of the smiling face exceeds the upper limit value. If it is "NO" in step S221, the process proceeds to step S225. On the other hand, if it is "YES" in step S221, in step S223, the clarity of the smiling face is set to the upper limit value, and the process proceeds to step S225.

[0225] In step S225, it is determined whether the size of the nodding exceeds the target value. If it is "NO" in step S225, in step S227, the size of the nodding is set to the target value, and the process proceeds to step S233 shown in FIG. 18. On the other hand, if it is "YES" in step S225, in step S229, it is determined whether the size of the nodding exceeds the upper limit value.

[0226] If it is "NO" in step S229, the process proceeds to step S233. On the other hand, if it is "YES" in step S229, in step S231, the size of the nodding is set to the upper limit value, and the process proceeds to step S233.

[0227] As shown in FIG. 18, in step S233, it is determined whether the expression of the avatar is a smiling face. If it is "YES" in step S233, in step S235, the operation data is updated with the set clarity of the smiling face and the size of the nodding, and the process proceeds to step S239. However, if the detected clarity of the smiling face and the size of the nodding have not been changed, the process of step S235 is not performed.

[0228] On the other hand, if it is "NO" in step S233, in step S237, the operation data is updated with the set size of the nodding, and the process proceeds to step S239. However, if the detected size of the nodding has not been changed, the process of step S235 is not performed.

[0229] In step S239, it is determined whether there is generation of the avatar's voice. If it is "NO" in step S239, that is, if there is no generation of the avatar's voice, in step S245, the operation data 404h is transmitted to the user terminal 12, and the process proceeds to step S247 shown in FIG. 19.

[0230] On the other hand, if it is "YES" in step S239, that is, if there is generation of the avatar voice, in step S241, the tone of the avatar voice is set to the target value, and in step S243, the operation data 404h and the avatar voice data with the tone set to the target value are transmitted to the user terminal 12, and the process proceeds to step S247.

[0231] As shown in FIG. 19, in step S247, it is determined whether data has been received from the user terminal 12. If it is "NO" in step S247, the process proceeds to step S255. On the other hand, if it is "YES" in step S247, in step S249, it is determined whether there is user voice data 404n.

[0232] If it is "NO" in step S249, the process proceeds to step S255. On the other hand, if it is "YES" in step S249, in step S251, the voice of the user is recognized by the dialogue unit, in step S253, the purpose of the user is interpreted by the dialogue unit, and the process proceeds to step S255.

[0233] In step S255, after updating (or setting) the target value and the upper limit value, it is determined whether the second predetermined time (for example, 5 seconds) has elapsed. If it is "NO" in step S255, the process returns to step S205.

[0234] On the other hand, if it is "YES" in step S255, in step S257, the target value and the upper limit value are set, that is, updated to the target value and the upper limit value described in the next number of the setting table, and the process returns to step S205.

[0235] Also in the second embodiment, similar to the first embodiment, the atmosphere of the conversation can be enlivened. Also in the second embodiment, similar to the first embodiment, when increasing the degree of clarity of the smiling face, the user who is interacting with the avatar naturally feels enlivened in the conversation, and as a result, the operator can smoothly proceed with the response to the user.

[0236] In addition, in the second embodiment as well, the modification contents described in the first embodiment can be appropriately adopted.

[0237] Also, as described above, the human operator and the dialogue unit can also control the avatar alternately. Whether the human operator or the dialogue unit controls the avatar is determined or decided by the human operator or the dialogue unit. When the human operator controls the avatar, as described in the first embodiment, the user terminal 12 and the operator terminal 16 are controlled. When the dialogue unit controls the avatar, as described in the second embodiment, the user terminal 12 and the operator terminal 16 are controlled.

[0238] In the above-described embodiment, the user terminal and the operator terminal perform P2P communication, but they may communicate via a server provided on the network.

[0239] Furthermore, in the case where the same results can be obtained for each step of the flowcharts shown in the above-described embodiments, the processing order can be changed.

[0240] Furthermore, all of the various screens and specific numerical values given in the above-described embodiments are merely examples and can be appropriately changed as needed.

Explanation of Reference Numerals

[0241] 10... Information processing system 12... User terminal 14... Network 16... Operator terminal 20, 50... CPU 22, 52... RAM 24, 54... Communication I / F 26, 56... Input / output I / F 28, 58... Input device 30, 60... Display device 32, 62... Microphone 34, 64... Speaker 36, 38, 66... Camera

Claims

1. A control device for an operator who responds to a user using an avatar, comprising: avatar control means for controlling the actions and speech of the avatar; The avatar control means increases the degree of at least one piece of non-verbal information of the avatar when the operator responds to the user from the start to the end of the dialogue with the user. A control device that increases as it goes from the start to the end of the dialogue.

2. The control device according to claim 1, wherein the non-verbal information includes the clarity of the avatar's facial expression, the tone of the avatar's voice, and the nodding of the avatar.

3. The control device according to claim 2, wherein the facial expression of the avatar is a smiling face.

4. Further comprising a microphone for detecting the voice of the operator, The avatar control means generates voice data for controlling the speech of the avatar based on the voice of the operator detected by the microphone, and increases the tone of the avatar's voice as it goes from the start to the end of the dialogue. The control device according to claim 2.

5. Further comprising a camera for photographing the operator, The avatar control means generates motion data for controlling an action including the facial expression of the avatar based on the video of the operator photographed by the camera, and increases the clarity of the avatar's smiling face as it goes from the start to the end of the dialogue. The control device according to claim 1 or 2.

6. Further comprising a camera for photographing the operator, The avatar control means generates motion data for controlling the action of the avatar based on the video of the operator photographed by the camera, and the avatar control means increases the size of the avatar's nodding as it goes from the start to the end of the dialogue. The control device according to claim 2.

7. The control device according to claim 2, wherein the avatar control means increases the rate of increasing the degree of each of the tone of the avatar's voice, the clarity of the avatar's facial expression, and the size of the avatar's nodding in the final stage compared to the initial stage of the dialogue.

8. The control device according to claim 1, wherein the operator is a human or a program using artificial intelligence.

9. A control program executed by a control device for an operator who responds to a user using an avatar, The control program causes the processor of the control device to control the actions and speech of the avatar, A control program that increases the degree of at least one piece of non-verbal information of the avatar when the operator responds to the user from the start to the end of the dialogue with the user.

10. A control method for a control device of an operator who responds to a user using an avatar, wherein a processor of the control device controls the movement and speech of the avatar, and increases the degree of at least one piece of non-verbal information of the avatar when the operator responds to the user from the start to the end of the dialogue with the user.

Citation Information

Patent Citations

  • Information processing device, program, and information processing method

    JP6937534B1